Researchers have developed AT-ViT, a novel dual-branch Vision Transformer designed to improve plant trait recognition from herbarium images. This model addresses the challenge of background noise and spurious correlations by employing a multi-scale, multi-view cross-attention fusion scheme. AT-ViT also incorporates a mask-guided patch weighting mechanism to focus on plant-relevant regions, leading to significant accuracy gains and improved robustness against background perturbations compared to existing models like CrossViT and ResNet101. AI
IMPACT This model's approach to handling noisy data could inform future computer vision applications in fields with similar data challenges.
RANK_REASON The item is a research paper detailing a new model architecture and its performance on a specific task. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- AT-ViT
- CatalyzeX
- Connected Papers
- CrossVit: enhancing canopy monitoring management practices in viticulture.
- DagsHub
- Gotit.pub
- Hugging Face
- Litmaps
- ResNet101
- ScienceCast
- scite Smart Citations
- vision transformer
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →