研究人员推出了一种新的文本引导开放词汇对象计数框架MambaCount,该框架利用空间稀疏状态空间对偶(S^4D)块来克服Transformer在处理密集场景和大尺度变化方面的局限性。MambaCount解决了Mamba中的因果建模问题和空间标记响应中的高熵问题,在线性复杂度下在FSC-147数据集上取得了最先进的性能。同时,RT-Counter为该任务提供了一个实时解决方案,通过视觉原型文本化模块和编织Transformer层来平衡准确性和效率,取得了具有竞争力的结果,同时速度更快、参数效率更高。此外,还提出了一个新的基准Robust-TOOC,用于评估在不利条件下的对象计数,以及Dual-TTT,一个旨在提高鲁棒性而不改变现有架构的测试时训练框架。
AI
arXiv:2606.17650v1 Announce Type: cross Abstract: Text-guided Open-vocabulary Object Counting (TOOC) aims to estimate the number of objects described by text prompts, which is particularly challenging in dense scenes with large scale variations. Existing TOOC approaches predomina…
Text-guided Open-vocabulary Object Counting (TOOC) aims to estimate the number of objects described by text prompts, which is particularly challenging in dense scenes with large scale variations. Existing TOOC approaches predominantly rely on Transformers, whose quadratic complex…
arXiv cs.CV
TIER_1English(EN)·Hao-Yuan Ma, Li Zhang, Zhiwei Zhu, Jie Gao·
arXiv:2606.17561v1 Announce Type: new Abstract: Text-guided open-vocabulary object counting (TOOC) aims to count objects belonging to the categories specified by natural language descriptions. Although vision-language pre-trained models have been successful applied to TOOC tasks,…
arXiv cs.CV
TIER_1English(EN)·Hao-Yuan Ma, Yuda Zou, Li Zhang, Yongchao Xu·
Text-guided Open-vocabulary Object Counting (TOOC) enables counting arbitrary object categories specified by text prompts, offering substantially greater flexibility than conventional closed-set counting. However, existing TOOC methods are developed and evaluated primarily on ide…
Text-guided open-vocabulary object counting (TOOC) aims to count objects belonging to the categories specified by natural language descriptions. Although vision-language pre-trained models have been successful applied to TOOC tasks, they still struggle with fine-grained spatial u…