PulseAugur
实时 09:44:40

新的Hi-Token方法提高了AI模型中的视觉基础准确性

研究人员开发了一种新颖的生成式视觉基础方法Hi-Token,通过分层标记坐标来提高边界框预测的准确性。该方法对百位、十位和个位数字进行编码,增加了结构并促进了视觉-语言模型中的标记重用。作为Hi-Token的补充,Hi-GAR在训练过程中使用基于几何的奖励来进一步提高定位精度。在多个视觉-语言模型骨干和基准上的实验表明,性能持续提升,其中Hi-R1与现有的专业方法相比取得了更优异的结果。 AI

影响 提高了AI模型在视觉基础任务中的定位准确性,可能增强需要通过文本描述精确识别对象的应用。

排序理由 该集群包含两篇arXiv论文,详细介绍了视觉基础领域的新研究和调查。

在 arXiv cs.CV 阅读 →

AI 生成摘要 · Google Gemini · 来自 2 个来源。 我们如何撰写摘要 →

新的Hi-Token方法提高了AI模型中的视觉基础准确性

报道来源 [2]

  1. arXiv cs.CV TIER_1 English(EN) · Xiuyuan Zhu, Ke Lu, Kun Dong, Siwen Jiao, Hao Wu, Zijin Du, Shun Mao, Dongming Zhang, Jian Xue ·

    Hi-Token: Hierarchical Coordinate Tokenization for Generative Visual Grounding

    arXiv:2608.03471v1 Announce Type: new Abstract: Generative Vision-Language Models (VLMs) commonly treat bounding-box coordinates as independent output symbols, leaving numerical order and axis semantics implicit. We identify this representation as an important source of error in …

  2. arXiv cs.CV TIER_1 English(EN) · Linhui Xiao, Xiaoshan Yang, Xiangyuan Lan, Yaowei Wang, Changsheng Xu ·

    Toward Visual Grounding: A Survey

    arXiv:2412.20206v4 Announce Type: replace Abstract: Visual Grounding, also known as Referring Expression Comprehension and Phrase Grounding, aims to ground the specific region(s) within the image(s) based on the given expression text. This task simulates the common referential re…