PulseAugur
中
实时 22:18:50
English(EN) TextGaze: Prompting Gaze Target Estimation with Textual Scene Cues

TextGaze 使用 LVLM 进行注视点估计 · arXiv cs.CV

研究人员推出了一种新颖的注视点估计架构 TextGaze,该架构利用大型视觉语言模型 (LVLM) 进行语义引导。该方法旨在克服现有方法的局限性,这些方法要么需要大量的标注,要么仅关注低级视觉显著性。TextGaze 提取视觉特征,并采用 LVLM 生成与注视对齐的文本线索,通过具有分层文本监督的基于 Transformer 的融合模块进行处理。该模型能够高效地预测注视热图以及帧内/帧外状态,在四个主流数据集上展示了具有竞争力的性能和鲁棒的跨数据集泛化能力,且无需额外微调。 AI

影响 这项研究为注视估计提供了一种更高效、更通用的方法,有望改善人机交互和辅助功能工具。

排序理由 该集群包含一篇详细介绍注视点估计新方法的学术论文。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.CV 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

TextGaze 使用 LVLM 进行注视点估计 · arXiv cs.CV

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
该集群包含一篇详细介绍注视点估计新方法的学术论文。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
79 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准。

报道来源 [1]

  1. arXiv cs.CV TIER_1 English(EN) · Junhui She, Fei Wang, Kun Li, Yiqi Nie, Yuxin Liu, Zhangling Duan, Xun Yang ·

    TextGaze:利用文本场景线索进行注视目标估计

    arXiv:2607.10130v1 Announce Type: new Abstract: Gaze target estimation aims to infer the position of a person's gaze within a scene. Within mainstream design logic, multi-branch methods require extra supervision and annotations, while streamlined designs prioritize low-level visu…