PulseAugur
EN
LIVE 03:13:43

New Gaze-to-Text Framework Uses LLMs to Decode Human Intentions

Researchers have introduced Gazette, a novel framework for decoding human gaze into natural language descriptions of goals. Unlike previous methods that relied on categorical decoding, Gazette uses multimodal large language models (MLLMs) to generate free-form text that captures the nuances of human intentions. The framework leverages synthesized "think-aloud" transcripts generated by a large language model to help Gazette learn goal-specific dynamics and filter individual gaze differences. This approach achieves state-of-the-art performance in gaze decoding across various tasks, enabling gaze to serve as a non-intrusive cue for inferring human goals. AI

IMPACT This research could lead to more intuitive human-computer interaction by enabling systems to better understand user intentions through gaze tracking.

RANK_REASON The cluster describes a novel framework and learning problem presented in an academic paper on arXiv. [lever_c_demoted from research: ic=1 ai=1.0]

Read on arXiv cs.CV →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

New Gaze-to-Text Framework Uses LLMs to Decode Human Intentions

COVERAGE [1]

  1. arXiv cs.CV TIER_1 English(EN) · Sounak Mondal, Dimitris Samaras, Gregory Zelinsky, Minh Hoai ·

    Gaze-to-text Generation: Beyond Categorical Decoding of Human Attention

    arXiv:2607.23917v1 Announce Type: new Abstract: We introduce a novel learning problem: decoding gaze into natural language descriptions of human goals across diverse visual tasks. Unlike prior work, which frames gaze decoding as a discriminative task over predefined categories, w…