Researchers have developed a new method to fine-tune vision-language models for generating more effective referring expressions. This approach uses estimated listener gaze data as a learning signal, transforming observations of incremental listener comprehension into rewards. Experiments show that models trained with this gaze-estimating listener produce significantly more pragmatic references, reducing word count from 15.4 to 4.0 while increasing success rates from 75.2% to 80.0%. This work highlights the potential for learning utterance generation through language-based interaction, incorporating implicit listener comprehension signals. AI
IMPACT This research could lead to more efficient and natural language generation in AI systems, improving human-computer interaction.
RANK_REASON The cluster contains a research paper detailing a new method for fine-tuning AI models. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- CatalyzeX
- Connected Papers
- DagsHub
- Gotit.pub
- Hugging Face
- Influence Flower
- Litmaps
- ScienceCast
- scite Smart Citations
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →