A new research paper explores the challenge of predicting human attention in text, establishing benchmarks for "floor" (naive truncation) and "ceiling" (split-half oracle) scores. The study found that current frontier language models achieve 35-53% of the gap between these bounds, with a state-of-the-art prompt compressor performing worse than random. However, an unweighted fusion of five frontier models significantly improved performance, and this gain was retained by distilling the fusion into a single 8B open-weight student model. AI
IMPACT Highlights limitations of current LLMs in understanding nuanced human attention and suggests fusion techniques as a path to improvement.
RANK_REASON Research paper published on arXiv detailing a new benchmark and findings on predicting human attention in text.
- Apify
- arXiv
- Floor, Ceiling, and the Fusion Gap: How Much of Crowd Reading Attention Can Machines Predict?
- Holm
- LLMLingua-2
AI-generated summary · Google Gemini · from 3 sources. How we write summaries →