A new research paper explores how Large Language Models (LLMs) in draft-verify-revise pipelines can struggle with deictic ambiguity, where context-dependent expressions like "previous" can refer to different things across stages. The study tested six models, finding that while GPT-5.2 improved significantly with increased reasoning effort, Gemini 3-Pro maintained high accuracy across all tested levels. The research suggests that context engineers should explicitly define referents at each stage to prevent deictic shifts. AI
IMPACT Highlights potential deictic shift issues in LLM orchestration, suggesting explicit referent definition for improved context handling.
RANK_REASON Research paper published on arXiv detailing LLM capabilities. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- CatalyzeX
- DagsHub
- Draft-Verify-Revise
- Gemini 3-Pro
- Gotit.pub
- GPT-5.2
- Hugging Face
- Obinna Ekekezie M.D.
- ScienceCast
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →