Researchers have developed new methods to improve the completeness and quality of answers generated by large language models (LLMs) for complex questions. Apple's research introduces DeepAmbigQA, a dataset and generation pipeline designed to test LLMs on questions requiring multi-hop reasoning and disambiguation of ambiguous entities, revealing that even advanced models like GPT-5 struggle with answer completeness. Separately, a new framework called QQ leverages the dual nature of question generation and answering to create more coherent multi-hop questions, showing significant improvements on datasets like HotpotQA and MuSiQue. AI
IMPACT Addresses limitations in LLM reasoning and answer completeness for complex queries, potentially leading to more reliable AI assistants.
RANK_REASON Two research papers introducing new datasets and frameworks for evaluating and improving multi-hop question answering in LLMs.
Read on Apple Machine Learning Research →
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →