Researchers have introduced EXACT, a novel supervision-allocation objective designed to improve long-context adaptation in language models. This method addresses a mismatch where packed training with document masking results in short effective contexts for target tokens. By assigning extra weight to long effective-context targets based on their frequency in the long tail, EXACT demonstrates significant improvements across various Qwen and LLaMA configurations on benchmarks like NoLiMa and RULER. The gains are particularly notable when evidence is located thousands of tokens away, while performance on shorter contexts remains stable, preserving standard QA and reasoning capabilities. AI
IMPACT Enhances long-context understanding in LLMs, potentially improving performance on tasks requiring retrieval of information from distant parts of a document.
RANK_REASON Academic paper detailing a new method for improving language model performance. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →