Researchers have developed a new method to evaluate in-context learning (ICL) in large language models, specifically focusing on how fine-tuning affects this ability. The study introduces "In-Context Sensitivity" (ICS) and "ICL-GAP" as metrics to measure attention-level changes and behavioral accuracy gaps, respectively. Experiments with Llama-2-7B demonstrated that optimizing for ICS can lead to a dissociation where attention patterns change significantly, but the model's actual task performance, measured by MMLU accuracy, declines. AI
IMPACT This research highlights potential pitfalls in evaluating LLM capabilities, suggesting that attention-based proxies for in-context learning may not accurately reflect behavioral performance after fine-tuning.
RANK_REASON Academic paper detailing a new methodology and experimental findings. [lever_c_demoted from research: ic=1 ai=1.0]
- Attention Sensitivity Is Not Enough: Dissociating Attention-Level and Behavioural In-Context Learning under Fine-Tuning
- In-context learning
- fine-tuning
- ICL-GAP
- In-Context Sensitivity
- large language models
- Llama-2-7B
- MMLU
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →