A new study published on arXiv investigates how supervised fine-tuning (SFT) affects the instruction sensitivity of large language models. Researchers found that SFT consistently reduces instruction sensitivity in smaller models like Qwen3 (1.7B and 4B parameters), leading to performance improvements across different instruction formulations. However, for larger models such as Qwen3-8B and Gemma-2-9B, the effect of SFT on sensitivity is less pronounced and can vary, with some models showing consistent directional contrasts while others do not. The study also highlights that the choice of evaluation method, such as free-generation versus forced-choice, can lead to different conclusions about model robustness. AI
IMPACT Understanding how fine-tuning affects model sensitivity is crucial for developing more robust and reliable LLMs.
RANK_REASON Academic paper on LLM behavior. [lever_c_demoted from research: ic=1 ai=1.0]
Read on arXiv cs.IR (Information Retrieval) →
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →