Researchers have identified two critical failure modes in post-training quantization (PTQ) for looped transformers, which reuse weights across recurrence steps. The first, termed 'feedback exposure,' occurs when a quantized layer perturbs the recurrent state without an identity path, leading to amplified errors. The second, 'calibration blindness,' arises from quantization methods that base their calculations on initial activations, neglecting later states in the recurrence. Experiments on models like Huginn-3.5B and Mamba demonstrated these issues, with solutions like accumulating Hessians across recurrence steps showing promise in recovering full accuracy. AI
IMPACT Identifies key challenges in optimizing looped transformer architectures for efficient deployment.
RANK_REASON The cluster contains an academic paper detailing novel research findings on model quantization techniques. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →