Fast and Stable Triangular Inversion for Delta-Rule Linear Transformers
Researchers have developed a new method for triangular inversion, a crucial operation in linear attention mechanisms used by advanced models like Qwen3.5/3.6 and Kimi Linear. This technique significantly improves the speed and numerical stability of this sub-routine, which is often a performance bottleneck. Experiments show up to a 4.3x speed-up on NPUs compared to existing implementations, leading to overall layer performance gains without sacrificing accuracy. AI
IMPACT Improves efficiency of linear attention mechanisms, potentially enabling faster and more accurate long-context models.