ByteDance's Seed team has identified a periodic performance issue in DeepSeek-V4 models, where the model's accuracy fluctuates based on the input token's position. This "phase sensitivity" is linked to DeepSeek's block KV cache compression technique, designed to improve efficiency with long contexts. Researchers found that the model's ability to recall information can vary by over 40 percentage points depending on where the information is placed within the compressed context window, with the cycle length correlating to the compression step size. AI
IMPACT Highlights potential pitfalls in long-context optimization techniques, suggesting a need for more nuanced evaluation beyond average performance.
RANK_REASON Research paper detailing a specific performance issue in a model due to an optimization technique. [lever_c_demoted from research: ic=1 ai=1.0]
- ByteDance
- DeepSeek
- DeepSeek-V4
- DeepSeek-V4.1-Flash-0910
- DeepSeek-V4-Flash-0731
- DeepSeek-V4-Flash-Base
- DeepSeek-V4-Pro-Base
- Qwen3 0.6B
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →