PulseAugur
EN
LIVE 21:03:10

Together AI's OSCAR slashes KV cache memory by 8x

Together AI has released OSCAR, an open-source 2-bit KV cache method that significantly reduces memory usage. Unlike previous 2-bit methods that failed at longer contexts, OSCAR maintains performance up to 128K tokens. This innovation was demonstrated using the Qwen3-8B model, showing an 8x reduction in KV cache memory. AI

IMPACT Reduces memory requirements for large language models, potentially enabling longer context windows and more efficient deployment.

RANK_REASON The cluster describes a new open-source technical method for improving AI model efficiency, which falls under research. [lever_c_demoted from research: ic=1 ai=1.0]

Read on Towards AI →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Together AI's OSCAR slashes KV cache memory by 8x

How we ranked this

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
The cluster describes a new open-source technical method for improving AI model efficiency, which falls under research. [lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
infra, model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
122 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

Full methodology in our editorial standards.

COVERAGE [1]

  1. Towards AI TIER_1 English(EN) · Chew Loong Nian - AI ENGINEER ·

    Together AI's OSCAR Killed KV Cache Memory 8x — The First 2-Bit That Doesn't Collapse at 128K

    <div class="medium-feed-item"><p class="medium-feed-snippet">Every 2-bit KV cache method I tried in 2025 collapsed past 32K context. Together AI&#x2019;s OSCAR, open-sourced on May 25, 2026, kept Qwen3&#x2013;8B&#x2026;</p><p class="medium-feed-link"><a href="https://pub.towardsa…