PulseAugur
EN
LIVE 00:04:36

Speech models achieve 64% KV cache reduction with new text compression

Researchers have developed a novel method for compressing key-value (KV) states in full-duplex speech models, significantly reducing memory usage during long-running interactions. This technique, termed acoustic-to-text KV compression, converts incoming speech into a compact textual format during processing "slack" time. Implemented with MiniCPM-o 4.5 and LoRA, the method achieved a 64.6% reduction in peak streaming KV-cache size on ten-minute sessions. The approach also demonstrated improvements in transcription, temporal question answering, and summarization, while maintaining comparable performance in pause-handling, turn-taking, and interruption. AI

IMPACT Reduces memory requirements for long-context speech models, potentially enabling more efficient real-time applications.

RANK_REASON Academic paper detailing a new technical method for improving AI model efficiency. [lever_c_demoted from research: ic=1 ai=1.0]

Read on arXiv cs.AI →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Speech models achieve 64% KV cache reduction with new text compression

COVERAGE [1]

  1. arXiv cs.AI TIER_1 English(EN) · Yejin Lee, Seungbeom Kim, Yongha Lee, Kyuhong Shim ·

    Acoustic-to-Text KV Compression for Full-Duplex Speech Models

    arXiv:2609.31224v1 Announce Type: cross Abstract: Full-duplex speech language models continuously accumulate acoustic key-value (KV) states, making long-running interactions memory-intensive. During listening, the model can finish processing an audio unit before the next arrives;…