Researchers have developed a new method called stride-k subsampling to reduce the number of audio tokens processed by OpenAI's Whisper model without requiring additional training. This technique involves selecting every k-th token, and a stride-2 configuration was found to cut audio tokens by 75% and GFLOPs by over 50% with minimal impact on Word Error Rate (WER) for most ASR benchmarks. The method also showed benefits for Whisper-based SpeechLMs, reducing latency by up to 27.4% with only modest accuracy drops. AI
IMPACT Reduces computational requirements for audio processing models, potentially lowering inference costs and latency.
RANK_REASON Academic paper detailing a new method for an existing model. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →