SpeechLMs
PulseAugur coverage of SpeechLMs — every cluster mentioning SpeechLMs across labs, papers, and developer communities, ranked by signal.
2 day(s) with sentiment data
-
Stride-k subsampling slashes Whisper audio tokens by 75% without retraining
Researchers have developed a new method called stride-k subsampling to reduce the number of audio tokens processed by OpenAI's Whisper model without requiring additional training. This technique involves selecting every…
-
New framework combats 'previous-belief contamination' in speech emotion models
Researchers have identified a phenomenon called previous-belief contamination (PBC) in streaming emotion understanding models, where the model's prior predictions can negatively impact its interpretation of current audi…
-
New benchmarks and platforms advance voice agent evaluation and development
New research introduces EVA-Bench, a comprehensive framework for evaluating voice agents, addressing challenges in simulating realistic conversations and measuring performance across various failure modes. Simultaneousl…