Researchers have developed a novel method for predicting text from low-pass filtered speech, a task previously overlooked in the field. By fine-tuning the Whisper model on only the lowest Mel bins (approximately 450Hz cutoff), they achieved a Word Error Rate (WER) of 36%. The study found that 10% of utterances were recovered perfectly, and 40% had a WER of 25% or lower, suggesting a strong correlation between low-frequency speech features and lexical content. This breakthrough could enable new applications, such as using prosody to guide text generation in large language models. AI
IMPACT This research could enable new applications for LLMs by allowing prosody to guide text generation.
RANK_REASON Academic paper detailing a new method for speech-to-text conversion. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- CatalyzeX
- Connected Papers
- DagsHub
- Gotit.pub
- Hugging Face
- Influence Flower
- Litmaps
- ScienceCast
- scite Smart Citations
- Whisper
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →