Researchers have introduced SoniSpeech, a new large-scale, open-vocabulary dataset designed for wearable silent speech interfaces (SSIs). This trimodal dataset synchronizes ultrasound echo profiles, voiced audio, and frontal video, capturing both voiced and silent speech. SoniSpeech contains 18,000 utterances totaling 34 hours, derived from the SODA dialogue dataset, and features contemporary conversational English with over 5,000 unique words and full phoneme coverage. A baseline model using CTC-based ResNet-34 achieved a 26.3% word error rate on the silent speech recognition task, establishing the first benchmark for this emerging field. AI
IMPACT This dataset could accelerate the development of more versatile and less intrusive silent speech interfaces.
RANK_REASON The cluster contains a research paper introducing a new dataset and benchmark for a specific AI task. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →