PulseAugur
EN
LIVE 08:23:52

New SoniSpeech dataset advances open-vocabulary silent speech recognition

Researchers have introduced SoniSpeech, a new large-scale, open-vocabulary dataset designed for wearable silent speech interfaces (SSIs). This trimodal dataset synchronizes ultrasound echo profiles, voiced audio, and frontal video, capturing both voiced and silent speech. SoniSpeech contains 18,000 utterances totaling 34 hours, derived from the SODA dialogue dataset, and features contemporary conversational English with over 5,000 unique words and full phoneme coverage. A baseline model using CTC-based ResNet-34 achieved a 26.3% word error rate on the silent speech recognition task, establishing the first benchmark for this emerging field. AI

IMPACT This dataset could accelerate the development of more versatile and less intrusive silent speech interfaces.

RANK_REASON The cluster contains a research paper introducing a new dataset and benchmark for a specific AI task. [lever_c_demoted from research: ic=1 ai=1.0]

Read on arXiv cs.LG →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

New SoniSpeech dataset advances open-vocabulary silent speech recognition

COVERAGE [1]

  1. arXiv cs.LG TIER_1 English(EN) · Ruidong Zhang, Jiacheng Liu, Fran\c{c}ois Guimbreti\`ere, Cheng Zhang ·

    SoniSpeech: A Large-Scale Open-Vocabulary Tri-Modal Dataset for Wearable Silent Speech Interfaces

    arXiv:2608.00803v1 Announce Type: cross Abstract: Wearable silent speech interfaces (SSIs) are limited to small, closed vocabularies. Approaches achieving larger vocabularies require obtrusive hardware such as facial electrodes. We present SoniSpeech, the first large-scale, open-…