Researchers have developed Arti-JEPA, a new joint embedding predictive architecture designed to model real-time MRI data of the vocal tract for speech analysis. This model was trained on approximately 62 hours of unlabeled vocal tract videos and evaluated on tasks including phoneme prediction, classification of fluent versus disfluent speech, and characterizing speech changes after surgery. The findings indicate that a temporal video prior is more effective than per-frame encoders, and domain adaptation is crucial for certain tasks like phoneme prediction, though it did not improve stuttering classification. Arti-JEPA demonstrated an ability to decode phoneme signals from post-operative speech, suggesting its potential as a reusable measurement tool for speech science. AI
IMPACT This research offers a new method for analyzing vocal tract dynamics using MRI, potentially advancing speech science and clinical applications.
RANK_REASON The cluster contains an academic paper detailing a new model and its evaluation. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →