LRS3
PulseAugur coverage of LRS3 — every cluster mentioning LRS3 across labs, papers, and developer communities, ranked by signal.
3 day(s) with sentiment data
-
SyncVoice framework enhances video dubbing with vision-augmented TTS
Researchers have developed SyncVoice, a novel framework for automatic video dubbing that enhances speech naturalness and temporal synchronization with visual content. By integrating a Text-Visual Fusion Module into a pr…
-
New Candor-LR dataset pushes audio-visual speech recognition toward natural conversation
Researchers have introduced Candor-LR, a new dataset designed to advance audio-visual speech recognition (AVSR) by simulating natural conversations. Unlike existing benchmarks like LRS3, which use scripted speech, Cando…
-
New AVSRBench benchmark reveals generalization gap in speech recognition
Researchers have developed AVSRBench, a new benchmark designed to evaluate Audio-Visual Speech Recognition (AVSR) systems across a variety of challenging conditions beyond standard broadcast speech. The study found that…
-
New method boosts LLM-based audio-visual speech recognition
Researchers have developed a new method called Attention-Guided Reliability Scaling (AGRS) to improve audio-visual speech recognition (AVSR) systems that use large language models. This technique adapts contrastive deco…
-
New DoubleHelix framework improves audio-visual speech recognition
Researchers have introduced DoubleHelix, a novel framework for audio-visual speech recognition (AVSR) that enhances the fusion of audio and visual data. Unlike previous methods that treat cross-modal interaction as a si…
-
New Framework Generates Speech from Facial Images
Researchers have developed a novel Face-to-Speech (F2S) framework capable of generating plausible voices from static facial images, addressing the limitation of text-to-speech (TTS) systems that require reference audio.…
-
New VSR method uses head pose to improve accuracy
Researchers have developed a new framework called HP-VSR-ResFiLM to improve visual speech recognition (VSR) by explicitly incorporating head-pose information. This method uses a pose-conditioned residual Feature-wise Li…