Vad
PulseAugur coverage of Vad — every cluster mentioning Vad across labs, papers, and developer communities, ranked by signal.
3 day(s) with sentiment data
-
Voice AI struggles with turn-taking, needs better endpointing models
Current voice AI systems often end conversations prematurely by relying solely on silence detection, leading to users being cut off mid-thought. Increasing the silence threshold to prevent interruptions introduces a wor…
-
New VAD method isolates visual evidence in multimodal AI distillation
Researchers have developed Visual Attribution Distillation (VAD), a novel method for multimodal on-policy distillation. VAD aims to isolate the visual evidence supporting a teacher model's corrections to a student model…
-
Visual Attribution Distillation (VAD) enhances multimodal knowledge transfer
Researchers have introduced Visual Attribution Distillation (VAD), a novel algorithm designed to improve multimodal on-policy distillation by isolating visual evidence in knowledge transfer. VAD works by reconstructing …
-
BEVLM framework enhances LLM reasoning for autonomous driving
Researchers have developed BEVLM, a new framework that integrates Large Language Models (LLMs) with Bird's-Eye View (BEV) representations for autonomous driving. This approach aims to overcome the limitations of current…
-
New REDDIT Framework Corrects Timestamp Drift in ASR Models
Researchers have developed REDDIT, a novel post-training framework designed to fix timestamp inaccuracies in autoregressive Automatic Speech Recognition (ASR) systems. This method addresses timestamp drift, where the de…
-
New REDDIT framework corrects ASR timestamp drift without model forgetting
Researchers have developed REDDIT, a novel post-training framework designed to correct timestamp drift in Automatic Speech Recognition (ASR) systems without causing catastrophic forgetting. This method uses a replay-bas…
-
Wan-Streamer v0.1: Unified model for real-time audio-visual interaction
Researchers have introduced Wan-Streamer v0.1, a novel end-to-end multimodal foundation model designed for real-time, low-latency audio-visual interaction. Unlike traditional cascaded systems, Wan-Streamer integrates la…
-
GraphBEV++ framework tackles feature misalignment in autonomous driving perception
Researchers have introduced GraphBEV++, a novel framework designed to tackle feature misalignment in Bird's-Eye View (BEV) perception for autonomous driving systems. The framework employs two main modules: LocalAlign-v2…
-
Voice AI paradox: Advanced chat, basic failures
Voice AI assistants like Yandex's Alisa exhibit a paradox of advanced conversational abilities alongside basic functional failures, stemming from their complex architecture. This hybrid system combines speech recognitio…
-
Unified Map Prior Encoder enhances autonomous driving mapping and planning
Researchers have developed a Unified Map Prior Encoder (UMPE) designed to integrate diverse map data, such as HD/SD vector maps, rasterized maps, and satellite imagery, into autonomous driving systems. This encoder addr…
-
New LLMs unify audio and language processing for full-duplex and medical applications
Researchers have developed UAF, a novel unified audio front-end LLM designed for full-duplex speech interaction. This model integrates diverse audio front-end tasks like voice activity detection and turn-taking into a s…