VoiceBench
PulseAugur coverage of VoiceBench — every cluster mentioning VoiceBench across labs, papers, and developer communities, ranked by signal.
-
KRAFTON releases bilingual speech model A.X K2 Raon-Speech
KRAFTON has released A.X K2 Raon-Speech, a bilingual English/Korean speech language model with 21.2 billion total parameters and 3.5 billion active parameters. This multimodal model integrates a text backbone from SK Te…
-
Alibaba releases Qwen-Audio-3.0-Realtime with enhanced voice AI capabilities
Alibaba has launched Qwen-Audio-3.0-Realtime, an upgraded real-time conversational voice model. This new version enhances intelligence, agent tool utilization, empathetic dialogue, and full-duplex interaction capabiliti…
-
ParaBridge method improves speech models' paralinguistic understanding
Researchers have developed ParaBridge, a novel on-policy self-distillation method designed to improve speech language models' ability to incorporate paralinguistic cues into dialogue. This technique trains models to bet…
-
New Shapley Value method explains multimodal AI models
Researchers have developed a novel extension of Shapley Values to explain the behavior of multimodal multilingual models (MLLMs). This framework addresses the challenges of integrating text and audio data by treating th…
-
NVIDIA launches Nemotron 3 Nano Omni, unifying multimodal AI for efficiency
NVIDIA has released Nemotron 3 Nano Omni, an open multimodal model capable of processing text, images, audio, and video. This model aims to unify these modalities into a single architecture, improving efficiency and ena…