Audio
PulseAugur coverage of Audio — every cluster mentioning Audio across labs, papers, and developer communities, ranked by signal.
4 day(s) with sentiment data
-
New benchmark Tri-PvP reveals modality bias in omni-modal LLMs
Researchers have developed Tri-PvP, a new benchmark designed to expose modality bias in omni-modal large language models (OLLMs). This benchmark addresses a limitation in previous evaluations by separating perceptual si…
-
Developers can estimate LLM API costs by understanding token pricing and cost-saving levers
Estimating the cost of using Large Language Model (LLM) APIs requires understanding that pricing is based on tokens, with output tokens being significantly more expensive than input tokens. Developers can calculate basi…
-
New research explores parallel drafting for speculative decoding in LLMs
Two new research papers explore advancements in speculative decoding for large language models, focusing on improving efficiency and coherence in parallel drafting. The first paper surveys the applicability of block-par…
-
New theory explains AI model representations across modalities
Researchers have developed a new theory that explains the internal workings of AI models across different modalities like vision, audio, and language. This theory posits that classification tasks create a shared represe…
-
MiniMax AI launches omni-modal generation model H3
MiniMax AI has launched its new omni-modal generation model, MiniMax H3. This model is capable of processing and generating content across multiple modalities including text, images, video, and audio. MiniMax H3 is bein…
-
Study finds multimodal LLMs perpetuate gender bias in musical instrument associations
A new study published on arXiv investigates gender bias in multimodal large language models (LLMs) by examining their associations with musical instruments. Researchers developed the Symphony-Bias dataset, which include…
-
JEPA models face challenges with language's conditional structure
A new paper explores the challenges of applying Joint-Embedding Predictive Architectures (JEPAs) to language processing, contrasting their effectiveness in image and audio domains with their limitations in text. The res…
-
New AHEAD framework boosts multi-class label aggregation accuracy
Researchers have developed AHEAD, a novel framework for multi-class label aggregation that improves the accuracy of inferring true labels from noisy crowdsourced annotations. AHEAD utilizes a graph neural network to lea…
-
New AHEAD framework enhances crowdsourced label aggregation accuracy
Researchers have developed AHEAD, a novel framework for multi-class label aggregation that improves the accuracy of inferring true labels from crowdsourced annotations. AHEAD utilizes a graph neural network to learn cro…
-
New PhysEditWorld dataset enables physics-editable game world models
Researchers have introduced PhysEditWorld, a large-scale dataset designed to enable physics-editable world models for game environments. This dataset focuses on gravity variations within 12 cinematic scenes rendered usi…
-
Volcanic Engine releases Doubao 2.1 Pro with enhanced AI capabilities · 1 source tracked
ByteDance's Volcanic Engine has released the Doubao large model 2.1, with the Pro version featuring enhanced capabilities in coding, agent technology, and visual language models. The company also announced new video, im…
-
New metric MultiMem quantifies memorization in multi-modal contrastive learning
Researchers have introduced MultiMem, a novel metric to quantify memorization in multi-modal contrastive learning, a field previously unexplored in this regard. Their analysis indicates that semantic misalignment betwee…
-
Edge ML Developers Debate Data Bottlenecks: Acquisition vs. Cleaning
A Reddit user on r/MachineLearning is seeking to identify the primary time sink for developers working with embedded/edge machine learning, specifically for time-series sensor data. The user is developing a hardware-agn…
-
New Omni-Fake dataset benchmarks multimodal deepfake detection on social media
Researchers have introduced Omni-Fake, a new benchmark dataset designed to improve the detection of multimodal deepfakes on social media. The dataset includes over 1 million samples across image, audio, video, and audio…