PulseAugur
中
实时 18:19:37
English(EN) Unified Audio Intelligence Without Regressing on Text Intelligence

Nemotron Labs 发布统一音频-文本大模型 Audex

Nemotron Labs 推出了 Audex (Nemotron-Labs-Audex-30B-A3B),一个统一的音频-文本大语言模型。Audex 基于 Nemotron-Cascade-2-30B-A3B 纯文本模型构建,使用单个 Transformer 解码器无缝处理和生成音频与文本。该模型在包含 1574 亿音频 token 和 3205 亿文本 token 的大型数据集上进行了训练,并结合了监督训练和强化学习技术。Audex 在理解、语音识别、翻译和生成等各种音频任务中取得了最先进的性能,同时保留了其文本模型前身的强大推理和知识能力。 AI

影响 这种统一的模型架构有望推动多模态 AI 能力的发展,从而在未来的应用中实现更无缝的音频和文本处理集成。

排序理由 该集群描述了一篇详细介绍新模型发布的论文。

在 Hugging Face Daily Papers 阅读 →

AI 生成摘要 · Google Gemini · 来自 5 个来源。 我们如何撰写摘要 →

Nemotron Labs 发布统一音频-文本大模型 Audex

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
该集群描述了一篇详细介绍新模型发布的论文。
Source corroboration
5 independent sources
Strong cross-source corroboration — multiple independent publishers covered this within the clustering window.
Topics
model release, paper
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
94 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.
Coverage growth since scoring
+1 source(s) since last score
New sources have picked up this story since our last re-score. Score will update on the next scoring pass.

完整方法见我们的编辑标准。

报道来源 [5]

  1. arXiv cs.AI TIER_1 English(EN) · Zhifeng Kong, Sang-gil Lee, Jaehyeon Kim, Boxin Wang, Zihan Liu, Sungwon Kim, Yang Chen, Arushi Goel, Rajarshi Roy, Wenliang Dai, Zhuolin Yang, Yangyi Chen, Dongfu Jiang, Sreyan Ghosh, Tuomas Rintamaki, Andrew Tao, Jonathan Raiman, Mohammad Shoeybi, Brya… ·

    统一音频智能,不牺牲文本智能

    arXiv:2607.05196v1 Announce Type: cross Abstract: Audio intelligence involves understanding, reasoning about, and generating both audio and speech. In this work, we introduce Nemotron-Labs-Audex-30B-A3B (Audex), a unified audio-text LLM built on Nemotron-Cascade-2-30B-A3B, a stro…

  2. arXiv cs.AI TIER_1 English(EN) · Wei Ping ·

    统一音频智能,不牺牲文本智能

    Audio intelligence involves understanding, reasoning about, and generating both audio and speech. In this work, we introduce Nemotron-Labs-Audex-30B-A3B (Audex), a unified audio-text LLM built on Nemotron-Cascade-2-30B-A3B, a strong text-only MoE LLM. Audex adopts a simple unifie…

  3. Hugging Face Daily Papers TIER_1 English(EN) ·

    统一音频智能,不牺牲文本智能

    Audio intelligence involves understanding, reasoning about, and generating both audio and speech. In this work, we introduce Nemotron-Labs-Audex-30B-A3B (Audex), a unified audio-text LLM built on Nemotron-Cascade-2-30B-A3B, a strong text-only MoE LLM. Audex adopts a simple unifie…

  4. arXiv cs.CL TIER_1 English(EN) · Wei Ping ·

    统一音频智能,不牺牲文本智能

    Audio intelligence involves understanding, reasoning about, and generating both audio and speech. In this work, we introduce Nemotron-Labs-Audex-30B-A3B (Audex), a unified audio-text LLM built on Nemotron-Cascade-2-30B-A3B, a strong text-only MoE LLM. Audex adopts a simple unifie…

  5. Hugging Face Daily Papers TIER_1 English(EN) ·

    统一音频智能,不牺牲文本智能

    A unified audio-text large language model is presented that integrates audio and text processing through a shared transformer decoder, achieving superior performance across multiple audio and speech tasks while maintaining strong text reasoning capabilities.