PulseAugur
中
实时 02:24:06
English(EN) Improved Gemini audio models for powerful voice experiences

Google DeepMind 增强 Gemini 音频模型,实现自然语音交互和翻译

Google DeepMind 发布了升级版的 Gemini 2.5 音频模型,增强了实时语音代理和文本到语音生成的能力。Gemini 2.5 Flash Native Audio 模型现在提供了改进的功能调用、指令遵循和对话上下文感知能力,在 ComplexFuncBench Audio 基准测试中取得了 71.5% 的分数。此外,Google Translate 应用正在推出新的实时语音翻译功能,支持实时语音到语音翻译,并能保留说话者的语调和音高。 AI

排序理由 前沿实验室模型发布,附带系统卡。

在 Google DeepMind 阅读 →

AI 生成摘要 · Google Gemini · 来自 2 个来源。 我们如何撰写摘要 →

Google DeepMind 增强 Gemini 音频模型,实现自然语音交互和翻译

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Frontier Release
前沿实验室模型发布,附带系统卡。
Source corroboration
2 independent sources
Multiple independent publishers reporting the same story raises confidence that it's real and newsworthy.
Topics
model release, product
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
492 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准。

报道来源 [2]

  1. Google DeepMind TIER_1 English(EN) ·

    改进的 Gemini 音频模型,带来强大的语音体验

  2. Google DeepMind TIER_1 English(EN) ·

    Gemini 2.5 的高级音频对话与生成

    Gemini 2.5 has new capabilities in AI-powered audio dialog and generation.