PulseAugur
实时 09:26:44
English(EN) Real-Time Voice AI Hears but Does Not Listen

研究发现:语音AI系统尽管能感知情绪,但未能据此采取行动 · 跟踪2个来源

一项新的研究论文评估了四种领先的实时语音AI系统——OpenAI的GPT Realtime 2、谷歌的Gemini 3.1 Flash Live,以及阿里巴巴的Qwen3.5 Omni Plus和Omni Flash——发现它们尽管能够感知到声音中的情绪或讽刺等线索,却常常未能据此采取行动。这种“情商差距”意味着系统优先考虑所说的字面意思,而不是语气和表达方式,导致在客户服务或金融交易等关键场景中出现不恰当的响应。虽然提示可以提供部分改进,但研究建议在使用这些系统处理对声音表达至关重要的内容时要谨慎。 AI

影响 强调了当前语音AI的一个关键差距,建议在声音语调和情绪对于理解至关重要的应用中要谨慎。

排序理由 在arXiv上发表的研究论文,详细介绍了AI模型能力的研究结果。

在 arXiv cs.CL 阅读 →

AI 生成摘要 · Google Gemini · 来自 2 个来源。 我们如何撰写摘要 →

研究发现:语音AI系统尽管能感知情绪,但未能据此采取行动 · 跟踪2个来源

报道来源 [2]

  1. arXiv cs.CL TIER_1 English(EN) · Martijn Bartelds, Federico Bianchi, James Zou ·

    实时语音AI能听见但不能理解

    arXiv:2606.26083v1 Announce Type: new Abstract: Speech conveys information through both words and vocal delivery. We evaluate four leading production realtime voice systems-OpenAI's GPT Realtime 2, Google's Gemini 3.1 Flash Live, and Alibaba's Qwen3.5 Omni Plus and Omni Flash-on …

  2. arXiv cs.CL TIER_1 English(EN) · James Zou ·

    实时语音AI能听见但不会倾听

    Speech conveys information through both words and vocal delivery. We evaluate four leading production realtime voice systems-OpenAI's GPT Realtime 2, Google's Gemini 3.1 Flash Live, and Alibaba's Qwen3.5 Omni Plus and Omni Flash-on tasks where the words and the delivery patterns …