PulseAugur
中
实时 10:29:58
English(EN) Output Language Confusion under Multilingual Prompt Contamination

新测试显示大型语言模型难以处理混合语言提示

一种名为多语言干扰(MDI)的新评估协议已被开发出来,用于评估大型语言模型如何处理包含不相关外语句子的提示。当在 Llama-3.1-8B 模型上使用印地语干扰进行测试时,该模型经常切换到梵文书写系统,尽管在精确匹配评分中这看起来像是幻觉,但通常保留了语义的正确性。其他模型在面对这种多语言干扰时倾向于不作答。该研究强调了在评估多语言大型语言模型可靠性时,区分脚本切换和语义错误的重要性,尤其是在检索增强生成等应用中。 AI

影响 突出了 RAG 等多语言大型语言模型应用中潜在的可靠性问题,并建议需要更细致的评估指标。

排序理由 介绍大型语言模型新评估方法的学术论文。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.CL 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

新测试显示大型语言模型难以处理混合语言提示

本文如何被排名

Signal score
11 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
介绍大型语言模型新评估方法的学术论文。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, safety
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Same-day
Cluster formed today. Ranking reflects the current source set at time of score.

完整方法见我们的编辑标准。

报道来源 [1]

  1. arXiv cs.CL TIER_1 English(EN) · Riju Marwah, Ritvik Garimella, Khusham Bansal, Atishay Jain, Amit Sheth ·

    多语言提示污染下的输出语言混淆

    arXiv:2610.02926v1 Announce Type: new Abstract: Standard factual benchmarks assume clean monolingual prompts and exact-match scoring, two assumptions that break simultaneously in real-world multilingual deployment, from retrieval-augmented generation pipelines returning mixed-lan…