English(EN)How Reliable Is Your Jailbreak Judge? Calibration and Adversarial Robustness of Automated ASR Scoring
新的ASR技术解决语音错误并提高判断可靠性
作者PulseAugur 编辑部·[7 个来源]·
研究人员正在开发先进的方法来改进自动语音识别(ASR)系统,特别是在低资源语言方面以及解决特定类型的错误。一种名为Error-Aware TF-IDF的方法使用一种新颖的算法,根据历史语音错误识别来优先处理更正文档,从而显著降低词错误率。另一种名为G-SPIN的方法将语音图模型与大型语言模型相结合,通过将搜索空间限制在合理的语音替代方案内来纠正语义关键错误。此外,一项研究质疑用于评估LLM越狱尝试的自动判断的可靠性,揭示了其准确性和鲁棒性方面的不一致和漏洞。
AI
arXiv:2606.24915v1 Announce Type: new Abstract: End-to-end automatic speech recognition systems frequently hallucinate rare entities and domain-specific terms, especially in low-resource languages. While retrieval-augmented generation frameworks can mitigate these errors using la…
arXiv:2606.24889v1 Announce Type: new Abstract: Automatic speech recognition (ASR) systems, despite low overall word error rates, produce residual lexical errors that disproportionately affect semantically critical tokens such as named entities, negations, and sentiment-bearing w…
arXiv:2606.25487v1 Announce Type: new Abstract: Almost every paper on LLM jailbreaks and prompt injection reports an attack-success rate (ASR), and that number is assigned not by people but by an automated judge: either a safety classifier trained for the task, or a general chat …
Almost every paper on LLM jailbreaks and prompt injection reports an attack-success rate (ASR), and that number is assigned not by people but by an automated judge: either a safety classifier trained for the task, or a general chat model prompted to grade. The judge is rarely che…
Almost every paper on LLM jailbreaks and prompt injection reports an attack-success rate (ASR), and that number is assigned not by people but by an automated judge: either a safety classifier trained for the task, or a general chat model prompted to grade. The judge is rarely che…
End-to-end automatic speech recognition systems frequently hallucinate rare entities and domain-specific terms, especially in low-resource languages. While retrieval-augmented generation frameworks can mitigate these errors using large language models, current architectures face …
<h4>A feasible framework for evaluating ASR models across semantic categories instead of a single aggregate metric</h4><figure><img alt="Introduction image showing decomposition of general WER into semantic categories, such as people, geography names, etc" src="https://cdn-images…