PulseAugur
实时 09:23:43
English(EN) Do Evaluation Metrics Detect Errors in Classical Chinese to English Translations?

古汉语到英语的AI翻译需要新指标

研究人员调查了当前自动评估指标在古汉语到英语翻译中的有效性,这项任务大型语言模型表现出惊人的熟练度,但缺乏可靠的评估。研究使用包含最小对的诊断框架来识别常见错误类型,发现现有指标存在显著盲点。尽管MetricX24在测试指标中表现最佳,但研究结果强调了为具有历史和文化差异的翻译背景量身定制更强大、更可解释的评估工具的必要性。 AI

影响 强调了在专业翻译任务中为LLM开发更好评估指标的必要性,可能影响数字人文领域未来模型的开发和部署。

排序理由 该集群包含一篇详细介绍特定翻译任务评估指标研究的学术论文。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

古汉语到英语的AI翻译需要新指标

报道来源 [1]

  1. arXiv cs.AI TIER_1 English(EN) · Osvaldo Quinjica, Eric Bennett, Xinchen Yang, Andrew Schonebaum, Marine Carpuat ·

    评估指标能否检测出古文英译的错误?

    arXiv:2608.08283v1 Announce Type: cross Abstract: Although large language models can translate some historical languages surprisingly well, their usefulness in digital humanities workflows is limited by the lack of reliable evaluation. We investigate whether existing automatic ev…