PulseAugur
实时 09:17:11
English(EN) Loanword or Switch? The Annotation Boundary, Not the Model, Drives Kazakh-Russian Code-Switching Identification

哈萨克语-俄语语码转换识别瓶颈在于注释而非模型

arXiv上的一篇新论文探讨了哈萨克语和俄语之间的语码转换识别问题,发现注释边界比所使用的模型更为关键。研究人员开发了一个黄金语言识别(LID)数据集,并为整合借词和从句级切换制定了具体指南。在测试中,FastText、Lingua和XLM-RoBERTa等模型表现各异,这表明注释定义而非模型类别本身,是准确识别的主要瓶颈。 AI

影响 强调了清晰的注释指南在自然语言处理(NLP)任务中的重要性,并提出在语码转换识别方面,改进数据标注可能比仅关注先进模型更具影响力。

排序理由 该集群包含一篇在arXiv上发表的学术论文。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.CL 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

哈萨克语-俄语语码转换识别瓶颈在于注释而非模型

报道来源 [1]

  1. arXiv cs.CL TIER_1 English(EN) · Bogdan Savelyev ·

    Loanword or Switch? The Annotation Boundary, Not the Model, Drives Kazakh-Russian Code-Switching Identification

    arXiv:2608.00581v1 Announce Type: new Abstract: Off-the-shelf LID and letter heuristics over-label Kazakh-Russian social text as mixed: Russian loanwords inside Kazakh look like code-switching under a shared Cyrillic script. We release a document-level gold LID set whose guidelin…