PulseAugur
实时 05:27:57
English(EN) From transcription to semantic corpus analysis: unsupervised learning of sentence representations for ancient languages

无监督方法为古语生成有效的句子嵌入

研究人员开发了两种无监督学习策略,TSDAE 和对比句子嵌入 (CSE),为古语生成有效的句子嵌入。这些方法仅使用原始文本来调整现有语言模型,在拉丁语和古希腊语文本的复用识别和对应检索等任务上,其表现优于多语言和专业基线。调整后的编码器即使在处理嘈杂的自动转录文本时也特别有效,并且提供了一个名为 Paraphrasis 的工具,供非专业人士使用完整的流程。 AI

影响 能够对历史文本进行语义分析,有可能在数字人文和语言学领域带来新的见解。

排序理由 学术论文,详细介绍了用于古语自然语言处理任务的新无监督学习方法。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.IR (Information Retrieval) 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

无监督方法为古语生成有效的句子嵌入

报道来源 [1]

  1. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Th{é}otime de la Selle ·

    从转录到语义语料库分析:古语言句子表示的无监督学习

    Automatic Text Recognition (ATR) now supplies digital humanities with large volumes of unstructured, heterogeneous, and often noisy text in ancient languages. Downstream semantic analysestext reuse identification, alignment, and semantic search-rely on sentence embeddings, yet ex…