PulseAugur
实时 06:33:35
English(EN) A large-scale corpus of religious radio broadcast transcripts from webstream recordings in the United States

美国宗教广播转录文本新语料库发布

研究人员创建了一个来自美国网络流媒体录音的宗教广播转录文本的大规模语料库。该数据集在2025年7月的一个月内收集,包含超过70万条录音和6000多万行语音,通过自动化流程进行转录和说话人区分,并使用大型语言模型标注了节目格式和主题。该语料库旨在促进对宗教广播、其对社会和政治问题的讨论以及该代表性不足领域语音处理研究的深入研究。 AI

影响 为宗教媒体中的语音处理和内容分析提供了一个新数据集。

排序理由 该条目是关于一个大规模语料库的数据描述论文。[lever_c_demoted from research: ic=1 ai=0.7]

在 arXiv cs.CL 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

美国宗教广播转录文本新语料库发布

报道来源 [1]

  1. arXiv cs.CL TIER_1 English(EN) · Samuel Bestvater, Athena Chapekis, Skyler Seets, Anna Lieb, Sono Shah, Aaron Smith ·

    美国网络流媒体录音的大规模宗教广播转录语料库

    arXiv:2607.26249v1 Announce Type: new Abstract: Religious radio is a widespread but understudied form of mass communication in the United States, and content-level analysis of it has been constrained by the absence of large-scale transcript data. This Data Descriptor presents a c…