PulseAugur
EN
LIVE 06:23:06

New corpus of US religious radio transcripts released

Researchers have created a large-scale corpus of religious radio broadcast transcripts from webstream recordings in the United States. This dataset, captured over a month in July 2025, includes over 700,000 recordings and more than 60 million lines of speech, transcribed and speaker-diarized using an automated pipeline and labeled with programming format and topic by a large language model. The corpus is designed to facilitate the study of religious broadcasting, its discussion of social and political issues, and speech-processing research in this underrepresented domain. AI

IMPACT Provides a new dataset for speech processing and content analysis in religious media.

RANK_REASON The item is a data descriptor paper for a large-scale corpus. [lever_c_demoted from research: ic=1 ai=0.7]

Read on arXiv cs.CL →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

New corpus of US religious radio transcripts released

COVERAGE [1]

  1. arXiv cs.CL TIER_1 English(EN) · Samuel Bestvater, Athena Chapekis, Skyler Seets, Anna Lieb, Sono Shah, Aaron Smith ·

    A large-scale corpus of religious radio broadcast transcripts from webstream recordings in the United States

    arXiv:2607.26249v1 Announce Type: new Abstract: Religious radio is a widespread but understudied form of mass communication in the United States, and content-level analysis of it has been constrained by the absence of large-scale transcript data. This Data Descriptor presents a c…