Researchers have created a large-scale corpus of religious radio broadcast transcripts from webstream recordings in the United States. This dataset, captured over a month in July 2025, includes over 700,000 recordings and more than 60 million lines of speech, transcribed and speaker-diarized using an automated pipeline and labeled with programming format and topic by a large language model. The corpus is designed to facilitate the study of religious broadcasting, its discussion of social and political issues, and speech-processing research in this underrepresented domain. AI
IMPACT Provides a new dataset for speech processing and content analysis in religious media.
RANK_REASON The item is a data descriptor paper for a large-scale corpus. [lever_c_demoted from research: ic=1 ai=0.7]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →