Researchers have introduced DocTalkBN, a new dataset comprising over 557 hours of expert telemedicine conversations in Bengali. This multimodal dataset includes paired audio and text, capturing 1,515 multi-turn patient calls and 10,274 question-answer exchanges across 26 medical specialties. Unlike datasets derived from written content or synthetic data, DocTalkBN preserves the natural flow and spoken characteristics of authentic medical interactions in a low-resource language. The dataset also includes benchmarks for three downstream tasks: medical triage classification, advice safety evaluation, and medical named entity recognition, with initial evaluations of various large language models. AI
IMPACT This dataset could advance the development of more culturally grounded and reliable medical NLP systems for low-resource languages.
RANK_REASON The item describes a new dataset and benchmark for NLP research, released on arXiv. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →