Researchers have developed a new corpus of 33 clinical cognitive assessment conversations, totaling 8,250 utterances, annotated with speaker roles and 56 dialogue acts. This dataset is designed to benchmark large language models (LLMs) in fine-grained dialogue-act classification and next-patient-utterance generation within a clinical context. Experiments using the LLaMA-3.1-8B model showed that instruction tuning improved performance, with reasoning-aware fine-tuning yielding the best classification results, although models still struggle with distinguishing closely related dialogue acts. AI
IMPACT This research provides a framework for developing more sophisticated AI agents capable of understanding and generating nuanced clinical dialogue.
RANK_REASON The cluster contains an academic paper detailing a new dataset and benchmark for LLM evaluation. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →