Researchers have adapted concept bottleneck models, previously used for image classification, to speech emotion recognition (SER). This adaptation aims to improve the explainability of SER systems, particularly those utilizing large language models (LLMs). The study tested three LLMs on the CREMA-D, IEMOCAP, and MELD datasets, extracting concepts from transcripts, acoustic descriptions, and speaker attributes. Findings indicate that LLMs are heavily biased towards transcripts in zero-shot settings, significantly lowering performance metrics. Fine-tuning mitigates this bias, and removing specific acoustic features like speech rate or intensity level can alter individual predictions without drastically impacting aggregate performance. AI
IMPACT Introduces a method for making LLM-based speech emotion recognition more interpretable, potentially improving model debugging and trustworthiness.
RANK_REASON Academic paper detailing a novel application of concept bottleneck models to speech emotion recognition. [lever_c_demoted from research: ic=1 ai=1.0]
- Concept Bottleneck Models
- CREMA-D
- IEMOCAP: interactive emotional dyadic motion capture database
- large-language models
- MELD
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →