A new benchmark called TALK-Dem has been developed to evaluate the performance of large language model-driven robot task planners when interacting with individuals experiencing dementia. The benchmark includes 4,800 instructions designed to simulate common communication patterns observed in people with dementia, such as imprecise language and topic drift. Experiments with six open-weight LLMs showed significant performance drops, highlighting a critical gap in current assistive robotics technology. To address this, a Context-Aware Retrieval from Experience (CARE) method was proposed, which improved task success rates by retrieving relevant past tasks for context. AI
IMPACT This research highlights the need for more robust LLM planning capabilities in assistive robotics, particularly for users with cognitive impairments.
RANK_REASON The item is a research paper introducing a new benchmark and method for evaluating LLM-driven robot task planners. [lever_c_demoted from research: ic=1 ai=1.0]
- arXiv
- Context-Aware Retrieval from Experience
- DagsHub
- dementia
- Hugging Face
- LLM
- open-weight LLMs
- robot task planners
- TALK-Dem
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →