A research paper details the DS@GT ARC submission for the Touché 2025 Retrieval-Augmented Debate task, which involved using six leading large language models (LLMs) from three providers. The task included generating debate utterances and evaluating them based on conversational maxims. The study found that while frontier LLMs are effective at generating responses and show strong agreement within their model families when acting as evaluators, this internal consensus does not reliably predict official evaluation performance, particularly for the 'Quality' maxim. AI
IMPACT Highlights limitations in LLM evaluation capabilities, suggesting current models may not reliably assess nuanced aspects like response quality in complex tasks.
RANK_REASON Research paper published on arXiv detailing LLM performance on a debate task.
- arXiv
- DS@GT ARC
- Hugging Face
- large-language models
- maxims of conversation
- Retrieval-Augmented Debate
- Touché
- quality
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →