Researchers have introduced InteractBench, a new benchmark designed to evaluate the algorithmic reasoning capabilities of large language models (LLMs) on interactive problems. These problems, common in competitive programming, require models to dynamically acquire information through queries rather than having all inputs provided upfront. The benchmark includes 322 problems from platforms like Codeforces and AtCoder, along with local interactors for offline evaluation. Initial results show a significant gap in performance, with current advanced models struggling with interactive tasks, often failing due to protocol violations or exceeding query budgets, in addition to algorithmic logic errors. AI
IMPACT Highlights a new frontier for LLM evaluation, pushing beyond static problem-solving to dynamic, interactive reasoning.
RANK_REASON The cluster contains a research paper introducing a new benchmark for evaluating LLMs. [lever_c_demoted from research: ic=1 ai=1.0]
- AtCoder
- Codeforces
- Hugging Face
- InteractBench
- International Collegiate Programming Contest
- International Olympiad in Informatics
- LLMs
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →