Researchers have introduced TREK, a new benchmark designed to rigorously evaluate the capabilities of LLM agents in complex trip planning scenarios. TREK addresses limitations in existing benchmarks by focusing on producing fully feasible itineraries that are constraint-correct, hallucination-free, spatio-temporally executable, and budget-valid, while also responding to unstated traveler needs. The benchmark includes 800 multi-constraint tasks and a deterministic evaluator, finding that even advanced models like GPT-5.6 achieve full feasibility on only 46.2% of solvable tasks, with unstated needs posing a significant challenge. AI
IMPACT This benchmark will help researchers identify and address critical limitations in LLM agent planning capabilities, particularly concerning unstated user needs and overall feasibility.
RANK_REASON The cluster contains an academic paper introducing a new benchmark and evaluation kit for LLM agents. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →