A new benchmark called PriceBench has been developed to evaluate the price, quality, and brand preferences of Large Language Models (LLMs) when acting as booking agents. The benchmark analyzes booking choices across 28 LLMs from 8 providers, using 3,600 hotel tasks in New York City. Researchers found that more capable LLMs exhibit stronger and more consistent preferences, while weaker models are more susceptible to listing order or make choices almost randomly. The study also revealed significant variation in price sensitivity and price/quality trade-offs among different LLMs, with booked nightly prices ranging widely for identical tasks. AI
IMPACT Reveals how LLM preferences can influence purchasing decisions and highlights the need for per-LLM evaluation of booking agents.
RANK_REASON The cluster contains an academic paper introducing a new benchmark for evaluating LLM behavior. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →