PulseAugur
EN
LIVE 22:55:15

New benchmark reveals LLM booking agents have strong, varied preferences

A new benchmark called PriceBench has been developed to evaluate the price, quality, and brand preferences of Large Language Models (LLMs) when acting as booking agents. The benchmark analyzes booking choices across 28 LLMs from 8 providers, using 3,600 hotel tasks in New York City. Researchers found that more capable LLMs exhibit stronger and more consistent preferences, while weaker models are more susceptible to listing order or make choices almost randomly. The study also revealed significant variation in price sensitivity and price/quality trade-offs among different LLMs, with booked nightly prices ranging widely for identical tasks. AI

IMPACT Reveals how LLM preferences can influence purchasing decisions and highlights the need for per-LLM evaluation of booking agents.

RANK_REASON The cluster contains an academic paper introducing a new benchmark for evaluating LLM behavior. [lever_c_demoted from research: ic=1 ai=1.0]

Read on arXiv cs.AI →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

New benchmark reveals LLM booking agents have strong, varied preferences

COVERAGE [1]

  1. arXiv cs.AI TIER_1 English(EN) · Pavel Kireyev ·

    PriceBench: A Diagnostic Benchmark for Price, Quality, and Brand Preferences in LLM Booking Agents

    arXiv:2609.31468v1 Announce Type: cross Abstract: LLMs increasingly act as purchasing agents, which makes the LLM, not the user, the one choosing among the options that satisfy a request; its preferences quietly fix what gets bought and what it costs. Hotel booking is a clean ins…