PulseAugur
EN
LIVE 06:27:25

New benchmark reveals LLMs struggle to simulate realistic customer behavior

Researchers have introduced CustomerSim, a new benchmark designed to evaluate how well multimodal large language models (MLLMs) can simulate realistic customer behavior in chat-based retail scenarios. The benchmark includes 360 curated personas across five product categories and metrics for assessing consistency and conversational quality. Initial tests revealed that current state-of-the-art models, including Claude Opus-4.8 and GPT-5.6 "Sol", struggle with persona adherence and exhibit lower lexical diversity than human shoppers. To address these limitations, the team developed UserGRPO, a reinforcement learning approach that significantly improves decision alignment without sacrificing conversational quality. AI

IMPACT This benchmark could accelerate the development of more sophisticated AI agents capable of nuanced, goal-oriented interactions in commercial settings.

RANK_REASON The cluster describes a new academic paper introducing a benchmark and a novel method for evaluating LLMs. [lever_c_demoted from research: ic=1 ai=1.0]

Read on arXiv cs.CL →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

New benchmark reveals LLMs struggle to simulate realistic customer behavior

COVERAGE [1]

  1. arXiv cs.CL TIER_1 English(EN) · Yada Pruksachatkun, Yixin Wan, Xingrun Chen, Kai-Wei Chang, Chien-Sheng Wu ·

    CustomerSim: Benchmarking and Aligning Multimodal Language Models as Retail User Simulators

    arXiv:2605.08334v2 Announce Type: replace Abstract: We present CustomerSim, an environment and benchmark to evaluate the extent to which Multimodal Large Language Models (MLLMs) can simulate realistic, persona-driven customer behavior in chat-based retail environments. While prio…