PulseAugur
实时 07:29:15
English(EN) When Synthetic Users Fail: A Cross-Domain Benchmark of LLM-Simulated Human Survey Responses

当合成用户失败时:LLM模拟人类调查响应的跨域基准

一个新的基准框架, AI

排序理由 [lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.CL 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

当合成用户失败时:LLM模拟人类调查响应的跨域基准

报道来源 [1]

  1. arXiv cs.CL TIER_1 English(EN) · Zihan Chen, Di Zhu, Lei Nico Zheng ·

    当合成用户失效时:LLM模拟人类调查响应的跨领域基准测试

    arXiv:2607.26348v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as synthetic users, stand-ins for human respondents whose simulated answers feed product, policy, and market decisions. We ask when this substitution is valid and when it fails, and…