PulseAugur
实时 20:35:53

新基准测试 LLM 从流式内容进行动态用户画像的能力

引入了一个新的基准测试 StreamProfileBench,用于评估大型语言模型 (LLM) 从连续到达的内容中推断用户画像的能力,这项任务是当前静态评估所忽略的。该基准测试包含一个数据集,其中包含来自 7,000 多名真实用户在五个平台上的 120,000 多条用户生成内容帖子。对 14 个 LLM 的实验显示,模型在连续更新用户画像方面存在困难,表现出保留旧兴趣的偏见,并且未能识别衰退的兴趣。 AI

影响 凸显了 LLM 在实时系统中适应不断变化的用��兴趣方面的局限性。

排序理由 该集群描述了一篇介绍用于评估 LLM 的基准测试和数据集的新学术论文。

在 Hugging Face Daily Papers 阅读 →

AI 生成摘要 · Google Gemini · 来自 3 个来源。 我们如何撰写摘要 →

新基准测试 LLM 从流式内容进行动态用户画像的能力

报道来源 [3]

  1. arXiv cs.CL TIER_1 English(EN) · Sizhe Wang, Feiyu Duan, Juelin Wang, Liwen Zhang, Feiyu Duan ·

    StreamProfileBench:真实流式场景下细粒度用户画像推断的基准测试

    arXiv:2605.25758v1 Announce Type: new Abstract: Large Language Models (LLMs) have reshaped user profiling, yet current evaluations mainly focus on static data snapshots. This paradigm overlooks the reality of personalized systems, where User-Generated Content (UGC) arrives contin…

  2. Hugging Face Daily Papers TIER_1 English(EN) ·

    StreamProfileBench:真实流式场景下细粒度用户画像推断的基准测试

    Large Language Models (LLMs) have reshaped user profiling, yet current evaluations mainly focus on static data snapshots. This paradigm overlooks the reality of personalized systems, where User-Generated Content (UGC) arrives continuously and fine-grained profile evolve rapidly. …

  3. arXiv cs.CL TIER_1 English(EN) · Feiyu Duan ·

    StreamProfileBench:真实流媒体场景下细粒度用户画像推断的基准测试

    Large Language Models (LLMs) have reshaped user profiling, yet current evaluations mainly focus on static data snapshots. This paradigm overlooks the reality of personalized systems, where User-Generated Content (UGC) arrives continuously and fine-grained profile evolve rapidly. …