PulseAugur
中
实时 18:45:44

新基准测试 LLM 从流式内容进行动态用户画像的能力

引入了一个新的基准测试 StreamProfileBench,用于评估大型语言模型 (LLM) 从连续到达的内容中推断用户画像的能力,这项任务是当前静态评估所忽略的。该基准测试包含一个数据集,其中包含来自 7,000 多名真实用户在五个平台上的 120,000 多条用户生成内容帖子。对 14 个 LLM 的实验显示,模型在连续更新用户画像方面存在困难,表现出保留旧兴趣的偏见,并且未能识别衰退的兴趣。 AI

影响 凸显了 LLM 在实时系统中适应不断变化的用��兴趣方面的局限性。

排序理由 该集群描述了一篇介绍用于评估 LLM 的基准测试和数据集的新学术论文。

在 Hugging Face Daily Papers 阅读 →

AI 生成摘要 · Google Gemini · 来自 3 个来源。 我们如何撰写摘要 →

新基准测试 LLM 从流式内容进行动态用户画像的能力

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
该集群描述了一篇介绍用于评估 LLM 的基准测试和数据集的新学术论文。
Source corroboration
3 independent sources
Multiple independent publishers reporting the same story raises confidence that it's real and newsworthy.
Topics
paper, other
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
136 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准。

报道来源 [3]

  1. arXiv cs.CL TIER_1 English(EN) · Sizhe Wang, Feiyu Duan, Juelin Wang, Liwen Zhang, Feiyu Duan ·

    StreamProfileBench:真实流式场景下细粒度用户画像推断的基准测试

    arXiv:2605.25758v1 Announce Type: new Abstract: Large Language Models (LLMs) have reshaped user profiling, yet current evaluations mainly focus on static data snapshots. This paradigm overlooks the reality of personalized systems, where User-Generated Content (UGC) arrives contin…

  2. Hugging Face Daily Papers TIER_1 English(EN) ·

    StreamProfileBench:真实流式场景下细粒度用户画像推断的基准测试

    Large Language Models (LLMs) have reshaped user profiling, yet current evaluations mainly focus on static data snapshots. This paradigm overlooks the reality of personalized systems, where User-Generated Content (UGC) arrives continuously and fine-grained profile evolve rapidly. …

  3. arXiv cs.CL TIER_1 English(EN) · Feiyu Duan ·

    StreamProfileBench:真实流媒体场景下细粒度用户画像推断的基准测试

    Large Language Models (LLMs) have reshaped user profiling, yet current evaluations mainly focus on static data snapshots. This paradigm overlooks the reality of personalized systems, where User-Generated Content (UGC) arrives continuously and fine-grained profile evolve rapidly. …