PulseAugur
EN
LIVE 17:44:01

New benchmark tests LLMs on dynamic user profiling from streaming content

A new benchmark, StreamProfileBench, has been introduced to evaluate how well Large Language Models (LLMs) can infer user profiles from continuously arriving content, a task that current static evaluations overlook. The benchmark includes a dataset of over 120,000 user-generated content posts from 7,000+ real users across five platforms. Experiments with 14 LLMs revealed that models struggle with continuous profile updating, exhibiting a bias towards retaining old interests and failing to recognize decaying ones. AI

IMPACT Highlights limitations in LLMs' ability to adapt to evolving user interests in real-time systems.

RANK_REASON The cluster describes a new academic paper introducing a benchmark and dataset for evaluating LLMs.

Read on Hugging Face Daily Papers →

AI-generated summary · Google Gemini · from 3 sources. How we write summaries →

New benchmark tests LLMs on dynamic user profiling from streaming content

COVERAGE [3]

  1. arXiv cs.CL TIER_1 English(EN) · Sizhe Wang, Feiyu Duan, Juelin Wang, Liwen Zhang, Feiyu Duan ·

    StreamProfileBench: A Benchmark for Fine-Grained User Profile Inference in Real-World Streaming Scenarios

    arXiv:2605.25758v1 Announce Type: new Abstract: Large Language Models (LLMs) have reshaped user profiling, yet current evaluations mainly focus on static data snapshots. This paradigm overlooks the reality of personalized systems, where User-Generated Content (UGC) arrives contin…

  2. Hugging Face Daily Papers TIER_1 English(EN) ·

    StreamProfileBench: A Benchmark for Fine-Grained User Profile Inference in Real-World Streaming Scenarios

    Large Language Models (LLMs) have reshaped user profiling, yet current evaluations mainly focus on static data snapshots. This paradigm overlooks the reality of personalized systems, where User-Generated Content (UGC) arrives continuously and fine-grained profile evolve rapidly. …

  3. arXiv cs.CL TIER_1 English(EN) · Feiyu Duan ·

    StreamProfileBench: A Benchmark for Fine-Grained User Profile Inference in Real-World Streaming Scenarios

    Large Language Models (LLMs) have reshaped user profiling, yet current evaluations mainly focus on static data snapshots. This paradigm overlooks the reality of personalized systems, where User-Generated Content (UGC) arrives continuously and fine-grained profile evolve rapidly. …