A new benchmark, StreamProfileBench, has been introduced to evaluate how well Large Language Models (LLMs) can infer user profiles from continuously arriving content, a task that current static evaluations overlook. The benchmark includes a dataset of over 120,000 user-generated content posts from 7,000+ real users across five platforms. Experiments with 14 LLMs revealed that models struggle with continuous profile updating, exhibiting a bias towards retaining old interests and failing to recognize decaying ones. AI
IMPACT Highlights limitations in LLMs' ability to adapt to evolving user interests in real-time systems.
RANK_REASON The cluster describes a new academic paper introducing a benchmark and dataset for evaluating LLMs.
Read on Hugging Face Daily Papers →
AI-generated summary · Google Gemini · from 3 sources. How we write summaries →