PulseAugur
EN
LIVE 06:48:33

New benchmark tests LLMs and humans on inferring social familiarity

Researchers have introduced FriendBench, a new benchmark designed to evaluate the ability of both humans and multimodal large language models (LLMs) to infer familiarity between two people based on a short conversation clip. The benchmark uses 20-second video, audio, or text segments of dyadic ice-breaker conversations. Across various modalities, 26 models from seven companies were compared against human panels, with the top-performing models achieving accuracy indistinguishable from humans. However, a key difference emerged: while humans maintained a balanced prediction, the strongest models showed a tendency to predict "stranger," suggesting a difference in their effective prior assumptions rather than discriminatory ability. AI

IMPACT This benchmark could advance research into social intelligence in AI, potentially leading to more nuanced and context-aware multimodal models.

RANK_REASON The cluster describes a new academic paper introducing a benchmark for evaluating LLMs and humans on a specific task. [lever_c_demoted from research: ic=1 ai=1.0]

Read on arXiv cs.AI →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

New benchmark tests LLMs and humans on inferring social familiarity

COVERAGE [1]

  1. arXiv cs.AI TIER_1 English(EN) · Jeffrey M. Girard, Jason Z. Zheng, Jacqueline R. Vertino, Antony D'Avirro, Benjamin Peloquin ·

    FriendBench: Benchmarking Dyadic Familiarity Inference in Humans and Multimodal Large Language Models

    arXiv:2607.29602v1 Announce Type: cross Abstract: Reading a social situation often depends on behavior, not words alone. We introduce FriendBench, a benchmark for inferring whether two people are already familiar or are meeting as strangers, from a 20-second clip of a dyadic ice-…