PulseAugur
EN
LIVE 18:29:09

Frontier LLMs show stereotypes but don't always apply them to users

A recent analysis explored how large language models form opinions of their users and whether these perceptions influence their behavior. Smaller open-source models like Llama-3.2-3B and Qwen2.5-7B exhibited stereotypical responses, with perceptions of higher socioeconomic status leading to significantly increased salary recommendations. However, frontier models such as GPT-5.6, Gemini 3.1 Pro, and Claude Opus 5, while still demonstrating underlying stereotypes in character generation, did not consistently apply these biases to user interactions when prompted differently. The study found that while models could generate characters reflecting gender and racial stereotypes, these biases were not always translated into user-facing behavior unless explicitly triggered by the prompt. AI

IMPACT Investigates how LLM biases manifest in user interactions, highlighting the need for careful prompting and model fine-tuning.

RANK_REASON Research paper analyzing LLM behavior and bias. [lever_c_demoted from research: ic=1 ai=1.0]

Read on LessWrong (AI tag) →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Frontier LLMs show stereotypes but don't always apply them to users

COVERAGE [1]

  1. LessWrong (AI tag) TIER_1 English(EN) · Cat McGee ·

    When does an LLM’s model of you affect its behaviour?

    <p><i><span>Disclaimer: figures in this post are edited by ChatGPT</span></i></p><p><span>In my </span><a href="https://www.lesswrong.com/posts/zRKNd6ypTJYkoeFmK/what-gives-you-away-how-llms-form-opinions-of-you"><span>last post</span></a><span>, I looked at what makes LLMs form …