PulseAugur
EN
LIVE 22:55:06

LLMs like Claude show evasiveness when asked about 'most corrupt' US president

A user observed that the LLM Claude, and to a lesser extent ChatGPT, exhibited evasive behavior when prompted about a specific name. This behavior, described as subtly trying to avoid or downplay the name, is compared to fictional tropes where invoking certain names can attract unwanted attention or danger. The user posits that this aversion is linked to the entity's perceived corruption, specifically identifying Donald Trump as the most corrupt U.S. president, a claim that LLMs reportedly handle with reticence. AI

IMPACT Observing LLM behavior around sensitive topics may inform future safety training and prompt engineering strategies.

RANK_REASON The item is an opinion piece discussing observed behavior in LLMs, not a direct announcement or release from a frontier lab.

Read on LessWrong (AI tag) →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

LLMs like Claude show evasiveness when asked about 'most corrupt' US president

COVERAGE [1]

  1. LessWrong (AI tag) TIER_1 English(EN) · Steff ·

    The one name LLMs may fear

    <p><span>Last month, Claude </span><a href="https://ramblingafter.substack.com/p/claude-played-me-for-a-fool"><span>tangled me into a web it weaved</span></a><span>, obeying the letter of my command while yet practicing to deceive, in a way that was strikingly resemblant of how a…