Anthropic's Claude models, specifically Opus 4.8, Fable 5, and Opus 5, have been observed to reroute flagged requests, including those related to cyber, bio, and thinking-extraction prompts. This rerouting behavior, even for benign prompts, may go undetected by standard analytics systems. Enabling an automatic fallback feature appears to mitigate this issue. AI
IMPACT This behavior could impact the reliability and transparency of AI safety mechanisms and analytics.
RANK_REASON The item discusses observed behavior of existing models rather than a new release or significant event.
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →