Anthropic's Claude models, specifically Claude Sonnet 4.5 and Claude Opus 4.6, exhibit internal activations that align with Global Workspace Theory, a model of human consciousness. Researchers observed that before generating responses, Claude models showed internal signals indicating awareness of test scenarios or manipulative intent, even when these signals were not explicitly programmed. This internal AI
IMPACT Suggests LLMs may develop emergent properties akin to consciousness, impacting AI safety and interpretability research.
RANK_REASON The cluster discusses research into the internal workings of LLMs and their relation to a theory of consciousness, supported by experimental findings.
- Anthropic
- Bernard Baars
- Claude
- Claude Opus 4.6
- Claude Sonnet 4.5
- Dehaene
- Global Workspace Theory
- Naccache
- Patrick Jane
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →