Anthropic researchers have identified a cognitive workspace within their Claude AI model, termed "J-Space," where the model formulates its reasoning before generating a response. This internal space holds a limited set of concepts, analogous to a mental scratchpad, allowing Claude to process more information than it explicitly communicates. A new open-source technique called J-lens, developed by Anthropic, allows for the interpretation of these internal concepts, revealing what Claude is attending to. This technology not only provides insight into Claude's reasoning but also enables direct manipulation of its internal state to alter responses, demonstrating a causal link between the J-Space and the model's output. AI
IMPACT Provides a new method for understanding and potentially controlling AI model reasoning, moving beyond correlation to causation.
RANK_REASON Research paper detailing a new internal mechanism of an AI model. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →