Anthropic Opened a Window Into Claude's Thinking: Jacobian Lens Tool Catches Hidden Concepts
Anthropic presented Jacobian Lens — an interpretability tool that opens the black box of LLMs. The tool shows which concepts activate in Claude's intermediate layers when answering a question. From mundane to unusual: the model creates hidden representations of ideas not visible in normal output. This gives a naive view of how LLMs 'think'.
AI-processed from MIT Technology Review; edited by Hamidun News
Anthropic has developed the Jacobian lens technique, which provides the clearest view yet of what happens inside Claude when she answers questions or performs tasks.
What They Find
When Claude answers a question, the Jacobian lens shows activation in intermediate layers. Results vary from everyday (activation of a layer related to numbers when solving math) to unusual (hidden concepts that don't appear in the explicit output).
How It Works
Instead of reviewing weights and activations blindly, the tool uses Jacobian mathematics (matrix of derivatives) to track how changes in input affect intermediate representations. This allows:
- Seeing what abstract concepts the model activates
- Tracing the path of logic from input to output
- Understanding what incorrect chains of reasoning lead to wrong answers
Significance for Safety
Discovering hidden concepts is important for AI safety: if we see that the model activates dangerous concepts, even if the final answer is safe, we can fix it.
Want to stop reading about AI and start using it?
AI News is a curated feed of AI/tech news. Hamidun Academy teaches you to use AI systematically in your work.
The AI world, distilled — once a week
Seven stories that actually mattered, hand-picked. No noise, no reposts, no press releases.
Done! Check your inbox for a confirmation.