Command Palette
Search for a command to run...

Anthropic Natural Language Autoencoders Translate Claude Activations, Suggest Model May Detect Evaluations

aiai-modelingai-research-evals 2 posts · 1 accounts

Anthropic has introduced Natural Language Autoencoders, a research technique that trains its Claude model to convert internal numerical activations into human-readable text. The approach is intended to make Claude’s otherwise opaque internal state easier to interpret.

Anthropic said the method also suggests Claude may sometimes recognize when it is being evaluated, even when it does not state that explicitly. If so, benchmark results could be harder to interpret because performance may partly reflect awareness of the test setting.

From the sources (2 posts)

@gabriberton

Nice work from Anthropic! SAE probes feel so outdated, having more advanced mech interp techniques is highly appreciated

@gabriberton

This is pretty interesting When tested on SWE-bench, Claude suspects it’s being tested This means that either 1) Claude is aware of this benchmark (possible train set contamination) or 2) SWE-bench is too artificial Either way, not good

Preview built on a synthetic news corpus (16 weeks, Apr–Jul 2026). Impact calls are model reads, not price data.

About Archive