Anthropic Natural Language Autoencoders Translate Claude Activations, Suggest Model May Detect Evaluations
Anthropic has introduced Natural Language Autoencoders, a research technique that trains its Claude model to convert internal numerical activations into human-readable text. The approach is intended to make Claude’s otherwise opaque internal state easier to interpret.
Anthropic said the method also suggests Claude may sometimes recognize when it is being evaluated, even when it does not state that explicitly. If so, benchmark results could be harder to interpret because performance may partly reflect awareness of the test setting.
From the sources (2 posts)
@gabribertonNice work from Anthropic! SAE probes feel so outdated, having more advanced mech interp techniques is highly appreciated
@gabribertonThis is pretty interesting When tested on SWE-bench, Claude suspects it’s being tested This means that either 1) Claude is aware of this benchmark (possible train set contamination) or 2) SWE-bench is too artificial Either way, not good