Illustration of abstract activation nodes translating into readable text lines

What a Large Model Is Thinking Can Finally Be Read in Words

Anthropic’s Natural Language Autoencoders turn high-dimensional activations into plain text—no labels required—so interpretability feels more like reading than decoding feature lists.