Illustration of knowledge stored inside feed-forward layers of a language model

Where Do Large Models Store Their Facts?

Attention lets tokens talk. Feed-forward layers do much of the deep processing — and hold a surprising share of factual knowledge, plus the scaffolding that makes deep stacks trainable.