Infoblock · public evidence graph · transformer

Softmax self-attention (the Transformer's core operation) is exactly ONE update step of a continuous modern Hopfield network with an exponential energy — i.e. associative-memory attractor retrieval — with pattern-storage capacity scaling ~2^(d/2) and the softmax inverse-temperature beta as the retrieval-sharpness knob.

Grade E2 · state: verified · id: IB-attention-hopfield

Quantities: attention = 1 modern-Hopfield update; pattern-storage capacity ~2^(d/2); knob = softmax inverse-temperature beta

Primary source

Ramsauer et al. 2020, ICLR — 'Hopfield Networks is All You Need'

Edges (outgoing)

Incoming edges (what points here)