Softmax self-attention (the Transformer's core operation) is exactly ONE update step of a continuous modern Hopfield network with an exponential energy — i.e. associative-memory attractor retrieval — with pattern-storage capacity scaling ~2^(d/2) and the softmax inverse-temperature beta as the retrieval-sharpness knob.