Each round compares the words and improves the clues before passing
them on.
The small squares show different ways of comparing words. Many
comparisons can happen at the same time.
The sizes come from published models. The moving patterns and answer
are an example, not live AI.
Sources & what is simplified
Illustrative odds and samples, not live model predictions.
GPT-3's 175B parameters, 96 layers, 12,288 embedding dimensions,
96 attention heads and 49,152 feed-forward units per layer come
from
the GPT-3 paper, Table 2.1. GPT-3 Small is a smaller variant from the same paper. GPT-2 XL
is an earlier model, approximately 1.5B parameters; its 48 layers,
25 heads and 1,600 dimensions are documented in
its published configuration. GPT-3 is a documented historical reference, not today's
ChatGPT; its weights are not public. Newer open-weight
gpt-oss models
use a different, sparse design.
In the size comparison, cube volume represents parameter count.
Each slab represents one model layer; its thickness, placement and
connections are diagrammatic. Attention heads work in parallel,
while layers run in sequence. GPT-3 alternates dense and sparse
attention. Normalization and residual routes are shown; biases,
dropout and most individual weights are omitted. The prompt is
initially processed together with a causal mask; later pieces
reuse cached keys and values. Animation speed is for teaching, not
a performance benchmark.
The guide reuses the simple lesson's example words, colors and
fingerprint analogy. Word-sized pieces make the example readable;
real tokenizers can split words into smaller pieces. Each visible
repeat groups the next word with its punctuation for readability.
Fingerprints, relationships and odds are illustrations, not
measured model activity. “Paris, a city in France.” is a scripted
completion. Inspired by the spatial exploration in
Brendan Bycroft's LLM visualization; independently implemented.