CEG5305 · Generative AI with Foundation Models

Home › Part 2 › Lecture 1

Part 2 - Lecture 1

Architecture of Large Language Models

Topic 4, Architecture and Training of LLMs. Session 01 covers the architecture; session 02 covers training.

Session 01 · 3 hours, three partsLecturer · Junyuan HongSlides · open the slidesNext · Session 02Homework · for practice, no credits
Why do LLMs work so well?

And the part of that question we can test today: is it the architecture, or the scale?

Part A

History of Language Models

See three models side by side, then trace the history from counting to attention.

Part B

What makes Transformers great?

One controlled experiment (ConvNeXt), then the six components inside nanoGPT.

Part C

Break a GPT

Remove components from a small GPT one at a time and measure. Then you do the same as homework.

Every diagram on this page is drawn in HTML and most are interactive. Where a diagram follows a paper's figure, the paper is named under it. Embedded demos belong to their authors and load only when you click.

PART A

History of Language Models

See the gap first, then trace where it came from.

Three models, side by side

Run three language models side by side and compare what comes back: a table of counts (1948), a small LSTM (1997), and GPT-2 small (2019), a Transformer with 124M parameters. All three take the same prompt. The n-gram model and the LSTM learn from the same book, Alice's Adventures in Wonderland; GPT-2 read about 8 million web pages (Radford et al., 2019). The difference is obvious before any theory. The rest of the lecture works out where it comes from.

All three designs have been used in products. In 2007 Google's machine translation used n-gram tables with up to 300 billion entries (Brants et al., 2007). In 2016 Google Translate replaced its phrase-based statistical system with an LSTM network (Wu et al., 2016). ChatGPT is fine-tuned from a much larger GPT model (OpenAI, 2022).

Build word or character n-grams from any text and generate one token at a time, with the top-5 next-token probabilities shown.
A two-layer LSTM (Hochreiter & Schmidhuber, 1997) that reads and writes one character at a time. It was trained on Alice's Adventures in Wonderland only, so it has seen the book the n-gram model counts, and nothing else.
GPT-2 small (Radford et al., 2019) runs on your own machine through Transformers.js. Nothing is sent to a server. It was trained only to predict the next token, so it continues your text rather than answering it.
Hypothesis 1 · today

Architecture

The Transformer can do something the older models structurally cannot. Testable by building models of matched size and removing pieces.

Hypothesis 2 · session 02

Scale

Data and compute. Testable with scaling laws: predict the loss before spending the money.

They are not rivals. Both are true, and the useful question is how much each one explains.

History of generative models: from n-grams to the Transformer

  1. 1948n-gram: Shannon approximates English with letter and word n-grams.
  2. 1990RNN: the simple recurrent network (Elman).
  3. 1997LSTM (Hochreiter & Schmidhuber).
  4. 2003NPLM, the neural probabilistic language model: learned word vectors and an MLP (Bengio et al.).
  5. 2010RNN language model (Mikolov et al.).
  6. 2014RNN encoder–decoder (seq2seq, Sutskever et al.), plus attention for translation (Bahdanau et al.): still built on an RNN.
  7. 2017Transformer: attention replaces the RNN (Vaswani et al.).
  8. 2018–20GPT, GPT-2, GPT-3: the decoder alone, pre-trained to predict the next token, then scaled up.

Each step below gets the same treatment: the structure, what it fixed, and what it left broken.

n-grams: counting

The Markov assumption keeps only the last \(n-1\) tokens: \[ p(x_t \mid x_1,\dots,x_{t-1}) \approx p(x_t \mid x_{t-n+1},\dots,x_{t-1}) = \frac{\mathrm{count}(x_{t-n+1},\dots,x_{t-1},\textcolor{#ef7c00}{x_t})}{\mathrm{count}(x_{t-n+1},\dots,x_{t-1})} \] There is nothing to learn. The table of counts is the whole model.

WORKEDFast and interpretable; ran speech recognition and translation for decades. Google’s 2007 translation system counted 5-grams in 2 trillion tokens of text (Brants et al., 2007).
BROKENo generalisation: “the cat sat on the” and “the dog sat on the” are unrelated rows. Most contexts are never seen.

NPLM (2003): learn word vectors instead of counting

The neural probabilistic language model of Bengio et al. (2003) gives every word a learned vector, look up the vectors for a fixed window of previous words, and feed them to an MLP. Similar words end up with similar vectors, so what is learned about “cat” carries over to “dog”.

The vector of word \(w\) is \(C(w)\), row \(w\) of a \(|V| \times m\) table \(C\) of learned numbers. For a window of \(n-1\) words the model computes (eq. 1 of the paper, without its optional direct connections)

\[ x = \big(C(w_{t-n+1}), \dots, C(w_{t-1})\big), \qquad h = \tanh(Hx + d), \qquad P(w_t = i \mid \text{context}) = \mathrm{softmax}(Uh + b)_i . \]

\(C, H, d, U, b\) are all trained together by gradient descent. \(H\) is \(h \times (n-1)m\): one block of columns for each position in the window. The window size is built into the weights, so a longer window is a different model, and a word outside the window has no way in.

FIXEDGeneralisation: words share structure through their vectors.
DID NOTThe window is still fixed in size, and each position has its own weights.

How neural models learn: gradient descent and backpropagation

An n-gram model is trained by counting. From the NPLM on, every model here has weights that are learned, and training repeats three steps. Compute the loss on some text: the cross-entropy of the true next token from Part 1 (lecture 02), \( \mathcal{L} = -\log p_\theta(x_{t+1} \mid x_{\le t}) \). Compute its gradient \(\nabla_\theta \mathcal{L}\), which says how the loss changes as each weight changes. Move every weight a small step against it, \( \theta \leftarrow \theta - \eta\,\nabla_\theta \mathcal{L} \), where \(\eta\) is the learning rate.

Backpropagation (Rumelhart et al., 1986) computes all the gradients with the chain rule, from the loss back towards the input. For a stack of \(L\) layers with outputs \(h_1, \dots, h_L\), \[ \frac{\partial \mathcal{L}}{\partial h_1} = \frac{\partial \mathcal{L}}{\partial h_L}\,\frac{\partial h_L}{\partial h_{L-1}} \cdots \frac{\partial h_2}{\partial h_1}, \] one factor per layer, or per step in an RNN. If the factors are mostly smaller than 1, the product shrinks towards zero and the early layers barely learn (vanishing gradients); if they are larger than 1, it explodes (Bengio et al., 1994). Several designs in this lecture exist because of this product: the LSTM's cell state and residual connections (C4) give the gradient a path on which the factors stay near 1, and LayerNorm (C5) keeps the size of each layer's input fixed.

Left: gradient descent on one weight, \(L(w) = \tfrac12 (w-1)^2\), one step per click; try a learning rate above 2. Right: one number \(g\) stands for the size of each factor \(\partial h_{k+1}/\partial h_k\). An illustration, not a trained model.

RNN: carry a hidden state forward

A recurrent neural network (Elman, 1990) updates one hidden state at every step, \(h_t = \tanh(W h_{t-1} + U x_t)\), so context is no longer limited to a window. The same weights are used at every step.

FIXEDUnbounded context, in principle.
DID NOTEarly inputs fade (or blow up) as the state is rewritten (Bengio et al., 1994), and training must run one step after another.

LSTM: gates that let gradients survive

The LSTM (Hochreiter & Schmidhuber, 1997) adds a cell state that is updated by addition, with gates deciding what to forget, what to write and what to output. When the forget gate (Gers et al., 2000) stays near 1, information and gradients pass through many steps.

An in-browser recurrent network that generates handwriting stroke by stroke, based on Graves (2013) — the same paper the cell diagram above follows. Turn the Speed slider up; the default is slow.
FIXEDMemory over hundreds of steps. LSTMs were the standard sequence model in NLP for years.
DID NOTNothing about the sequential bottleneck: step \(t\) still waits for step \(t-1\).

Seq2seq + attention (2014–15): a patch

In seq2seq translation (Sutskever et al., 2014) the encoder compresses the whole source sentence into one vector. Attention (Bahdanau et al., 2015) was introduced to let the decoder look back at every encoder state. It was a fix on top of the RNN, not a replacement.

Transformer (2017): keep attention, drop recurrence

The Transformer (Vaswani et al., 2017) is attention and an MLP, stacked, with position added explicitly. GPT-style models use only the decoder half. Click through the diagram; each component has its own section in Part B.

FIXEDBoth problems at once: any token reaches any earlier token in one step, and all positions train in parallel.
DID NOTAttention compares every pair of tokens, so time and memory grow as \(n^2\).
Type a prompt and press Generate: embeddings, Q/K/V, attention, MLP and the next-token probabilities animate for the new token. Temperature and top-k are adjustable.
A guided 3D walk through a GPT: nano-gpt (85,584 parameters), GPT-2 and GPT-3 side by side. Use full screen: the view needs a wide window.

RNN vs Transformer: two numbers explain it

Five architectures, four questions

ModelContextPath between two tokensParallel trainingWhy it lost / won
n-gramlast \(n-1\) tokens—n/a (counting)No generalisation
NPLM (2003)fixed window1 step, inside the windowyesThe window cannot grow
RNNunbounded in principle\(O(n)\)noForgets; sequential training
LSTMlong\(O(n)\)noMemory fixed; bottleneck not
Transformerthe full context window\(O(1)\)yesBoth at once, at \(O(n^2)\) cost

Transformers beyond language

The Transformer did not stay in language. Cut the input into tokens and the same Transformer applies; only the two ends change: how the input becomes tokens, and what comes out. One paper per domain:

Images

ViT · Dosovitskiy et al., 2021

One token per 16×16 image patch.

Image generation

DiT · Peebles & Xie, 2023

A diffusion model whose denoising network is a Transformer over latent patches, in place of the usual U-Net.

Video

ViViT · Arnab et al., 2021

One token per space-time tube of the video.

Speech

Whisper · Radford et al., 2023

An encoder–decoder, like the 2017 Transformer, from spectrogram to text; trained on 680,000 hours of audio.

Time series

PatchTST · Nie et al., 2023

One token per window of the series; forecasts what comes next.

Proteins

ESM-2 · Lin et al., 2023

One token per amino acid; a language model of up to 15 billion parameters, from which 3D structure is predicted.

The sketches are simplified redrawings, not the papers' figures. Examples: a DiT-XL/2 sample from facebookresearch/DiT (CC BY-NC 4.0), shown in pixels although DiT works on a compressed latent; Eadweard Muybridge, The Horse in Motion (1878, public domain, via Wikimedia Commons); ubiquitin drawn from the coordinates of PDB 1UBQ; the eagle is a U.S. Fish and Wildlife Service photo and the speech is Neil Armstrong’s “That’s one small step for man” (Apollo 11, NASA), both public domain, via Wikimedia Commons.

Where that leaves us. Two structural changes: information travels in one hop, and training parallelises. That is a satisfying answer. Part B shows it is also incomplete.
PART B

What makes Transformers great?

One controlled experiment, then the components themselves.

The comparison we just made was not controlled

A Transformer differs from a ConvNet or an RNN in many ways at once: attention, normalisation, activation, optimiser, training recipe, data augmentation, depth ratios. When one model beats another on a dozen axes, the result does not say which axis mattered.

Liu et al. (2022), A ConvNet for the 2020s, ran the controlled version. Start from a ResNet-50 (He et al., 2016), apply the design decisions that came with vision Transformers one at a time, and measure after every change.

It shows

Credit for the ConvNet-vs-Transformer gap had gone partly to the wrong place. The training recipe alone was worth +2.7 points, and twelve design changes added up to a large one.

It does not show

That attention is useless. This is ImageNet classification: fixed-size images, no long-range dependency across a sentence. Language is where one-hop paths pay off.

The method outlasts the facts. Ablation — change one thing, hold everything else fixed, measure — is how credit is assigned. Part C does it to a GPT, and the homework asks you to do it too.

The six components inside nanoGPT

nanoGPT implements a GPT-2-style model in about 300 lines of model.py. It has six kinds of component, C1–C6. Tokens enter at the bottom, become vectors (C1), pass through n_layer identical blocks (C2–C5), and leave as next-token probabilities (C6). Each code link below opens that part of model.py.

What is it actually trained to do?

Maximum likelihood over the training text, factorised token by token: \[ \theta^* = \arg\max_\theta \sum_i \log p_\theta(x^{(i)}),\qquad \log p_\theta(x) = \sum_t \log p_\theta(x_t \mid x_{<t}) \] Negate and average over tokens to get the loss, and exponentiate to get perplexity: \[ \mathcal{L} = -\frac{1}{T}\sum_t \log p_\theta(x_t \mid x_{<t}),\qquad \mathrm{PPL} = e^{\mathcal{L}} \] Nothing in this lecture changes the objective. The architectures above, and the six components below, are different ways to compute \(p_\theta\).

In training, every position of a sequence predicts its next token at once, and the loss is the average of \(-\log p\) over those predictions: the cross-entropy with the true next token. Gradient descent lowers it (how neural models learn).

C1 · Token and position embeddings

\( x = E[\text{idx}] + P[0,\dots,T-1] \). The token table is the 2003 idea unchanged. The position table is needed because attention has no built-in notion of order.

C2 · Causal self-attention

\[ \mathrm{Attn}(Q,K,V) = \mathrm{softmax}\!\left(\frac{QK^\top}{\sqrt{d_k}} + M\right)V,\qquad Q = XW_Q,\; K = XW_K,\; V = XW_V \] This is the only place in the model where tokens exchange information. Remove it and every position is processed independently.

C3 · The MLP, with GELU

\( \mathrm{MLP}(x) = W_2\,\mathrm{GELU}(W_1 x) \) with a hidden width of \(4d\) and the GELU activation (Hendrycks & Gimpel, 2016). It works on each position separately. Most people guess attention holds most of the parameters; the MLP does.

Some nonlinearity is needed, or \(W_2 W_1 x\) collapses to one matrix. Which one is an empirical choice: the 2017 Transformer used ReLU, GPT-1 (Radford et al., 2018) and BERT (Devlin et al., 2019) switched to GELU, and Llama (Touvron et al., 2023) uses SwiGLU (Shazeer, 2020). In ConvNeXt, swapping ReLU for GELU left accuracy unchanged at 80.6% (Liu et al., 2022).

The difference from ReLU is in the slope, which is what a gradient is multiplied by on its way back through the unit. For \(x < 0\), ReLU's slope is exactly 0, so no gradient passes; GELU, \(x\,\Phi(x)\) with \(\Phi\) the standard normal CDF, is smooth, and its slope there is small but not 0 (except at its minimum, near \(x = -0.75\)). Hendrycks & Gimpel reported that it beat ReLU on all the vision, language and speech tasks they tested; GPT-1 adopted it, and BERT notes that it uses GELU “following OpenAI GPT”.

C4 · Residual connections

\( x \leftarrow x + \mathrm{Sublayer}(x) \). Unrolled, \( x_N = x_0 + \sum_n \mathrm{Sublayer}_n(x_n) \): there is a direct path from the input to the output, so in backpropagation the gradient reaches the first layer without being multiplied by one factor per layer (see above). Residual connections came from ResNet (He et al., 2016) and are still in every large model.

C5 · LayerNorm, and where to put it

LayerNorm (Ba et al., 2016): \( \mathrm{LN}(x) = \gamma \odot \dfrac{x-\mu}{\sqrt{\sigma^2+\epsilon}} + \beta \), where \(\mu\) and \(\sigma^2\) are the mean and variance of the \(d\) numbers in one token's vector, and \(\gamma\), \(\beta\) are learned. The 2017 Transformer (Vaswani et al., 2017) normalised after each residual addition; GPT-2 (Radford et al., 2019) and nanoGPT normalise before each sub-layer, which leaves the residual path untouched and makes deep stacks easier to train.

C6 · The output head and weight tying

\( \text{logits} = \mathrm{LN}(x)\,E^\top \). nanoGPT uses the embedding matrix again as the output layer (weight tying, Press & Wolf, 2017): the same space in, the same space out, and \(V\times d\) fewer parameters.

What nanoGPT does not have — yet

ComponentWhat it changes
RoPERotary position embedding: relative position, no learned table
RMSNormLayerNorm without subtracting the mean; cheaper, about as good
SwiGLUA gated MLP; the hidden width shrinks to \(8/3\cdot d\) to keep the parameter count
Grouped-query attentionFewer key/value heads than query heads; a smaller memory cache at inference
Mixture of expertsMany MLPs, only a few active for each token

All five are refinements of C1–C6 found in current open models. Session 02 returns to them.

PART C

Break a GPT

The ConvNeXt method, on a model small enough to run on a laptop.

Ablation, live

What we break. nanoGPT is a minimal GPT: a model file of about 300 lines with the six components above. We use a one-file copy with the same six parts, made small on purpose: TinyShakespeare (char-rnn), 1.1 MB of text with one character per token, and a model of 818,241 parameters, about 200,000 times fewer than GPT-3 (Brown et al., 2020). One training run takes about a minute on a laptop. Unlike nanoGPT, the copy does not tie the output head to the token embedding and has no dropout.

The baseline's training loss every 10 steps, mean of 3 seeds, from our laptop-CPU run of 29 Sep 2026. It starts near ln 65, the loss of a uniform guess over 65 characters. The parameters take 3.3 MB (818,241 × 4 bytes); GPT-3's would take 350 GB at 2 bytes each.
Held fixed, for every run

Data: TinyShakespeare, one character per token (65 characters); first 90% to train, last 10% to validate. Model: 4 blocks, 4 heads, width 128, context 64 characters, 818,241 parameters. Training: 3,000 steps of 32 sequences; AdamW, learning rate 3 × 10⁻⁴; gradients clipped at 1.0. Seeds: 1337, 1338, 1339. Measure: validation loss, averaged over 20 batches.

Changed, one at a time

Remove LayerNorm · remove residual connections · remove positional embeddings · collapse multi-head attention to one head.

Results. Validation loss for all five variants is added here after the session.

Reading an ablation honestly

  1. Train each version more than once. Use a different random seed each time. One run can be lucky or unlucky, so report the average and the range.
  2. Train every version the same way. Same number of steps, same data. Otherwise a difference may come from the training, not from the change.
  3. Say what else the change removed. Taking out a part can also take out parameters, so the model is smaller as well as different.
  4. Report changes that made no difference. “It made no measurable difference” is a real finding.

Homework, for practice

Homework is for practice. Free to do. No credits.

Run your own ablation on the small GPT. The goal: find out what one component does, by training the model with it and without it.

  1. Run the baseline: the base code below as given, 3 seeds.
  2. Change one component, nothing else: remove it or replace it, with the same data, steps and seeds. Choose one of today's four in more depth (does removing LayerNorm still help at 12 layers?), another (MLP width, depth, GELU → ReLU, tied vs untied embeddings), or add one nanoGPT lacks (RMSNorm, RoPE).
  3. Compare and explain: is the change in validation loss larger than the spread across seeds? What does the component do that explains it?

What you produce:

  • Your code change: a diff against ablation_base.py, with its fixed settings unchanged.
  • A table: validation loss, mean and range over 3 seeds, for the baseline and your version.
  • One plot: the training-loss curves of both.
  • A few sentences: what the component does, why the loss changed or did not, and whether you predicted it.

Before you conclude, check your runs against the four rules above, and check that the loss at step 0 is near ln 65 = 4.17.

Base code: ablation_base.py, the model and settings behind the class table: a nanoGPT-shaped Transformer (Karpathy, 2022) in one file, with the fixed settings at the top and the four class variants as switches. Add your variant as a new switch, then run python ablation_base.py --variant baseline and --variant your_variant; each runs the three seeds. The output file also holds the training loss every 10 steps, for your plot.

python -m venv gpt && source gpt/bin/activate   # optional: a clean environment
pip install torch                               # Python 3.9+; no GPU needed
curl -O https://costa-nus.github.io/CEG5305_GenAI/part2/lecture-1/code/ablation_base.py
python ablation_base.py --out baseline.json     # 3 seeds; downloads the data
# your change: add one entry to VARIANTS in ablation_base.py, e.g. "deeper": {"n_layer": 12}
python ablation_base.py --variant deeper --out mine.json

Where to run it: open the notebook in Colab (choose a T4 GPU runtime; about 35 s per run), or run it locally with Python 3.9+ and PyTorch. No GPU is needed: the run uses under 0.5 GB of RAM and takes about 75 s per run on a laptop CPU (measured on an Apple M5 Pro), so about 20 minutes for all 15 runs of the class table.

Resources

Reading

  • Vaswani et al., Attention Is All You Need, NeurIPS 2017. §3, Model Architecture; read §3.5 on positional encoding carefully. Read it as a historical document: most current models differ from it in several places this page points out.
  • Liu et al., A ConvNet for the 2020s, CVPR 2022. §2, the modernisation roadmap.
  • Raschka, Build a Large Language Model (From Scratch), Manning 2024. Chapters 2–4. Code: rasbt/LLMs-from-scratch.

The official reading list for the module is on Canvas.

Code

More demos

Not a language model: it classifies 2D points. Useful for seeing how an MLP moves information, with edge thickness showing weight size.
Animated state machines and an editable transition matrix. The n-gram model is a Markov chain over tokens.
  • AttentionViz — query and key vectors of trained models across layers and heads. Heavy: takes up to a minute to load.

References

Every work cited on this page, in APA style with shortened author lists. Titles link to the paper.

    Next session: how the model is actually trained

    The objective at scale

    What it takes to minimise the next-token prediction loss over trillions of tokens.

    Pre-training data

    Where the tokens come from, and why most of them are thrown away.

    Scaling laws

    Predicting that loss from model size, data and compute, before spending the compute.

    Adapting without retraining

    LoRA and parameter-efficient fine-tuning.

    Go to the session 02 page →