What operations run inside a transformer decoder layer?
A transformer decoder layer updates token representations through attention and a feed-forward network. Attention combines information from visible tokens. The feed-forward network transforms each token’s representation separately. Normalization and residual additions support the repeated computation through many layers.
Follow one Llama-style layer
A common pre-norm decoder runs this sequence:
- Normalize the input representation.
- Project queries, keys, and values and apply causal attention.
- Project the attention result and add it to the original input.
- Normalize that updated representation.
- Run the feed-forward network and add its output back.
The residual additions preserve an input path through each sub-block. RMSNorm rescales representations without LayerNorm’s mean subtraction. Models differ in normalization placement, attention type, activation, and bias terms; this sequence describes a Llama-style block, rather than every transformer.
The Transformer paper defines attention as a softmax-weighted combination of values. Causal masking lets a position attend to itself and earlier positions. Multi-head projections let different learned subspaces contribute to the result, without assigning a fixed linguistic role to each head.
Count the actual projections
Meta’s Llama 3 8B configuration has hidden dimension 4,096, feed-forward dimension 14,336, 32 query heads, eight KV heads, and head dimension 128.
| Projection group | Parameter calculation | Per-layer count |
|---|---|---|
| Queries and attention output | 2 × 4,096² | 33,554,432 |
| Keys and values | 2 × 4,096 × (8 × 128) | 8,388,608 |
| SwiGLU feed-forward matrices | 3 × 4,096 × 14,336 | 176,160,768 |
SwiGLU uses two input projections, combines one activated result element by element with the other, and applies an output projection. GLU Variants Improve Transformer defines that formulation.
The feed-forward matrices account for approximately 81% of these projection weights. That is not 81% of total execution time or the whole model: embeddings, the vocabulary output head, normalization, attention computation, and memory traffic also matter.
When estimating parameters, use the exact FFN width and KV-head count. A generic 12 × layers × hidden_dimension² approximation assumes a different conventional block and misses important GQA/SwiGLU details.
Engineering Guide: transformer architecture includes the layer diagram and parameter estimate.