What operations run inside a transformer decoder layer?

A transformer decoder layer updates token representations through attention and a feed-forward network. Attention combines information from visible tokens. The feed-forward network transforms each token’s representation separately. Normalization and residual additions support the repeated computation through many layers.

Follow one Llama-style layer

A common pre-norm decoder runs this sequence:

  1. Normalize the input representation.
  2. Project queries, keys, and values and apply causal attention.
  3. Project the attention result and add it to the original input.
  4. Normalize that updated representation.
  5. Run the feed-forward network and add its output back.

The residual additions preserve an input path through each sub-block. RMSNorm rescales representations without LayerNorm’s mean subtraction. Models differ in normalization placement, attention type, activation, and bias terms; this sequence describes a Llama-style block, rather than every transformer.

The Transformer paper defines attention as a softmax-weighted combination of values. Causal masking lets a position attend to itself and earlier positions. Multi-head projections let different learned subspaces contribute to the result, without assigning a fixed linguistic role to each head.

Count the actual projections

Meta’s Llama 3 8B configuration has hidden dimension 4,096, feed-forward dimension 14,336, 32 query heads, eight KV heads, and head dimension 128.

Projection groupParameter calculationPer-layer count
Queries and attention output2 × 4,096²33,554,432
Keys and values2 × 4,096 × (8 × 128)8,388,608
SwiGLU feed-forward matrices3 × 4,096 × 14,336176,160,768

SwiGLU uses two input projections, combines one activated result element by element with the other, and applies an output projection. GLU Variants Improve Transformer defines that formulation.

The feed-forward matrices account for approximately 81% of these projection weights. That is not 81% of total execution time or the whole model: embeddings, the vocabulary output head, normalization, attention computation, and memory traffic also matter.

When estimating parameters, use the exact FFN width and KV-head count. A generic 12 × layers × hidden_dimension² approximation assumes a different conventional block and misses important GQA/SwiGLU details.

Engineering Guide: transformer architecture includes the layer diagram and parameter estimate.