Encoder vs decoder models: which architecture fits?
Use a task-trained encoder for representations or predictions over an input. Use an encoder-decoder model when a trained system maps one sequence to another. Use a decoder-only language model for autoregressive generation and prompt-based tasks. The checkpoint’s training and output interface matter as much as the architecture name.
These are useful starting choices for engineers selecting models, rather than limits on what an adapted architecture can ever do.
Compare what each model returns
| Family | Attention pattern | Typical output and example |
|---|---|---|
| Encoder-only | Tokens see the full input | Token representations, labels, or spans; BERT |
| Encoder-decoder | Full-input encoder; causal output decoder with cross-attention | Generated output sequence; T5 |
| Decoder-only | Causal attention over the token sequence | Next-token distributions and generated continuations; Llama |
BERT learns bidirectional representations using masked-language pretraining. A suitable output head and task training turn those representations into classifications, spans, or other predictions. Standard BERT is not configured as an ordinary left-to-right text generator. Also, an arbitrary encoder’s pooled output is not automatically a good retrieval embedding: that use needs suitable training and evaluation.
T5 uses an encoder to represent an input and a decoder to generate an output while attending to those representations. It can express classification tasks as text generation as well as sequence transformation. Encoder-decoder does not mean translation only.
A decoder-only model uses the same causal sequence for instructions, examples, and the generated continuation. Each position can attend to itself and earlier positions. Its earlier KV states remain valid as generation appends tokens, enabling incremental decoding. The Transformer paper explains the masking and cross-attention distinction.
Choose the checkpoint for the task
For fixed-label classification or extraction, compare a task-trained smaller model with a prompted generator. For retrieval, compare trained embedding or reranking checkpoints, rather than raw architecture labels. For flexible answers, code, or tool arguments, evaluate a generation model on the output format and errors your application can tolerate.
Check input limits, supported languages, required output head, quality, latency, and deployment runtime. A family label predicts the computation pattern; it does not guarantee task quality or make a larger model the better choice.
Engineering Guide: model families explains the generation and caching properties together.