Embedding vs Generative Models: What Does Each Output?
An embedding model represents an input as a vector for similarity search or other downstream tasks. A generative model produces a sequence, such as an answer or structured output. In RAG, an embedding model helps retrieve evidence; a generator uses that evidence to answer.
The roles can use related architectures, but their training objectives and outputs differ.
A vector is a representation, not an answer
| Property | Embedding model | Generative model |
|---|---|---|
| Typical output | Fixed-dimensional vector | Variable-length token sequence |
| Example job | Find passages similar to a query | Explain or extract from passages |
| Evaluation | Retrieval relevance and ranking | Answer correctness and evidence use |
Text embedding models may pool token representations from an encoder or a decoder-derived architecture. The pooling and representation training make the vector useful; averaging arbitrary hidden states does not establish retrieval quality.
For example, Qwen3-Embedding-8B is a text-embedding model with configurable output dimensions and task instructions. Follow the model’s query and document conventions. Compatible representations need not use identical input prefixes.
Keep the index compatible
Store the encoder identity, dimension, and preprocessing used to index documents. Queries must use the corresponding query encoder and instructions. Vectors from unrelated model versions generally do not belong in one similarity comparison, even if their dimensions match.
Changing the generator can leave the vector index usable. Changing the embedding model usually requires re-embedding documents or another explicitly tested migration.
Some models support shortened vectors through Matryoshka training. OpenAI’s embedding announcement reports that a 256-dimensional text-embedding-3-large representation beat its 1,536-dimensional ada-002 comparison on MTEB. This is a model-specific benchmark, not proof that fewer dimensions always preserve quality.
For your corpus, test retrieval with labeled queries at supported dimensions. Record relevant-document recall, ranking, storage, and search latency. A stronger answer generator cannot reliably recover a source passage the retrieval stage never supplied.
The embedding-model section of the LLM Engineering Guide compares example models and explains dimension truncation.