LLM Distillation: How Do Small Models Learn From Teachers?

LLM distillation trains a student model using signals produced by a teacher model. Logit-based methods match output distributions. Data-based methods train on teacher-generated examples. Consider distillation when a smaller model could meet a defined task threshold at lower serving cost.

Distillation does not guarantee that the student retains every teacher capability. Data coverage and the evaluation task determine what transfers.

Two ways to transfer behavior

MethodTeacher informationMain requirement
Logit-based distillationToken probabilities or logitsAccess to distributions and a compatible training objective
Data-based distillationGenerated input/output examplesReviewed or filtered examples and student fine-tuning

An API-only teacher can support data-based distillation even when it does not expose the distributions needed by your logit-training method. Different tokenizers or architectures can make direct distribution matching more involved.

The DeepSeek-R1 report describes fine-tuning smaller Qwen and Llama models on a curated mixture of roughly 800,000 examples. Its distilled models performed well on the paper’s reasoning benchmarks. Those results establish that transfer worked in that setup; they do not show that a small student can replace any teacher on any workload.

Build examples around a measurable task

For a support classifier, provide real question types and have the teacher propose labels and explanations. Check labels against a reviewed taxonomy; reject unsupported classifications. Include rare classes and ambiguous cases rather than producing many easy paraphrases of common requests.

Separate teacher-generated training examples from the held-out evaluation set. Otherwise the student can appear strong because the test reuses examples from training rather than measuring unseen cases. Verify that the teacher’s terms permit the intended training use before generating the dataset.

Compare the student with both the teacher and its own untuned base model. Measure task quality, error types, output length, latency, and cost at the intended concurrency. Include fallback calls to the teacher when estimating savings.

Use teacher explanations only when they improve the student’s measured behavior; longer training targets also cost more to generate and learn.

The distillation section of the LLM Engineering Guide describes the published R1 comparisons and their limits.