Synthetic Training Data: How Do I Generate Useful Examples?
Generate synthetic examples to address specific gaps in real training data. Start from reviewed cases, vary the inputs deliberately, verify the outputs, and test the resulting model on real held-out examples. A larger generated dataset is useful only when it improves the task you measure.
Use synthetic data for coverage you can check, such as missing labels, rare formats, or examples with known answers.
Define what each generated example adds
For an extraction task, first name the field and failure you need to cover:
| Real failure | Proposed generated variation | Output check |
|---|---|---|
| Missing field causes an invented value | Remove that field from otherwise valid documents | Expected field is empty |
| Reordered fields break extraction | Change layout while retaining the values | Values still match the source |
| Rare label is misclassified | Generate diverse cases for that reviewed label | Label passes review |
Retain the source case, generation prompt, teacher identity, and validation result. This lets you remove or regenerate a subset when a labeling rule changes.
Self-Instruct generates and filters instruction examples from a seed set. Phi-4’s technical report describes more involved generation, critique, revision, and instruction-reversal methods. These are options for producing data; their published results do not certify your examples.
Filter before increasing volume
Check duplicates, unsupported answers, label balance, language, and task difficulty. Use deterministic verifiers where the correct result is known, and human review where it is not. Keep evaluation cases out of the generation prompts and training data.
Recursive training on generated data can lose rare patterns. The model-collapse study examines this risk. A separate study of accumulating real and synthetic data shows that the data-retention strategy changes the outcome. Synthetic examples do not inevitably cause collapse; indiscriminate replacement and repeated generation require particular scrutiny.
Train with and without the synthetic subset under the same budget. Keep it when real held-out quality improves without unacceptable regressions, especially on rare cases.
The synthetic-data section of the LLM Engineering Guide explains the main generation approaches.