BF16 vs FP16 vs FP8: Which Format Should I Train In?
Use BF16 as a starting point when the GPU and training stack support it. Use FP16 with loss scaling when BF16 is unavailable or the measured workload favors FP16. Test FP8 through a supported mixed-precision implementation when training throughput justifies the additional numerical checks.
Mixed precision uses different formats for different operations and states. Selecting BF16 or FP8 does not mean every tensor uses that format.
Range and precision solve different problems
The exponent controls the range of values. The fraction bits control how closely neighboring values can be represented. BF16 keeps the same exponent width as FP32 but has fewer fraction bits. FP16 has finer precision than BF16 within its narrower range.
| Format | Exponent / fraction bits | Practical consideration |
|---|---|---|
| BF16 | 8 / 7 | Wide range; usually avoids loss scaling |
| FP16 | 5 / 10 | Small gradients can underflow; use loss scaling |
| FP8 E4M3 | 4 / 3 | More precision than E5M2, less range |
| FP8 E5M2 | 5 / 2 | Wider range for tensors such as gradients |
With FP16, loss scaling multiplies the loss before backward computation and unscales gradients before the optimizer update. This helps small gradients remain representable. Overflow still needs handling.
NVIDIA’s Transformer Engine documentation describes a conventional FP8 hybrid recipe: E4M3 for forward tensors and E5M2 for backward tensors, with scaling chosen for the tensors. Blackwell’s block-scaled formats add other recipes; hardware support alone does not choose one for your model.
Establish a numerical baseline
Run a representative training segment with the same data order and optimizer settings. Compare peak memory, time per step, loss behavior, non-finite values, and held-out quality. Keep numerically sensitive operations and optimizer states at the precision the implementation requires.
A faster run that diverges or misses the quality requirement does not save training cost. Conversely, a stable BF16 baseline may be sufficient without adding FP8 work.
The mixed-precision section of the LLM Engineering Guide includes the format ranges and published training results.