BF16 vs FP16 vs FP8: Which Format Should I Train In?

Use BF16 as a starting point when the GPU and training stack support it. Use FP16 with loss scaling when BF16 is unavailable or the measured workload favors FP16. Test FP8 through a supported mixed-precision implementation when training throughput justifies the additional numerical checks.

Mixed precision uses different formats for different operations and states. Selecting BF16 or FP8 does not mean every tensor uses that format.

Range and precision solve different problems

The exponent controls the range of values. The fraction bits control how closely neighboring values can be represented. BF16 keeps the same exponent width as FP32 but has fewer fraction bits. FP16 has finer precision than BF16 within its narrower range.

FormatExponent / fraction bitsPractical consideration
BF168 / 7Wide range; usually avoids loss scaling
FP165 / 10Small gradients can underflow; use loss scaling
FP8 E4M34 / 3More precision than E5M2, less range
FP8 E5M25 / 2Wider range for tensors such as gradients

With FP16, loss scaling multiplies the loss before backward computation and unscales gradients before the optimizer update. This helps small gradients remain representable. Overflow still needs handling.

NVIDIA’s Transformer Engine documentation describes a conventional FP8 hybrid recipe: E4M3 for forward tensors and E5M2 for backward tensors, with scaling chosen for the tensors. Blackwell’s block-scaled formats add other recipes; hardware support alone does not choose one for your model.

Establish a numerical baseline

Run a representative training segment with the same data order and optimizer settings. Compare peak memory, time per step, loss behavior, non-finite values, and held-out quality. Keep numerically sensitive operations and optimizer states at the precision the implementation requires.

A faster run that diverges or misses the quality requirement does not save training cost. Conversely, a stable BF16 baseline may be sufficient without adding FP8 work.

The mixed-precision section of the LLM Engineering Guide includes the format ranges and published training results.