• Decrease Text SizeIncrease Text Size

How does Mixed Precision Training work?

NVIDIA's Tensor Cores (Volta, Turing, Ampere, Hopper) and AMD's Matrix Cores (CDNA) execute lower-precision matrix multiplications 2x to 16x faster than FP32 equivalents, and the reduced memory footprint enables larger batches or larger models. BF16 (bfloat16) has become the dominant choice for LLM training because its wider exponent range matches FP32, eliminating the loss-scaling complexity required for FP16 stability. The platform stren