• Decrease Text SizeIncrease Text Size

How does Gradient Checkpointing work?

Standard training stores every layer's activations to enable gradient computation through backpropagation, which can require hundreds of gigabytes for large LLM fine-tuning. Gradient checkpointing strategically discards intermediate activations and recomputes them when needed, typically reducing memory by 50%-80% at the cost of 20%-30% additional compute. The platform strengt