How does Gradient Checkpointing work?
Standard training stores every layer's activations to enable gradient computation through backpropagation, which can require hundreds of gigabytes for large LLM fine-tuning. Gradient checkpointing strategically discards intermediate activations and recomputes them when needed, typically reducing memory by 50%-80% at the cost of 20%-30% additional compute. The platform strengt