How does Gradient Accumulation work?
With gradient accumulation steps of 8, a per-device batch size of 4 yields an effective batch of 32 — equivalent in optimization dynamics to processing all 32 examples at once. The technique trades wall-clock training time for memory: each accumulation step requires its own forward and backward pass, so total training time is approximately proportional to total examples processed. The platform strength