Gradient Accumulation Can Slow Training
A TraceML benchmark using Qwen3-1.7B with LoRA found that equivalent effective batch sizes can have substantially different training times depending on physical batch size and accumulation steps. Across 100 optimizer updates, 4×1 was 17% faster than 1×4 on a T4 and 41% faster on an L4.



