
1.3x MXFP8 MoE Training Speedup vs BF16
PyTorch achieves 30.2% training speedup for Llama4 Scout MoE model using MXFP8 primitives in TorchAO, matching BF16 convergence. Demonstrated on GB200 cluster with TorchTitan framework. Reaches ~81% of theoretical peak performance.





