MXFP8 GEMM Matches 99% cuBLAS Speed
Meta/PyTorch engineer details MXFP8 GEMM kernel design in CUDA + PTX, achieving up to 99% of cuBLAS performance despite FP8 constraints. Includes deep dives into challenges and links to PyTorch integrations for DeepSeek-V3 training on B200 GPUs.





