FP8 Quantization: Prefill Latency vs. Decoding Speed Trade-offs
This benchmark analyzes the performance impact of FP8 quantization on Gemma 2 9B models running on NVIDIA L4 GPUs. It reveals that while FP8 improves decoding speed, it introduces a significant 'prefill tax' that increases time-to-first-token (TTFT).





