Search

Few direct matches — filled in with the latest updates.

Tag: #fp83 results

⚙️

MXFP8 GEMM Matches 99% cuBLAS Speed

Meta/PyTorch engineer details MXFP8 GEMM kernel design in CUDA + PTX, achieving up to 99% of cuBLAS performance despite FP8 constraints. Includes deep dives into challenges and links to PyTorch integrations for DeepSeek-V3 training on B200 GPUs.

Reddit r/MachineLearningCommunityMar 30#fp8#gemm#nvidia-b200
⚙️

Moore Threads Hits Nvidia FP8 Parity

Moore Threads achieves systemic FP8 breakthrough, syncing with Nvidia globally as one of few domestic GPU makers. MTT S5000 delivers 1000 TFLOPS FP8 AI compute with 80GB VRAM and full precision support. MUSA platform enables compatibility with PyTorch, vLLM, and other AI frameworks.

IT之家MediaMar 13#fp8#gpu-acceleration#ai-hardware
⚙️

FP8 Inference on Older GPUs via Software

Feather emulates FP8 inference on Ampere GPUs like RTX 3050 using custom Triton kernels and bit-packing, achieving 1.5x speedup over FP32 for TinyLlama with minimal accuracy loss. It targets memory bandwidth optimization without native hardware support. Future plans include CUDA Graphs, block quantization, and Llama support; accepted at PyTorch Conference Europe 2026.

Reddit r/MachineLearningCommunityFeb 26#fp8#quantization
📰

When AI Diagnoses Clash With Doctors

Patients are increasingly bringing answers from Doubao and other AI assistants into hospitals, sometimes challenging doctors' diagnoses and requesting unnecessary tests. The article documents how AI hallucinations, incomplete symptom interpretation, and misplaced patient trust are increasing consultation time and worsening doctor-patient friction.

Kling AI Becomes Kuaishou’s Growth Engine

Kling AI Becomes Kuaishou’s Growth Engine

Kling AI generated more than RMB 850 million in Q2 revenue, up 240% year over year, making it the standout growth driver in Kuaishou’s otherwise slowing business. Kuaishou is prioritizing AI investment, spinning Kling AI out for independent financing at an implied valuation of US$18 billion despite near-term profit pressure.

Rapidus Bets on 2nm Without Fighting TSMC

Rapidus Bets on 2nm Without Fighting TSMC

Rapidus is pursuing mass production of advanced 2nm semiconductors through a large-scale Japanese national project. Instead of matching TSMC’s scale, the company plans to compete with RUMS, an integrated one-building production model designed for diverse, lower-volume manufacturing.

ITmedia AI+ (日本)Media2h ago#2nm#semiconductors#chip-manufacturing