
Cursor Rewrites MoE GPU Execution with MoK
Cursor open-sourced Mixture-of-Kittens (MoK), a GPU kernel design that combines token scheduling, cross-GPU communication, and expert computation for Mixture-of-Experts workloads. Its Pull-based dispatch reduced signaling latency from about 103μs to 18μs in a multi-node microbenchmark, while improving NVLink utilization by up to 29% in imbalanced workloads.
雷峰网 · 20d ago


















