MoE Model Training Speed Surges 41% via Cursor's MoK Pipeline Optimization
Cursor open-sources MoK to merge GPU data transfer and computation, boosting throughput by 41% on GB300 clusters. The tool targets large-scale MoE training efficiency.
Woofun AI reports that Cursor has open-sourced the Mixture-of-Kittens (MoK) framework to accelerate Mixture-of-Experts (MoE) large-scale model training. The system integrates previously distinct GPU data transfer and computation processes into a single kernel, allowing simultaneous execution to minimize idle time. In tests utilizing 512 GB300 GPUs, overall training throughput increased by 41%, while MoE layer performance reached 2.37 times the speed of public baseline solutions. Currently deployed across tens of thousands of GPUs within Composer training environments, MoK is released under the Apache 2.0 license. Support is limited to NVIDIA Blackwell GPUs, specifically targeting institutions operating GB200 and GB300 NVL72 clusters.
Comments
No comments yet.