LMSYS Says Nvidia GB300 NVL72 Beats B200 by Up to 6.5x
LMSYS and SemiAnalysis said Nvidia's Dynamo with SGLang disaggregation on GB300 NVL72 systems supplied by CoreWeave achieved up to 6.5 times the performance of the B200 on DeepSeek-v4 Pro 1.6T.
They said the high-throughput configuration used DeepSeek's MegaMoe kernels, which fuse and overlap EP dispatch, EP combine and GEMM operations into a single kernel. The performance work was credited to engineers at Radix, LMSYS and Nvidia, while CoreWeave supplied temporary GB300 NVL72 racks for the open-source optimization effort.
From the sources (2 posts)
@semianalysis_GB300 NVL72 Rack Scale Dynamo SGLang disaggregation has up to 6.5x better performance than B200 on DeepSeekv4 Pro 1.6T 🚀 The high throughput configuration uses @deepseek_ai 's MegaMoe kernels which fully fuses & overlaps EP dispatch & EP
@lmsysorgBy leveraging @NVIDIAAI Dynamo + SGLang disaggregation on the GB300 NVL72 (powered by @CoreWeave), we're seeing a massive 6.5x performance leap over B200 for DeepSeek-v4 Pro 1.6T 📈