Command Palette
Search for a command to run...

LMSYS Says Nvidia GB300 NVL72 Beats B200 by Up to 6.5x

aitech 2 posts · 2 accounts

LMSYS and SemiAnalysis said Nvidia's Dynamo with SGLang disaggregation on GB300 NVL72 systems supplied by CoreWeave achieved up to 6.5 times the performance of the B200 on DeepSeek-v4 Pro 1.6T.

They said the high-throughput configuration used DeepSeek's MegaMoe kernels, which fuse and overlap EP dispatch, EP combine and GEMM operations into a single kernel. The performance work was credited to engineers at Radix, LMSYS and Nvidia, while CoreWeave supplied temporary GB300 NVL72 racks for the open-source optimization effort.

From the sources (2 posts)

@semianalysis_

GB300 NVL72 Rack Scale Dynamo SGLang disaggregation has up to 6.5x better performance than B200 on DeepSeekv4 Pro 1.6T 🚀 The high throughput configuration uses @deepseek_ai 's MegaMoe kernels which fully fuses & overlaps EP dispatch & EP

@lmsysorg

By leveraging @NVIDIAAI Dynamo + SGLang disaggregation on the GB300 NVL72 (powered by @CoreWeave), we're seeing a massive 6.5x performance leap over B200 for DeepSeek-v4 Pro 1.6T 📈

Preview built on a synthetic news corpus (16 weeks, Apr–Jul 2026). Impact calls are model reads, not price data.

About Archive