Command Palette
Search for a command to run...

AMD $1.1 Million Hackathon Yields MI355X Optimizations That Double Inference Speed and Beat Nvidia B200

aiai-infrastructureai-compute-chipsai-inference-platforms 3 posts · 2 accounts

A $1.1 million GPU_MODE hackathon funded by AMD yielded kernel optimizations that more than doubled the MI355X AI processor’s end-to-end inference performance. The winning development team reworked the processor software to run Kimi artificial-intelligence models across every tested setting faster than Nvidia’s B200 chip.

The benchmarking code has been merged into the main branch of AMD’s AITER kernel library and the ATOM inference engine. Developers aim to integrate the updates into the open-source vLLM framework next, a move that would allow the open-source vLLM framework to run on AMD hardware at speeds matching the industry standard built for Nvidia chips.

From the sources (3 posts)

@semianalysis_

@GPU_MODE This work has been upstreamed to the main branch of @AIatAMD’s AITER kernel library and to the ATOM inference engine. We hope this work will also be upstreamed to vLLM so that AMD vLLM can reach performance parity with CUDA vLLM.

@semianalysis_

GREAT WORK BY @GPU_MODE 🚨 FOR LAUNCHING THE $1.1mil AMD KERNEL HACKATHON. The GPUMODE Readonflow Team’s kernels improved end-to-end MI355X performance by over 2x. We explain the optimizations below. 1/4🧵

@gpu_mode

Kinda wild that our community got together and collectively made Kimi inference on MI355X faster than B200 across all settings. Congrats to the winners team RadeonFlow and thanks again to AMD for working so closely with our community https:

Preview built on a synthetic news corpus (16 weeks, Apr–Jul 2026). Impact calls are model reads, not price data.

About Archive