Artificial Analysis Launches AgentPerf, Nvidia GB300 Leads Initial Test at 61,354 Agents/MW
Artificial Analysis launched AgentPerf, which it described as the first benchmark built for agentic inference, and published initial DeepSeek V4 Pro results across Nvidia Blackwell, Hopper and AMD hardware. At a service level of 20 tokens per second and 10 seconds time-to-first-token, the benchmark’s Agents per Megawatt metric put Nvidia’s rack-scale GB300 at 61,354, versus 21,053 for Nvidia’s B300, 3,551 for AMD’s MI355X and 2,594 for Nvidia’s H200.
AgentPerf is designed to measure long-running agent workloads rather than a single chat completion, replaying coding-agent trajectories that can run up to 200 turns and exceed 100,000 tokens while allowing production techniques such as KV cache reuse and speculative decoding. Artificial Analysis said the benchmark will be updated on a rolling basis, and cautioned that its MI355X configuration was older than its Blackwell setup and couldn’t stably use speculative decoding.
From the sources (5 posts)
@tftc21The first benchmark built specifically for agentic AI is here. Artificial Analysis just launched AgentPerf, a new benchmark designed to measure how AI infrastructure handles agentic workloads, where a single task chains dozens to hundreds
@artificialanlysToday we're releasing the first results for AA-AgentPerf, our new agentic inference benchmark: initially covering DeepSeek V4 Pro across NVIDIA Blackwell, Hopper, and AMD. AA-AgentPerf is the first benchmark built for agentic inference. We
@artificialanlysAs more hardware & config combinations are tested, we will be able to display how optimal each combination is at each performance target (SLO). So far, we have tested with 20 tokens/s and 60 tokens/s per-user SLOs for DeepSeek V4 Pro, and w
@rohanpaul_aiNVIDIA just posted the first agentic AI benchmark results where GB300 NVL72 runs up to 20x more coding agents per megawatt than H200. Older inference benchmarks mostly ask how fast a system can produce tokens after one prompt. AgentPerf f
@rohanpaul_aiAgentPerf treats agentic AI as a long-running systems workload, where memory, networking, batching, MoE routing, and latency control all matter together. It includes long context windows, with requests ranging from 5K to 131K tokens and an