Command Palette
Search for a command to run...

Alibaba Qwen Audio Leads Speech Benchmark at 84.1 Percent, Beating OpenAI Model

aiai-modelingai-research-evalsai-model-releasesai-productsai-generative-media 3 posts · 1 accounts

Alibaba's Qwen Audio 3.0 Realtime Plus variant became the leader on the Artificial Analysis Speech to Speech Index, scoring 84.1 percent. The mark beats OpenAI's GPT Realtime 2.1 High, which sits at 79.1 percent, and covers higher performance across the index's three sub benchmarks covering speech reasoning, conversational dynamics, and agentic performance.

The release includes a Flash variant that placed fourth on the index at 76.3 percent. Testing on Aliyun endpoints showed the Plus model requires 4.02 seconds to generate the first audio response, lagging behind the 1.10 second startup time of the best Open AI alternative. The Plus model costs $4.42 per hour of input audio, pricing it below the newer Open AI tier at $10.75 an hour while running slightly more expensive than the standard GPT Realtime 2 High.

From the sources (3 posts)

@artificialanlys

Alibaba has released Qwen Audio 3.0 Realtime, with the Plus variant debuting as the new #1 model on the Artificial Analysis Speech to Speech Index at 84.1%, ahead of GPT-Realtime-2.1 High at 79.1% Released earlier this month, Qwen Audio 3.

@artificialanlys

Qwen Audio 3.0 Realtime Plus leads all three components of our Speech to Speech Index. On Big Bench Audio it achieves 99.2%, the highest score we have measured, with the Flash variant achieving 96.1%. On our Full Duplex Bench subset the plu

@artificialanlys

Plus costs $4.42 per hour of input audio on our Big Bench Audio subset, more expensive than GPT-Realtime-2 High ($4.14) and ~2.4x cheaper than GPT-Realtime-2.1 High ($10.75). The Flash variant measures $4.77, appearing more costly than Plus

Preview built on a synthetic news corpus (16 weeks, Apr–Jul 2026). Impact calls are model reads, not price data.

About Archive