Command Palette
Search for a command to run...

OpenAI Launches GPT-5.6 Luna and Terra Model Variants to Agent Arena With Benchmark Gains Up to 10 Percent

aiai-modelingai-model-releasesai-research-evals 2 posts · 1 accounts

OpenAI introduced the GPT-5.6 Luna and Terra model variants to the Agent Arena leaderboard, posting positive performance gains across all new releases. Luna, operating at a higher reasoning setting, edged out cheaper Terra configurations by delivering a 3.3% score increase, while the Sol variant ranked second on the leaderboard with a 10.1% gain.

The benchmark tracks millions of real-world agentic tasks from a global user community, measuring how models use web search and terminal tools to complete complex workflows. Test-time scaling proved highly effective on the smaller Luna model, which saw substantial gains moving its reasoning effort from low to medium settings.

From the sources (2 posts)

@arena

GPT-5.6 Sol vs. Terra vs. Luna - all at xHigh, in Agent Arena. As previously reported, Sol is top 2 on the leaderboard (+10.1%), but Terra (+4.0%) and Luna (+3.3%) both post positive net improvement too, landing right next to Claude Opus 4

@arena

GPT-5.6 Terra and Luna (xHigh variants) have joined Sol in Agent Arena! Built by @OpenAI, they rank #15 and #17 on the leaderboard, based on real-world agentic sessions from our global community. In comparing all variants in Agent Arena, a

Preview built on a synthetic news corpus (16 weeks, Apr–Jul 2026). Impact calls are model reads, not price data.

About Archive