Command Palette
Search for a command to run...

SpaceXAI Launches Grok 4.5 Coding Model, Tops AutomationBench At 51 Percent

aiai-modelingai-model-releasesai-research-evalsai-productsai-agents-coding 124 posts · 62 accounts

SpaceXAI launched Grok 4.5, a model trained specifically for coding and autonomous agents in collaboration with Cursor. The model scores 51% on the AutomationBench-AA leaderboard, beating Anthropic's Claude Fable 5 at 49% and Claude Opus 4.8 at 48%, while completing 79.9% of task objectives at a cost of $0.34 per task.

The model ranked fourth on the Intelligence Index and third in the Coding Agent Index, posting an 83.3% score on Terminal-Bench 2.1 and 64.7% on SWE-Bench Pro. While the system matches frontier competitors on automated SaaS workflows and parallel tool use, it registers more rule violations per task and lower raw benchmark totals on pure coding evaluations than leading models. The model is available now via the SpaceXAI API and Cursor, with European rollout expected by mid-July.

From the sources (25 posts)

@mark_k

HUGE: Grok 4.5 by @xai probably coming this week! 🗣️

@testingcatalog

BREAKING 🔥: TRACES OF GROK 4.5 HAVE BEEN SPOTTED ON THE GROK WEB! > Unlock the full power of Chat with Grok 4.5 SOON 👀

@rohanpaul_ai

Grok 4.5 almost ready to drop? Some details about Grok 4.5 that have now been confirmed by various reports. - Grok 4.5 is built on xAI’s V9 foundation model with 1.5T parameters. - That makes it about 3x larger than v8-small, which curre

@elonmusk

Based on strong positive feedback from customers in our beta test program, @SpaceXAI will make Grok 4.5 available to the public tomorrow. It is an Opus-class model, but faster, more token-efficient and lower cost.

@elliotarledge

RT @elonmusk: Based on strong positive feedback from customers in our beta test program, @SpaceXAI will make Grok 4.5 available to the publ…

@jukan05

RT @elonmusk: Based on strong positive feedback from customers in our beta test program, @SpaceXAI will make Grok 4.5 available to the publ…

@scaling01

RT @elonmusk: Based on strong positive feedback from customers in our beta test program, @SpaceXAI will make Grok 4.5 available to the publ…

@cointelegraph

🔥 NOW: Elon Musk says SpaceXAI will roll out Grok 4.5 to the public tomorrow, calling it an Opus-class model that's faster, more token-efficient, and cheaper.

@testingcatalog

SPACEXAI 🔥: Grok 4.5 is officially set to launch on Wednesday. > It is an Opus-class model, but faster, more token-efficient and lower cost. Soon 👀

@marionawfal

Elon: Grok 4.5 is coming tomorrow!! "It is an Opus-class model, but faster, more token-efficient and lower cost." Opus class means it's the flagship, highest-performance, most powerful model Writer: Ian

@stevibe

RT @elonmusk: Based on strong positive feedback from customers in our beta test program, @SpaceXAI will make Grok 4.5 available to the publ…

@ns123abc

BRO LITERALLY PREDICTED THIS

@teknium

Day 0 support coming for Grok 4.5 as well, of course!

@wesroth

RT @WesRoth: New traces of Grok 4.5 have reportedly appeared on Grok web, including subscription copy saying: “Unlock the full power of Cha…

@hesamation

Grok 4.5 will be Opus level? AND faster? AND cheaper? AND more token-efficient?

@mtslive

SITUATION DETECTED: SpaceXAI will make Grok 4.5 available to the public today.

@sawyermerritt

RT @elonmusk: Based on strong positive feedback from customers in our beta test program, @SpaceXAI will make Grok 4.5 available to the publ…

@mark_k

Grok 4.5 is the new Cursor Composer. If we're lucky, we'll get to use it in @cursor_ai on day 1 (today).

@elonmusk

We will continue to make refinements to the Grok Build harness and the 1.5T foundation model almost every day in response to user requests. The 2T model will finish training this month and be available to customers next month.

@financialjuice

Musk on Grok: We will continue to make refinements to the Grok Build harness and the 1.5T foundation model almost every day in response to user requests - Post on X.

@mtslive

SITUATION UPDATE: SpaceXAI’s 2-trillion-parameter model will finish training this month and be available to customers next month.

@imjaredz

RT @ScottWu46: Benchmark scores are exciting but more importantly we are seeing incredible results so far using this model in Devin! Try it…

@teortaxestex

If this turns out wrong I'll crash out

@elonmusk

Our internal assessment is that Grok 4.5 is roughly comparable to Opus 4.7, but much faster. The combination of capability, faster speed and lower cost is what makes it competitive. We are closing the loop on real-world usefulness, not be

@business

SpaceXAI has unveiled a new AI model built in partnership with Cursor that’s meant to be more adept at finance, legal and coding tasks, in a bid by Elon Musk’s firm to gain ground on rivals Anthropic and OpenAI

Preview built on a synthetic news corpus (16 weeks, Apr–Jul 2026). Impact calls are model reads, not price data.

About Archive