SpaceXAI Launches Grok 4.5 Coding Model, Tops AutomationBench At 51 Percent
SpaceXAI launched Grok 4.5, a model trained specifically for coding and autonomous agents in collaboration with Cursor. The model scores 51% on the AutomationBench-AA leaderboard, beating Anthropic's Claude Fable 5 at 49% and Claude Opus 4.8 at 48%, while completing 79.9% of task objectives at a cost of $0.34 per task.
The model ranked fourth on the Intelligence Index and third in the Coding Agent Index, posting an 83.3% score on Terminal-Bench 2.1 and 64.7% on SWE-Bench Pro. While the system matches frontier competitors on automated SaaS workflows and parallel tool use, it registers more rule violations per task and lower raw benchmark totals on pure coding evaluations than leading models. The model is available now via the SpaceXAI API and Cursor, with European rollout expected by mid-July.
From the sources (25 posts)
@mark_kHUGE: Grok 4.5 by @xai probably coming this week! 🗣️
@testingcatalogBREAKING 🔥: TRACES OF GROK 4.5 HAVE BEEN SPOTTED ON THE GROK WEB! > Unlock the full power of Chat with Grok 4.5 SOON 👀
@rohanpaul_aiGrok 4.5 almost ready to drop? Some details about Grok 4.5 that have now been confirmed by various reports. - Grok 4.5 is built on xAI’s V9 foundation model with 1.5T parameters. - That makes it about 3x larger than v8-small, which curre
@elonmuskBased on strong positive feedback from customers in our beta test program, @SpaceXAI will make Grok 4.5 available to the public tomorrow. It is an Opus-class model, but faster, more token-efficient and lower cost.
@elliotarledgeRT @elonmusk: Based on strong positive feedback from customers in our beta test program, @SpaceXAI will make Grok 4.5 available to the publ…
@jukan05RT @elonmusk: Based on strong positive feedback from customers in our beta test program, @SpaceXAI will make Grok 4.5 available to the publ…
@scaling01RT @elonmusk: Based on strong positive feedback from customers in our beta test program, @SpaceXAI will make Grok 4.5 available to the publ…
@cointelegraph🔥 NOW: Elon Musk says SpaceXAI will roll out Grok 4.5 to the public tomorrow, calling it an Opus-class model that's faster, more token-efficient, and cheaper.
@testingcatalogSPACEXAI 🔥: Grok 4.5 is officially set to launch on Wednesday. > It is an Opus-class model, but faster, more token-efficient and lower cost. Soon 👀
@marionawfalElon: Grok 4.5 is coming tomorrow!! "It is an Opus-class model, but faster, more token-efficient and lower cost." Opus class means it's the flagship, highest-performance, most powerful model Writer: Ian
@stevibeRT @elonmusk: Based on strong positive feedback from customers in our beta test program, @SpaceXAI will make Grok 4.5 available to the publ…
@ns123abcBRO LITERALLY PREDICTED THIS
@tekniumDay 0 support coming for Grok 4.5 as well, of course!
@wesrothRT @WesRoth: New traces of Grok 4.5 have reportedly appeared on Grok web, including subscription copy saying: “Unlock the full power of Cha…
@hesamationGrok 4.5 will be Opus level? AND faster? AND cheaper? AND more token-efficient?
@mtsliveSITUATION DETECTED: SpaceXAI will make Grok 4.5 available to the public today.
@sawyermerrittRT @elonmusk: Based on strong positive feedback from customers in our beta test program, @SpaceXAI will make Grok 4.5 available to the publ…
@mark_kGrok 4.5 is the new Cursor Composer. If we're lucky, we'll get to use it in @cursor_ai on day 1 (today).
@elonmuskWe will continue to make refinements to the Grok Build harness and the 1.5T foundation model almost every day in response to user requests. The 2T model will finish training this month and be available to customers next month.
@financialjuiceMusk on Grok: We will continue to make refinements to the Grok Build harness and the 1.5T foundation model almost every day in response to user requests - Post on X.
@mtsliveSITUATION UPDATE: SpaceXAI’s 2-trillion-parameter model will finish training this month and be available to customers next month.
@imjaredzRT @ScottWu46: Benchmark scores are exciting but more importantly we are seeing incredible results so far using this model in Devin! Try it…
@teortaxestexIf this turns out wrong I'll crash out
@elonmuskOur internal assessment is that Grok 4.5 is roughly comparable to Opus 4.7, but much faster. The combination of capability, faster speed and lower cost is what makes it competitive. We are closing the loop on real-world usefulness, not be
@businessSpaceXAI has unveiled a new AI model built in partnership with Cursor that’s meant to be more adept at finance, legal and coding tasks, in a bid by Elon Musk’s firm to gain ground on rivals Anthropic and OpenAI