Command Palette
Search for a command to run...

Meta Muse Spark 1.1 Beats Opus, Grok 4.5 and Gemini On New Theoretical Computer Science Benchmark

aiai-modelingai-research-evals 25 posts · 15 accounts

Meta Muse Spark 1.1 topped a new theoretical computer science benchmark, outperforming rival models including Anthropic Opus, X.AI Grok 4.5 and Google Gemini on a test of graph induction reasoning. Researcher Dimitris Batzoglou posted the results Friday, noting the evaluation comprises 64 problems where models must generate logical formulas to describe designated nodes across multiple graphs.

The performance extends Muse Spark 1.1’s recent rankings on scientific coding and agent evaluation suites. Batzoglou stated the benchmark, described as a spotlight paper at the ICML 2026 conference, is now publicly available. The results follow days of independent testing that placed the model near or ahead of earlier frontier releases.

From the sources (25 posts)

@charlesrollet1

RT @alexeheath: Meta AI chief @alexandr_wang told me the company sees its new API for selling AI models as a real business, not just a lear…

@garrytan

RT @alexeheath: Meta AI chief @alexandr_wang told me the company sees its new API for selling AI models as a real business, not just a lear…

@alexeheath

Meta AI chief @alexandr_wang told me the company sees its new API for selling AI models as a real business, not just a learning exercise. Tokens are “gigantic and growing very, very quickly,” he told me, and capturing even a slice of that

@kimmonismus

I do not think anyone saw this massive comeback from Meta coming. And now they are even receiving high praise from SemiAnalysis. Just to be clear: I have the utmost respect for SemiAnalysis and Dylan Patel. They do truly outstanding resear

@alexandr_wang

RT @robertcourson: I am very impressed with Muse Spark 1.1 for UI/UX. Especially for the price. This was the result of 1 prompt to make a…

@alexandr_wang

RT @mweinbach: Let's see how Muse Spark 1.1 is!

@alexandr_wang

RT @emdashsh: Muse Spark 1.1 is now available in Emdash with OpenCode.

@alexandr_wang

RT @eashish93: I was testing Muse Spark 1.1 and honestly, the result is pretty impressive. Built from planning to completion in a single p…

@mweinbach

Just because this got retweeted, Muse Spark 1.1 impressions are it's actually pretty good This fills the gap for me that Kimi K2.6 or GLM 5.2 would, good enough for most things, pretty quick, very cheap (more efficient than those models)

@alexandr_wang

RT @mweinbach: Just because this got retweeted, Muse Spark 1.1 impressions are it's actually pretty good This fills the gap for me that Ki…

@wesroth

Vals AI’s latest results place Meta Muse Spark 1.1 among the strongest new agent models. It is now ranked #1 on MedScribe and TaxEval, taking the top spot from Fable 5, while reportedly being 10x cheaper and twice as fast. Muse Spark 1.1

@therundownai

RT @rowancheung: Many people were doubting Meta's position in the AI race. Yesterday, they dropped Muse Spark 1.1, now one of the stronges…

@rohanpaul_ai

RT @rohanpaul_ai: Meta is so back in the AI coding race with Muse Spark 1.1, using cut-rate pricing to pressure OpenAI and Anthropic in age…

@alexandr_wang

muse spark is able to do end-to-end tasks based on short video instructions

@alexandr_wang

RT @aicodeking: Meta Muse Spark 1.1 is a great model. Awesome at frontend and everything. It still lacks at some places like overwriting fi…

@alexandr_wang

RT @CuriousRefuge: So @Meta has officially entered the AI image and video space, and we're excited to share our early test results. Muse Vi…

@teortaxestex

Muse Spark 1.1 is surprisingly close to Grok 4.5 on many high-signal evals This is the current top on CritPT

@yunta_tsai

It is interesting that Muse Spark 1.1 has 1M context length but did not perform as well as Grok 4.5 at 500K context length. Sometimes I feel people chasing context length without scrutinizing every token inside the context are wasting memo

@alexandr_wang

muse spark 1.1 can be even better than opus 4.8 at 20% of the cost

@denny_zhou

muse spark 1.1 passed elon bench

@ofirpress

RT @alexandr_wang: muse spark 1.1 is ahead of gpt-5.6 on SciCode

@alexandr_wang

@teortaxesTex this isn’t true we’re going to increase access to muse spark incl openrouter

@alexandr_wang

muse spark 1.1 outperforms opus, grok 4.5, and gemini on a new challenging finite model theory / theoretical cs eval

@jack_w_rae

RT @alexandr_wang: muse spark 1.1 outperforms opus, grok 4.5, and gemini on a new challenging finite model theory / theoretical cs eval

@s_batzoglou

I benchmarked the new models (Sol, Terra, Luna, Fable 5, Meta Muse Spark 1.1, Grok 4.5) on an induction reasoning task. This is an updated table from yesterday, and my benchmark is described as spotlight in ICML '26. How to read the tabl

Preview built on a synthetic news corpus (16 weeks, Apr–Jul 2026). Impact calls are model reads, not price data.

About Archive