Meta Muse Spark 1.1 Beats GPT-5.6 and Gemini 3.1 on Radiology Benchmark, Nears Human Performance
Meta Muse Spark 1.1 posted state-of-the-art results on the Radiologists Last Exam Handover Readiness Index, outperforming OpenAI’s GPT-5.6 and Google’s Gemini 3.1 on a clinical radiology benchmark and reaching performance levels close to human experts.
The model does not yet surpass the Fable benchmark on this specific medical task, though developers note human experts still hold a significant lead. The latest benchmark adds a health-focused evaluation to a model that recently demonstrated strength in coding, theoretical computer science, and long-context reasoning across prior public tests.
From the sources (16 posts)
@rohanpaul_aiRT @rohanpaul_ai: Meta is so back in the AI coding race with Muse Spark 1.1, using cut-rate pricing to pressure OpenAI and Anthropic in age…
@alexandr_wangmuse spark is able to do end-to-end tasks based on short video instructions
@alexandr_wangRT @aicodeking: Meta Muse Spark 1.1 is a great model. Awesome at frontend and everything. It still lacks at some places like overwriting fi…
@alexandr_wangRT @CuriousRefuge: So @Meta has officially entered the AI image and video space, and we're excited to share our early test results. Muse Vi…
@teortaxestexMuse Spark 1.1 is surprisingly close to Grok 4.5 on many high-signal evals This is the current top on CritPT
@yunta_tsaiIt is interesting that Muse Spark 1.1 has 1M context length but did not perform as well as Grok 4.5 at 500K context length. Sometimes I feel people chasing context length without scrutinizing every token inside the context are wasting memo
@alexandr_wangmuse spark 1.1 can be even better than opus 4.8 at 20% of the cost
@denny_zhoumuse spark 1.1 passed elon bench
@ofirpressRT @alexandr_wang: muse spark 1.1 is ahead of gpt-5.6 on SciCode
@alexandr_wang@teortaxesTex this isn’t true we’re going to increase access to muse spark incl openrouter
@alexandr_wangmuse spark 1.1 outperforms opus, grok 4.5, and gemini on a new challenging finite model theory / theoretical cs eval
@jack_w_raeRT @alexandr_wang: muse spark 1.1 outperforms opus, grok 4.5, and gemini on a new challenging finite model theory / theoretical cs eval
@s_batzoglouI benchmarked the new models (Sol, Terra, Luna, Fable 5, Meta Muse Spark 1.1, Grok 4.5) on an induction reasoning task. This is an updated table from yesterday, and my benchmark is described as spotlight in ICML '26. How to read the tabl
@alexandr_wangMuse Spark 1.1 is SOTA on the Radiologists Last Exam Handover Readiness Index (RadLE-H), nearing human expert performance.
@_jasonweiMuse Spark 1.1 outperforms GPT-5.6 Sol and Gemini 3.1 on Radiology's Last Exam. We don't beat Fable (yet). And Humans are still a lot better, but we are working on closing the gap!
@jack_w_raeMuse Spark 1.1 is really strong on health topics! I have found this to be the case from personal usage but it’s cool to see it show up on benchmarks also.