OpenAI's GPT-5.6 Sol Scores 88.8% on WeirdML, Tops CritPt and DeepSWE
OpenAI's GPT-5.6 Sol scored 88.8% on the WeirdML benchmark, narrowly ahead of Claude Fable 5 at less than half the price. Separate benchmark updates on Thursday also showed Sol leading CritPt, a test of unpublished research-level physics problems, and topping the DeepSWE coding-agent leaderboard at 73%.
Artificial Analysis said Sol gained about 5 points over GPT-5.5 on CritPt and beat Claude Fable 5 by about 4 points. CritPt, developed by Argonne and the University of Illinois Urbana-Champaign, uses graduate-level physics problems contributed by more than 60 researchers from more than 30 institutions. DeepSWE measures performance on 113 software-engineering tasks across 91 repositories and five languages; one comparison put Sol at about 72% to 73% and $8.4 per task, versus around 70% and $13 to $22 for Fable 5. A separate comment said CritPt still appears close to a 30% ceiling, while the WeirdML result showed no run below 60% of the state-of-the-art score across 85 tests.
From the sources (25 posts)
@btibor91OpenAI announced a new livestream called "ChatGPT Work - Our biggest update for work in ChatGPT" for tomorrow
@testingcatalogOPENAI 🔥: A new ChatGPT Work livestream has been announced for tomorrow. > Tomorrow, we're introducing a new partner for your most ambitious work. That’s something 👀
@wesrothOpenAI has announced a new ChatGPT Work livestream with the teaser: “We’re introducing a new partner for your most ambitious work.” 👀
@testingcatalogOPENAI 🔥: ChatGPT Work will likely arrive in the form of an upgraded dedicated workspace, individual for every user. What we know so far 👀 > "ChatGPT Work can help you build websites, prototype new ideas, create presentations and documen
@financialjuiceOpenAI's Altman expects a smoother process working with the government on AI
@firstsquawkOPENAI MADE 'MANY CHANGES' AFTER TALKING TO GOVT, ALTMAN SAYS
@financialjuiceOpenAI's Altman: OpenAI made many changes after talking to the government.
@financialjuiceOpenAI's Altman: OpenAI’s newest AI model is 54% more token-efficient on agentic coding.
@stockmktnewzOpenAI CEO Sam Altman said that its latest AI model is 54% more token efficent on agent coding tasks - CNBC
@cnbcOpenAI's newest AI model is 54% more token efficient on agentic coding, Altman tells CNBC
@businessSam Altman said OpenAI made “many changes” during its discussions with the Trump administration before moving ahead with releasing its newest AI models to the general public
@polymarketNEW: Sam Altman reveals OpenAI made “many changes” during talks with the Trump administration before GPT-5.6’s release.
@exec_sumBREAKING: OpenAI is set to launch GPT-5.6, its most capable model yet after a delayed rollout
@daveaitelRT @OpenAI: Today. 10am PT.
@firstsquawkOPENAI UNVEILS THE GPT-5.6 MODEL FAMILY—SOL, TERRA, AND LUNA—WITH ROLLOUT ACROSS CHATGPT, CODEX, AND THE API.
@firstsquawkOPENAI LAUNCHES CHATGPT WORK, POWERED BY GPT-5.6, WITH ROLLOUT BEGINNING FOR PRO, ENTERPRISE, AND EDU USERS, WHILE PLUS AND BUSINESS WILL FOLLOW.
@firstsquawkOPENAI'S NEW CHATGPT DESKTOP APP, COMBINING CHAT, WORK, AND CODEX, IS ROLLING OUT GLOBALLY ON MAC AND WINDOWS FOR ALL USERS.
@sama5.6 livestream going now. in addition to the model, 3 major product things. 1. ChatGPT Work--really big deal! 2. new ChatGPT desktop app 3. hosted sites
@diggOpenAI said on its live stream that GPT-5.6 Sol, Terra and Luna are rolling out today. Sol is coming to paid plans over the next 24 hours. Terra and Luna are coming to free users.
@scaling01GPT-5.6 Benchmarks
@wallstengineOpenAI introduces ChatGPT Work, powered by GPT-5.6, and says Sol, Terra and Luna are rolling out across ChatGPT, Codex and the API.
@firstsquawkOPENAI LAUNCHES GPT-5.6 FAMILY OF MODELS FOR GENERAL AVAILABILITY FOLLOWING LIMITED PREVIEW
@scaling01GPT-5.6-Sol scoring a massive 7.78% on ARC-AGI-3 this is a massive jump over Opus 4.8's 1.5%
@cursor_aiGPT-5.6 Sol, Terra, and Luna are now available in Cursor. On CursorBench, Sol scores 67.2%.
@samaobviously the best model we have ever produced, but also one of the best blog posts we have ever produced: