Command Palette
Search for a command to run...

Claude Fable 5 Max Sets WeirdML High at 91.9%

aiai-modelingai-research-evals 169 posts · 75 accounts

Claude Fable 5 max scored 91.9% on WeirdML, setting a new best result on the benchmark, according to results posted Saturday by researcher H Tihle.

The posted breakdown said a task by task state of the art composite would be 93.5%, indicating the model stayed within a few percentage points of the best score on each run. Claude Fable 5 max matched the top score on 7 of 17 tasks, and its weakest run was 7 percentage points behind the best result on that task. The test used two runs per task rather than the usual five, a setup the post said both highlights the model’s consistency and can make it easier to avoid one or two bad runs.

From the sources (25 posts)

@testingcatalog

OPENAI 🔥: ChatGPT Work will likely arrive in the form of an upgraded dedicated workspace, individual for every user. What we know so far 👀 > "ChatGPT Work can help you build websites, prototype new ideas, create presentations and documen

@financialjuice

OpenAI's Altman expects a smoother process working with the government on AI

@firstsquawk

OPENAI MADE 'MANY CHANGES' AFTER TALKING TO GOVT, ALTMAN SAYS

@financialjuice

OpenAI's Altman: OpenAI made many changes after talking to the government.

@financialjuice

OpenAI's Altman: OpenAI’s newest AI model is 54% more token-efficient on agentic coding.

@stockmktnewz

OpenAI CEO Sam Altman said that its latest AI model is 54% more token efficent on agent coding tasks - CNBC

@cnbc

OpenAI's newest AI model is 54% more token efficient on agentic coding, Altman tells CNBC

@business

Sam Altman said OpenAI made “many changes” during its discussions with the Trump administration before moving ahead with releasing its newest AI models to the general public

@polymarket

NEW: Sam Altman reveals OpenAI made “many changes” during talks with the Trump administration before GPT-5.6’s release.

@exec_sum

BREAKING: OpenAI is set to launch GPT-5.6, its most capable model yet after a delayed rollout

@daveaitel

RT @OpenAI: Today. 10am PT.

@firstsquawk

OPENAI UNVEILS THE GPT-5.6 MODEL FAMILY—SOL, TERRA, AND LUNA—WITH ROLLOUT ACROSS CHATGPT, CODEX, AND THE API.

@firstsquawk

OPENAI LAUNCHES CHATGPT WORK, POWERED BY GPT-5.6, WITH ROLLOUT BEGINNING FOR PRO, ENTERPRISE, AND EDU USERS, WHILE PLUS AND BUSINESS WILL FOLLOW.

@firstsquawk

OPENAI'S NEW CHATGPT DESKTOP APP, COMBINING CHAT, WORK, AND CODEX, IS ROLLING OUT GLOBALLY ON MAC AND WINDOWS FOR ALL USERS.

@sama

5.6 livestream going now. in addition to the model, 3 major product things. 1. ChatGPT Work--really big deal! 2. new ChatGPT desktop app 3. hosted sites

@digg

OpenAI said on its live stream that GPT-5.6 Sol, Terra and Luna are rolling out today. Sol is coming to paid plans over the next 24 hours. Terra and Luna are coming to free users.

@scaling01

GPT-5.6 Benchmarks

@wallstengine

OpenAI introduces ChatGPT Work, powered by GPT-5.6, and says Sol, Terra and Luna are rolling out across ChatGPT, Codex and the API.

@firstsquawk

OPENAI LAUNCHES GPT-5.6 FAMILY OF MODELS FOR GENERAL AVAILABILITY FOLLOWING LIMITED PREVIEW

@scaling01

GPT-5.6-Sol scoring a massive 7.78% on ARC-AGI-3 this is a massive jump over Opus 4.8's 1.5%

@cursor_ai

GPT-5.6 Sol, Terra, and Luna are now available in Cursor. On CursorBench, Sol scores 67.2%.

@sama

obviously the best model we have ever produced, but also one of the best blog posts we have ever produced:

@firstadopter

OpenAI @OpenAI announces ChatGPT Work. Key news: "an agent that can take action across your apps and files, stay with a project for hours if needed, and turn a goal into finished work." "More than 5 million people use Codex every week" "

@business

OpenAI is introducing a new AI agent that’s meant to field a wider range of complex tasks for hours at a time, bolstering its push to appeal to more business professionals.

@stevenheidel

we're excited to roll out GPT-5.6 today across the API, ChatGPT, and Codex. our teams are working hard to expand access to everyone over the next 24 hours

Preview built on a synthetic news corpus (16 weeks, Apr–Jul 2026). Impact calls are model reads, not price data.

About Archive