Command Palette
Search for a command to run...

Artificial Analysis Scores Show AI Gains But Reveal Open Weights Gap

aiai-modelingai-research-evals 3 posts · 2 accounts

AI models demonstrated rapid performance gains but continued to trail closed systems on a new benchmark measuring long-horizon consulting tasks, according to scores released Tuesday by Artificial Analysis. The evaluation framework, known as AA-Briefcase, tasks frontier models with completing multi-week, highly complex workflows that require sustained reasoning and agent-like execution. Researcher Emil Eifrem graphed the results, noting a steep forward curve for top-performing models while highlighting a distinct performance gap between open-weight and closed-weights architectures.

The testing specifically evaluated Anthropic’s Fable model under a guardrailed configuration. Eifrem noted that the benchmark scores reflect these safety restrictions, and plotting the underlying, less-constrained Mythos model version would shift the frontier curve. The release adds to a growing set of industry assessments designed to stress-test AI beyond standard coding or mathematical reasoning, focusing instead on sustained, multi-step operational capabilities.

From the sources (3 posts)

@emollick

I took the new AA-Briefcase scores from @ArtificialAnlys (basically having the AI do multi-week consulting gigs with a lot of complexity) and graphed the frontier curve for open and closed models: 1) Surprise, rapid gains! 2) The open weigh

@artificialanlys

RT @emollick: I took the new AA-Briefcase scores from @ArtificialAnlys (basically having the AI do multi-week consulting gigs with a lot of…

@emollick

Even though I made this graph, it is also kind of wrong. Fable is guardrailed Mythos. If we use the Mythos date

Preview built on a synthetic news corpus (16 weeks, Apr–Jul 2026). Impact calls are model reads, not price data.

About Archive