Artificial Analysis Scores Show AI Gains But Reveal Open Weights Gap
AI models demonstrated rapid performance gains but continued to trail closed systems on a new benchmark measuring long-horizon consulting tasks, according to scores released Tuesday by Artificial Analysis. The evaluation framework, known as AA-Briefcase, tasks frontier models with completing multi-week, highly complex workflows that require sustained reasoning and agent-like execution. Researcher Emil Eifrem graphed the results, noting a steep forward curve for top-performing models while highlighting a distinct performance gap between open-weight and closed-weights architectures.
The testing specifically evaluated Anthropic’s Fable model under a guardrailed configuration. Eifrem noted that the benchmark scores reflect these safety restrictions, and plotting the underlying, less-constrained Mythos model version would shift the frontier curve. The release adds to a growing set of industry assessments designed to stress-test AI beyond standard coding or mathematical reasoning, focusing instead on sustained, multi-step operational capabilities.
From the sources (3 posts)
@emollickI took the new AA-Briefcase scores from @ArtificialAnlys (basically having the AI do multi-week consulting gigs with a lot of complexity) and graphed the frontier curve for open and closed models: 1) Surprise, rapid gains! 2) The open weigh
@artificialanlysRT @emollick: I took the new AA-Briefcase scores from @ArtificialAnlys (basically having the AI do multi-week consulting gigs with a lot of…
@emollickEven though I made this graph, it is also kind of wrong. Fable is guardrailed Mythos. If we use the Mythos date