Command Palette
Search for a command to run...

Artificial Analysis Launches Harvey Legal Benchmark, Claude Fable 5 Leads at 14.2% Pass Rate

aiai-modelingai-research-evals 7 posts · 1 accounts

Artificial Analysis launched the Harvey LAB-AA benchmark on Monday, evaluating 28 large language models on a private set of 120 real-world legal tasks spanning 24 practice areas. The benchmark measures the all-pass rate, or the share of tasks where a model fully satisfies every rubric criterion. Claude Fable 5 from Anthropic led the evaluation at 14.2%, roughly double the 7.5% achieved by Anthropic’s Claude Opus 4.8 and Zai Group’s GLM-5.2.

The evaluation highlights that while frontier models handle most individual requirements, they rarely complete whole legal matters end-to-end. Thirteen of the 28 models passed zero tasks entirely, underscoring that the industry is far from automating professional legal deliverables. Artificial Analysis noted the test is an independent reimplementation of a previous Harvey evaluation, using a different grading framework and excluding the original company's custom document-generation scripts. Costs to run the evaluation spanned roughly 950 times, with Claude Fable 5 averaging about $19 per task compared to $1.30 for GLM-5.2.

From the sources (7 posts)

@artificialanlys

After our announcement last month, Artificial Analysis is now launching Harvey LAB-AA (Legal Agent Benchmark), our implementation of Harvey's new agentic legal benchmark that evaluates language models on real-world legal work across 24 prac

@artificialanlys

Models handle most requirements of a legal matter, but rarely all of them. Leading models satisfy over 90% of individual rubric criteria, but pass only 0% to 14.2% of tasks in full, and only 4 models score above 90% criteria pass rate (Clau

@artificialanlys

Topping the leaderboard is expensive. Claude Fable 5 leads all-pass at 14.2% and is the most expensive model run as of this post at ~$18.9 per task, ahead of Claude Sonnet 5 at ~$11.8 per task. Claude Opus 4.8 costs ~$8.2 per task, while De

@artificialanlys

All-pass rate is generally tied to how much work a model does per task. Claude Fable 5 generates ~117k output tokens per task to reach 14.2% all-pass, while Claude Opus 4.8 and GLM-5.2 generate ~111k and ~78k output tokens per task respecti

@artificialanlys

Stronger models tend to spend longer per task. Claude Fable 5 averages ~16.9 minutes per task, Claude Opus 4.8 ~18.5 minutes, and Claude Sonnet 5 ~22.8 minutes. GLM-5.2 is the exception, matching Claude Opus 4.8's all-pass rate at only ~5.0

@artificialanlys

The leaders work in long agentic loops. Claude Fable 5 averages ~64 turns per task for its 14.2% all-pass, with Claude Opus 4.8, GLM-5.2, and MiniMax-M3 in the 56-64 turn range. Claude Sonnet 5 runs the longest loops of any model at ~161 tu

@artificialanlys

Harvey LAB-AA is our independent reimplementation of Harvey’s evaluation, and there are several key differences to the original version: ➤ Models are run on our Stirrup agent harness, enabling features such as context compaction rather tha

Preview built on a synthetic news corpus (16 weeks, Apr–Jul 2026). Impact calls are model reads, not price data.

About Archive