Artificial Analysis Launches Harvey Legal Benchmark, Claude Fable 5 Leads at 14.2% Pass Rate
Artificial Analysis launched the Harvey LAB-AA benchmark on Monday, evaluating 28 large language models on a private set of 120 real-world legal tasks spanning 24 practice areas. The benchmark measures the all-pass rate, or the share of tasks where a model fully satisfies every rubric criterion. Claude Fable 5 from Anthropic led the evaluation at 14.2%, roughly double the 7.5% achieved by Anthropic’s Claude Opus 4.8 and Zai Group’s GLM-5.2.
The evaluation highlights that while frontier models handle most individual requirements, they rarely complete whole legal matters end-to-end. Thirteen of the 28 models passed zero tasks entirely, underscoring that the industry is far from automating professional legal deliverables. Artificial Analysis noted the test is an independent reimplementation of a previous Harvey evaluation, using a different grading framework and excluding the original company's custom document-generation scripts. Costs to run the evaluation spanned roughly 950 times, with Claude Fable 5 averaging about $19 per task compared to $1.30 for GLM-5.2.
From the sources (7 posts)
@artificialanlysAfter our announcement last month, Artificial Analysis is now launching Harvey LAB-AA (Legal Agent Benchmark), our implementation of Harvey's new agentic legal benchmark that evaluates language models on real-world legal work across 24 prac
@artificialanlysModels handle most requirements of a legal matter, but rarely all of them. Leading models satisfy over 90% of individual rubric criteria, but pass only 0% to 14.2% of tasks in full, and only 4 models score above 90% criteria pass rate (Clau
@artificialanlysTopping the leaderboard is expensive. Claude Fable 5 leads all-pass at 14.2% and is the most expensive model run as of this post at ~$18.9 per task, ahead of Claude Sonnet 5 at ~$11.8 per task. Claude Opus 4.8 costs ~$8.2 per task, while De
@artificialanlysAll-pass rate is generally tied to how much work a model does per task. Claude Fable 5 generates ~117k output tokens per task to reach 14.2% all-pass, while Claude Opus 4.8 and GLM-5.2 generate ~111k and ~78k output tokens per task respecti
@artificialanlysStronger models tend to spend longer per task. Claude Fable 5 averages ~16.9 minutes per task, Claude Opus 4.8 ~18.5 minutes, and Claude Sonnet 5 ~22.8 minutes. GLM-5.2 is the exception, matching Claude Opus 4.8's all-pass rate at only ~5.0
@artificialanlysThe leaders work in long agentic loops. Claude Fable 5 averages ~64 turns per task for its 14.2% all-pass, with Claude Opus 4.8, GLM-5.2, and MiniMax-M3 in the 56-64 turn range. Claude Sonnet 5 runs the longest loops of any model at ~161 tu
@artificialanlysHarvey LAB-AA is our independent reimplementation of Harvey’s evaluation, and there are several key differences to the original version: ➤ Models are run on our Stirrup agent harness, enabling features such as context compaction rather tha