Command Palette
Search for a command to run...

Nvidia, Harvey and Trajectory Post-Train Nemotron 3 Super on Legal Benchmark, Cite 25% Gain

aiai-modelingai-research-evalsai-open-models 6 posts · 4 accounts

Harvey, Trajectory and Nvidia said they post-trained Nvidia's open-weight Nemotron 3 Super on Harvey's Legal Agent Benchmark, or LAB, and that initial results show the model can match closed-source frontier systems on complex legal tasks. Trajectory said the work lifted rubric-pass criteria 25% in hours and matched GPT 5.5, while Nvidia said LAB spans more than 1,200 end-to-end tasks across 24 practice areas.

The companies cast the project as a step toward continual learning and sovereign legal AI, saying open-weight models offer auditability, security, provenance and data sovereignty in regulated work. Artificial Analysis has said it is working with Harvey to launch a full LAB leaderboard, and Nvidia said the partners plan to test the more powerful Nemotron 3 Ultra when available.

From the sources (6 posts)

@baseten

Our research team partnered with @Harvey and showed that post-trained open models can compete at the frontier on LAB, their legal agent benchmark.

@baseten

RT @brian_a_burns: Promising early results from Harvey's collaboration with @baseten Research on training custom open-weight agents for leg…

@baseten

RT @mudithj: LAB is such a useful foundation to explore what model behaviours lead to frontier performance and also to hill climb on. We v…

@artificialanlys

We're excited to work with Harvey to launch the full leaderboard for Legal Agent Benchmark - coming soon to Artificial Analysis!

@rronak_

5 days of Trajectory. Welcome to Day 2. At Trajectory, we are building the platform for continual learning. We unlock the real usage, edits, retries, that happen in products to improve model capabilities. Today, we are announcing our res

@nvidiaai

This is a great read on post-training and open models. @harvey & @trajectorylabs post-trained Nemotron 3 Super on complex legal tasks with some very impressive initial results. All with auditable weights, real security, and clear prove

Preview built on a synthetic news corpus (16 weeks, Apr–Jul 2026). Impact calls are model reads, not price data.

About Archive