Command Palette
Search for a command to run...

Arena Launches AutoEval AI Ranking System to Cut Model Evaluation Time to Hours With Over 90% Accuracy

aiai-modelingai-research-evalsai-products 5 posts · 1 accounts

Arena launched AutoEval, an automated evaluation system that ranks newly released artificial intelligence models in hours rather than days by applying a reward model trained on millions of user preference votes. The preliminary scores appear directly on the platform’s public leaderboard and cover text, vision, image generation, and code categories.

The methodology provides developers an early signal to compare model checkpoints while waiting for human validation. In a temporal holdout test, the system produced a ranking correlation greater than 0.98 with subsequent live scores and correctly ordered competing models separated by at least 10 Arena points in more than 90% of cases. A leaderboard entry displaying an estimated score will not receive an official rank until sufficient live human votes confirm the ranking.

From the sources (5 posts)

@arena

Today we’re launching AutoEval: a new evaluation methodology that ranks models using reward models based on millions of real Arena user preferences. Highlights: - High-quality evaluation signals calibrated on real preference data - Strong

@arena

How does AutoEval produce Arena Score? For a new model to be evaluated, we generate responses on prompts sampled from live Arena evaluations. A pointwise reward model on millions of pairwise human votes to score each response and converts

@arena

We train our reward model to handle challenges like preference shifts, increasingly subtle model differences, ambiguous tie votes, and longer and more context-dependent interactions. To evaluate it fairly, we use a temporal holdout: train o

@arena

AutoEval scores provide an early signal for comparing candidate model checkpoints before full live results become available. We compared AutoEval’s early ordering with the final live leaderboard. When two models were separated by at least

@arena

The same approach extends beyond text. Across vision, image generation, and code arena, we train modality-specific reward models on live Arena preferences. For example, on Text-to-Image, our reward model was trained on more than 3 million

Preview built on a synthetic news corpus (16 weeks, Apr–Jul 2026). Impact calls are model reads, not price data.

About Archive