Command Palette
Search for a command to run...

Vals AI Launches Benchmark Showing Exa Search Tool Beats Native LLM Capabilities at 48.5% Score

aiai-modelingai-research-evals 6 posts · 1 accounts

Vals AI launched the Web Search Index benchmark to compare third-party search tools against built-in model capabilities, recording a top score of 48.5% for Claude Fable 5 paired with Exa search. The methodology isolates variables by providing each model with the same setup and stripping out preloaded domain tools like SEC EDGAR and CourtListener, ensuring performance differences trace directly to the search provider.

Exa outperformed native search across the benchmark's eight model-tool combinations, though native search remained cheaper on average, with Gemini 3.5 Flash pricing tasks at $0.40 each. Results varied by sector: Exa led finance tasks by 6 percentage points in market and earnings analysis, while native search held a statistically insignificant edge in legal applications except for regulatory framework interpretation, where Exa scored 54% versus 48%. Vals AI plans to add additional search providers to the benchmark in the coming months.

From the sources (6 posts)

@valsai

We are excited to launch Web Search Index! Web Search Index compares native provider search against independent web-search tools on real world tasks.

@valsai

The core design choice is isolation. We give the same model the same setup every time and change only one thing: the web search tool it uses. We also strip out the domain-specific tools the benchmarks ship with, like SEC EDGAR and CourtList

@valsai

Claude Fable 5 with @ExaAILabs tops the overall Web Search Index at 48.5%, the highest of all eight model-tool combinations. Overall, Exa performs better than the native search tool, whereas native search cost less than Exa on average, with

@valsai

We found that the winner is domain dependent. On finance, Exa scores about 6 percentage points higher than native once we control for task difficulty (p < 0.001). The margin is widest where a task depends heavily on retrieving data: Market

@valsai

On legal, native is ahead by less than a percentage point, a gap that is not statistically significant (p = 0.86). Native search performs best on tasks requiring the application of precedent to a fact pattern (about 35% vs. 32% for Exa) and

@valsai

We are expecting to add more search providers over time, if you have any requests let us know. For the full methodology and results, visit:

Preview built on a synthetic news corpus (16 weeks, Apr–Jul 2026). Impact calls are model reads, not price data.

About Archive