Frontier AI Models Beat OpenEvidence, UpToDate in 3 Evaluations, Nature Medicine Paper Finds
General AI frontier models from Google, OpenAI and Anthropic outperformed specialized clinical tools including OpenEvidence and UpToDate in medical-information evaluations, according to a Nature Medicine paper. The paper said 12 US clinicians assessed the systems in randomized, blinded testing and found frontier large language models led in all three evaluations.
The findings suggest general-purpose models are increasingly competitive in clinician-facing information tasks. The paper also said clinical AI tools performed comparably to auto-enabled Google Search AI Overview on the RCQ.
From the sources (4 posts)
@jeffdeanRT @EricTopol: For medical information, general AI frontier models (Google, OpenAI, Anthropic) outperformed specialized @EvidenceOpen and…
@emollickThere has been a push to use OpenEvidence AI for doctors. But this paper suggests general models are much better: “Frontier LLMs outperformed clinical AI tools in all three evaluations. Clinical AI tools performed comparably to auto-enabled
@andrewcurran_RT @emollick: There has been a push to use OpenEvidence AI for doctors. But this paper suggests general models are much better: “Frontier L…
@erictopolFor medical information, general AI frontier models (Google, OpenAI, Anthropic) outperformed specialized @EvidenceOpen and @UpToDate as assessed by 12 US clinicians, randomized and blinded to which model and extensive testing/benchmarks.