OpenEvidence launches four medical AI models

04-09-2026

OpenEvidence released Osler, Sackett, Snow and Darwin on 3 September 2026. Darwin is the first AI to score 100% on MedQA and is application only.

Written by:

Jorick van Weelie

Marketing Lead at DataNorth | AI Enthusiast & Tech Storyteller

openevidence model family darwin scores 100% on medqa
Sign up for our Newsletter

Published 4 September 2026

OpenEvidence released a family of four medical AI models on 3 September 2026. Three of them, Osler, Sackett and Snow, are free to verified clinicians from today. The fourth, OpenEvidence Darwin, is available by application only and is reported as the first AI model to score 100% on MedQA, the leading independent benchmark of medical AI.

What are Osler, Sackett and Snow?

The three production models differ in one variable: how long they think. OpenEvidence holds all of them to the same standard of clinical accuracy and lets the clinician pick based on the time available.

  • Osler: about 5 seconds to answer, built for the pace of a patient encounter, now the default model
  • Sackett: about 30 seconds, asks the clinician for more context and weighs the evidence
  • Snow: about 5 minutes, runs a full investigation of the literature before writing
  • Where: openevidence.com and the OpenEvidence iOS and Android apps
  • Price: free to verified clinicians
  • Darwin: research preview, by application only

The names are the argument. William Osler moved medicine to the bedside, David Sackett founded evidence-based medicine, and John Snow traced the 1854 cholera outbreak to a single water pump. Osler replaces the model that previously powered OpenEvidence answers, and Snow succeeds the Deep Consult feature.

How good is OpenEvidence Darwin?

Darwin is the reason this release matters, and its scores are the only ones OpenEvidence published.

BenchmarkOpenEvidence DarwinNamed rivals
MedQA (standard medical exam questions)100%First model in history to reach it
HealthBench Professional (clinical judgement)82.7%Ahead of Claude Fable 5 and Gemini 3.7
NOHARM (safety of clinical advice)87.2%Ahead of Claude Fable 5 and Gemini 3.7
MedXpertQA (expert-level medical reasoning)72.8%Ahead of Claude Fable 5 and Gemini 3.7
Answer speed, fastest production modelOsler: about 5 secondsSnow: about 5 minutes
Price for verified cliniciansFreeDarwin: application only

These are OpenEvidence’s own figures. The company names Claude Fable 5 and Gemini 3.7 as the next-best models but does not publish their scores, so you can see that Darwin leads without seeing by how much. It does say it is publishing the MedQA results alongside Darwin’s explanation for every answer, which is more than a bare score and makes the claim checkable in a way most benchmark announcements are not.

A perfect score is worth reading carefully. It usually signals that a benchmark has been saturated rather than that a problem has been solved. MedQA is built from medical licensing exam questions, which is a narrower thing than clinical practice.

Why is Darwin restricted?

OpenEvidence gives a direct reason: capability in medicine is dual-use. A model reasoning at the frontier of virology, immunology and human genetics could accelerate bioweapons-relevant research or germline editing outside mainstream oversight. So Darwin goes to institutional partners such as the National Organization for Rare Disorders, research collaborators, and accredited academic researchers benchmarking clinical AI safety.

The plan is that Darwin’s capabilities flow into Osler, Sackett and Snow as its safeguards are validated. This is now the standard shape of a frontier launch. OpenAI did the same thing with GPT-6 Astra and its Daybreak programme on the same day, and Google did it with the Fairwind-gated Gemini 3.8 Flash Cyber the day before.

What this means

If you build clinical software, this changes your default. OpenEvidence says more American physicians use it than all other AI platforms combined, and it is free to verified US clinicians, so the three production models are about to become the baseline your product is compared against. The speed tiers are the practical detail: a five-second answer fits inside a consultation, a five-minute one does not, and matching the model to the moment is a product decision rather than a model decision.

Treat the Darwin numbers as a claim rather than a finding, for now. Every figure is first-party, the rival scores are absent, and no independent group has published a replication. That is not a reason to dismiss it, particularly since OpenEvidence is publishing per-answer explanations, but it is a reason to wait for the academic partners it named to report back. If you work outside healthcare, the interesting part is not the medicine. It is that a fourth vendor in three days gated its best model behind an application process on dual-use grounds. That is now how frontier releases work.

For more information, visit the official announcement of the OpenEvidence Model Family on the OpenEvidence newsroom.

Add DataNorth AI to your Google favorites