Pymander Health clinical AI achieved a 98.2% traceable score on the USMLE — the gold standard benchmark for complex medical reasoning.

August 3, 2026

What we measured

Two-tier cascade, cost-routed$1.70
91.4%
Single model with retrieval (corpus + live literature)$23.36
97.1%
Single model, no retrieval (shipped)$9.30 per run
98.2%
88%92%96%100%

Figure 1. Accuracy and cost by configuration, measured on the same 244-item set. Horizontal scale is truncated at 88%.

About the benchmark

The USMLE is the three-step examination required for medical licensure in the United States: Step 1 on the basic science of medicine, Step 2 CK on clinical knowledge and patient management, Step 3 on practising unsupervised. It is the standard benchmark for medical AI because its questions demand reasoning rather than recall, and because an official answer key exists — which makes a claimed score checkable instead of merely asserted.

The evaluation set comprised 244 text-only questions drawn from the official published sample examinations, excluding items dependent on an image. Each question was answered once, by a single model, without retrieval. The full set was administered twice as independent runs, yielding 479 correct responses of 488.

What this does not mean

It does not mean the system can practise medicine. The USMLE is a closed problem: the relevant facts are contained in the vignette, the differential is bounded, and one of the listed options is correct. Performance on a closed-form examination does not, without further evidence, extend to open-ended clinical care.

Why this is the beginning

What an examination can establish is that the reasoning underneath is sound. That is a precondition for the work, not the work itself. The harder measurement is the one Pymander is built on: continuous care, sustained over weeks, months, and years, and coupled with protocols, laboratory testing, and every other part of a person's healthcare.

For methodology questions or access to the underlying run data: hello@pymander.app