Which AI is best for clinical questions?

Not a benchmark — a verdict. The same case and the same question go to every model; the answers are judged blind. Licensed clinicians judge real de-identified charts inside CareOS. You can judge realistic sample cases right here, free.

Use it on real charts
Clinician verdictsCommunity verdicts Always judged blind Zero patient data

The leaderboard

Win rate = share of blinded head-to-heads this model's answer was crowned best.

No blind clinician verdicts in this category yet

This board fills as clinicians across CareOS judge blinded comparisons on real cases.

Licensed clinicians judging real de-identified charts inside CareOS. Ratings are 1–5.

Try the arena — free, no account

Pick a case, ask the question, judge the answers blind. Your verdict joins the community board above. Cases are fictional; the models are real.

Models answer as A / B / C — you judge first, then they're revealed.

The contenders

List prices as published by each vendor — the arena measures judgment; this is what the judgment costs.

ModelVendorAccessInput $/M tokensOutput $/M tokens
Claude Sonnet 5AnthropicAPI — runs automatically$3.00$15.00
Claude Opus 4.8AnthropicAPI — runs automatically$15.00$75.00
GPT-5OpenAIAPI — runs automatically$1.25$10.00
Gemini 3.1 ProGoogleAPI — runs automatically$1.25$10.00
Gemini 3.6 FlashGoogleAPI — runs automatically$0.30$2.50
OpenEvidenceOpenEvidenceNo public API — pasted in
FunctionalMindJohn Snow LabsNo public API — pasted in

How the arena works

  1. 1

    One case, every model

    The identical context and question go to every contender — here a fictional sample case; inside CareOS, the real chart with identity stripped server-side. Same input, or the comparison means nothing.

  2. 2

    Blind answers

    Responses come back as Model A, B, C — no logos, no vendor tells. Judge them the way you'd weigh a colleague's curbside opinion: does it use the actual data, does it catch the unsafe thing, could you act on it?

  3. 3

    A verdict, then the reveal

    Crown the best answer, then see who wrote what — often a surprise. Clinician verdicts and community verdicts build separate boards, and only blind votes count.

Frequently asked

How are the models judged?

Every model receives the identical case context and question. Answers come back blinded — Model A, B, C — and the judge crowns the best one before identities are revealed. Inside CareOS, licensed clinicians run this on real de-identified charts; those verdicts build the clinician board. On this page, anyone can judge fictional sample cases; those verdicts build the community board. The two never mix.

Is any patient data on this page?

No. The try-it cases are entirely fictional, written to read like real charts. The leaderboards are aggregate vote counts and rating averages. And inside CareOS, patient identity — name, date of birth, contact details — is stripped server-side before a question reaches any model.

Why does OpenEvidence have fewer votes?

OpenEvidence has no public API, so it can't run in the automated arena. Inside CareOS, clinicians paste its answer into a run by hand; pasted answers are identified rather than blinded, and identified verdicts are excluded from the public boards — so its sample grows more slowly.

Can I run this on my own patients?

Yes — Model Arena ships inside CareOS, the AI-native EHR for functional medicine, wellness, hormone, and longevity clinics. Assemble a case from the actual chart (intake, notes, transcripts, labs, meds, wearables), fan the question out to every connected model, and judge blind. Unlimited runs, your own data, your verdicts remembered.

How many free runs do I get here?

A handful per day — each run calls the real vendor APIs on our dime. Clinics on CareOS run unlimited comparisons with their own model credentials.

The arena is the demo. The EHR is the product.

CareOS is the AI-native EHR for functional medicine, wellness, hormone, and longevity clinics — charts, scribing, labs, and every frontier model working from the same record. Providers run the arena on real patients; patients get care teams whose AI is actually audited.