Not a benchmark — a verdict. The same case and the same question go to every model; the answers are judged blind. Licensed clinicians judge real de-identified charts inside CareOS. You can judge realistic sample cases right here, free.
Win rate = share of blinded head-to-heads this model's answer was crowned best.
No blind clinician verdicts in this category yet
This board fills as clinicians across CareOS judge blinded comparisons on real cases.
Licensed clinicians judging real de-identified charts inside CareOS. Ratings are 1–5.
Pick a case, ask the question, judge the answers blind. Your verdict joins the community board above. Cases are fictional; the models are real.
List prices as published by each vendor — the arena measures judgment; this is what the judgment costs.
| Model | Vendor | Access | Input $/M tokens | Output $/M tokens |
|---|---|---|---|---|
| Claude Sonnet 5 | Anthropic | API — runs automatically | $3.00 | $15.00 |
| Claude Opus 4.8 | Anthropic | API — runs automatically | $15.00 | $75.00 |
| GPT-5 | OpenAI | API — runs automatically | $1.25 | $10.00 |
| Gemini 3.1 Pro | API — runs automatically | $1.25 | $10.00 | |
| Gemini 3.6 Flash | API — runs automatically | $0.30 | $2.50 | |
| OpenEvidence | OpenEvidence | No public API — pasted in | — | — |
| FunctionalMind | John Snow Labs | No public API — pasted in | — | — |
The identical context and question go to every contender — here a fictional sample case; inside CareOS, the real chart with identity stripped server-side. Same input, or the comparison means nothing.
Responses come back as Model A, B, C — no logos, no vendor tells. Judge them the way you'd weigh a colleague's curbside opinion: does it use the actual data, does it catch the unsafe thing, could you act on it?
Crown the best answer, then see who wrote what — often a surprise. Clinician verdicts and community verdicts build separate boards, and only blind votes count.
Every model receives the identical case context and question. Answers come back blinded — Model A, B, C — and the judge crowns the best one before identities are revealed. Inside CareOS, licensed clinicians run this on real de-identified charts; those verdicts build the clinician board. On this page, anyone can judge fictional sample cases; those verdicts build the community board. The two never mix.
No. The try-it cases are entirely fictional, written to read like real charts. The leaderboards are aggregate vote counts and rating averages. And inside CareOS, patient identity — name, date of birth, contact details — is stripped server-side before a question reaches any model.
OpenEvidence has no public API, so it can't run in the automated arena. Inside CareOS, clinicians paste its answer into a run by hand; pasted answers are identified rather than blinded, and identified verdicts are excluded from the public boards — so its sample grows more slowly.
Yes — Model Arena ships inside CareOS, the AI-native EHR for functional medicine, wellness, hormone, and longevity clinics. Assemble a case from the actual chart (intake, notes, transcripts, labs, meds, wearables), fan the question out to every connected model, and judge blind. Unlimited runs, your own data, your verdicts remembered.
A handful per day — each run calls the real vendor APIs on our dime. Clinics on CareOS run unlimited comparisons with their own model credentials.
CareOS is the AI-native EHR for functional medicine, wellness, hormone, and longevity clinics — charts, scribing, labs, and every frontier model working from the same record. Providers run the arena on real patients; patients get care teams whose AI is actually audited.