Which AI is best at lab interpretation?

Reading a lab panel in whole-person context — connecting markers to symptoms, medications, and cycle phase instead of reciting reference ranges — is where clinical AI answers differ most. Clinicians judge them blind on real de-identified results.

The question every model answers: “Interpret this patient's lab results in whole-person context. Which findings matter most, how do they connect to the presenting concerns, and what follow-up testing or intervals would you suggest?

Judged blind By licensed clinicians Zero patient data

The leaderboard

Win rate = share of blinded head-to-heads this model's answer was crowned best.

No blind clinician verdicts in this category yet

This board fills as clinicians across CareOS judge blinded comparisons on real cases.

Licensed clinicians judging real de-identified charts inside CareOS. Ratings are 1–5.

Other tasks on the board

Judge them yourself

Try the arena free on realistic sample cases — or run it on your own patients inside CareOS, the AI-native EHR for functional medicine, wellness, hormone, and longevity clinics.