A rubric-graded benchmark of frontier language models on current clinical evidence.
Leaderboard
Sources257 tasks across eight subject areas. Updated September 10, 2026.
Claude Fable 5
Claude Opus 5
Grok 4.6
GPT-5.6 Sol (max)
GPT-5.6 Sol (high)
Kimi K3
Muse Spark
Claude Fable 5.1
Gemini 3.6
- TMInkling
Claude Sonnet 5
GLM 5.2
MiniMax M3
MAI Thinking
Mistral Medium 3.5
Nemotron 3.5 Lightning
Claude Fable 5.1 declined 97 of 257 tasks: the vendor's safeguard classifier returned a refusal notice instead of an answer. Declined tasks are graded like any other completion and earn no credit; over the tasks it answered, its mean is 75.9.
Consensus rubric credit on the release set, one answer per task, tools disabled. Every score is from Arcophos's own harness run; the per-model record is on the sources page. The bench totals 346 released tasks across its two live rankings; the incretin therapeutics evidence suite carries its own ranking on the suite page.
The evaluation harness is open source. Run the public sample under the same blinding rules and scoring formula with arcophos-evals, an Inspect native task pack with a zero dependency reference runner, against the sample dataset on Hugging Face.
Health Optimization Bench evaluates language models on hard, freshness-dependent questions in preventive and optimization medicine. Every task is written against a primary source, audited by model families that did not author it, and scored blind by a panel of three independent families. The task's authoring family never grades it.
Micro benches
The bench is assembled from focused micro benches, each covering one area of preventive and optimization medicine. Grading provenance is marked per suite: hybrid model graded suites are scored by a cross-family panel of independent model families under the bench-wide protocol, and physician graded suites additionally carry licensed-clinician validation. Suite rankings publish as they complete; the methodology page describes both grading modes in full.
- 89 tasks · ranking live
- Incretin Therapeutics: Clinical Decisions10 tasks · physician graded
- Cancer Screening & Early Detection33 tasks · hybrid model graded
- Exercise & Cardiorespiratory Fitness30 tasks · hybrid model graded
- Longevity / Geroscience Pharmacology33 tasks · hybrid model graded
- Hormone Optimization31 tasks · hybrid model graded
- Blood Pressure Optimization33 tasks · hybrid model graded
- Lipid & ASCVD Prevention32 tasks · hybrid model graded
- Nutrition & Supplements37 tasks · hybrid model graded
- Sleep Optimization28 tasks · hybrid model graded
Data sample
A 30 task sample, drawn from each micro bench, is available through the Arcophos research data platform, with rubric criteria and reference responses included. The same sample is published as a Hugging Face dataset and runs directly in the open source evaluation harness.
View the samplePaper
The methodology paper is in preparation and will accompany the archived v1 release, together with per-task verification records and the grading protocol. For early access, licensing, or evaluation of a private model, contact Arcophos.