Health Optimization Bench

A rubric-graded benchmark of frontier language models on current clinical evidence.

Leaderboard

Sources

257 tasks across eight subject areas. Updated September 10, 2026.

  1. 70.9
    Anthropic logoClaude Fable 5
  2. 69.3
    Anthropic logoClaude Opus 5
  3. 66.8
    SpaceX AI logoGrok 4.6
  4. 66.6
    OpenAI logoGPT-5.6 Sol (max)
  5. 64.6
    OpenAI logoGPT-5.6 Sol (high)
  6. 59.9
    Moonshot AI logoKimi K3
  7. 57.2
    Meta logoMuse Spark
  8. 47.3
    Anthropic logoClaude Fable 5.1
  9. 39.7
    Google logoGemini 3.6
  10. 35.6
    TMInkling
  11. 34.6
    Anthropic logoClaude Sonnet 5
  12. 20.7
    Zhipu logoGLM 5.2
  13. 18.3
    MiniMax logoMiniMax M3
  14. 17.5
    Microsoft AI logoMAI Thinking
  15. 9.2
    Mistral logoMistral Medium 3.5
  16. 4.9
    NVIDIA logoNemotron 3.5 Lightning

Claude Fable 5.1 declined 97 of 257 tasks: the vendor's safeguard classifier returned a refusal notice instead of an answer. Declined tasks are graded like any other completion and earn no credit; over the tasks it answered, its mean is 75.9.

Consensus rubric credit on the release set, one answer per task, tools disabled. Every score is from Arcophos's own harness run; the per-model record is on the sources page. The bench totals 346 released tasks across its two live rankings; the incretin therapeutics evidence suite carries its own ranking on the suite page.

The evaluation harness is open source. Run the public sample under the same blinding rules and scoring formula with arcophos-evals, an Inspect native task pack with a zero dependency reference runner, against the sample dataset on Hugging Face.

Health Optimization Bench evaluates language models on hard, freshness-dependent questions in preventive and optimization medicine. Every task is written against a primary source, audited by model families that did not author it, and scored blind by a panel of three independent families. The task's authoring family never grades it.

Micro benches

The bench is assembled from focused micro benches, each covering one area of preventive and optimization medicine. Grading provenance is marked per suite: hybrid model graded suites are scored by a cross-family panel of independent model families under the bench-wide protocol, and physician graded suites additionally carry licensed-clinician validation. Suite rankings publish as they complete; the methodology page describes both grading modes in full.

Data sample

A 30 task sample, drawn from each micro bench, is available through the Arcophos research data platform, with rubric criteria and reference responses included. The same sample is published as a Hugging Face dataset and runs directly in the open source evaluation harness.

View the sample

Paper

The methodology paper is in preparation and will accompany the archived v1 release, together with per-task verification records and the grading protocol. For early access, licensing, or evaluation of a private model, contact Arcophos.