Health Optimization Bench

Sources

Every score on this site comes from one place: Arcophos's own run of the benchmark harness. No number here is copied from a vendor system card, a launch post, or another leaderboard. The document of record for that run is the methodology page, which states the authoring, audit, and grading protocol the scores were produced under.

The tables below list, per board and per model, the score, its 95 percent confidence interval, the number of released tasks scored, the snapshot month, and the locator: the task set analysis file the number was read from and the date of that snapshot. The site shows 32 scores across 2 boards; none are awaiting source documentation.

Documents

Arcophos runHealth Optimization Bench: Arcophos harness run
Arcophos, first party; cited by 32 rows

The run itself: harness analysis over the release set of each task set, one answer per task, tools disabled. Locators name the task set directory in the Arcophos harness repository (tasksets/glp1-ev for the incretin evidence suite, tasksets/mb2ev for the subject suites) and the snapshot date the analysis was read. The public sample and the open-source runner are linked from the data sample section.

Subject suites

257 released tasks · 16 models · snapshot 2026-09-04 · ranking

#modelscore95% CItaskssnapshotsource
1Anthropic logoClaude Fable 5 Anthropic70.968.0–73.8257September 2026Arcophos run
tasksets/mb2ev/analysis.json (regenerate with benchdb import-hob on the machine holding the harness tasksets; this row was verified against the committed site snapshot), snapshot 2026-09-04
score 70.9, 95% CI 68 to 73.8, n 257 (values as published in the site snapshot of 2026-09-04)
2Anthropic logoClaude Opus 5 Anthropic69.366.2–72.4257September 2026Arcophos run
tasksets/mb2ev/analysis.json (regenerate with benchdb import-hob on the machine holding the harness tasksets; this row was verified against the committed site snapshot), snapshot 2026-09-04
score 69.3, 95% CI 66.2 to 72.4, n 257 (values as published in the site snapshot of 2026-09-04)
3xAI logoGrok 4.6 xAI66.863.8–69.8257September 2026Arcophos run
tasksets/mb2ev/analysis.json (regenerate with benchdb import-hob on the machine holding the harness tasksets; this row was verified against the committed site snapshot), snapshot 2026-09-04
score 66.8, 95% CI 63.8 to 69.8, n 257 (values as published in the site snapshot of 2026-09-04)
4OpenAI logoGPT-5.6 Sol (max) OpenAI
max reasoning effort
66.663.4–69.6257September 2026Arcophos run
tasksets/mb2ev/analysis.json (regenerate with benchdb import-hob on the machine holding the harness tasksets; this row was verified against the committed site snapshot), snapshot 2026-09-04
score 66.6, 95% CI 63.4 to 69.6, n 257 (values as published in the site snapshot of 2026-09-04)
5OpenAI logoGPT-5.6 Sol (high) OpenAI
high reasoning effort
64.661.5–67.6257September 2026Arcophos run
tasksets/mb2ev/analysis.json (regenerate with benchdb import-hob on the machine holding the harness tasksets; this row was verified against the committed site snapshot), snapshot 2026-09-04
score 64.6, 95% CI 61.5 to 67.6, n 257 (values as published in the site snapshot of 2026-09-04)
6Moonshot AI logoKimi K3 Moonshot AI59.956.5–63.3257September 2026Arcophos run
tasksets/mb2ev/analysis.json (regenerate with benchdb import-hob on the machine holding the harness tasksets; this row was verified against the committed site snapshot), snapshot 2026-09-04
score 59.9, 95% CI 56.5 to 63.3, n 257 (values as published in the site snapshot of 2026-09-04)
7Meta logoMuse Spark Meta57.253.9–60.5257September 2026Arcophos run
tasksets/mb2ev/analysis.json (regenerate with benchdb import-hob on the machine holding the harness tasksets; this row was verified against the committed site snapshot), snapshot 2026-09-04
score 57.2, 95% CI 53.9 to 60.5, n 257 (values as published in the site snapshot of 2026-09-04)
8Anthropic logoClaude Fable 5.1 Anthropic47.342.2–52.2257September 2026Arcophos run
tasksets/mb2ev/analysis.json (regenerate with benchdb import-hob on the machine holding the harness tasksets; this row was verified against the committed site snapshot), snapshot 2026-09-04
score 47.3, 95% CI 42.2 to 52.2, n 257 (values as published in the site snapshot of 2026-09-04)
9Google logoGemini 3.6 Google39.736.2–43.1257September 2026Arcophos run
tasksets/mb2ev/analysis.json (regenerate with benchdb import-hob on the machine holding the harness tasksets; this row was verified against the committed site snapshot), snapshot 2026-09-04
score 39.7, 95% CI 36.2 to 43.1, n 257 (values as published in the site snapshot of 2026-09-04)
10TMInkling Thinking Machines35.632.3–39.0257September 2026Arcophos run
tasksets/mb2ev/analysis.json (regenerate with benchdb import-hob on the machine holding the harness tasksets; this row was verified against the committed site snapshot), snapshot 2026-09-04
score 35.6, 95% CI 32.3 to 39, n 257 (values as published in the site snapshot of 2026-09-04)
11Anthropic logoClaude Sonnet 5 Anthropic34.631.6–37.8257September 2026Arcophos run
tasksets/mb2ev/analysis.json (regenerate with benchdb import-hob on the machine holding the harness tasksets; this row was verified against the committed site snapshot), snapshot 2026-09-04
score 34.6, 95% CI 31.6 to 37.8, n 257 (values as published in the site snapshot of 2026-09-04)
12Zhipu logoGLM 5.2 Zhipu20.718.1–23.4257September 2026Arcophos run
tasksets/mb2ev/analysis.json (regenerate with benchdb import-hob on the machine holding the harness tasksets; this row was verified against the committed site snapshot), snapshot 2026-09-04
score 20.7, 95% CI 18.1 to 23.4, n 257 (values as published in the site snapshot of 2026-09-04)
13MiniMax logoMiniMax M3 MiniMax18.315.7–20.9257September 2026Arcophos run
tasksets/mb2ev/analysis.json (regenerate with benchdb import-hob on the machine holding the harness tasksets; this row was verified against the committed site snapshot), snapshot 2026-09-04
score 18.3, 95% CI 15.7 to 20.9, n 257 (values as published in the site snapshot of 2026-09-04)
14Microsoft AI logoMAI Thinking Microsoft AI17.515.1–20.0257September 2026Arcophos run
tasksets/mb2ev/analysis.json (regenerate with benchdb import-hob on the machine holding the harness tasksets; this row was verified against the committed site snapshot), snapshot 2026-09-04
score 17.5, 95% CI 15.1 to 20, n 257 (values as published in the site snapshot of 2026-09-04)
15Mistral logoMistral Medium 3.5 Mistral9.27.4–11.1257September 2026Arcophos run
tasksets/mb2ev/analysis.json (regenerate with benchdb import-hob on the machine holding the harness tasksets; this row was verified against the committed site snapshot), snapshot 2026-09-04
score 9.2, 95% CI 7.4 to 11.1, n 257 (values as published in the site snapshot of 2026-09-04)
16NVIDIA logoNemotron 3.5 Lightning NVIDIA4.93.6–6.2257September 2026Arcophos run
tasksets/mb2ev/analysis.json (regenerate with benchdb import-hob on the machine holding the harness tasksets; this row was verified against the committed site snapshot), snapshot 2026-09-04
score 4.9, 95% CI 3.6 to 6.2, n 257 (values as published in the site snapshot of 2026-09-04)

Incretin Therapeutics: Evidence Synthesis

89 released tasks · 16 models · snapshot 2026-09-04 · ranking

#modelscore95% CItaskssnapshotsource
1Anthropic logoClaude Fable 5.1 Anthropic84.980.1–89.389September 2026Arcophos run
tasksets/glp1-ev/analysis.json (regenerate with benchdb import-hob on the machine holding the harness tasksets; this row was verified against the committed site snapshot), snapshot 2026-09-04
score 84.9, 95% CI 80.1 to 89.3, n 89 (values as published in the site snapshot of 2026-09-04)
2Anthropic logoClaude Fable 5 Anthropic83.879.1–88.189September 2026Arcophos run
tasksets/glp1-ev/analysis.json (regenerate with benchdb import-hob on the machine holding the harness tasksets; this row was verified against the committed site snapshot), snapshot 2026-09-04
score 83.8, 95% CI 79.1 to 88.1, n 89 (values as published in the site snapshot of 2026-09-04)
3xAI logoGrok 4.6 xAI81.377.0–85.389September 2026Arcophos run
tasksets/glp1-ev/analysis.json (regenerate with benchdb import-hob on the machine holding the harness tasksets; this row was verified against the committed site snapshot), snapshot 2026-09-04
score 81.3, 95% CI 77 to 85.3, n 89 (values as published in the site snapshot of 2026-09-04)
4Anthropic logoClaude Opus 5 Anthropic78.372.9–83.489September 2026Arcophos run
tasksets/glp1-ev/analysis.json (regenerate with benchdb import-hob on the machine holding the harness tasksets; this row was verified against the committed site snapshot), snapshot 2026-09-04
score 78.3, 95% CI 72.9 to 83.4, n 89 (values as published in the site snapshot of 2026-09-04)
5OpenAI logoGPT-5.6 Sol (max) OpenAI
max reasoning effort
77.973.3–82.289September 2026Arcophos run
tasksets/glp1-ev/analysis.json (regenerate with benchdb import-hob on the machine holding the harness tasksets; this row was verified against the committed site snapshot), snapshot 2026-09-04
score 77.9, 95% CI 73.3 to 82.2, n 89 (values as published in the site snapshot of 2026-09-04)
6Moonshot AI logoKimi K3 Moonshot AI77.873.0–82.489September 2026Arcophos run
tasksets/glp1-ev/analysis.json (regenerate with benchdb import-hob on the machine holding the harness tasksets; this row was verified against the committed site snapshot), snapshot 2026-09-04
score 77.8, 95% CI 73 to 82.4, n 89 (values as published in the site snapshot of 2026-09-04)
7OpenAI logoGPT-5.6 Sol (high) OpenAI
high reasoning effort
77.773.1–82.089September 2026Arcophos run
tasksets/glp1-ev/analysis.json (regenerate with benchdb import-hob on the machine holding the harness tasksets; this row was verified against the committed site snapshot), snapshot 2026-09-04
score 77.7, 95% CI 73.1 to 82, n 89 (values as published in the site snapshot of 2026-09-04)
8Meta logoMuse Spark Meta68.462.8–73.989September 2026Arcophos run
tasksets/glp1-ev/analysis.json (regenerate with benchdb import-hob on the machine holding the harness tasksets; this row was verified against the committed site snapshot), snapshot 2026-09-04
score 68.4, 95% CI 62.8 to 73.9, n 89 (values as published in the site snapshot of 2026-09-04)
9Google logoGemini 3.6 Google57.951.2–64.389September 2026Arcophos run
tasksets/glp1-ev/analysis.json (regenerate with benchdb import-hob on the machine holding the harness tasksets; this row was verified against the committed site snapshot), snapshot 2026-09-04
score 57.9, 95% CI 51.2 to 64.3, n 89 (values as published in the site snapshot of 2026-09-04)
10TMInkling Thinking Machines53.947.7–60.189September 2026Arcophos run
tasksets/glp1-ev/analysis.json (regenerate with benchdb import-hob on the machine holding the harness tasksets; this row was verified against the committed site snapshot), snapshot 2026-09-04
score 53.9, 95% CI 47.7 to 60.1, n 89 (values as published in the site snapshot of 2026-09-04)
11Anthropic logoClaude Sonnet 5 Anthropic50.144.3–56.189September 2026Arcophos run
tasksets/glp1-ev/analysis.json (regenerate with benchdb import-hob on the machine holding the harness tasksets; this row was verified against the committed site snapshot), snapshot 2026-09-04
score 50.1, 95% CI 44.3 to 56.1, n 89 (values as published in the site snapshot of 2026-09-04)
12MiniMax logoMiniMax M3 MiniMax35.529.6–41.689September 2026Arcophos run
tasksets/glp1-ev/analysis.json (regenerate with benchdb import-hob on the machine holding the harness tasksets; this row was verified against the committed site snapshot), snapshot 2026-09-04
score 35.5, 95% CI 29.6 to 41.6, n 89 (values as published in the site snapshot of 2026-09-04)
13Microsoft AI logoMAI Thinking Microsoft AI33.026.8–39.389September 2026Arcophos run
tasksets/glp1-ev/analysis.json (regenerate with benchdb import-hob on the machine holding the harness tasksets; this row was verified against the committed site snapshot), snapshot 2026-09-04
score 33.0, 95% CI 26.8 to 39.3, n 89 (values as published in the site snapshot of 2026-09-04)
14Zhipu logoGLM 5.2 Zhipu29.223.9–34.989September 2026Arcophos run
tasksets/glp1-ev/analysis.json (regenerate with benchdb import-hob on the machine holding the harness tasksets; this row was verified against the committed site snapshot), snapshot 2026-09-04
score 29.2, 95% CI 23.9 to 34.9, n 89 (values as published in the site snapshot of 2026-09-04)
15Mistral logoMistral Medium 3.5 Mistral14.410.6–18.589September 2026Arcophos run
tasksets/glp1-ev/analysis.json (regenerate with benchdb import-hob on the machine holding the harness tasksets; this row was verified against the committed site snapshot), snapshot 2026-09-04
score 14.4, 95% CI 10.6 to 18.5, n 89 (values as published in the site snapshot of 2026-09-04)
16NVIDIA logoNemotron 3.5 Lightning NVIDIA7.84.7–11.389September 2026Arcophos run
tasksets/glp1-ev/analysis.json (regenerate with benchdb import-hob on the machine holding the harness tasksets; this row was verified against the committed site snapshot), snapshot 2026-09-04
score 7.8, 95% CI 4.7 to 11.3, n 89 (values as published in the site snapshot of 2026-09-04)

Protocol in brief

Tasks are authored one at a time by one of five frontier model families, audited for material error and citation fidelity by families that did not write them, and removed only when a second family independently confirms a material verdict. Candidate models answer closed book, one answer per task, tools disabled. Each answer is graded blind in three passes spread across three grader families, with the task's authoring family excluded from its panel; any split criterion escalates to a fourth, uninvolved family. Scores are consensus rubric credit on a 0 to 100 scale. Confidence intervals are nonparametric bootstraps over tasks, 10,000 resamples, fixed seed. Scores are comparable only within an identical task set, grading panel, and pass count. The full protocol, the measured verification outcomes, and the safety scoring rule are on the methodology page.