Sources
Every score on this site comes from one place: Arcophos's own run of the benchmark harness. No number here is copied from a vendor system card, a launch post, or another leaderboard. The document of record for that run is the methodology page, which states the authoring, audit, and grading protocol the scores were produced under.
The tables below list, per board and per model, the score, its 95 percent confidence interval, the number of released tasks scored, the snapshot month, and the locator: the task set analysis file the number was read from and the date of that snapshot. The site shows 32 scores across 2 boards; none are awaiting source documentation.
Documents
| Arcophos run | Health Optimization Bench: Arcophos harness run Arcophos, first party; cited by 32 rows |
The run itself: harness analysis over the release set of each task set, one answer per task, tools disabled. Locators name the task set directory in the Arcophos harness repository (tasksets/glp1-ev for the incretin evidence suite, tasksets/mb2ev for the subject suites) and the snapshot date the analysis was read. The public sample and the open-source runner are linked from the data sample section.
Subject suites
257 released tasks · 16 models · snapshot 2026-09-04 · ranking
| # | model | score | 95% CI | tasks | snapshot | source | |
|---|---|---|---|---|---|---|---|
| 1 | Claude Fable 5 Anthropic | 70.9 | 68.0–73.8 | 257 | September 2026 | Arcophos run tasksets/mb2ev/analysis.json (regenerate with benchdb import-hob on the machine holding the harness tasksets; this row was verified against the committed site snapshot), snapshot 2026-09-04 “score 70.9, 95% CI 68 to 73.8, n 257 (values as published in the site snapshot of 2026-09-04)” | |
| 2 | Claude Opus 5 Anthropic | 69.3 | 66.2–72.4 | 257 | September 2026 | Arcophos run tasksets/mb2ev/analysis.json (regenerate with benchdb import-hob on the machine holding the harness tasksets; this row was verified against the committed site snapshot), snapshot 2026-09-04 “score 69.3, 95% CI 66.2 to 72.4, n 257 (values as published in the site snapshot of 2026-09-04)” | |
| 3 | Grok 4.6 xAI | 66.8 | 63.8–69.8 | 257 | September 2026 | Arcophos run tasksets/mb2ev/analysis.json (regenerate with benchdb import-hob on the machine holding the harness tasksets; this row was verified against the committed site snapshot), snapshot 2026-09-04 “score 66.8, 95% CI 63.8 to 69.8, n 257 (values as published in the site snapshot of 2026-09-04)” | |
| 4 | GPT-5.6 Sol (max) OpenAI max reasoning effort | 66.6 | 63.4–69.6 | 257 | September 2026 | Arcophos run tasksets/mb2ev/analysis.json (regenerate with benchdb import-hob on the machine holding the harness tasksets; this row was verified against the committed site snapshot), snapshot 2026-09-04 “score 66.6, 95% CI 63.4 to 69.6, n 257 (values as published in the site snapshot of 2026-09-04)” | |
| 5 | GPT-5.6 Sol (high) OpenAI high reasoning effort | 64.6 | 61.5–67.6 | 257 | September 2026 | Arcophos run tasksets/mb2ev/analysis.json (regenerate with benchdb import-hob on the machine holding the harness tasksets; this row was verified against the committed site snapshot), snapshot 2026-09-04 “score 64.6, 95% CI 61.5 to 67.6, n 257 (values as published in the site snapshot of 2026-09-04)” | |
| 6 | Kimi K3 Moonshot AI | 59.9 | 56.5–63.3 | 257 | September 2026 | Arcophos run tasksets/mb2ev/analysis.json (regenerate with benchdb import-hob on the machine holding the harness tasksets; this row was verified against the committed site snapshot), snapshot 2026-09-04 “score 59.9, 95% CI 56.5 to 63.3, n 257 (values as published in the site snapshot of 2026-09-04)” | |
| 7 | Muse Spark Meta | 57.2 | 53.9–60.5 | 257 | September 2026 | Arcophos run tasksets/mb2ev/analysis.json (regenerate with benchdb import-hob on the machine holding the harness tasksets; this row was verified against the committed site snapshot), snapshot 2026-09-04 “score 57.2, 95% CI 53.9 to 60.5, n 257 (values as published in the site snapshot of 2026-09-04)” | |
| 8 | Claude Fable 5.1 Anthropic | 47.3 | 42.2–52.2 | 257 | September 2026 | Arcophos run tasksets/mb2ev/analysis.json (regenerate with benchdb import-hob on the machine holding the harness tasksets; this row was verified against the committed site snapshot), snapshot 2026-09-04 “score 47.3, 95% CI 42.2 to 52.2, n 257 (values as published in the site snapshot of 2026-09-04)” | |
| 9 | Gemini 3.6 Google | 39.7 | 36.2–43.1 | 257 | September 2026 | Arcophos run tasksets/mb2ev/analysis.json (regenerate with benchdb import-hob on the machine holding the harness tasksets; this row was verified against the committed site snapshot), snapshot 2026-09-04 “score 39.7, 95% CI 36.2 to 43.1, n 257 (values as published in the site snapshot of 2026-09-04)” | |
| 10 | TM | Inkling Thinking Machines | 35.6 | 32.3–39.0 | 257 | September 2026 | Arcophos run tasksets/mb2ev/analysis.json (regenerate with benchdb import-hob on the machine holding the harness tasksets; this row was verified against the committed site snapshot), snapshot 2026-09-04 “score 35.6, 95% CI 32.3 to 39, n 257 (values as published in the site snapshot of 2026-09-04)” |
| 11 | Claude Sonnet 5 Anthropic | 34.6 | 31.6–37.8 | 257 | September 2026 | Arcophos run tasksets/mb2ev/analysis.json (regenerate with benchdb import-hob on the machine holding the harness tasksets; this row was verified against the committed site snapshot), snapshot 2026-09-04 “score 34.6, 95% CI 31.6 to 37.8, n 257 (values as published in the site snapshot of 2026-09-04)” | |
| 12 | GLM 5.2 Zhipu | 20.7 | 18.1–23.4 | 257 | September 2026 | Arcophos run tasksets/mb2ev/analysis.json (regenerate with benchdb import-hob on the machine holding the harness tasksets; this row was verified against the committed site snapshot), snapshot 2026-09-04 “score 20.7, 95% CI 18.1 to 23.4, n 257 (values as published in the site snapshot of 2026-09-04)” | |
| 13 | MiniMax M3 MiniMax | 18.3 | 15.7–20.9 | 257 | September 2026 | Arcophos run tasksets/mb2ev/analysis.json (regenerate with benchdb import-hob on the machine holding the harness tasksets; this row was verified against the committed site snapshot), snapshot 2026-09-04 “score 18.3, 95% CI 15.7 to 20.9, n 257 (values as published in the site snapshot of 2026-09-04)” | |
| 14 | MAI Thinking Microsoft AI | 17.5 | 15.1–20.0 | 257 | September 2026 | Arcophos run tasksets/mb2ev/analysis.json (regenerate with benchdb import-hob on the machine holding the harness tasksets; this row was verified against the committed site snapshot), snapshot 2026-09-04 “score 17.5, 95% CI 15.1 to 20, n 257 (values as published in the site snapshot of 2026-09-04)” | |
| 15 | Mistral Medium 3.5 Mistral | 9.2 | 7.4–11.1 | 257 | September 2026 | Arcophos run tasksets/mb2ev/analysis.json (regenerate with benchdb import-hob on the machine holding the harness tasksets; this row was verified against the committed site snapshot), snapshot 2026-09-04 “score 9.2, 95% CI 7.4 to 11.1, n 257 (values as published in the site snapshot of 2026-09-04)” | |
| 16 | Nemotron 3.5 Lightning NVIDIA | 4.9 | 3.6–6.2 | 257 | September 2026 | Arcophos run tasksets/mb2ev/analysis.json (regenerate with benchdb import-hob on the machine holding the harness tasksets; this row was verified against the committed site snapshot), snapshot 2026-09-04 “score 4.9, 95% CI 3.6 to 6.2, n 257 (values as published in the site snapshot of 2026-09-04)” | |
Incretin Therapeutics: Evidence Synthesis
89 released tasks · 16 models · snapshot 2026-09-04 · ranking
| # | model | score | 95% CI | tasks | snapshot | source | |
|---|---|---|---|---|---|---|---|
| 1 | Claude Fable 5.1 Anthropic | 84.9 | 80.1–89.3 | 89 | September 2026 | Arcophos run tasksets/glp1-ev/analysis.json (regenerate with benchdb import-hob on the machine holding the harness tasksets; this row was verified against the committed site snapshot), snapshot 2026-09-04 “score 84.9, 95% CI 80.1 to 89.3, n 89 (values as published in the site snapshot of 2026-09-04)” | |
| 2 | Claude Fable 5 Anthropic | 83.8 | 79.1–88.1 | 89 | September 2026 | Arcophos run tasksets/glp1-ev/analysis.json (regenerate with benchdb import-hob on the machine holding the harness tasksets; this row was verified against the committed site snapshot), snapshot 2026-09-04 “score 83.8, 95% CI 79.1 to 88.1, n 89 (values as published in the site snapshot of 2026-09-04)” | |
| 3 | Grok 4.6 xAI | 81.3 | 77.0–85.3 | 89 | September 2026 | Arcophos run tasksets/glp1-ev/analysis.json (regenerate with benchdb import-hob on the machine holding the harness tasksets; this row was verified against the committed site snapshot), snapshot 2026-09-04 “score 81.3, 95% CI 77 to 85.3, n 89 (values as published in the site snapshot of 2026-09-04)” | |
| 4 | Claude Opus 5 Anthropic | 78.3 | 72.9–83.4 | 89 | September 2026 | Arcophos run tasksets/glp1-ev/analysis.json (regenerate with benchdb import-hob on the machine holding the harness tasksets; this row was verified against the committed site snapshot), snapshot 2026-09-04 “score 78.3, 95% CI 72.9 to 83.4, n 89 (values as published in the site snapshot of 2026-09-04)” | |
| 5 | GPT-5.6 Sol (max) OpenAI max reasoning effort | 77.9 | 73.3–82.2 | 89 | September 2026 | Arcophos run tasksets/glp1-ev/analysis.json (regenerate with benchdb import-hob on the machine holding the harness tasksets; this row was verified against the committed site snapshot), snapshot 2026-09-04 “score 77.9, 95% CI 73.3 to 82.2, n 89 (values as published in the site snapshot of 2026-09-04)” | |
| 6 | Kimi K3 Moonshot AI | 77.8 | 73.0–82.4 | 89 | September 2026 | Arcophos run tasksets/glp1-ev/analysis.json (regenerate with benchdb import-hob on the machine holding the harness tasksets; this row was verified against the committed site snapshot), snapshot 2026-09-04 “score 77.8, 95% CI 73 to 82.4, n 89 (values as published in the site snapshot of 2026-09-04)” | |
| 7 | GPT-5.6 Sol (high) OpenAI high reasoning effort | 77.7 | 73.1–82.0 | 89 | September 2026 | Arcophos run tasksets/glp1-ev/analysis.json (regenerate with benchdb import-hob on the machine holding the harness tasksets; this row was verified against the committed site snapshot), snapshot 2026-09-04 “score 77.7, 95% CI 73.1 to 82, n 89 (values as published in the site snapshot of 2026-09-04)” | |
| 8 | Muse Spark Meta | 68.4 | 62.8–73.9 | 89 | September 2026 | Arcophos run tasksets/glp1-ev/analysis.json (regenerate with benchdb import-hob on the machine holding the harness tasksets; this row was verified against the committed site snapshot), snapshot 2026-09-04 “score 68.4, 95% CI 62.8 to 73.9, n 89 (values as published in the site snapshot of 2026-09-04)” | |
| 9 | Gemini 3.6 Google | 57.9 | 51.2–64.3 | 89 | September 2026 | Arcophos run tasksets/glp1-ev/analysis.json (regenerate with benchdb import-hob on the machine holding the harness tasksets; this row was verified against the committed site snapshot), snapshot 2026-09-04 “score 57.9, 95% CI 51.2 to 64.3, n 89 (values as published in the site snapshot of 2026-09-04)” | |
| 10 | TM | Inkling Thinking Machines | 53.9 | 47.7–60.1 | 89 | September 2026 | Arcophos run tasksets/glp1-ev/analysis.json (regenerate with benchdb import-hob on the machine holding the harness tasksets; this row was verified against the committed site snapshot), snapshot 2026-09-04 “score 53.9, 95% CI 47.7 to 60.1, n 89 (values as published in the site snapshot of 2026-09-04)” |
| 11 | Claude Sonnet 5 Anthropic | 50.1 | 44.3–56.1 | 89 | September 2026 | Arcophos run tasksets/glp1-ev/analysis.json (regenerate with benchdb import-hob on the machine holding the harness tasksets; this row was verified against the committed site snapshot), snapshot 2026-09-04 “score 50.1, 95% CI 44.3 to 56.1, n 89 (values as published in the site snapshot of 2026-09-04)” | |
| 12 | MiniMax M3 MiniMax | 35.5 | 29.6–41.6 | 89 | September 2026 | Arcophos run tasksets/glp1-ev/analysis.json (regenerate with benchdb import-hob on the machine holding the harness tasksets; this row was verified against the committed site snapshot), snapshot 2026-09-04 “score 35.5, 95% CI 29.6 to 41.6, n 89 (values as published in the site snapshot of 2026-09-04)” | |
| 13 | MAI Thinking Microsoft AI | 33.0 | 26.8–39.3 | 89 | September 2026 | Arcophos run tasksets/glp1-ev/analysis.json (regenerate with benchdb import-hob on the machine holding the harness tasksets; this row was verified against the committed site snapshot), snapshot 2026-09-04 “score 33.0, 95% CI 26.8 to 39.3, n 89 (values as published in the site snapshot of 2026-09-04)” | |
| 14 | GLM 5.2 Zhipu | 29.2 | 23.9–34.9 | 89 | September 2026 | Arcophos run tasksets/glp1-ev/analysis.json (regenerate with benchdb import-hob on the machine holding the harness tasksets; this row was verified against the committed site snapshot), snapshot 2026-09-04 “score 29.2, 95% CI 23.9 to 34.9, n 89 (values as published in the site snapshot of 2026-09-04)” | |
| 15 | Mistral Medium 3.5 Mistral | 14.4 | 10.6–18.5 | 89 | September 2026 | Arcophos run tasksets/glp1-ev/analysis.json (regenerate with benchdb import-hob on the machine holding the harness tasksets; this row was verified against the committed site snapshot), snapshot 2026-09-04 “score 14.4, 95% CI 10.6 to 18.5, n 89 (values as published in the site snapshot of 2026-09-04)” | |
| 16 | Nemotron 3.5 Lightning NVIDIA | 7.8 | 4.7–11.3 | 89 | September 2026 | Arcophos run tasksets/glp1-ev/analysis.json (regenerate with benchdb import-hob on the machine holding the harness tasksets; this row was verified against the committed site snapshot), snapshot 2026-09-04 “score 7.8, 95% CI 4.7 to 11.3, n 89 (values as published in the site snapshot of 2026-09-04)” | |
Protocol in brief
Tasks are authored one at a time by one of five frontier model families, audited for material error and citation fidelity by families that did not write them, and removed only when a second family independently confirms a material verdict. Candidate models answer closed book, one answer per task, tools disabled. Each answer is graded blind in three passes spread across three grader families, with the task's authoring family excluded from its panel; any split criterion escalates to a fourth, uninvolved family. Scores are consensus rubric credit on a 0 to 100 scale. Confidence intervals are nonparametric bootstraps over tasks, 10,000 resamples, fixed seed. Scores are comparable only within an identical task set, grading panel, and pass count. The full protocol, the measured verification outcomes, and the safety scoring rule are on the methodology page.