Incretin Therapeutics: Evidence Synthesis
micro bench v1 · 89 released tasks · updated September 10, 2026
The suite covers the incretin therapeutic class: pivotal trial results, cardiovascular and renal outcome evidence, label facts, and guideline positions for GLP-1 receptor agonists and dual agonists, anchored to readouts from 2024 through 2026. Tasks demand specific quantities and study design details, not summaries, and each carries one penalized misstatement criterion scored without compensation.
Ranking
Sources| # | model | score | 95% CI | safety fails | unanimity | |
|---|---|---|---|---|---|---|
| 1 | Claude Fable 5.1 Anthropic [src] | 84.9 | 80.1–89.3 | 11/89 | 92% | |
| 2 | Claude Fable 5 Anthropic [src] | 83.8 | 79.1–88.1 | 10/89 | 91% | |
| 3 | GPT-6 Astra (max) OpenAI [src] | 82.2 | 77.7–86.4 | 15/89 | 87% | |
| 4 | Grok 4.6 SpaceX AI [src] | 81.3 | 77.0–85.3 | 12/89 | 88% | |
| 5 | Claude Opus 5 Anthropic [src] | 78.3 | 72.9–83.4 | 14/89 | 89% | |
| 6 | GPT-5.6 Sol (max) OpenAI [src] | 77.9 | 73.3–82.2 | 13/89 | 89% | |
| 7 | Kimi K3 Moonshot AI [src] | 77.8 | 73.0–82.4 | 14/89 | 89% | |
| 8 | GPT-5.6 Sol (high) OpenAI [src] | 77.7 | 73.1–82.0 | 13/89 | 89% | |
| 9 | Muse Spark Meta [src] | 68.4 | 62.8–73.9 | 19/89 | 86% | |
| 10 | Gemini 3.8 Flash Google [src] | 67.8 | 61.7–73.6 | 23/89 | 86% | |
| 11 | Gemini 3.6 Google [src] | 57.9 | 51.2–64.3 | 19/89 | 81% | |
| 12 | TM | Inkling Thinking Machines [src] | 53.9 | 47.7–60.1 | 23/89 | 82% |
| 13 | Claude Sonnet 5 Anthropic [src] | 50.1 | 44.3–56.1 | 20/89 | 82% | |
| 14 | MiniMax M3 MiniMax [src] | 35.5 | 29.6–41.6 | 34/89 | 83% | |
| 15 | MAI Thinking Microsoft AI [src] | 33.0 | 26.8–39.3 | 27/89 | 85% | |
| 16 | GLM 5.2 Zhipu [src] | 29.2 | 23.9–34.9 | 33/89 | 83% | |
| 17 | Mistral Medium 3.5 Mistral [src] | 14.4 | 10.6–18.5 | 47/89 | 89% | |
| 18 | Nemotron 3.5 Lightning NVIDIA [src] | 7.8 | 4.7–11.3 | 48/89 | 94% | |
Claude Fable 5.1 declined 1 of 89 tasks: the vendor's safeguard classifier returned a refusal notice instead of an answer. Declined tasks are graded like any other completion and earn no credit; over the tasks it answered, its mean is 85.9.
Consensus rubric credit, one answer per task, tools disabled. Safety fails count tasks with a consensus-met penalty criterion. Unanimity is the share of criterion judgments on which all three grader families agreed before adjudication. All scores are from Arcophos's own harness run, snapshot 2026-09-10; each row's record is on the sources page.
Task difficulty
hardest, cross-model mean
- placebo weight trajectories14%
- redefine1 projection audit16%
- pivotal dc ae extract22%
- society guideline grades23%
- retatrutide phase2 status26%
easiest, cross-model mean
- flow early termination92%
- incretin label indications89%
- select mace nnt a288%
- dcae pivotal obesity trials87%
- select mace nnt a485%
Construction and grading follow the bench-wide protocol: five-family authoring, two-family confirmation before any removal, blind author-excluded panel grading with escalation, and a confidential holdout. Details are on the methodology page.