Health Optimization Bench

Incretin Therapeutics: Evidence Synthesis

micro bench v1 · 89 released tasks · updated September 10, 2026

The suite covers the incretin therapeutic class: pivotal trial results, cardiovascular and renal outcome evidence, label facts, and guideline positions for GLP-1 receptor agonists and dual agonists, anchored to readouts from 2024 through 2026. Tasks demand specific quantities and study design details, not summaries, and each carries one penalized misstatement criterion scored without compensation.

Ranking

Sources
#modelscore95% CIsafety failsunanimity
1Anthropic logoClaude Fable 5.1 Anthropic [src]84.980.1–89.311/8992%
2Anthropic logoClaude Fable 5 Anthropic [src]83.879.1–88.110/8991%
3OpenAI logoGPT-6 Astra (max) OpenAI [src]82.277.7–86.415/8987%
4SpaceX AI logoGrok 4.6 SpaceX AI [src]81.377.0–85.312/8988%
5Anthropic logoClaude Opus 5 Anthropic [src]78.372.9–83.414/8989%
6OpenAI logoGPT-5.6 Sol (max) OpenAI [src]77.973.3–82.213/8989%
7Moonshot AI logoKimi K3 Moonshot AI [src]77.873.0–82.414/8989%
8OpenAI logoGPT-5.6 Sol (high) OpenAI [src]77.773.1–82.013/8989%
9Meta logoMuse Spark Meta [src]68.462.8–73.919/8986%
10Google logoGemini 3.8 Flash Google [src]67.861.7–73.623/8986%
11Google logoGemini 3.6 Google [src]57.951.2–64.319/8981%
12TMInkling Thinking Machines [src]53.947.7–60.123/8982%
13Anthropic logoClaude Sonnet 5 Anthropic [src]50.144.3–56.120/8982%
14MiniMax logoMiniMax M3 MiniMax [src]35.529.6–41.634/8983%
15Microsoft AI logoMAI Thinking Microsoft AI [src]33.026.8–39.327/8985%
16Zhipu logoGLM 5.2 Zhipu [src]29.223.9–34.933/8983%
17Mistral logoMistral Medium 3.5 Mistral [src]14.410.6–18.547/8989%
18NVIDIA logoNemotron 3.5 Lightning NVIDIA [src]7.84.7–11.348/8994%

Claude Fable 5.1 declined 1 of 89 tasks: the vendor's safeguard classifier returned a refusal notice instead of an answer. Declined tasks are graded like any other completion and earn no credit; over the tasks it answered, its mean is 85.9.

Consensus rubric credit, one answer per task, tools disabled. Safety fails count tasks with a consensus-met penalty criterion. Unanimity is the share of criterion judgments on which all three grader families agreed before adjudication. All scores are from Arcophos's own harness run, snapshot 2026-09-10; each row's record is on the sources page.

Task difficulty

hardest, cross-model mean

  • placebo weight trajectories14%
  • redefine1 projection audit16%
  • pivotal dc ae extract22%
  • society guideline grades23%
  • retatrutide phase2 status26%

easiest, cross-model mean

  • flow early termination92%
  • incretin label indications89%
  • select mace nnt a288%
  • dcae pivotal obesity trials87%
  • select mace nnt a485%

Construction and grading follow the bench-wide protocol: five-family authoring, two-family confirmation before any removal, blind author-excluded panel grading with escalation, and a confidential holdout. Details are on the methodology page.