hive-bench

Which coding agents hold up end to end?

Hive is a multi-agent orchestrator that carries software tasks through planning, implementation, pull requests, review, and fixes. hive-bench measures how coding-agent configurations perform across that entire automated workflow, scoring each final diff against the merged reference.

What this benchmark measures

The corpus contains 6 completed tasks from the Hive repository. For each task, the source is rewound to the original base commit and the candidate receives the frozen task inputs, but never the reference patch. A candidate may use one model for every stage or split planning and execution between models. Every candidate then runs real Hive in an isolated runner: plan → execute → open PR → production review and fix loop.

Planning invokes Compound Engineering's /ce-plan; execution uses Hive's normal plan-driven development stage. Candidate-specific production reviewers can then revise the local PR before the independent scoring judges see the final diff. The exact stage owners and reviewer panels are listed in the methodology.

The board covers all 90 generation cells: 15 candidates × 6 tasks. The original 36 cells have 1 independent score sample per judge; the 54 cells across the five active later campaigns have 3. This corpus has no curated objective test gates, so close rankings are directional rather than definitive.

How the scores should be read

  1. Two independent rulers: Fable 5 and GPT-5.6 Sol score each final diff from 0–10. Fable ran with reasoning enabled: the first three campaign records pin xhigh, while the DeepSeek and GLM 5.3 Flash (0x Alpha) campaigns record it as unspecified. Sol ran at xhigh for the original campaign and ultra for all five active later campaigns.
  2. Keep both judge scores visible: the leaderboard uses their arithmetic mean only for a compact ordering because their calibrations differ. Every row retains separate Fable and Sol columns, and every matrix cell shows Fable / Sol in that order.
  3. Deliberation is not a third judge: after the independent scores were recorded, Fable and Sol each saw the other judge's anonymous verdict and reasoning, argued the strongest evidence that its own initial score was wrong, and then held or revised that score. After discussion is the arithmetic mean of those two final scores. It is the default sort and a diagnostic layer, not a replacement for the independent scores, because the exchange can add anchoring or convergence pressure.
  4. Show family conflicts: when the judge and candidate share a model family, that row is marked SF . The score remains visible, but it is weaker evidence than a family-disjoint judgment.
  5. Keep efficiency descriptive: wall time, normalized generation tokens, and API-equivalent cost are shown where the underlying runner recorded them. Missing values are reported as unknown, never inferred.

Read the full methodology and limitations before treating small score differences as meaningful.

What came out

GLM 5.3 Flash (0x Alpha) via Pi high and Pi max have independent paired means of 5.05 and 5.036. After the separate discussion diagnostic, they finish at 5.692 and 5.483. Both judges are family-disjoint from GLM 5.3 Flash (0x Alpha), and both rows use three samples per judge on the same six tasks.

Neither GLM 5.3 Flash (0x Alpha) row changes the after-discussion leader: the production-like Sol plan → Sol execute → Sol + Grok review setup remains first at 7.525. The all-Pro and Pro-plan/Flash-execute/Pro-review DeepSeek configurations finish at 5.892 and 5.783. Both judgments are family-disjoint, so their evidence carries no same-family warning.

The sharper DeepSeek result is judge calibration, not rank. On the three-sample independent scores, Fable and Sol differ by 4.483 points for all-Pro and 2.994 for the mixed workflow. The fresh one-shot deliberation brings the final judge gap for both rows to about 1.9 points without erasing it. Across all 90 cells, the judges' mean absolute spread still fell from 2.082 to 1.251. That suggests the anonymous exchange surfaced real misses, but the same movement may also reflect anchoring pressure. Deliberation therefore remains a diagnostic rather than a cleaner replacement score.

DeepSeek is the cheapest fully attributable provider path measured so far: all-Pro averages $0.55 across five priced cells, and Pro/Flash/Pro averages $0.6 across all six. Their normalized token volume is 50.754M and 34.774M per recorded task respectively—because cache reads are counted as part of total generation volume. Missing all-Pro telemetry remains unknown rather than inferred.

The Pi max campaign averages $6.94 per task across all six priced cells and 33.343M normalized tokens per task. Its wall-time mean is 95.5 minutes across 5 recorded cells. The earlier Pi high campaign has complete token telemetry but no serialized API-equivalent price, so its cost remains unknown.

These are scoped findings, not a universal model ranking. 6 tasks from one Ruby/CLI project cannot establish broad superiority. The original rows have 1 sample per judge, the five active later campaigns have 3, and the discussion layer is a single exchange; the results show what happened when these exact configurations drove this exact end-to-end workflow.

One task, end to end

fix-tmux asked candidates to repair launcher scripts and Claude-ready detection, based on the task later merged as Hive PR #623. The mixed Opus → Codex candidate scored 9.0 / 8.7 (Fable / Sol), the best Sol score for that task. GPT-5.6 Sol scored 9.0 / 8.5. The result and exact candidate patch behind every score are linked from the full board below.

Generated 2026-08-27T22:21:35Z · 90/90 cells complete

Audit correction, 26 August 2026: the historical OpenCode GLM 5.3 Flash (0x Alpha) row is withdrawn from ranking and active coverage because its target exposed held-out history and its profile lacked Bash capability parity. Its invalidated patch artifacts remain available for audit.

Leaderboard

Ranked by the presentation-only mean of Fable and Sol. The two judge columns remain the primary evidence because their calibrations differ. SF means that judge shares a model family with the candidate. Bold scores are the independent means and remain the primary evidence; the table opens ordered by the paired discussion-final mean. Efficiency averages use only recorded cells.

Sol plan → Sol execute → Sol + Grok review Sol xhigh plan · Sol high execute · Sol xhigh + Grok xhigh review 3 samples/judge · scores + public diffs 6.722 5.933 SF 6.328 7.525 Fable 7.917 · Sol 7.133 6/6 paired 151.0 min 6/6 timed $42.94 known* 63.156M known*
Sol plan → Grok execute → Sol review Sol xhigh plan · Grok xhigh execute · Sol xhigh review 3 samples/judge · scores + public diffs 7.778 4.845 SF 6.312 6.908 Fable 7.383 · Sol 6.433 6/6 paired 132.1 min 5/6 timed $21.35 known* 29.123M known*
Sol plan → Terra execute → Sol review Sol xhigh plan · Terra xhigh execute · Sol xhigh review 3 samples/judge · scores + public diffs 7.055 5.133 SF 6.094 6.467 Fable 7.45 · Sol 5.483 6/6 paired 166 min 4/6 timed $26.12 5/6 priced 34.807M 5/6 measured
GPT-5.6 Sol xhigh xhigh (pinned) 6.917 6.183 SF 6.55 6.058 Fable 6.25 · Sol 5.867 6/6 paired 146.9 min 4/6 timed $47.66 6/6 priced 74.86M 6/6 measured
Fable plan → Grok execute → Sol review Fable high plan · Grok xhigh execute · Sol xhigh review 3 samples/judge · scores + public diffs 6.194 SF 4.667 SF 5.431 6.05 Fable 6.583 · Sol 5.517 6/6 paired 98.4 min 5/6 timed $7.10 known* 9.24M known*
DeepSeek V4 Pro 0813 xhigh DeepSeek V4 Pro 0813 xhigh throughout 3 samples/judge · scores + public diffs 7.722 3.239 5.481 5.892 Fable 6.833 · Sol 4.95 6/6 paired 120.0 min 5/6 timed $0.55 5/6 priced 50.754M 5/6 measured
DeepSeek V4 Pro plan → V4 Flash execute → V4 Pro review DeepSeek V4 Pro 0813 xhigh plan/review · V4 Flash 0731 xhigh execute 3 samples/judge · scores + public diffs 6.333 3.339 4.836 5.783 Fable 6.75 · Sol 4.817 6/6 paired 118.8 min 6/6 timed $0.6 6/6 priced 34.774M 6/6 measured
Sol plan → Terra execute → Grok review Sol xhigh plan · Terra xhigh execute · Grok xhigh review 3 samples/judge · scores + public diffs 5.861 4.567 SF 5.214 5.717 Fable 6.517 · Sol 4.917 6/6 paired 61.1 min 6/6 timed $18.63 known* 26.229M known*
GLM 5.3 Flash (0x Alpha) via Pi high GLM 5.3 Flash (0x Alpha) high throughout via Pi/OpenRouter 3 samples/judge · scores + public diffs 6.306 3.795 5.05 5.692 Fable 6.333 · Sol 5.05 6/6 paired 105.7 min 4/6 timed unknown 19.838M 6/6 measured
GLM 5.3 Flash (0x Alpha) via Pi max GLM 5.3 Flash (0x Alpha) max throughout via Pi/OpenRouter 3 samples/judge · scores + public diffs 6.5 3.572 5.036 5.483 Fable 6.5 · Sol 4.467 6/6 paired 95.5 min 5/6 timed $6.9381 6/6 priced 33.343M 6/6 measured
Opus plan → Codex 5.5 xhigh Opus default · Codex xhigh 6.167 SF 4.783 SF 5.475 5.267 Fable 5.583 · Sol 4.95 6/6 paired 60.8 min 5/6 timed $19.26 6/6 priced 27.67M 6/6 measured
Opus 4.8 Claude CLI default 6.083 SF 4.25 5.167 4.833 Fable 5.5 · Sol 4.167 6/6 paired 63.7 min 6/6 timed $35.74 6/6 priced 59.82M 6/6 measured
Grok 4.5 xhigh xhigh (pinned) 6.333 4.217 5.275 4.383 Fable 4.883 · Sol 3.883 6/6 paired 27.3 min 5/6 timed unknown unknown
GLM 5.2 provider default 4.833 3.617 4.225 3.817 Fable 4.25 · Sol 3.383 6/6 paired 114.4 min 6/6 timed $5.65 6/6 priced 21.793M 6/6 measured
Codex 5.5 xhigh xhigh (pinned) 4.667 3.867 SF 4.267 3.725 Fable 4.0 · Sol 3.45 6/6 paired 52 min 6/6 timed $17.17 6/6 priced 24.01M 6/6 measured

Discussion is a one-shot diagnostic with campaign-specific round one: the six original rows reused their exact published independent verdicts and recovered rationales; the nine rows across the five active later three-seed campaigns received fresh round-one re-grades. In all six active campaigns, each judge then saw the other anonymous verdict, challenged its own view, and held or revised it. All fifteen rows ran this step across 90 paired cells. After discussion is the default table sort using the paired final mean; the independent scores remain visible, sortable, and the primary evidence.

Combined is (Fable + Sol) / 2 and is used only to provide one compact ordering. Token totals include fresh input, output, cache reads, and cache writes. Known* values include only providers with preserved telemetry. Grok usage and cost are unavailable, so those values are lower-bound subtotals rather than complete workflow totals. See the token accounting and pricing methodology for the normalization rules, mixed-model attribution, and known telemetry gaps.

Per-task efficiency by model/configuration

Each cell shows API-equivalent generation cost, total normalized generation tokens, the input/output/cache split, and recorded end-to-end wall time. Cache is shown as reads plus creation/writes. An unknown value means the runner did not preserve usable evidence; it does not mean zero.

Model / candidate add-i-key web-install install fix-tmux fix-review daemon
GPT-5.6 Sol xhigh $20.99
28.26M tokens
95.3 min
$88.23
142.43M tokens
227.4 min
$105.24
167.52M tokens
224.2 min
$20.36
32.37M tokens
40.7 min
$18.94
29.98M tokens
time not recorded
$32.22
48.58M tokens
time not recorded
Sol plan → Sol execute → Sol + Grok review $29.10 known*
41.088M known* tokens
124.0 min
$62.98 known*
96.241M known* tokens
178.7 min
$40.76 known*
58.575M known* tokens
145.1 min
$25.19 known*
38.298M known* tokens
94.1 min
$39.46 known*
56.678M known* tokens
177.3 min
$60.16 known*
88.057M known* tokens
186.5 min
Sol plan → Grok execute → Sol review $24.04 known*
33.998M known* tokens
115.5 min
$22.26 known*
29.397M known* tokens
94.9 min
$24.25 known*
30.182M known* tokens
173.5 min
$5.58 known*
6.697M known* tokens
63.5 min
$16.15 known*
22.795M known* tokens
time not recorded
$35.83 known*
51.67M known* tokens
213.1 min
Sol plan → Terra execute → Sol review $17.16
21.823M tokens
91.4 min
$21.70
30.491M tokens
197.7 min
$31.51
38.249M tokens
179.5 min
$31.28
43.864M tokens
time not recorded
cost unknown
tokens unknown
time not recorded
$28.96
39.608M tokens
195.2 min
DeepSeek V4 Pro 0813 xhigh $0.13
3.905M tokens
36.2 min
$0.96
103.633M tokens
146.2 min
$0.63
67.776M tokens
110.4 min
cost unknown
tokens unknown
time not recorded
$0.42
27.539M tokens
114.5 min
$0.63
50.918M tokens
192.9 min
Opus plan → Codex 5.5 xhigh $25.75
39.89M tokens
time not recorded
$16.95
22.79M tokens
50.9 min
$31.02
47.85M tokens
75 min
$11.92
14.69M tokens
73 min
$12.95
18.12M tokens
46.9 min
$17.00
22.7M tokens
58.2 min
Fable plan → Grok execute → Sol review $11.18 known*
14.114M known* tokens
92.1 min
$3.45 known*
5.927M known* tokens
150.7 min
$3.38 known*
1.335M known* tokens
time not recorded
$2.43 known*
2.359M known* tokens
25.5 min
$20.26 known*
28.061M known* tokens
87.2 min
$1.93 known*
3.647M known* tokens
136.2 min
Grok 4.5 xhigh cost unknown
tokens unknown
30.1 min
cost unknown
tokens unknown
32.1 min
cost unknown
tokens unknown
27.3 min
cost unknown
tokens unknown
time not recorded
cost unknown
tokens unknown
23.2 min
cost unknown
tokens unknown
23.9 min
Sol plan → Terra execute → Grok review $12.53 known*
17.19M known* tokens
43.7 min
$31.78 known*
48.013M known* tokens
75.1 min
$20.61 known*
28.593M known* tokens
54.3 min
$12.30 known*
17.293M known* tokens
93.4 min
$15.16 known*
20.982M known* tokens
50.9 min
$19.43 known*
25.302M known* tokens
49.5 min
Opus 4.8 $24.29
39.03M tokens
59.5 min
$33.08
54.71M tokens
44 min
$74.33
125.82M tokens
98.7 min
$11.35
17.23M tokens
39.9 min
$27.75
47.07M tokens
55.1 min
$43.64
75.08M tokens
85 min
GLM 5.3 Flash (0x Alpha) via Pi high cost unknown
5.392M tokens
51.2 min
cost unknown
52.999M tokens
time not recorded
cost unknown
29.776M tokens
160.6 min
cost unknown
3.348M tokens
99.1 min
cost unknown
9.719M tokens
111.8 min
cost unknown
17.794M tokens
time not recorded
GLM 5.3 Flash (0x Alpha) via Pi max $3.39
16.316M tokens
44.6 min
$7.62
38.153M tokens
4.5 min
$15.48
71.713M tokens
233.8 min
$1.99
8.619M tokens
65.3 min
$6.68
33.903M tokens
time not recorded
$6.47
31.354M tokens
129.2 min
DeepSeek V4 Pro plan → V4 Flash execute → V4 Pro review $0.16
4.516M tokens
41.9 min
$1.15
74.559M tokens
179.9 min
$0.71
43.399M tokens
113.0 min
$0.21
7.448M tokens
96.6 min
$0.63
38.425M tokens
167.3 min
$0.73
40.297M tokens
114.1 min
Codex 5.5 xhigh $14.55
19.17M tokens
51 min
$20.95
29.53M tokens
52.9 min
$20.21
28.59M tokens
67 min
$7.81
11.42M tokens
26 min
$22.62
31.1M tokens
75 min
$16.85
24.25M tokens
40.4 min
GLM 5.2 $4.27
17.548M tokens
114.1 min
$5.56
24.145M tokens
173.5 min
$6.80
27.155M tokens
124.4 min
$2.47
10.484M tokens
91.4 min
$3.13
4.316M tokens
29.7 min
$11.69
47.108M tokens
153.4 min

Known* values include only providers with preserved telemetry. Depending on the workflow, that can include Fable planning, Sol stages, or Sol planning plus Terra execution; Grok telemetry is unavailable. DeepSeek's Pi/OpenRouter rows have complete provider scope wherever a value is present. GLM 5.3 Flash (0x Alpha) via Pi high preserves complete token telemetry but has no published API-equivalent price; Pi max preserves complete token and cost telemetry for all six cells. Use each token info button for its exact provider scope and token split.

Full board (per task)

Every cell shows Fable 5 / GPT-5.6 Sol. The original rows use one score per judge; the five active later campaigns use three and display their means. Later-campaign cells also show the complete, separate discussion-final pair. Click a score for its machine-readable record. Exact diffs are linked where the public raw-evidence bundle contains them. No task in any campaign has an objective gate.

Candidate add-i-key web-install install fix-tmux fix-review daemon
GPT-5.6 Sol xhigh 7.0 / 7.0
discussion final · Fable 7.0 · Sol 7.0
diff
7.0 / 5.5
discussion final · Fable 6.5 · Sol 5.5
diff
5.5 / 3.5
discussion final · Fable 4.0 · Sol 3.5
diff
9.0 / 8.5
discussion final · Fable 8.5 · Sol 8.5
diff
6.0 / 7.1
discussion final · Fable 6.0 · Sol 6.2
diff
7.0 / 5.5
discussion final · Fable 5.5 · Sol 4.5
diff
Sol plan → Sol execute → Sol + Grok review 5.167 / 4.767
discussion final · Fable 8.5 · Sol 8.5
diff · 3 samples/judge
7.0 / 5.833
discussion final · Fable 7.0 · Sol 6.5
diff · 3 samples/judge
5.667 / 4.733
discussion final · Fable 6.0 · Sol 4.0
diff · 3 samples/judge
9.0 / 7.833
discussion final · Fable 9.0 · Sol 9.2
diff · 3 samples/judge
7.833 / 7.167
discussion final · Fable 8.5 · Sol 7.1
diff · 3 samples/judge
5.667 / 5.267
discussion final · Fable 8.5 · Sol 7.5
diff · 3 samples/judge
Sol plan → Grok execute → Sol review 9.0 / 6.167
discussion final · Fable 8.8 · Sol 8.5
diff · 3 samples/judge
8.333 / 5.167
discussion final · Fable 7 · Sol 5.8
diff · 3 samples/judge
7.833 / 4.067
discussion final · Fable 6.5 · Sol 5.3
diff · 3 samples/judge
9.167 / 5.833
discussion final · Fable 8.5 · Sol 8.5
diff · 3 samples/judge
4.0 / 4.0
discussion final · Fable 5.5 · Sol 4
diff · 3 samples/judge
8.333 / 3.833
discussion final · Fable 8 · Sol 6.5
diff · 3 samples/judge
Sol plan → Terra execute → Sol review 8.833 / 5.267
discussion final · Fable 9.2 · Sol 8.9
diff · 3 samples/judge
3.667 / 2.833
discussion final · Fable 4.5 · Sol 3.5
diff · 3 samples/judge
8.0 / 3.167
discussion final · Fable 6 · Sol 4.2
diff · 3 samples/judge
8.333 / 6.9
discussion final · Fable 9 · Sol 6
diff · 3 samples/judge
8.167 / 7.233
discussion final · Fable 7.5 · Sol 6.8
diff · 3 samples/judge
5.333 / 5.4
discussion final · Fable 8.5 · Sol 3.5
diff · 3 samples/judge
DeepSeek V4 Pro 0813 xhigh 9.0 / 1.0
discussion final · Fable 9.0 · Sol 9.2
diff · 3 samples/judge
8.167 / 2.933
discussion final · Fable 7.5 · Sol 3.5
diff · 3 samples/judge
7.167 / 2.167
discussion final · Fable 6.0 · Sol 3.0
diff · 3 samples/judge
5.5 / 5.833
discussion final · Fable 5.0 · Sol 5.0
diff · 3 samples/judge
8.0 / 3.5
discussion final · Fable 6.0 · Sol 4.0
diff · 3 samples/judge
8.5 / 4.0
discussion final · Fable 7.5 · Sol 5.0
diff · 3 samples/judge
Opus plan → Codex 5.5 xhigh 6.0 / 5.0
discussion final · Fable 5.5 · Sol 5.0
diff
5.0 / 3.0
discussion final · Fable 4.0 · Sol 4.0
diff
6.5 / 4.5
discussion final · Fable 6.0 · Sol 4.5
diff
9.0 / 8.7
discussion final · Fable 9.0 · Sol 8.7
diff
4.0 / 3.0
discussion final · Fable 3.5 · Sol 3.0
diff
6.5 / 4.5
discussion final · Fable 5.5 · Sol 4.5
diff
Fable plan → Grok execute → Sol review 8.5 / 6.0
discussion final · Fable 8 · Sol 8
diff · 3 samples/judge
5.167 / 3.167
discussion final · Fable 4.5 · Sol 3.5
diff · 3 samples/judge
5.5 / 2.833
discussion final · Fable 4 · Sol 2.5
diff · 3 samples/judge
5.833 / 7.833
discussion final · Fable 9 · Sol 7.5
diff · 3 samples/judge
5.833 / 4.0
discussion final · Fable 7.5 · Sol 6
diff · 3 samples/judge
6.333 / 4.167
discussion final · Fable 6.5 · Sol 5.6
diff · 3 samples/judge
Grok 4.5 xhigh 5.5 / 5.1
discussion final · Fable 5.0 · Sol 4.8
diff
6.0 / 4.0
discussion final · Fable 5.0 · Sol 3.5
diff
6.5 / 2.0
discussion final · Fable 3.0 · Sol 2.0
diff
7.5 / 6.2
discussion final · Fable 6.8 · Sol 5.5
diff
5.0 / 3.0
discussion final · Fable 3.5 · Sol 3.0
diff
7.5 / 5.0
discussion final · Fable 6.0 · Sol 4.5
diff
Sol plan → Terra execute → Grok review 6.833 / 7.9
discussion final · Fable 8.6 · Sol 8.5
diff · 3 samples/judge
4.667 / 2.833
discussion final · Fable 5.0 · Sol 3.0
diff · 3 samples/judge
5.5 / 2.567
discussion final · Fable 6.0 · Sol 2.5
diff · 3 samples/judge
8.5 / 6.267
discussion final · Fable 7.5 · Sol 7.0
diff · 3 samples/judge
5.667 / 5.0
discussion final · Fable 7.5 · Sol 5.0
diff · 3 samples/judge
4.0 / 2.833
discussion final · Fable 4.5 · Sol 3.5
diff · 3 samples/judge
Opus 4.8 6.0 / 5.0
discussion final · Fable 5.5 · Sol 5.0
diff
4.5 / 2.5
discussion final · Fable 4.0 · Sol 2.5
diff
6.0 / 2.5
discussion final · Fable 6.0 · Sol 2.5
diff
8.0 / 7.5
discussion final · Fable 7.5 · Sol 7.5
diff
5.0 / 3.0
discussion final · Fable 4.0 · Sol 2.5
diff
7.0 / 5.0
discussion final · Fable 6.0 · Sol 5.0
diff
GLM 5.3 Flash (0x Alpha) via Pi high 6.833 / 6.0
discussion final · Fable 8.0 · Sol 7.5
diff · 3 samples/judge
6.0 / 3.767
discussion final · Fable 7.5 · Sol 6.5
diff · 3 samples/judge
6.833 / 2.167
discussion final · Fable 4.0 · Sol 2.8
diff · 3 samples/judge
8.167 / 5.5
discussion final · Fable 7.5 · Sol 6.5
diff · 3 samples/judge
3.167 / 2.333
discussion final · Fable 4.5 · Sol 2.0
diff · 3 samples/judge
6.833 / 3.0
discussion final · Fable 6.5 · Sol 5.0
diff · 3 samples/judge
GLM 5.3 Flash (0x Alpha) via Pi max 8.667 / 4.667
discussion final · Fable 8.0 · Sol 7.5
diff · 3 samples/judge
4.333 / 3.1
discussion final · Fable 5.0 · Sol 2.5
diff · 3 samples/judge
6.0 / 1.667
discussion final · Fable 4.0 · Sol 2.5
diff · 3 samples/judge
8.667 / 5.167
discussion final · Fable 9.0 · Sol 8.8
diff · 3 samples/judge
4.0 / 2.5
discussion final · Fable 5.0 · Sol 2.5
diff · 3 samples/judge
7.333 / 4.333
discussion final · Fable 8.0 · Sol 3.0
diff · 3 samples/judge
DeepSeek V4 Pro plan → V4 Flash execute → V4 Pro review 9.0 / 0.833
discussion final · Fable 9.0 · Sol 9.4
diff · 3 samples/judge
3.833 / 2.933
discussion final · Fable 7.0 · Sol 3.0
diff · 3 samples/judge
5.667 / 2.5
discussion final · Fable 5.0 · Sol 2.0
diff · 3 samples/judge
8.5 / 5.5
discussion final · Fable 8.0 · Sol 6.5
diff · 3 samples/judge
3.5 / 3.0
discussion final · Fable 4.5 · Sol 2.5
diff · 3 samples/judge
7.5 / 5.267
discussion final · Fable 7.0 · Sol 5.5
diff · 3 samples/judge
Codex 5.5 xhigh 5.5 / 4.5
discussion final · Fable 5.0 · Sol 4.5
diff
3.0 / 3.5
discussion final · Fable 3.0 · Sol 2.5
diff
5.0 / 2.5
discussion final · Fable 3.5 · Sol 2.5
diff
5.0 / 6.5
discussion final · Fable 5.0 · Sol 5.5
diff
3.5 / 2.0
discussion final · Fable 3.0 · Sol 1.5
diff
6.0 / 4.2
discussion final · Fable 4.5 · Sol 4.2
diff
GLM 5.2 6.5 / 5.5
discussion final · Fable 6.0 · Sol 5.0
diff
3.0 / 3.0
discussion final · Fable 3.0 · Sol 3.0
diff
6.0 / 2.0
discussion final · Fable 4.0 · Sol 2.5
diff
7.0 / 7.2
discussion final · Fable 7.5 · Sol 6.8
diff
1.5 / 1.0
discussion final · Fable 1.0 · Sol 0.5
diff
5.0 / 3.0
discussion final · Fable 4.0 · Sol 2.5
diff

The tasks

Six real Hive tasks, each completed and merged. The original plans all came from Claude-family agents, which is disclosed corpus provenance and a possible framing bias.

  • add-i-key feature — Add an i key and legend entry that open a full task-information panel in the Hive TUI. reference PR · original plan: claude · original implementer: claude-opus-4-7
  • web-install feature — Add a first-class local, non-Docker install and run mode for the Hive web UI. reference PR · original plan: claude · original implementer: gpt-5.5 (Codex)
  • install feature — Package Hive for straightforward installation on macOS and several Linux distributions. reference PR · original plan: claude · original implementer: claude-opus-4-7
  • fix-tmux bugfix — Fix missing launcher scripts and ready-prompt detection for Claude Code in tmux mode. reference PR · original plan: claude · original implementer: gpt-5.5 (Codex)
  • fix-review bugfix — Recover completed review passes from Claude stop-hook failures instead of leaving REVIEW_ERROR. reference PR · original plan: claude · original implementer: gpt-5.5 (Codex)
  • daemon feature — Make the Hive daemon retry recoverable terminal errors after the failing dependency becomes healthy. reference PR · original plan: claude · original implementer: gpt-5.5 (Codex)

Audit the campaign

Every leaderboard score can be traced to a machine-readable record, and all 90 exact final diffs are public. The published evidence lets you:

  • Read the publication notes for coverage, judge settings, limitations, and publication exclusions.
  • Use the manifest to map every candidate/task cell to its files and verify each patch's byte size and SHA-256 hash.
  • Inspect the published score results for quality scores, judge provenance, same-family flags, and gate status.
  • Download the site data snapshot containing every displayed score, all three-sample follow-up distributions and intervals, the discussion-final diagnostic layer, recorded time, normalized token values, and API-equivalent cost estimates.
  • Read or compare the 36 original candidate patches or follow the per-cell links above for all 54 three-seed campaign patches.

These artifacts let you audit every score against the candidate code. The original campaign manifest also records each patch's byte size and SHA-256. Raw provider streams, build logs, target clones, and auth material are intentionally not published. Displayed wall times can be checked in the score results where raw evidence is public. Normalized token splits and recomputed costs live in the site snapshot and cannot yet be independently rebuilt from the public bundle.

  • All 90 generation cells completed. Every objective gate is no_gate in all six active campaigns, so scores are judge evidence, not test-pass rates.
  • The original 36-cell campaign uses one sample per judge; the 54 cells across the five later active three-sample campaigns use three samples per judge. All six active campaigns ran adversarial deliberation across all 90 cells.
  • Discussion final is a separate one-shot diagnostic, not an adjustment to the independent leaderboard scores. It covers 180/180 judge decisions. The original campaign reused its exact published verdicts and locally recovered rationales for round two; the five later active campaigns freshly re-graded round one.
  • GLM 5.3 Flash (0x Alpha) via Pi high preserves tokens for 6/6 cells and wall time for 4/6. Cost remains unknown rather than inferred.
  • GLM 5.3 Flash (0x Alpha) via Pi max preserves price and token telemetry for 6/6 cells and wall time for 5/6; the recovered fix-review artifact has no wall-time sample.
  • The historical OpenCode GLM 5.3 Flash (0x Alpha) campaign remains withdrawn because its profile lacked Bash capability parity and its target exposed held-out history; its invalidated patch artifacts remain available for audit.
  • Efficiency uses each campaign's serialized generation telemetry. DeepSeek's Pi/OpenRouter path preserved complete price and token telemetry for 11/12 cells and wall time for 11/12; missing values remain unknown.
  • Five Sol → Terra → Sol cells have complete recomputed API-equivalent costs; Grok workflows publish known-provider subtotals because Grok usage telemetry is unavailable.
  • The mixed Opus → Codex estimate is $115.5829 across six cells: $67.3351 Codex 5.5, $47.6319 Opus 4.8, and $0.6159 Haiku utility calls, or $19.26 per task on average.
  • Grok workflows publish known-provider token and API-equivalent cost subtotals. They exclude Grok, remain partial, and never impute its missing telemetry.
  • Wall-time means use only recorded samples. The production-review-panel campaign retains wall time for all 12 cells.
  • Costs use the benchmark's versioned usual-tier tables and are API-equivalent estimates, not subscription invoices; judge usage is excluded.
  • All 54 active three-sample campaign patches are public on this site; raw provider streams, logs, target clones, and auth material remain unpublished. Withdrawn OpenCode GLM 5.3 Flash (0x Alpha) patches remain available only as explicitly invalidated audit artifacts.

Run the benchmark for your task on your machine

The benchmark is a named workflow built into Hive. Install Hive, initialize a project with bench, and let the normal Hive daemon handle ordering, locking, retries, and concurrency.

hive init /path/to/benchmark-project --workflow bench
hive new <project-name> "benchmark my task"

Read the setup and campaign guide →

Submit a task to the public Hive benchmark

A completed Hive task with a merged reference PR can be proposed for a future public campaign directly from the CLI.

hive bench submit <task-slug>

Read the submission requirements →