hive-bench

Which coding agents hold up end to end?

Hive is a multi-agent orchestrator that carries software tasks through planning, implementation, pull requests, review, and fixes. hive-bench measures how coding-agent configurations perform across that entire automated workflow, scoring each final diff against the merged reference.

What this benchmark measures

The corpus contains 6 completed tasks from the Hive repository. For each task, the source is rewound to the original base commit and the candidate receives the frozen task inputs, but never the reference patch. A candidate may use one model for every stage or split planning and execution between models. Every candidate then runs real Hive in an isolated runner: plan → execute → open PR → production review and fix loop.

Planning invokes Compound Engineering's /ce-plan; execution uses Hive's normal plan-driven development stage. Candidate-specific production reviewers can then revise the local PR before the independent scoring judges see the final diff. The exact stage owners and reviewer panels are listed in the methodology.

The board covers all 66 generation cells: 11 candidates × 6 tasks. The original 36 cells have 1 independent score sample per judge; the 30 cells across the two later campaigns have 3. This corpus has no curated objective test gates, so close rankings are directional rather than definitive.

How the scores should be read

  1. Two independent rulers: Fable 5 and GPT-5.6 Sol score each final diff from 0–10. Fable 5 ran at xhigh in all three campaigns. Sol ran at xhigh for the original campaign and ultra for both three-seed follow-ups.
  2. Keep both judge scores visible: the leaderboard uses their arithmetic mean only for a compact ordering because their calibrations differ. Every row retains separate Fable and Sol columns, and every matrix cell shows Fable / Sol in that order.
  3. Deliberation is not a third judge: after the independent scores were recorded, Fable and Sol each saw the other judge's anonymous verdict and reasoning, argued the strongest evidence that its own initial score was wrong, and then held or revised that score. After discussion is the arithmetic mean of those two final scores. It is the default sort and a diagnostic layer, not a replacement for the independent scores, because the exchange can add anchoring or convergence pressure.
  4. Show family conflicts: when the judge and candidate share a model family, that row is marked SF . The score remains visible, but it is weaker evidence than a family-disjoint judgment.
  5. Keep efficiency descriptive: wall time, normalized generation tokens, and API-equivalent cost are shown where the underlying runner recorded them. Missing values are reported as unknown, never inferred.

Read the full methodology and limitations before treating small score differences as meaningful.

What came out

The independent scores put the all-Sol configuration first at 6.55: 6.917 by Fable and 6.183 by Sol. The production-like Sol plan → Sol execute → Sol + Grok review setup is second at 6.328: 6.722 by Fable and 5.933 by Sol. After deliberation, that production-like setup leads at 7.525, with 7.917 by Fable and 7.133 by Sol. Sol plan → Grok execute → Sol review follows at 6.908, then Sol plan → Terra execute → Sol review at 6.467, then all-Sol at 6.058.

The most curious result is the new winner: it ranks second under the three-sample independent mean but first under the one-shot after-discussion mean. The two later campaigns freshly re-graded round one before the judges exchanged verdicts, so this is not a mathematical adjustment to the independent mean. Across all 66 cells, the judges' mean absolute spread still fell from 1.77 to 1.039. That suggests the anonymous exchange surfaced real misses, but the same movement may also reflect anchoring pressure. Deliberation therefore remains a diagnostic rather than a cleaner replacement score.

Grok 4.5 remains fastest on its recorded timed cells at 27.3 minutes on average. The production-like winner reports $42.94 and 63.156M known Sol-stage tokens per task. Sol plan → Terra execute → Grok review reports $18.63 and 26.229M known Sol/Terra tokens. Both cover all six cells, but both remain lower-bound subtotals because Grok reviewer telemetry is unavailable.

These are scoped findings, not a universal model ranking. 6 tasks from one Ruby/CLI project cannot establish broad superiority. The original rows have 1 sample per judge, the two later campaigns have 3, and the discussion layer is a single exchange; the results show what happened when these exact configurations drove this exact end-to-end workflow.

One task, end to end

fix-tmux asked candidates to repair launcher scripts and Claude-ready detection, based on the task later merged as Hive PR #623. The mixed Opus → Codex candidate scored 9.0 / 8.7 (Fable / Sol), the best Sol score for that task. GPT-5.6 Sol scored 9.0 / 8.5. The result and exact candidate patch behind every score are linked from the full board below.

Generated 2026-07-26T01:10:35Z · 66/66 cells complete

Leaderboard

Ranked by the presentation-only mean of Fable and Sol. The two judge columns remain the primary evidence because their calibrations differ. SF means that judge shares a model family with the candidate. Bold scores are the independent means and remain the primary evidence; the table opens ordered by the paired discussion-final mean. Efficiency averages use only recorded cells.

Sol plan → Sol execute → Sol + Grok review Sol xhigh plan · Sol high execute · Sol xhigh + Grok xhigh review 3 samples/judge · scores + public diffs 6.722 5.933 SF 6.328 7.525 Fable 7.917 · Sol 7.133 6/6 paired 151.0 min 6/6 timed $42.94 known* 63.156M known*
Sol plan → Grok execute → Sol review Sol xhigh plan · Grok xhigh execute · Sol xhigh review 3 samples/judge · scores + public diffs 7.778 4.845 SF 6.312 6.908 Fable 7.383 · Sol 6.433 6/6 paired 132.1 min 5/6 timed $21.35 known* 29.123M known*
Sol plan → Terra execute → Sol review Sol xhigh plan · Terra xhigh execute · Sol xhigh review 3 samples/judge · scores + public diffs 7.055 5.133 SF 6.094 6.467 Fable 7.45 · Sol 5.483 6/6 paired 166 min 4/6 timed $26.12 5/6 priced 34.807M 5/6 measured
GPT-5.6 Sol xhigh xhigh (pinned) 6.917 6.183 SF 6.55 6.058 Fable 6.25 · Sol 5.867 6/6 paired 146.9 min 4/6 timed $47.66 6/6 priced 74.86M 6/6 measured
Fable plan → Grok execute → Sol review Fable high plan · Grok xhigh execute · Sol xhigh review 3 samples/judge · scores + public diffs 6.194 SF 4.667 SF 5.431 6.05 Fable 6.583 · Sol 5.517 6/6 paired 98.4 min 5/6 timed $7.10 known* 9.24M known*
Sol plan → Terra execute → Grok review Sol xhigh plan · Terra xhigh execute · Grok xhigh review 3 samples/judge · scores + public diffs 5.861 4.567 SF 5.214 5.717 Fable 6.517 · Sol 4.917 6/6 paired 61.1 min 6/6 timed $18.63 known* 26.229M known*
Opus plan → Codex 5.5 xhigh Opus default · Codex xhigh 6.167 SF 4.783 SF 5.475 5.267 Fable 5.583 · Sol 4.95 6/6 paired 60.8 min 5/6 timed $19.26 6/6 priced 27.67M 6/6 measured
Opus 4.8 Claude CLI default 6.083 SF 4.25 5.167 4.833 Fable 5.5 · Sol 4.167 6/6 paired 63.7 min 6/6 timed $35.74 6/6 priced 59.82M 6/6 measured
Grok 4.5 xhigh xhigh (pinned) 6.333 4.217 5.275 4.383 Fable 4.883 · Sol 3.883 6/6 paired 27.3 min 5/6 timed unknown unknown
GLM 5.2 provider default 4.833 3.617 4.225 3.817 Fable 4.25 · Sol 3.383 6/6 paired 114.4 min 6/6 timed $21.97 6/6 priced 83.65M 6/6 measured
Codex 5.5 xhigh xhigh (pinned) 4.667 3.867 SF 4.267 3.725 Fable 4.0 · Sol 3.45 6/6 paired 52 min 6/6 timed $17.17 6/6 priced 24.01M 6/6 measured

Discussion is a one-shot diagnostic with campaign-specific round one: the six original rows reused their exact published independent verdicts and recovered rationales; the five rows across the two later three-seed campaigns received fresh round-one re-grades. In all three campaigns, each judge then saw the other anonymous verdict, challenged its own view, and held or revised it. All eleven rows ran this step across 66 paired cells. After discussion is the default table sort using the paired final mean; the independent scores remain visible, sortable, and the primary evidence.

Combined is (Fable + Sol) / 2 and is used only to provide one compact ordering. Token totals include fresh input, output, cache reads, and cache writes. Known* values include only providers with preserved telemetry. Grok usage and cost are unavailable, so those values are lower-bound subtotals rather than complete workflow totals. See the token accounting and pricing methodology for the normalization rules, mixed-model attribution, and known telemetry gaps.

Per-task efficiency by model/configuration

Each cell shows API-equivalent generation cost, total normalized generation tokens, the input/output/cache split, and recorded end-to-end wall time. Cache is shown as reads plus creation/writes. An unknown value means the runner did not preserve usable evidence; it does not mean zero.

Model / candidate add-i-key web-install install fix-tmux fix-review daemon
GPT-5.6 Sol xhigh $20.99
28.26M tokens
95.3 min
$88.23
142.43M tokens
227.4 min
$105.24
167.52M tokens
224.2 min
$20.36
32.37M tokens
40.7 min
$18.94
29.98M tokens
time not recorded
$32.22
48.58M tokens
time not recorded
Sol plan → Sol execute → Sol + Grok review $29.10 known*
41.088M known* tokens
124.0 min
$62.98 known*
96.241M known* tokens
178.7 min
$40.76 known*
58.575M known* tokens
145.1 min
$25.19 known*
38.298M known* tokens
94.1 min
$39.46 known*
56.678M known* tokens
177.3 min
$60.16 known*
88.057M known* tokens
186.5 min
Sol plan → Grok execute → Sol review $24.04 known*
33.998M known* tokens
115.5 min
$22.26 known*
29.397M known* tokens
94.9 min
$24.25 known*
30.182M known* tokens
173.5 min
$5.58 known*
6.697M known* tokens
63.5 min
$16.15 known*
22.795M known* tokens
time not recorded
$35.83 known*
51.67M known* tokens
213.1 min
Sol plan → Terra execute → Sol review $17.16
21.823M tokens
91.4 min
$21.70
30.491M tokens
197.7 min
$31.51
38.249M tokens
179.5 min
$31.28
43.864M tokens
time not recorded
cost unknown
tokens unknown
time not recorded
$28.96
39.608M tokens
195.2 min
Opus plan → Codex 5.5 xhigh $25.75
39.89M tokens
time not recorded
$16.95
22.79M tokens
50.9 min
$31.02
47.85M tokens
75 min
$11.92
14.69M tokens
73 min
$12.95
18.12M tokens
46.9 min
$17.00
22.7M tokens
58.2 min
Fable plan → Grok execute → Sol review $11.18 known*
14.114M known* tokens
92.1 min
$3.45 known*
5.927M known* tokens
150.7 min
$3.38 known*
1.335M known* tokens
time not recorded
$2.43 known*
2.359M known* tokens
25.5 min
$20.26 known*
28.061M known* tokens
87.2 min
$1.93 known*
3.647M known* tokens
136.2 min
Grok 4.5 xhigh cost unknown
tokens unknown
30.1 min
cost unknown
tokens unknown
32.1 min
cost unknown
tokens unknown
27.3 min
cost unknown
tokens unknown
time not recorded
cost unknown
tokens unknown
23.2 min
cost unknown
tokens unknown
23.9 min
Sol plan → Terra execute → Grok review $12.53 known*
17.19M known* tokens
43.7 min
$31.78 known*
48.013M known* tokens
75.1 min
$20.61 known*
28.593M known* tokens
54.3 min
$12.30 known*
17.293M known* tokens
93.4 min
$15.16 known*
20.982M known* tokens
50.9 min
$19.43 known*
25.302M known* tokens
49.5 min
Opus 4.8 $24.29
39.03M tokens
59.5 min
$33.08
54.71M tokens
44 min
$74.33
125.82M tokens
98.7 min
$11.35
17.23M tokens
39.9 min
$27.75
47.07M tokens
55.1 min
$43.64
75.08M tokens
85 min
Codex 5.5 xhigh $14.55
19.17M tokens
51 min
$20.95
29.53M tokens
52.9 min
$20.21
28.59M tokens
67 min
$7.81
11.42M tokens
26 min
$22.62
31.1M tokens
75 min
$16.85
24.25M tokens
40.4 min
GLM 5.2 $15.85
64.59M tokens
114.1 min
$22.09
93.86M tokens
173.5 min
$26.15
103.24M tokens
124.4 min
$9.80
40.87M tokens
91.4 min
$12.39
17.05M tokens
29.7 min
$45.53
182.31M tokens
153.4 min

Known* values include only providers with preserved telemetry. Depending on the workflow, that can include Fable planning, Sol stages, or Sol planning plus Terra execution; Grok telemetry is unavailable. Use each token info button for its exact provider scope and token split.

Full board (per task)

Every cell shows Fable 5 / GPT-5.6 Sol. The original rows use one score per judge; the mixed-workflow rows use three and display their means. Mixed-workflow cells also show the complete, separate discussion-final pair. Click a score for its machine-readable record. Exact diffs are linked where the public raw-evidence bundle contains them. No task in either campaign has an objective gate.

Candidate add-i-key web-install install fix-tmux fix-review daemon
GPT-5.6 Sol xhigh 7.0 / 7.0
discussion final · Fable 7.0 · Sol 7.0
diff
7.0 / 5.5
discussion final · Fable 6.5 · Sol 5.5
diff
5.5 / 3.5
discussion final · Fable 4.0 · Sol 3.5
diff
9.0 / 8.5
discussion final · Fable 8.5 · Sol 8.5
diff
6.0 / 7.1
discussion final · Fable 6.0 · Sol 6.2
diff
7.0 / 5.5
discussion final · Fable 5.5 · Sol 4.5
diff
Sol plan → Sol execute → Sol + Grok review 5.167 / 4.767
discussion final · Fable 8.5 · Sol 8.5
diff · 3 samples/judge
7.0 / 5.833
discussion final · Fable 7.0 · Sol 6.5
diff · 3 samples/judge
5.667 / 4.733
discussion final · Fable 6.0 · Sol 4.0
diff · 3 samples/judge
9.0 / 7.833
discussion final · Fable 9.0 · Sol 9.2
diff · 3 samples/judge
7.833 / 7.167
discussion final · Fable 8.5 · Sol 7.1
diff · 3 samples/judge
5.667 / 5.267
discussion final · Fable 8.5 · Sol 7.5
diff · 3 samples/judge
Sol plan → Grok execute → Sol review 9.0 / 6.167
discussion final · Fable 8.8 · Sol 8.5
diff · 3 samples/judge
8.333 / 5.167
discussion final · Fable 7 · Sol 5.8
diff · 3 samples/judge
7.833 / 4.067
discussion final · Fable 6.5 · Sol 5.3
diff · 3 samples/judge
9.167 / 5.833
discussion final · Fable 8.5 · Sol 8.5
diff · 3 samples/judge
4.0 / 4.0
discussion final · Fable 5.5 · Sol 4
diff · 3 samples/judge
8.333 / 3.833
discussion final · Fable 8 · Sol 6.5
diff · 3 samples/judge
Sol plan → Terra execute → Sol review 8.833 / 5.267
discussion final · Fable 9.2 · Sol 8.9
diff · 3 samples/judge
3.667 / 2.833
discussion final · Fable 4.5 · Sol 3.5
diff · 3 samples/judge
8.0 / 3.167
discussion final · Fable 6 · Sol 4.2
diff · 3 samples/judge
8.333 / 6.9
discussion final · Fable 9 · Sol 6
diff · 3 samples/judge
8.167 / 7.233
discussion final · Fable 7.5 · Sol 6.8
diff · 3 samples/judge
5.333 / 5.4
discussion final · Fable 8.5 · Sol 3.5
diff · 3 samples/judge
Opus plan → Codex 5.5 xhigh 6.0 / 5.0
discussion final · Fable 5.5 · Sol 5.0
diff
5.0 / 3.0
discussion final · Fable 4.0 · Sol 4.0
diff
6.5 / 4.5
discussion final · Fable 6.0 · Sol 4.5
diff
9.0 / 8.7
discussion final · Fable 9.0 · Sol 8.7
diff
4.0 / 3.0
discussion final · Fable 3.5 · Sol 3.0
diff
6.5 / 4.5
discussion final · Fable 5.5 · Sol 4.5
diff
Fable plan → Grok execute → Sol review 8.5 / 6.0
discussion final · Fable 8 · Sol 8
diff · 3 samples/judge
5.167 / 3.167
discussion final · Fable 4.5 · Sol 3.5
diff · 3 samples/judge
5.5 / 2.833
discussion final · Fable 4 · Sol 2.5
diff · 3 samples/judge
5.833 / 7.833
discussion final · Fable 9 · Sol 7.5
diff · 3 samples/judge
5.833 / 4.0
discussion final · Fable 7.5 · Sol 6
diff · 3 samples/judge
6.333 / 4.167
discussion final · Fable 6.5 · Sol 5.6
diff · 3 samples/judge
Grok 4.5 xhigh 5.5 / 5.1
discussion final · Fable 5.0 · Sol 4.8
diff
6.0 / 4.0
discussion final · Fable 5.0 · Sol 3.5
diff
6.5 / 2.0
discussion final · Fable 3.0 · Sol 2.0
diff
7.5 / 6.2
discussion final · Fable 6.8 · Sol 5.5
diff
5.0 / 3.0
discussion final · Fable 3.5 · Sol 3.0
diff
7.5 / 5.0
discussion final · Fable 6.0 · Sol 4.5
diff
Sol plan → Terra execute → Grok review 6.833 / 7.9
discussion final · Fable 8.6 · Sol 8.5
diff · 3 samples/judge
4.667 / 2.833
discussion final · Fable 5.0 · Sol 3.0
diff · 3 samples/judge
5.5 / 2.567
discussion final · Fable 6.0 · Sol 2.5
diff · 3 samples/judge
8.5 / 6.267
discussion final · Fable 7.5 · Sol 7.0
diff · 3 samples/judge
5.667 / 5.0
discussion final · Fable 7.5 · Sol 5.0
diff · 3 samples/judge
4.0 / 2.833
discussion final · Fable 4.5 · Sol 3.5
diff · 3 samples/judge
Opus 4.8 6.0 / 5.0
discussion final · Fable 5.5 · Sol 5.0
diff
4.5 / 2.5
discussion final · Fable 4.0 · Sol 2.5
diff
6.0 / 2.5
discussion final · Fable 6.0 · Sol 2.5
diff
8.0 / 7.5
discussion final · Fable 7.5 · Sol 7.5
diff
5.0 / 3.0
discussion final · Fable 4.0 · Sol 2.5
diff
7.0 / 5.0
discussion final · Fable 6.0 · Sol 5.0
diff
Codex 5.5 xhigh 5.5 / 4.5
discussion final · Fable 5.0 · Sol 4.5
diff
3.0 / 3.5
discussion final · Fable 3.0 · Sol 2.5
diff
5.0 / 2.5
discussion final · Fable 3.5 · Sol 2.5
diff
5.0 / 6.5
discussion final · Fable 5.0 · Sol 5.5
diff
3.5 / 2.0
discussion final · Fable 3.0 · Sol 1.5
diff
6.0 / 4.2
discussion final · Fable 4.5 · Sol 4.2
diff
GLM 5.2 6.5 / 5.5
discussion final · Fable 6.0 · Sol 5.0
diff
3.0 / 3.0
discussion final · Fable 3.0 · Sol 3.0
diff
6.0 / 2.0
discussion final · Fable 4.0 · Sol 2.5
diff
7.0 / 7.2
discussion final · Fable 7.5 · Sol 6.8
diff
1.5 / 1.0
discussion final · Fable 1.0 · Sol 0.5
diff
5.0 / 3.0
discussion final · Fable 4.0 · Sol 2.5
diff

The tasks

Six real Hive tasks, each completed and merged. The original plans all came from Claude-family agents, which is disclosed corpus provenance and a possible framing bias.

  • add-i-key feature — Add an i key and legend entry that open a full task-information panel in the Hive TUI. reference PR · original plan: claude · original implementer: claude-opus-4-7
  • web-install feature — Add a first-class local, non-Docker install and run mode for the Hive web UI. reference PR · original plan: claude · original implementer: gpt-5.5 (Codex)
  • install feature — Package Hive for straightforward installation on macOS and several Linux distributions. reference PR · original plan: claude · original implementer: claude-opus-4-7
  • fix-tmux bugfix — Fix missing launcher scripts and ready-prompt detection for Claude Code in tmux mode. reference PR · original plan: claude · original implementer: gpt-5.5 (Codex)
  • fix-review bugfix — Recover completed review passes from Claude stop-hook failures instead of leaving REVIEW_ERROR. reference PR · original plan: claude · original implementer: gpt-5.5 (Codex)
  • daemon feature — Make the Hive daemon retry recoverable terminal errors after the failing dependency becomes healthy. reference PR · original plan: claude · original implementer: gpt-5.5 (Codex)

Audit the campaign

Every leaderboard score can be traced to a machine-readable record, and all 66 exact final diffs are public. The published evidence lets you:

  • Read the publication notes for coverage, judge settings, limitations, and publication exclusions.
  • Use the manifest to map every candidate/task cell to its files and verify each patch's byte size and SHA-256 hash.
  • Inspect the published score results for quality scores, judge provenance, same-family flags, and gate status.
  • Download the site data snapshot containing every displayed score, all three-sample follow-up distributions and intervals, the discussion-final diagnostic layer, recorded time, normalized token values, and API-equivalent cost estimates.
  • Read or compare the 36 original candidate patches or follow the per-cell links above for all 30 three-seed follow-up patches.

These artifacts let you audit every score against the candidate code. The original campaign manifest also records each patch's byte size and SHA-256. Raw provider streams, build logs, target clones, and auth material are intentionally not published. Displayed wall times can be checked in the score results where raw evidence is public. Normalized token splits and recomputed costs live in the site snapshot and cannot yet be independently rebuilt from the public bundle.

  • All 66 generation cells completed. Every objective gate is no_gate in all three campaigns, so scores are judge evidence, not test-pass rates.
  • The original 36-cell campaign uses one sample per judge; the 30 cells across the two three-seed follow-ups use three samples per judge. All three campaigns ran adversarial deliberation across all 66 cells.
  • Discussion final is a separate one-shot diagnostic, not an adjustment to the independent leaderboard scores. It covers 132/132 judge decisions. The original campaign reused its exact published verdicts and locally recovered rationales for round two; the follow-ups freshly re-graded round one.
  • Efficiency uses each campaign's serialized generation telemetry. Five Sol → Terra → Sol cells have complete recomputed API-equivalent costs; Grok workflows publish known-provider subtotals because Grok usage telemetry is unavailable.
  • The mixed Opus → Codex estimate is $115.5829 across six cells: $67.3351 Codex 5.5, $47.6319 Opus 4.8, and $0.6159 Haiku utility calls, or $19.26 per task on average.
  • Grok workflows publish known-provider token and API-equivalent cost subtotals. They exclude Grok, remain partial, and never impute its missing telemetry.
  • Wall-time means use only recorded samples. The production-review-panel campaign retains wall time for all 12 cells.
  • Costs use the benchmark's versioned 2026-06 usual-tier table and are API-equivalent estimates, not subscription invoices; judge usage is excluded.
  • All 30 follow-up candidate patches are public on this site; raw provider streams, logs, target clones, and auth material remain unpublished.

Run the benchmark for your task on your machine

The benchmark is a named workflow built into Hive. Install Hive, initialize a project with bench, and let the normal Hive daemon handle ordering, locking, retries, and concurrency.

hive init /path/to/benchmark-project --workflow bench
hive new <project-name> "benchmark my task"

Read the setup and campaign guide →

Submit a task to the public Hive benchmark

A completed Hive task with a merged reference PR can be proposed for a future public campaign directly from the CLI.

hive bench submit <task-slug>

Read the submission requirements →