hive-bench
Which coding agents hold up end to end?
Hive is a multi-agent orchestrator that carries software tasks through planning, implementation, pull requests, review, and fixes. hive-bench measures how coding-agent configurations perform across that entire automated workflow, scoring each final diff against the merged reference.
What this benchmark measures
The corpus contains 6 completed tasks from the Hive repository. For each task, the source is rewound to the original base commit and the candidate receives the frozen task inputs, but never the reference patch. A candidate may use one model for every stage or split planning and execution between models. Every candidate then runs real Hive in an isolated runner: plan → execute → open PR → production review and fix loop.
Planning invokes
Compound Engineering's
/ce-plan; execution uses Hive's normal plan-driven development
stage. Candidate-specific production reviewers can then revise the local
PR before the independent scoring judges see the final diff. The exact
stage owners and reviewer panels are listed in the
methodology.
The board covers all 66 generation cells: 11 candidates × 6 tasks. The original 36 cells have 1 independent score sample per judge; the 30 cells across the two later campaigns have 3. This corpus has no curated objective test gates, so close rankings are directional rather than definitive.
How the scores should be read
- Two independent rulers: Fable 5 and GPT-5.6 Sol
score each final diff from 0–10. Fable 5 ran at
xhighin all three campaigns. Sol ran atxhighfor the original campaign andultrafor both three-seed follow-ups. - Keep both judge scores visible: the leaderboard uses
their arithmetic mean only for a compact ordering because their
calibrations differ. Every row retains separate Fable and Sol columns,
and every matrix cell shows
Fable / Solin that order. - Deliberation is not a third judge: after the independent scores were recorded, Fable and Sol each saw the other judge's anonymous verdict and reasoning, argued the strongest evidence that its own initial score was wrong, and then held or revised that score. After discussion is the arithmetic mean of those two final scores. It is the default sort and a diagnostic layer, not a replacement for the independent scores, because the exchange can add anchoring or convergence pressure.
- Show family conflicts: when the judge and candidate share a model family, that row is marked SF . The score remains visible, but it is weaker evidence than a family-disjoint judgment.
- Keep efficiency descriptive: wall time, normalized generation tokens, and API-equivalent cost are shown where the underlying runner recorded them. Missing values are reported as unknown, never inferred.
Read the full methodology and limitations before treating small score differences as meaningful.
What came out
The independent scores put the all-Sol configuration first at 6.55: 6.917 by Fable and 6.183 by Sol. The production-like Sol plan → Sol execute → Sol + Grok review setup is second at 6.328: 6.722 by Fable and 5.933 by Sol. After deliberation, that production-like setup leads at 7.525, with 7.917 by Fable and 7.133 by Sol. Sol plan → Grok execute → Sol review follows at 6.908, then Sol plan → Terra execute → Sol review at 6.467, then all-Sol at 6.058.
The most curious result is the new winner: it ranks second under the three-sample independent mean but first under the one-shot after-discussion mean. The two later campaigns freshly re-graded round one before the judges exchanged verdicts, so this is not a mathematical adjustment to the independent mean. Across all 66 cells, the judges' mean absolute spread still fell from 1.77 to 1.039. That suggests the anonymous exchange surfaced real misses, but the same movement may also reflect anchoring pressure. Deliberation therefore remains a diagnostic rather than a cleaner replacement score.
Grok 4.5 remains fastest on its recorded timed cells at 27.3 minutes on average. The production-like winner reports $42.94 and 63.156M known Sol-stage tokens per task. Sol plan → Terra execute → Grok review reports $18.63 and 26.229M known Sol/Terra tokens. Both cover all six cells, but both remain lower-bound subtotals because Grok reviewer telemetry is unavailable.
These are scoped findings, not a universal model ranking. 6 tasks from one Ruby/CLI project cannot establish broad superiority. The original rows have 1 sample per judge, the two later campaigns have 3, and the discussion layer is a single exchange; the results show what happened when these exact configurations drove this exact end-to-end workflow.
One task, end to end
fix-tmux asked candidates to repair launcher scripts and
Claude-ready detection, based on the task later merged as
Hive PR #623.
The mixed Opus → Codex candidate scored 9.0 / 8.7
(Fable / Sol), the best Sol score for that task. GPT-5.6 Sol scored
9.0 / 8.5. The result and exact candidate patch behind
every score are linked from the full board below.
Leaderboard
Sol plan → Sol execute → Sol + Grok review
Sol xhigh plan · Sol high execute · Sol xhigh + Grok xhigh review
3 samples/judge · scores + public diffs
|
6.722 | 5.933 SF | 6.328 | 7.525 Fable 7.917 · Sol 7.133 6/6 paired | 151.0 min 6/6 timed | $42.94 known* | 63.156M known* |
|---|---|---|---|---|---|---|---|
Sol plan → Grok execute → Sol review
Sol xhigh plan · Grok xhigh execute · Sol xhigh review
3 samples/judge · scores + public diffs
|
7.778 | 4.845 SF | 6.312 | 6.908 Fable 7.383 · Sol 6.433 6/6 paired | 132.1 min 5/6 timed | $21.35 known* | 29.123M known* |
Sol plan → Terra execute → Sol review
Sol xhigh plan · Terra xhigh execute · Sol xhigh review
3 samples/judge · scores + public diffs
|
7.055 | 5.133 SF | 6.094 | 6.467 Fable 7.45 · Sol 5.483 6/6 paired | 166 min 4/6 timed | $26.12 5/6 priced | 34.807M 5/6 measured |
GPT-5.6 Sol xhigh
xhigh (pinned)
|
6.917 | 6.183 SF | 6.55 | 6.058 Fable 6.25 · Sol 5.867 6/6 paired | 146.9 min 4/6 timed | $47.66 6/6 priced | 74.86M 6/6 measured |
Fable plan → Grok execute → Sol review
Fable high plan · Grok xhigh execute · Sol xhigh review
3 samples/judge · scores + public diffs
|
6.194 SF | 4.667 SF | 5.431 | 6.05 Fable 6.583 · Sol 5.517 6/6 paired | 98.4 min 5/6 timed | $7.10 known* | 9.24M known* |
Sol plan → Terra execute → Grok review
Sol xhigh plan · Terra xhigh execute · Grok xhigh review
3 samples/judge · scores + public diffs
|
5.861 | 4.567 SF | 5.214 | 5.717 Fable 6.517 · Sol 4.917 6/6 paired | 61.1 min 6/6 timed | $18.63 known* | 26.229M known* |
Opus plan → Codex 5.5 xhigh
Opus default · Codex xhigh
|
6.167 SF | 4.783 SF | 5.475 | 5.267 Fable 5.583 · Sol 4.95 6/6 paired | 60.8 min 5/6 timed | $19.26 6/6 priced | 27.67M 6/6 measured |
Opus 4.8
Claude CLI default
|
6.083 SF | 4.25 | 5.167 | 4.833 Fable 5.5 · Sol 4.167 6/6 paired | 63.7 min 6/6 timed | $35.74 6/6 priced | 59.82M 6/6 measured |
Grok 4.5 xhigh
xhigh (pinned)
|
6.333 | 4.217 | 5.275 | 4.383 Fable 4.883 · Sol 3.883 6/6 paired | 27.3 min 5/6 timed | unknown | unknown |
GLM 5.2
provider default
|
4.833 | 3.617 | 4.225 | 3.817 Fable 4.25 · Sol 3.383 6/6 paired | 114.4 min 6/6 timed | $21.97 6/6 priced | 83.65M 6/6 measured |
Codex 5.5 xhigh
xhigh (pinned)
|
4.667 | 3.867 SF | 4.267 | 3.725 Fable 4.0 · Sol 3.45 6/6 paired | 52 min 6/6 timed | $17.17 6/6 priced | 24.01M 6/6 measured |
Per-task efficiency by model/configuration
| Model / candidate | add-i-key | web-install | install | fix-tmux | fix-review | daemon |
|---|---|---|---|---|---|---|
GPT-5.6 Sol xhigh |
$20.99 28.26M tokens 95.3 min |
$88.23 142.43M tokens 227.4 min |
$105.24 167.52M tokens 224.2 min |
$20.36 32.37M tokens 40.7 min |
$18.94 29.98M tokens time not recorded |
$32.22 48.58M tokens time not recorded |
Sol plan → Sol execute → Sol + Grok review |
$29.10 known* 41.088M known* tokens 124.0 min |
$62.98 known* 96.241M known* tokens 178.7 min |
$40.76 known* 58.575M known* tokens 145.1 min |
$25.19 known* 38.298M known* tokens 94.1 min |
$39.46 known* 56.678M known* tokens 177.3 min |
$60.16 known* 88.057M known* tokens 186.5 min |
Sol plan → Grok execute → Sol review |
$24.04 known* 33.998M known* tokens 115.5 min |
$22.26 known* 29.397M known* tokens 94.9 min |
$24.25 known* 30.182M known* tokens 173.5 min |
$5.58 known* 6.697M known* tokens 63.5 min |
$16.15 known* 22.795M known* tokens time not recorded |
$35.83 known* 51.67M known* tokens 213.1 min |
Sol plan → Terra execute → Sol review |
$17.16 21.823M tokens 91.4 min |
$21.70 30.491M tokens 197.7 min |
$31.51 38.249M tokens 179.5 min |
$31.28 43.864M tokens time not recorded |
cost unknown tokens unknown time not recorded |
$28.96 39.608M tokens 195.2 min |
Opus plan → Codex 5.5 xhigh |
$25.75 39.89M tokens time not recorded |
$16.95 22.79M tokens 50.9 min |
$31.02 47.85M tokens 75 min |
$11.92 14.69M tokens 73 min |
$12.95 18.12M tokens 46.9 min |
$17.00 22.7M tokens 58.2 min |
Fable plan → Grok execute → Sol review |
$11.18 known* 14.114M known* tokens 92.1 min |
$3.45 known* 5.927M known* tokens 150.7 min |
$3.38 known* 1.335M known* tokens time not recorded |
$2.43 known* 2.359M known* tokens 25.5 min |
$20.26 known* 28.061M known* tokens 87.2 min |
$1.93 known* 3.647M known* tokens 136.2 min |
Grok 4.5 xhigh |
cost unknown tokens unknown 30.1 min |
cost unknown tokens unknown 32.1 min |
cost unknown tokens unknown 27.3 min |
cost unknown tokens unknown time not recorded |
cost unknown tokens unknown 23.2 min |
cost unknown tokens unknown 23.9 min |
Sol plan → Terra execute → Grok review |
$12.53 known* 17.19M known* tokens 43.7 min |
$31.78 known* 48.013M known* tokens 75.1 min |
$20.61 known* 28.593M known* tokens 54.3 min |
$12.30 known* 17.293M known* tokens 93.4 min |
$15.16 known* 20.982M known* tokens 50.9 min |
$19.43 known* 25.302M known* tokens 49.5 min |
Opus 4.8 |
$24.29 39.03M tokens 59.5 min |
$33.08 54.71M tokens 44 min |
$74.33 125.82M tokens 98.7 min |
$11.35 17.23M tokens 39.9 min |
$27.75 47.07M tokens 55.1 min |
$43.64 75.08M tokens 85 min |
Codex 5.5 xhigh |
$14.55 19.17M tokens 51 min |
$20.95 29.53M tokens 52.9 min |
$20.21 28.59M tokens 67 min |
$7.81 11.42M tokens 26 min |
$22.62 31.1M tokens 75 min |
$16.85 24.25M tokens 40.4 min |
GLM 5.2 |
$15.85 64.59M tokens 114.1 min |
$22.09 93.86M tokens 173.5 min |
$26.15 103.24M tokens 124.4 min |
$9.80 40.87M tokens 91.4 min |
$12.39 17.05M tokens 29.7 min |
$45.53 182.31M tokens 153.4 min |
Full board (per task)
| Candidate | add-i-key | web-install | install | fix-tmux | fix-review | daemon |
|---|---|---|---|---|---|---|
GPT-5.6 Sol xhigh |
7.0 / 7.0 discussion final · Fable 7.0 · Sol 7.0 diff |
7.0 / 5.5 discussion final · Fable 6.5 · Sol 5.5 diff |
5.5 / 3.5 discussion final · Fable 4.0 · Sol 3.5 diff |
9.0 / 8.5 discussion final · Fable 8.5 · Sol 8.5 diff |
6.0 / 7.1 discussion final · Fable 6.0 · Sol 6.2 diff |
7.0 / 5.5 discussion final · Fable 5.5 · Sol 4.5 diff |
Sol plan → Sol execute → Sol + Grok review |
5.167 / 4.767 discussion final · Fable 8.5 · Sol 8.5 diff · 3 samples/judge |
7.0 / 5.833 discussion final · Fable 7.0 · Sol 6.5 diff · 3 samples/judge |
5.667 / 4.733 discussion final · Fable 6.0 · Sol 4.0 diff · 3 samples/judge |
9.0 / 7.833 discussion final · Fable 9.0 · Sol 9.2 diff · 3 samples/judge |
7.833 / 7.167 discussion final · Fable 8.5 · Sol 7.1 diff · 3 samples/judge |
5.667 / 5.267 discussion final · Fable 8.5 · Sol 7.5 diff · 3 samples/judge |
Sol plan → Grok execute → Sol review |
9.0 / 6.167 discussion final · Fable 8.8 · Sol 8.5 diff · 3 samples/judge |
8.333 / 5.167 discussion final · Fable 7 · Sol 5.8 diff · 3 samples/judge |
7.833 / 4.067 discussion final · Fable 6.5 · Sol 5.3 diff · 3 samples/judge |
9.167 / 5.833 discussion final · Fable 8.5 · Sol 8.5 diff · 3 samples/judge |
4.0 / 4.0 discussion final · Fable 5.5 · Sol 4 diff · 3 samples/judge |
8.333 / 3.833 discussion final · Fable 8 · Sol 6.5 diff · 3 samples/judge |
Sol plan → Terra execute → Sol review |
8.833 / 5.267 discussion final · Fable 9.2 · Sol 8.9 diff · 3 samples/judge |
3.667 / 2.833 discussion final · Fable 4.5 · Sol 3.5 diff · 3 samples/judge |
8.0 / 3.167 discussion final · Fable 6 · Sol 4.2 diff · 3 samples/judge |
8.333 / 6.9 discussion final · Fable 9 · Sol 6 diff · 3 samples/judge |
8.167 / 7.233 discussion final · Fable 7.5 · Sol 6.8 diff · 3 samples/judge |
5.333 / 5.4 discussion final · Fable 8.5 · Sol 3.5 diff · 3 samples/judge |
Opus plan → Codex 5.5 xhigh |
6.0 / 5.0 discussion final · Fable 5.5 · Sol 5.0 diff |
5.0 / 3.0 discussion final · Fable 4.0 · Sol 4.0 diff |
6.5 / 4.5 discussion final · Fable 6.0 · Sol 4.5 diff |
9.0 / 8.7 discussion final · Fable 9.0 · Sol 8.7 diff |
4.0 / 3.0 discussion final · Fable 3.5 · Sol 3.0 diff |
6.5 / 4.5 discussion final · Fable 5.5 · Sol 4.5 diff |
Fable plan → Grok execute → Sol review |
8.5 / 6.0 discussion final · Fable 8 · Sol 8 diff · 3 samples/judge |
5.167 / 3.167 discussion final · Fable 4.5 · Sol 3.5 diff · 3 samples/judge |
5.5 / 2.833 discussion final · Fable 4 · Sol 2.5 diff · 3 samples/judge |
5.833 / 7.833 discussion final · Fable 9 · Sol 7.5 diff · 3 samples/judge |
5.833 / 4.0 discussion final · Fable 7.5 · Sol 6 diff · 3 samples/judge |
6.333 / 4.167 discussion final · Fable 6.5 · Sol 5.6 diff · 3 samples/judge |
Grok 4.5 xhigh |
5.5 / 5.1 discussion final · Fable 5.0 · Sol 4.8 diff |
6.0 / 4.0 discussion final · Fable 5.0 · Sol 3.5 diff |
6.5 / 2.0 discussion final · Fable 3.0 · Sol 2.0 diff |
7.5 / 6.2 discussion final · Fable 6.8 · Sol 5.5 diff |
5.0 / 3.0 discussion final · Fable 3.5 · Sol 3.0 diff |
7.5 / 5.0 discussion final · Fable 6.0 · Sol 4.5 diff |
Sol plan → Terra execute → Grok review |
6.833 / 7.9 discussion final · Fable 8.6 · Sol 8.5 diff · 3 samples/judge |
4.667 / 2.833 discussion final · Fable 5.0 · Sol 3.0 diff · 3 samples/judge |
5.5 / 2.567 discussion final · Fable 6.0 · Sol 2.5 diff · 3 samples/judge |
8.5 / 6.267 discussion final · Fable 7.5 · Sol 7.0 diff · 3 samples/judge |
5.667 / 5.0 discussion final · Fable 7.5 · Sol 5.0 diff · 3 samples/judge |
4.0 / 2.833 discussion final · Fable 4.5 · Sol 3.5 diff · 3 samples/judge |
Opus 4.8 |
6.0 / 5.0 discussion final · Fable 5.5 · Sol 5.0 diff |
4.5 / 2.5 discussion final · Fable 4.0 · Sol 2.5 diff |
6.0 / 2.5 discussion final · Fable 6.0 · Sol 2.5 diff |
8.0 / 7.5 discussion final · Fable 7.5 · Sol 7.5 diff |
5.0 / 3.0 discussion final · Fable 4.0 · Sol 2.5 diff |
7.0 / 5.0 discussion final · Fable 6.0 · Sol 5.0 diff |
Codex 5.5 xhigh |
5.5 / 4.5 discussion final · Fable 5.0 · Sol 4.5 diff |
3.0 / 3.5 discussion final · Fable 3.0 · Sol 2.5 diff |
5.0 / 2.5 discussion final · Fable 3.5 · Sol 2.5 diff |
5.0 / 6.5 discussion final · Fable 5.0 · Sol 5.5 diff |
3.5 / 2.0 discussion final · Fable 3.0 · Sol 1.5 diff |
6.0 / 4.2 discussion final · Fable 4.5 · Sol 4.2 diff |
GLM 5.2 |
6.5 / 5.5 discussion final · Fable 6.0 · Sol 5.0 diff |
3.0 / 3.0 discussion final · Fable 3.0 · Sol 3.0 diff |
6.0 / 2.0 discussion final · Fable 4.0 · Sol 2.5 diff |
7.0 / 7.2 discussion final · Fable 7.5 · Sol 6.8 diff |
1.5 / 1.0 discussion final · Fable 1.0 · Sol 0.5 diff |
5.0 / 3.0 discussion final · Fable 4.0 · Sol 2.5 diff |
The tasks
-
add-i-keyfeature — Add an i key and legend entry that open a full task-information panel in the Hive TUI. reference PR · original plan: claude · original implementer:claude-opus-4-7 -
web-installfeature — Add a first-class local, non-Docker install and run mode for the Hive web UI. reference PR · original plan: claude · original implementer:gpt-5.5 (Codex) -
installfeature — Package Hive for straightforward installation on macOS and several Linux distributions. reference PR · original plan: claude · original implementer:claude-opus-4-7 -
fix-tmuxbugfix — Fix missing launcher scripts and ready-prompt detection for Claude Code in tmux mode. reference PR · original plan: claude · original implementer:gpt-5.5 (Codex) -
fix-reviewbugfix — Recover completed review passes from Claude stop-hook failures instead of leaving REVIEW_ERROR. reference PR · original plan: claude · original implementer:gpt-5.5 (Codex) -
daemonfeature — Make the Hive daemon retry recoverable terminal errors after the failing dependency becomes healthy. reference PR · original plan: claude · original implementer:gpt-5.5 (Codex)
Audit the campaign
Every leaderboard score can be traced to a machine-readable record, and all 66 exact final diffs are public. The published evidence lets you:
- Read the publication notes for coverage, judge settings, limitations, and publication exclusions.
- Use the manifest to map every candidate/task cell to its files and verify each patch's byte size and SHA-256 hash.
- Inspect the published score results for quality scores, judge provenance, same-family flags, and gate status.
- Download the site data snapshot containing every displayed score, all three-sample follow-up distributions and intervals, the discussion-final diagnostic layer, recorded time, normalized token values, and API-equivalent cost estimates.
- Read or compare the 36 original candidate patches or follow the per-cell links above for all 30 three-seed follow-up patches.
These artifacts let you audit every score against the candidate code. The original campaign manifest also records each patch's byte size and SHA-256. Raw provider streams, build logs, target clones, and auth material are intentionally not published. Displayed wall times can be checked in the score results where raw evidence is public. Normalized token splits and recomputed costs live in the site snapshot and cannot yet be independently rebuilt from the public bundle.
- All 66 generation cells completed. Every objective gate is no_gate in all three campaigns, so scores are judge evidence, not test-pass rates.
- The original 36-cell campaign uses one sample per judge; the 30 cells across the two three-seed follow-ups use three samples per judge. All three campaigns ran adversarial deliberation across all 66 cells.
- Discussion final is a separate one-shot diagnostic, not an adjustment to the independent leaderboard scores. It covers 132/132 judge decisions. The original campaign reused its exact published verdicts and locally recovered rationales for round two; the follow-ups freshly re-graded round one.
- Efficiency uses each campaign's serialized generation telemetry. Five Sol → Terra → Sol cells have complete recomputed API-equivalent costs; Grok workflows publish known-provider subtotals because Grok usage telemetry is unavailable.
- The mixed Opus → Codex estimate is $115.5829 across six cells: $67.3351 Codex 5.5, $47.6319 Opus 4.8, and $0.6159 Haiku utility calls, or $19.26 per task on average.
- Grok workflows publish known-provider token and API-equivalent cost subtotals. They exclude Grok, remain partial, and never impute its missing telemetry.
- Wall-time means use only recorded samples. The production-review-panel campaign retains wall time for all 12 cells.
- Costs use the benchmark's versioned 2026-06 usual-tier table and are API-equivalent estimates, not subscription invoices; judge usage is excluded.
- All 30 follow-up candidate patches are public on this site; raw provider streams, logs, target clones, and auth material remain unpublished.
Run the benchmark for your task on your machine
The benchmark is a named workflow built into Hive. Install Hive, initialize
a project with bench, and let the normal Hive daemon handle
ordering, locking, retries, and concurrency.
hive init /path/to/benchmark-project --workflow bench
hive new <project-name> "benchmark my task"
Read the setup and campaign guide →
Submit a task to the public Hive benchmark
A completed Hive task with a merged reference PR can be proposed for a future public campaign directly from the CLI.
hive bench submit <task-slug>