hive-bench
Which coding agents hold up end to end?
Hive is a multi-agent orchestrator that carries software tasks through planning, implementation, pull requests, review, and fixes. hive-bench measures how coding-agent configurations perform across that entire automated workflow, scoring each final diff against the merged reference.
What this benchmark measures
The corpus contains 6 completed tasks from the Hive repository. For each task, the source is rewound to the original base commit and the candidate receives the frozen task inputs, but never the reference patch. A candidate may use one model for every stage or split planning and execution between models. Every candidate then runs real Hive in an isolated runner: plan → execute → open PR → production review and fix loop.
Planning invokes
Compound Engineering's
/ce-plan; execution uses Hive's normal plan-driven development
stage. Candidate-specific production reviewers can then revise the local
PR before the independent scoring judges see the final diff. The exact
stage owners and reviewer panels are listed in the
methodology.
The board covers all 90 generation cells: 15 candidates × 6 tasks. The original 36 cells have 1 independent score sample per judge; the 54 cells across the five active later campaigns have 3. This corpus has no curated objective test gates, so close rankings are directional rather than definitive.
How the scores should be read
- Two independent rulers: Fable 5 and GPT-5.6 Sol
score each final diff from 0–10. Fable ran with reasoning enabled:
the first three campaign records pin
xhigh, while the DeepSeek and GLM 5.3 Flash (0x Alpha) campaigns record it as unspecified. Sol ran atxhighfor the original campaign andultrafor all five active later campaigns. - Keep both judge scores visible: the leaderboard uses
their arithmetic mean only for a compact ordering because their
calibrations differ. Every row retains separate Fable and Sol columns,
and every matrix cell shows
Fable / Solin that order. - Deliberation is not a third judge: after the independent scores were recorded, Fable and Sol each saw the other judge's anonymous verdict and reasoning, argued the strongest evidence that its own initial score was wrong, and then held or revised that score. After discussion is the arithmetic mean of those two final scores. It is the default sort and a diagnostic layer, not a replacement for the independent scores, because the exchange can add anchoring or convergence pressure.
- Show family conflicts: when the judge and candidate share a model family, that row is marked SF . The score remains visible, but it is weaker evidence than a family-disjoint judgment.
- Keep efficiency descriptive: wall time, normalized generation tokens, and API-equivalent cost are shown where the underlying runner recorded them. Missing values are reported as unknown, never inferred.
Read the full methodology and limitations before treating small score differences as meaningful.
What came out
GLM 5.3 Flash (0x Alpha) via Pi high and Pi max have independent paired means of 5.05 and 5.036. After the separate discussion diagnostic, they finish at 5.692 and 5.483. Both judges are family-disjoint from GLM 5.3 Flash (0x Alpha), and both rows use three samples per judge on the same six tasks.
Neither GLM 5.3 Flash (0x Alpha) row changes the after-discussion leader: the production-like Sol plan → Sol execute → Sol + Grok review setup remains first at 7.525. The all-Pro and Pro-plan/Flash-execute/Pro-review DeepSeek configurations finish at 5.892 and 5.783. Both judgments are family-disjoint, so their evidence carries no same-family warning.
The sharper DeepSeek result is judge calibration, not rank. On the three-sample independent scores, Fable and Sol differ by 4.483 points for all-Pro and 2.994 for the mixed workflow. The fresh one-shot deliberation brings the final judge gap for both rows to about 1.9 points without erasing it. Across all 90 cells, the judges' mean absolute spread still fell from 2.082 to 1.251. That suggests the anonymous exchange surfaced real misses, but the same movement may also reflect anchoring pressure. Deliberation therefore remains a diagnostic rather than a cleaner replacement score.
DeepSeek is the cheapest fully attributable provider path measured so far: all-Pro averages $0.55 across five priced cells, and Pro/Flash/Pro averages $0.6 across all six. Their normalized token volume is 50.754M and 34.774M per recorded task respectively—because cache reads are counted as part of total generation volume. Missing all-Pro telemetry remains unknown rather than inferred.
The Pi max campaign averages $6.94 per task across all six priced cells and 33.343M normalized tokens per task. Its wall-time mean is 95.5 minutes across 5 recorded cells. The earlier Pi high campaign has complete token telemetry but no serialized API-equivalent price, so its cost remains unknown.
These are scoped findings, not a universal model ranking. 6 tasks from one Ruby/CLI project cannot establish broad superiority. The original rows have 1 sample per judge, the five active later campaigns have 3, and the discussion layer is a single exchange; the results show what happened when these exact configurations drove this exact end-to-end workflow.
One task, end to end
fix-tmux asked candidates to repair launcher scripts and
Claude-ready detection, based on the task later merged as
Hive PR #623.
The mixed Opus → Codex candidate scored 9.0 / 8.7
(Fable / Sol), the best Sol score for that task. GPT-5.6 Sol scored
9.0 / 8.5. The result and exact candidate patch behind
every score are linked from the full board below.
Leaderboard
Sol plan → Sol execute → Sol + Grok review
Sol xhigh plan · Sol high execute · Sol xhigh + Grok xhigh review
3 samples/judge · scores + public diffs
|
6.722 | 5.933 SF | 6.328 | 7.525 Fable 7.917 · Sol 7.133 6/6 paired | 151.0 min 6/6 timed | $42.94 known* | 63.156M known* |
|---|---|---|---|---|---|---|---|
Sol plan → Grok execute → Sol review
Sol xhigh plan · Grok xhigh execute · Sol xhigh review
3 samples/judge · scores + public diffs
|
7.778 | 4.845 SF | 6.312 | 6.908 Fable 7.383 · Sol 6.433 6/6 paired | 132.1 min 5/6 timed | $21.35 known* | 29.123M known* |
Sol plan → Terra execute → Sol review
Sol xhigh plan · Terra xhigh execute · Sol xhigh review
3 samples/judge · scores + public diffs
|
7.055 | 5.133 SF | 6.094 | 6.467 Fable 7.45 · Sol 5.483 6/6 paired | 166 min 4/6 timed | $26.12 5/6 priced | 34.807M 5/6 measured |
GPT-5.6 Sol xhigh
xhigh (pinned)
|
6.917 | 6.183 SF | 6.55 | 6.058 Fable 6.25 · Sol 5.867 6/6 paired | 146.9 min 4/6 timed | $47.66 6/6 priced | 74.86M 6/6 measured |
Fable plan → Grok execute → Sol review
Fable high plan · Grok xhigh execute · Sol xhigh review
3 samples/judge · scores + public diffs
|
6.194 SF | 4.667 SF | 5.431 | 6.05 Fable 6.583 · Sol 5.517 6/6 paired | 98.4 min 5/6 timed | $7.10 known* | 9.24M known* |
DeepSeek V4 Pro 0813 xhigh
DeepSeek V4 Pro 0813 xhigh throughout
3 samples/judge · scores + public diffs
|
7.722 | 3.239 | 5.481 | 5.892 Fable 6.833 · Sol 4.95 6/6 paired | 120.0 min 5/6 timed | $0.55 5/6 priced | 50.754M 5/6 measured |
DeepSeek V4 Pro plan → V4 Flash execute → V4 Pro review
DeepSeek V4 Pro 0813 xhigh plan/review · V4 Flash 0731 xhigh execute
3 samples/judge · scores + public diffs
|
6.333 | 3.339 | 4.836 | 5.783 Fable 6.75 · Sol 4.817 6/6 paired | 118.8 min 6/6 timed | $0.6 6/6 priced | 34.774M 6/6 measured |
Sol plan → Terra execute → Grok review
Sol xhigh plan · Terra xhigh execute · Grok xhigh review
3 samples/judge · scores + public diffs
|
5.861 | 4.567 SF | 5.214 | 5.717 Fable 6.517 · Sol 4.917 6/6 paired | 61.1 min 6/6 timed | $18.63 known* | 26.229M known* |
GLM 5.3 Flash (0x Alpha) via Pi high
GLM 5.3 Flash (0x Alpha) high throughout via Pi/OpenRouter
3 samples/judge · scores + public diffs
|
6.306 | 3.795 | 5.05 | 5.692 Fable 6.333 · Sol 5.05 6/6 paired | 105.7 min 4/6 timed | unknown | 19.838M 6/6 measured |
GLM 5.3 Flash (0x Alpha) via Pi max
GLM 5.3 Flash (0x Alpha) max throughout via Pi/OpenRouter
3 samples/judge · scores + public diffs
|
6.5 | 3.572 | 5.036 | 5.483 Fable 6.5 · Sol 4.467 6/6 paired | 95.5 min 5/6 timed | $6.9381 6/6 priced | 33.343M 6/6 measured |
Opus plan → Codex 5.5 xhigh
Opus default · Codex xhigh
|
6.167 SF | 4.783 SF | 5.475 | 5.267 Fable 5.583 · Sol 4.95 6/6 paired | 60.8 min 5/6 timed | $19.26 6/6 priced | 27.67M 6/6 measured |
Opus 4.8
Claude CLI default
|
6.083 SF | 4.25 | 5.167 | 4.833 Fable 5.5 · Sol 4.167 6/6 paired | 63.7 min 6/6 timed | $35.74 6/6 priced | 59.82M 6/6 measured |
Grok 4.5 xhigh
xhigh (pinned)
|
6.333 | 4.217 | 5.275 | 4.383 Fable 4.883 · Sol 3.883 6/6 paired | 27.3 min 5/6 timed | unknown | unknown |
GLM 5.2
provider default
|
4.833 | 3.617 | 4.225 | 3.817 Fable 4.25 · Sol 3.383 6/6 paired | 114.4 min 6/6 timed | $5.65 6/6 priced | 21.793M 6/6 measured |
Codex 5.5 xhigh
xhigh (pinned)
|
4.667 | 3.867 SF | 4.267 | 3.725 Fable 4.0 · Sol 3.45 6/6 paired | 52 min 6/6 timed | $17.17 6/6 priced | 24.01M 6/6 measured |
Per-task efficiency by model/configuration
| Model / candidate | add-i-key | web-install | install | fix-tmux | fix-review | daemon |
|---|---|---|---|---|---|---|
GPT-5.6 Sol xhigh |
$20.99 28.26M tokens 95.3 min |
$88.23 142.43M tokens 227.4 min |
$105.24 167.52M tokens 224.2 min |
$20.36 32.37M tokens 40.7 min |
$18.94 29.98M tokens time not recorded |
$32.22 48.58M tokens time not recorded |
Sol plan → Sol execute → Sol + Grok review |
$29.10 known* 41.088M known* tokens 124.0 min |
$62.98 known* 96.241M known* tokens 178.7 min |
$40.76 known* 58.575M known* tokens 145.1 min |
$25.19 known* 38.298M known* tokens 94.1 min |
$39.46 known* 56.678M known* tokens 177.3 min |
$60.16 known* 88.057M known* tokens 186.5 min |
Sol plan → Grok execute → Sol review |
$24.04 known* 33.998M known* tokens 115.5 min |
$22.26 known* 29.397M known* tokens 94.9 min |
$24.25 known* 30.182M known* tokens 173.5 min |
$5.58 known* 6.697M known* tokens 63.5 min |
$16.15 known* 22.795M known* tokens time not recorded |
$35.83 known* 51.67M known* tokens 213.1 min |
Sol plan → Terra execute → Sol review |
$17.16 21.823M tokens 91.4 min |
$21.70 30.491M tokens 197.7 min |
$31.51 38.249M tokens 179.5 min |
$31.28 43.864M tokens time not recorded |
cost unknown tokens unknown time not recorded |
$28.96 39.608M tokens 195.2 min |
DeepSeek V4 Pro 0813 xhigh |
$0.13 3.905M tokens 36.2 min |
$0.96 103.633M tokens 146.2 min |
$0.63 67.776M tokens 110.4 min |
cost unknown tokens unknown time not recorded |
$0.42 27.539M tokens 114.5 min |
$0.63 50.918M tokens 192.9 min |
Opus plan → Codex 5.5 xhigh |
$25.75 39.89M tokens time not recorded |
$16.95 22.79M tokens 50.9 min |
$31.02 47.85M tokens 75 min |
$11.92 14.69M tokens 73 min |
$12.95 18.12M tokens 46.9 min |
$17.00 22.7M tokens 58.2 min |
Fable plan → Grok execute → Sol review |
$11.18 known* 14.114M known* tokens 92.1 min |
$3.45 known* 5.927M known* tokens 150.7 min |
$3.38 known* 1.335M known* tokens time not recorded |
$2.43 known* 2.359M known* tokens 25.5 min |
$20.26 known* 28.061M known* tokens 87.2 min |
$1.93 known* 3.647M known* tokens 136.2 min |
Grok 4.5 xhigh |
cost unknown tokens unknown 30.1 min |
cost unknown tokens unknown 32.1 min |
cost unknown tokens unknown 27.3 min |
cost unknown tokens unknown time not recorded |
cost unknown tokens unknown 23.2 min |
cost unknown tokens unknown 23.9 min |
Sol plan → Terra execute → Grok review |
$12.53 known* 17.19M known* tokens 43.7 min |
$31.78 known* 48.013M known* tokens 75.1 min |
$20.61 known* 28.593M known* tokens 54.3 min |
$12.30 known* 17.293M known* tokens 93.4 min |
$15.16 known* 20.982M known* tokens 50.9 min |
$19.43 known* 25.302M known* tokens 49.5 min |
Opus 4.8 |
$24.29 39.03M tokens 59.5 min |
$33.08 54.71M tokens 44 min |
$74.33 125.82M tokens 98.7 min |
$11.35 17.23M tokens 39.9 min |
$27.75 47.07M tokens 55.1 min |
$43.64 75.08M tokens 85 min |
GLM 5.3 Flash (0x Alpha) via Pi high |
cost unknown 5.392M tokens 51.2 min |
cost unknown 52.999M tokens time not recorded |
cost unknown 29.776M tokens 160.6 min |
cost unknown 3.348M tokens 99.1 min |
cost unknown 9.719M tokens 111.8 min |
cost unknown 17.794M tokens time not recorded |
GLM 5.3 Flash (0x Alpha) via Pi max |
$3.39 16.316M tokens 44.6 min |
$7.62 38.153M tokens 4.5 min |
$15.48 71.713M tokens 233.8 min |
$1.99 8.619M tokens 65.3 min |
$6.68 33.903M tokens time not recorded |
$6.47 31.354M tokens 129.2 min |
DeepSeek V4 Pro plan → V4 Flash execute → V4 Pro review |
$0.16 4.516M tokens 41.9 min |
$1.15 74.559M tokens 179.9 min |
$0.71 43.399M tokens 113.0 min |
$0.21 7.448M tokens 96.6 min |
$0.63 38.425M tokens 167.3 min |
$0.73 40.297M tokens 114.1 min |
Codex 5.5 xhigh |
$14.55 19.17M tokens 51 min |
$20.95 29.53M tokens 52.9 min |
$20.21 28.59M tokens 67 min |
$7.81 11.42M tokens 26 min |
$22.62 31.1M tokens 75 min |
$16.85 24.25M tokens 40.4 min |
GLM 5.2 |
$4.27 17.548M tokens 114.1 min |
$5.56 24.145M tokens 173.5 min |
$6.80 27.155M tokens 124.4 min |
$2.47 10.484M tokens 91.4 min |
$3.13 4.316M tokens 29.7 min |
$11.69 47.108M tokens 153.4 min |
Full board (per task)
| Candidate | add-i-key | web-install | install | fix-tmux | fix-review | daemon |
|---|---|---|---|---|---|---|
GPT-5.6 Sol xhigh |
7.0 / 7.0 discussion final · Fable 7.0 · Sol 7.0 diff |
7.0 / 5.5 discussion final · Fable 6.5 · Sol 5.5 diff |
5.5 / 3.5 discussion final · Fable 4.0 · Sol 3.5 diff |
9.0 / 8.5 discussion final · Fable 8.5 · Sol 8.5 diff |
6.0 / 7.1 discussion final · Fable 6.0 · Sol 6.2 diff |
7.0 / 5.5 discussion final · Fable 5.5 · Sol 4.5 diff |
Sol plan → Sol execute → Sol + Grok review |
5.167 / 4.767 discussion final · Fable 8.5 · Sol 8.5 diff · 3 samples/judge |
7.0 / 5.833 discussion final · Fable 7.0 · Sol 6.5 diff · 3 samples/judge |
5.667 / 4.733 discussion final · Fable 6.0 · Sol 4.0 diff · 3 samples/judge |
9.0 / 7.833 discussion final · Fable 9.0 · Sol 9.2 diff · 3 samples/judge |
7.833 / 7.167 discussion final · Fable 8.5 · Sol 7.1 diff · 3 samples/judge |
5.667 / 5.267 discussion final · Fable 8.5 · Sol 7.5 diff · 3 samples/judge |
Sol plan → Grok execute → Sol review |
9.0 / 6.167 discussion final · Fable 8.8 · Sol 8.5 diff · 3 samples/judge |
8.333 / 5.167 discussion final · Fable 7 · Sol 5.8 diff · 3 samples/judge |
7.833 / 4.067 discussion final · Fable 6.5 · Sol 5.3 diff · 3 samples/judge |
9.167 / 5.833 discussion final · Fable 8.5 · Sol 8.5 diff · 3 samples/judge |
4.0 / 4.0 discussion final · Fable 5.5 · Sol 4 diff · 3 samples/judge |
8.333 / 3.833 discussion final · Fable 8 · Sol 6.5 diff · 3 samples/judge |
Sol plan → Terra execute → Sol review |
8.833 / 5.267 discussion final · Fable 9.2 · Sol 8.9 diff · 3 samples/judge |
3.667 / 2.833 discussion final · Fable 4.5 · Sol 3.5 diff · 3 samples/judge |
8.0 / 3.167 discussion final · Fable 6 · Sol 4.2 diff · 3 samples/judge |
8.333 / 6.9 discussion final · Fable 9 · Sol 6 diff · 3 samples/judge |
8.167 / 7.233 discussion final · Fable 7.5 · Sol 6.8 diff · 3 samples/judge |
5.333 / 5.4 discussion final · Fable 8.5 · Sol 3.5 diff · 3 samples/judge |
DeepSeek V4 Pro 0813 xhigh |
9.0 / 1.0 discussion final · Fable 9.0 · Sol 9.2 diff · 3 samples/judge |
8.167 / 2.933 discussion final · Fable 7.5 · Sol 3.5 diff · 3 samples/judge |
7.167 / 2.167 discussion final · Fable 6.0 · Sol 3.0 diff · 3 samples/judge |
5.5 / 5.833 discussion final · Fable 5.0 · Sol 5.0 diff · 3 samples/judge |
8.0 / 3.5 discussion final · Fable 6.0 · Sol 4.0 diff · 3 samples/judge |
8.5 / 4.0 discussion final · Fable 7.5 · Sol 5.0 diff · 3 samples/judge |
Opus plan → Codex 5.5 xhigh |
6.0 / 5.0 discussion final · Fable 5.5 · Sol 5.0 diff |
5.0 / 3.0 discussion final · Fable 4.0 · Sol 4.0 diff |
6.5 / 4.5 discussion final · Fable 6.0 · Sol 4.5 diff |
9.0 / 8.7 discussion final · Fable 9.0 · Sol 8.7 diff |
4.0 / 3.0 discussion final · Fable 3.5 · Sol 3.0 diff |
6.5 / 4.5 discussion final · Fable 5.5 · Sol 4.5 diff |
Fable plan → Grok execute → Sol review |
8.5 / 6.0 discussion final · Fable 8 · Sol 8 diff · 3 samples/judge |
5.167 / 3.167 discussion final · Fable 4.5 · Sol 3.5 diff · 3 samples/judge |
5.5 / 2.833 discussion final · Fable 4 · Sol 2.5 diff · 3 samples/judge |
5.833 / 7.833 discussion final · Fable 9 · Sol 7.5 diff · 3 samples/judge |
5.833 / 4.0 discussion final · Fable 7.5 · Sol 6 diff · 3 samples/judge |
6.333 / 4.167 discussion final · Fable 6.5 · Sol 5.6 diff · 3 samples/judge |
Grok 4.5 xhigh |
5.5 / 5.1 discussion final · Fable 5.0 · Sol 4.8 diff |
6.0 / 4.0 discussion final · Fable 5.0 · Sol 3.5 diff |
6.5 / 2.0 discussion final · Fable 3.0 · Sol 2.0 diff |
7.5 / 6.2 discussion final · Fable 6.8 · Sol 5.5 diff |
5.0 / 3.0 discussion final · Fable 3.5 · Sol 3.0 diff |
7.5 / 5.0 discussion final · Fable 6.0 · Sol 4.5 diff |
Sol plan → Terra execute → Grok review |
6.833 / 7.9 discussion final · Fable 8.6 · Sol 8.5 diff · 3 samples/judge |
4.667 / 2.833 discussion final · Fable 5.0 · Sol 3.0 diff · 3 samples/judge |
5.5 / 2.567 discussion final · Fable 6.0 · Sol 2.5 diff · 3 samples/judge |
8.5 / 6.267 discussion final · Fable 7.5 · Sol 7.0 diff · 3 samples/judge |
5.667 / 5.0 discussion final · Fable 7.5 · Sol 5.0 diff · 3 samples/judge |
4.0 / 2.833 discussion final · Fable 4.5 · Sol 3.5 diff · 3 samples/judge |
Opus 4.8 |
6.0 / 5.0 discussion final · Fable 5.5 · Sol 5.0 diff |
4.5 / 2.5 discussion final · Fable 4.0 · Sol 2.5 diff |
6.0 / 2.5 discussion final · Fable 6.0 · Sol 2.5 diff |
8.0 / 7.5 discussion final · Fable 7.5 · Sol 7.5 diff |
5.0 / 3.0 discussion final · Fable 4.0 · Sol 2.5 diff |
7.0 / 5.0 discussion final · Fable 6.0 · Sol 5.0 diff |
GLM 5.3 Flash (0x Alpha) via Pi high |
6.833 / 6.0 discussion final · Fable 8.0 · Sol 7.5 diff · 3 samples/judge |
6.0 / 3.767 discussion final · Fable 7.5 · Sol 6.5 diff · 3 samples/judge |
6.833 / 2.167 discussion final · Fable 4.0 · Sol 2.8 diff · 3 samples/judge |
8.167 / 5.5 discussion final · Fable 7.5 · Sol 6.5 diff · 3 samples/judge |
3.167 / 2.333 discussion final · Fable 4.5 · Sol 2.0 diff · 3 samples/judge |
6.833 / 3.0 discussion final · Fable 6.5 · Sol 5.0 diff · 3 samples/judge |
GLM 5.3 Flash (0x Alpha) via Pi max |
8.667 / 4.667 discussion final · Fable 8.0 · Sol 7.5 diff · 3 samples/judge |
4.333 / 3.1 discussion final · Fable 5.0 · Sol 2.5 diff · 3 samples/judge |
6.0 / 1.667 discussion final · Fable 4.0 · Sol 2.5 diff · 3 samples/judge |
8.667 / 5.167 discussion final · Fable 9.0 · Sol 8.8 diff · 3 samples/judge |
4.0 / 2.5 discussion final · Fable 5.0 · Sol 2.5 diff · 3 samples/judge |
7.333 / 4.333 discussion final · Fable 8.0 · Sol 3.0 diff · 3 samples/judge |
DeepSeek V4 Pro plan → V4 Flash execute → V4 Pro review |
9.0 / 0.833 discussion final · Fable 9.0 · Sol 9.4 diff · 3 samples/judge |
3.833 / 2.933 discussion final · Fable 7.0 · Sol 3.0 diff · 3 samples/judge |
5.667 / 2.5 discussion final · Fable 5.0 · Sol 2.0 diff · 3 samples/judge |
8.5 / 5.5 discussion final · Fable 8.0 · Sol 6.5 diff · 3 samples/judge |
3.5 / 3.0 discussion final · Fable 4.5 · Sol 2.5 diff · 3 samples/judge |
7.5 / 5.267 discussion final · Fable 7.0 · Sol 5.5 diff · 3 samples/judge |
Codex 5.5 xhigh |
5.5 / 4.5 discussion final · Fable 5.0 · Sol 4.5 diff |
3.0 / 3.5 discussion final · Fable 3.0 · Sol 2.5 diff |
5.0 / 2.5 discussion final · Fable 3.5 · Sol 2.5 diff |
5.0 / 6.5 discussion final · Fable 5.0 · Sol 5.5 diff |
3.5 / 2.0 discussion final · Fable 3.0 · Sol 1.5 diff |
6.0 / 4.2 discussion final · Fable 4.5 · Sol 4.2 diff |
GLM 5.2 |
6.5 / 5.5 discussion final · Fable 6.0 · Sol 5.0 diff |
3.0 / 3.0 discussion final · Fable 3.0 · Sol 3.0 diff |
6.0 / 2.0 discussion final · Fable 4.0 · Sol 2.5 diff |
7.0 / 7.2 discussion final · Fable 7.5 · Sol 6.8 diff |
1.5 / 1.0 discussion final · Fable 1.0 · Sol 0.5 diff |
5.0 / 3.0 discussion final · Fable 4.0 · Sol 2.5 diff |
The tasks
-
add-i-keyfeature — Add an i key and legend entry that open a full task-information panel in the Hive TUI. reference PR · original plan: claude · original implementer:claude-opus-4-7 -
web-installfeature — Add a first-class local, non-Docker install and run mode for the Hive web UI. reference PR · original plan: claude · original implementer:gpt-5.5 (Codex) -
installfeature — Package Hive for straightforward installation on macOS and several Linux distributions. reference PR · original plan: claude · original implementer:claude-opus-4-7 -
fix-tmuxbugfix — Fix missing launcher scripts and ready-prompt detection for Claude Code in tmux mode. reference PR · original plan: claude · original implementer:gpt-5.5 (Codex) -
fix-reviewbugfix — Recover completed review passes from Claude stop-hook failures instead of leaving REVIEW_ERROR. reference PR · original plan: claude · original implementer:gpt-5.5 (Codex) -
daemonfeature — Make the Hive daemon retry recoverable terminal errors after the failing dependency becomes healthy. reference PR · original plan: claude · original implementer:gpt-5.5 (Codex)
Audit the campaign
Every leaderboard score can be traced to a machine-readable record, and all 90 exact final diffs are public. The published evidence lets you:
- Read the publication notes for coverage, judge settings, limitations, and publication exclusions.
- Use the manifest to map every candidate/task cell to its files and verify each patch's byte size and SHA-256 hash.
- Inspect the published score results for quality scores, judge provenance, same-family flags, and gate status.
- Download the site data snapshot containing every displayed score, all three-sample follow-up distributions and intervals, the discussion-final diagnostic layer, recorded time, normalized token values, and API-equivalent cost estimates.
- Read or compare the 36 original candidate patches or follow the per-cell links above for all 54 three-seed campaign patches.
These artifacts let you audit every score against the candidate code. The original campaign manifest also records each patch's byte size and SHA-256. Raw provider streams, build logs, target clones, and auth material are intentionally not published. Displayed wall times can be checked in the score results where raw evidence is public. Normalized token splits and recomputed costs live in the site snapshot and cannot yet be independently rebuilt from the public bundle.
- All 90 generation cells completed. Every objective gate is no_gate in all six active campaigns, so scores are judge evidence, not test-pass rates.
- The original 36-cell campaign uses one sample per judge; the 54 cells across the five later active three-sample campaigns use three samples per judge. All six active campaigns ran adversarial deliberation across all 90 cells.
- Discussion final is a separate one-shot diagnostic, not an adjustment to the independent leaderboard scores. It covers 180/180 judge decisions. The original campaign reused its exact published verdicts and locally recovered rationales for round two; the five later active campaigns freshly re-graded round one.
- GLM 5.3 Flash (0x Alpha) via Pi high preserves tokens for 6/6 cells and wall time for 4/6. Cost remains unknown rather than inferred.
- GLM 5.3 Flash (0x Alpha) via Pi max preserves price and token telemetry for 6/6 cells and wall time for 5/6; the recovered fix-review artifact has no wall-time sample.
- The historical OpenCode GLM 5.3 Flash (0x Alpha) campaign remains withdrawn because its profile lacked Bash capability parity and its target exposed held-out history; its invalidated patch artifacts remain available for audit.
- Efficiency uses each campaign's serialized generation telemetry. DeepSeek's Pi/OpenRouter path preserved complete price and token telemetry for 11/12 cells and wall time for 11/12; missing values remain unknown.
- Five Sol → Terra → Sol cells have complete recomputed API-equivalent costs; Grok workflows publish known-provider subtotals because Grok usage telemetry is unavailable.
- The mixed Opus → Codex estimate is $115.5829 across six cells: $67.3351 Codex 5.5, $47.6319 Opus 4.8, and $0.6159 Haiku utility calls, or $19.26 per task on average.
- Grok workflows publish known-provider token and API-equivalent cost subtotals. They exclude Grok, remain partial, and never impute its missing telemetry.
- Wall-time means use only recorded samples. The production-review-panel campaign retains wall time for all 12 cells.
- Costs use the benchmark's versioned usual-tier tables and are API-equivalent estimates, not subscription invoices; judge usage is excluded.
- All 54 active three-sample campaign patches are public on this site; raw provider streams, logs, target clones, and auth material remain unpublished. Withdrawn OpenCode GLM 5.3 Flash (0x Alpha) patches remain available only as explicitly invalidated audit artifacts.
Run the benchmark for your task on your machine
The benchmark is a named workflow built into Hive. Install Hive, initialize
a project with bench, and let the normal Hive daemon handle
ordering, locking, retries, and concurrency.
hive init /path/to/benchmark-project --workflow bench
hive new <project-name> "benchmark my task"
Read the setup and campaign guide →
Submit a task to the public Hive benchmark
A completed Hive task with a merged reference PR can be proposed for a future public campaign directly from the CLI.
hive bench submit <task-slug>