Published Harbor datasets¶
Results cover the six native pipelines, Tasksmith and all 14 research recipes. Historical evidence and the newer publication checks are reported separately below.
Browse the HuggingEnvs collection. The native datasets retain their existing owners.
Native pipelines¶
600 task entries across 6 earlier datasets. Recovered from cached Hub manifests and local publication stagings on 2026-09-15. These are historical snapshots, not a fresh Hub recount. Older revisions and duplicate stagings are excluded; cross-pipeline content is not deduplicated.
| Pipeline | Tasks | Recovered validation evidence | Dataset and evidence |
|---|---|---|---|
| pr_diff | 181 | 181 listed; no per-task validation in this snapshot. Earlier 100-task release reported oracle passes. | Dataset · Manifest |
| pr_runtime | 100 | 100 oracle/tracked passes; 88 clean commands; 87 with a regression guard. | Dataset · Manifest |
| commit_runtime | 100 | 100 generation-time verified stamps. Full 100-task oracle gate not recovered; older 52-task gate is separate. | Dataset · Manifest |
| code_instruct | 100 | 100 emitted after generation checks; no separate 100-task Harbor gate recovered. | Dataset · Local evidence |
| equivalence_tests | 100 | 100 emitted after generation checks; no separate 100-task Harbor gate recovered. | Dataset · Local evidence |
| cve_patches | 19 | 19 listed; validation fields absent. A 19-task oracle or independent quality gate was not recovered. | Dataset · Manifest |
These tasks have historical evidence scopes, not retrospectively assigned verified labels. In particular, the earlier 52-task commit-runtime gate does not validate the later 100-task dataset. Read the native results and solver samples.
Tasksmith and research recipes¶
1,330 tasks across 15 datasets. Each dataset contains complete Harbor task directories, archives and a registry pinned to its artifact revision.
| Pipeline | Tasks | Evaluation labels | Dataset and pinned manifest |
|---|---|---|---|
| swe-smith | 100 | 100 unverified | Dataset · Manifest |
| r2e | 100 | 100 unverified | Dataset · Manifest |
| swe-gen | 100 | 100 unverified | Dataset · Manifest |
| swe-next | 100 | 100 unverified | Dataset · Manifest |
| r2e-gym | 100 | 100 unverified | Dataset · Manifest |
| scaler | 100 | 100 unverified | Dataset · Manifest |
| endless-terminals | 100 | 100 unverified | Dataset · Manifest |
| cli-gym | 25 | 25 unverified | Dataset · Manifest |
| swe-flow | 100 | 2 needs repair; 98 unverified | Dataset · Manifest |
| seta-seed2synth | 100 | 100 unverified | Dataset · Manifest |
| seta-evol | 100 | 100 unverified | Dataset · Manifest |
| tmax | 55 | 55 unverified | Dataset · Manifest |
| terminalworld | 100 | 3 needs repair; 97 unverified | Dataset · Manifest |
| dataarc | 100 | 100 unverified | Dataset · Manifest |
| tasksmith | 50 | 50 verified | Dataset · Manifest |
What the labels establish¶
The Tasksmith and research-recipe release contains 50 verified, 5 needing repair and 1,275 unverified tasks. These totals exclude the historical native inventories above. Tasksmith's verified cohort came from an assisted campaign; this does not claim unattended conversion. Two SWE-flow instruction issues and three TerminalWorld verifier gaps remain explicitly diagnosed. Each dataset manifest supplies task-level labels, diagnostics and evidence scope.
Publication checks for those 15 datasets compared 214,097 file identities and parsed every selected task with Harbor. This establishes artifact integrity and format, not semantic quality of every task. See evaluation labels, yield and cost, and how to publish.