Published Harbor datasets
Published results cover the six native pipelines, Tasksmith and 14 research recipes. Local collections are reported separately and are not included in published totals.
Browse every dataset in the Repo2RLEnv collection on the Hugging Face Hub.
Native pipelines
600 task entries across 6 earlier datasets. Recovered from cached Hub manifests and local publication stagings on 2026-09-15. These are historical snapshots, not a fresh Hub recount. Older revisions and duplicate stagings are excluded; cross-pipeline content is not deduplicated.
| Pipeline | Tasks | Recovered validation evidence | Dataset and evidence |
|---|---|---|---|
| pr_diff | 181 | 181 listed; no per-task validation in this snapshot. Earlier 100-task release reported oracle passes. | Dataset · Manifest |
| pr_runtime | 100 | 100 oracle/tracked passes; 88 clean commands; 87 with a regression guard. | Dataset · Manifest |
| commit_runtime | 100 | 100 generation-time verified stamps. Full 100-task oracle gate not recovered; older 52-task gate is separate. | Dataset · Manifest |
| code_instruct | 100 | 100 emitted after generation checks; no separate 100-task Harbor gate recovered. | Dataset · Local evidence |
| equivalence_tests | 100 | 100 emitted after generation checks; no separate 100-task Harbor gate recovered. | Dataset · Local evidence |
| cve_patches | 19 | 19 listed; validation fields absent. A 19-task oracle or independent quality gate was not recovered. | Dataset · Manifest |
These tasks have historical evidence scopes, not retrospectively assigned verified labels. In particular, the earlier 52-task commit-runtime gate does not validate the later 100-task dataset. Read the native results and solver samples.
Tasksmith and research recipes
1,330 tasks across 15 datasets. Each dataset contains complete Harbor task directories, archives and a registry pinned to its artifact revision.
| Pipeline | Tasks | Evaluation labels | Dataset and pinned manifest |
|---|---|---|---|
| swe-smith | 100 | 100 unverified | Dataset · Manifest |
| r2e | 100 | 100 unverified | Dataset · Manifest |
| swe-gen | 100 | 100 unverified | Dataset · Manifest |
| swe-next | 100 | 100 unverified | Dataset · Manifest |
| r2e-gym | 100 | 100 unverified | Dataset · Manifest |
| scaler | 100 | 100 unverified | Dataset · Manifest |
| endless-terminals | 100 | 100 unverified | Dataset · Manifest |
| cli-gym | 25 | 25 unverified | Dataset · Manifest |
| swe-flow | 100 | 2 needs repair; 98 unverified | Dataset · Manifest |
| seta-seed2synth | 100 | 100 unverified | Dataset · Manifest |
| seta-evol | 100 | 100 unverified | Dataset · Manifest |
| tmax | 55 | 55 unverified | Dataset · Manifest |
| terminalworld | 100 | 3 needs repair; 97 unverified | Dataset · Manifest |
| dataarc | 100 | 100 unverified | Dataset · Manifest |
| tasksmith | 50 | 50 verified | Dataset · Manifest |
Local collections awaiting publication
CodeMidas has 100 staged Harbor tasks, separate from the published totals above. The 2026-09-25 campaign generated 128 exports; 101 passed ordinary solver review, 26 had demonstrated verifier/instruction defects and 1 remained unresolved.
The curated collection has 200 baseline failures and 400 oracle passes, plus four reviewed Luna attempts and four Sol screens per task. All 100 retain blocked labels because adversarial checks could not run. No full-method acceptance is claimed. See the release notes and audit and source/difficulty breakdown.
FrontierSmith has 100 local construction-checked Harbor tasks across 20 problem families, measured 2026-09-29. All retain unverified labels; 21/100 final bundles have a completed blind OpenAI rollout. These artifacts have not been published and are excluded from the totals above.
What the labels establish
The Tasksmith and research-recipe release contains 50 verified, 5 needing repair and 1,275 unverified tasks. These totals exclude the historical native inventories above. Tasksmith's verified cohort came from an assisted campaign; this does not claim unattended conversion. Two SWE-flow instruction issues and three TerminalWorld verifier gaps remain explicitly diagnosed. Each dataset manifest supplies task-level labels, diagnostics and evidence scope.
Publication checks for those 15 datasets compared 214,097 file identities and parsed every selected task with Harbor. This establishes artifact integrity and format, not semantic quality of every task. See evaluation labels, yield and cost, and how to publish.