Repo2RLEnv

Published Harbor datasets

Edit on GitHub

Published results cover the six native pipelines, Tasksmith and 14 research recipes. Local collections are reported separately and are not included in published totals.

Browse every dataset in the Repo2RLEnv collection on the Hugging Face Hub.

Native pipelines

600 task entries across 6 earlier datasets. Recovered from cached Hub manifests and local publication stagings on 2026-09-15. These are historical snapshots, not a fresh Hub recount. Older revisions and duplicate stagings are excluded; cross-pipeline content is not deduplicated.

PipelineTasksRecovered validation evidenceDataset and evidence
pr_diff181181 listed; no per-task validation in this snapshot. Earlier 100-task release reported oracle passes.Dataset · Manifest
pr_runtime100100 oracle/tracked passes; 88 clean commands; 87 with a regression guard.Dataset · Manifest
commit_runtime100100 generation-time verified stamps. Full 100-task oracle gate not recovered; older 52-task gate is separate.Dataset · Manifest
code_instruct100100 emitted after generation checks; no separate 100-task Harbor gate recovered.Dataset · Local evidence
equivalence_tests100100 emitted after generation checks; no separate 100-task Harbor gate recovered.Dataset · Local evidence
cve_patches1919 listed; validation fields absent. A 19-task oracle or independent quality gate was not recovered.Dataset · Manifest

These tasks have historical evidence scopes, not retrospectively assigned verified labels. In particular, the earlier 52-task commit-runtime gate does not validate the later 100-task dataset. Read the native results and solver samples.

Tasksmith and research recipes

1,330 tasks across 15 datasets. Each dataset contains complete Harbor task directories, archives and a registry pinned to its artifact revision.

PipelineTasksEvaluation labelsDataset and pinned manifest
swe-smith100100 unverifiedDataset · Manifest
r2e100100 unverifiedDataset · Manifest
swe-gen100100 unverifiedDataset · Manifest
swe-next100100 unverifiedDataset · Manifest
r2e-gym100100 unverifiedDataset · Manifest
scaler100100 unverifiedDataset · Manifest
endless-terminals100100 unverifiedDataset · Manifest
cli-gym2525 unverifiedDataset · Manifest
swe-flow1002 needs repair; 98 unverifiedDataset · Manifest
seta-seed2synth100100 unverifiedDataset · Manifest
seta-evol100100 unverifiedDataset · Manifest
tmax5555 unverifiedDataset · Manifest
terminalworld1003 needs repair; 97 unverifiedDataset · Manifest
dataarc100100 unverifiedDataset · Manifest
tasksmith5050 verifiedDataset · Manifest

Local collections awaiting publication

CodeMidas has 100 staged Harbor tasks, separate from the published totals above. The 2026-09-25 campaign generated 128 exports; 101 passed ordinary solver review, 26 had demonstrated verifier/instruction defects and 1 remained unresolved.

The curated collection has 200 baseline failures and 400 oracle passes, plus four reviewed Luna attempts and four Sol screens per task. All 100 retain blocked labels because adversarial checks could not run. No full-method acceptance is claimed. See the release notes and audit and source/difficulty breakdown.

FrontierSmith has 100 local construction-checked Harbor tasks across 20 problem families, measured 2026-09-29. All retain unverified labels; 21/100 final bundles have a completed blind OpenAI rollout. These artifacts have not been published and are excluded from the totals above.

What the labels establish

The Tasksmith and research-recipe release contains 50 verified, 5 needing repair and 1,275 unverified tasks. These totals exclude the historical native inventories above. Tasksmith's verified cohort came from an assisted campaign; this does not claim unattended conversion. Two SWE-flow instruction issues and three TerminalWorld verifier gaps remain explicitly diagnosed. Each dataset manifest supplies task-level labels, diagnostics and evidence scope.

Publication checks for those 15 datasets compared 214,097 file identities and parsed every selected task with Harbor. This establishes artifact integrity and format, not semantic quality of every task. See evaluation labels, yield and cost, and how to publish.

On this page