Pipelines
Every way Repo2RLEnv turns repositories, pull requests and task seeds into Harbor tasks, grouped by the kind of task you get.
A pipeline turns source material, such as a repository, its pull requests or an existing task, into Harbor tasks. They're grouped here by the kind of task they produce. Each card shows the reward, whether the pipeline runs on your machine or on a remote worker, and its status. If you're not sure which to use, start with Choose a pipeline.
Tasksmith
Point an agent at a merged pull request and get a verified Harbor environment back. Tasksmith investigates the repository, builds its environment, designs the task, then reviews and repairs its own work until the controls pass.
- 01
Investigate
Reads the PR, its diff and the repository around it.
- 02
Bootstrap
Builds the repository’s environment on a remote worker.
- 03
Design
Writes the instruction and a private verifier.
- 04
Construct
Assembles the Harbor task; the merged code is the oracle.
- 05
Review and repair
Runs the controls and a blind solver, then fixes what fails.
Repository repair
7 pipelinesFix real code in a real repository. The repository’s own tests decide the reward. Tasksmith, above, is the agentic route.
Fix the issue a merged PR fixed; the PR’s own tests stay hidden
Fix the bug a commit fixed, from an LLM-written symptom report
Patch a published vulnerability from its stripped advisory
Fix a seeded single-site defect in a healthy Python repository
Repair a historical change mined from merged PRs
Repair a historical change mined from first-parent commits
Implementation and reconstruction
6 pipelinesWrite or restore functionality, graded by hidden or differential tests.
Write a module for an LLM-authored problem anchored in the repository’s API
Reimplement a stubbed real function to match its reference
Implement a documented function against differential tests
Reimplement functions removed from a working Python repository
Restore a merged PR’s source change, from explicit PR URLs
Rebuild a removed feature from its written behavioral contract
Patch similarity
1 pipelineReproduce a real change. The patch is scored against the merged one, no test suite needed.
Terminal tasks
7 pipelinesWork in a shell to reach a state a verifier can check.
Terminal task synthesized from a question-and-answer seed
Evolved variant of an existing Harbor task
Augmented variant of a Harbor seed task
Terminal task combining three to five sampled skills
Terminal task from a sampled category, complexity and scenario
Task reconstructed from a real terminal recording
Repair a deliberately broken development environment
Reasoning and optimization
2 pipelinesProblems without a repository: exact answers, or open objectives with graded scores.
How you run each kind
| Kind | Command | Where it runs |
|---|---|---|
| Native pipeline | repo2rlenv generate --repo <repo> --pipeline <name> | Your machine. pr_diff needs neither Docker nor an LLM; the others need Docker and --llm to bootstrap the repository once. |
| Tasksmith | repo2rlenv tasksmith run <panel> --campaign … --output … --runtime-wheel … | A Modal or Daytona worker. See Tasksmith. |
| Research recipe | repo2rlenv generate --config <file> with pipeline.name and pipeline.recipe set | A Modal or Daytona worker; there's no local fallback. See Run research recipes. |
The native datasets are historical inventories, and native results records what was and wasn't validated. The recipe and Tasksmith datasets hold 1,330 tasks in 15 datasets, each with per-task evaluation labels. CodeMidas and FrontierSmith collections are staged locally and reported separately. See published datasets and yield and cost.
Planned or deferred
| Proposal | What it would do | Status |
|---|---|---|
SEC-bench (cve_patches recipe) | Reconstruct a vulnerable build and reproduction input, then emit an isolated repair task | Deferred; not implemented |
RFC 0007: native pr_to_env | One Harbor environment per PR URL you supply, verified like pr_runtime | Draft |
RFC 0008: env_setup | Make a bare repository's test suite build and run from scratch | Draft |
RFC 0009: test_synthesis | Write the test that catches a broken variant of a function | Draft |
RFC 0010: issue_runtime | Mine closed issues linked to fixing PRs, graded by the PR's tests | Draft |
Where to go next
- Tasks and Rewards explain the task layout and reward kinds every pipeline shares.
- Run with Harbor runs any of these tasks with the oracle, a no-op agent or a real agent.
- Quality and verification covers how a generated task becomes a verified one.
- Add a pipeline walks through the pipeline contract, options model, tests and doc page.