Repo2RLEnv

Choose a pipeline

Match your source material, reward and infrastructure to the right generator.

Edit on GitHub

Repo2RLEnv has six native pipelines, Tasksmith, and 16 research recipes. Which one fits depends on what you start from, what kind of reward you need, and where you can run containers.

Four questions

  1. What's your source material? A merged PR, a repository's commit history, a repository's own functions and APIs, a security advisory, or seeds such as question/answer pairs, existing Harbor tasks, terminal recordings or problem families.
  2. Do you need a test-based reward? Diff similarity scores an agent's patch against the real one without running any code. Test-based rewards run the repository's tests, so generation has to build an environment where they pass.
  3. Where can you run containers? Native pipelines build and check environments with Docker on your machine. Tasksmith and research recipes run on remote Modal or Daytona workers.
  4. How much churn can you absorb? pr_diff, pr_runtime and commit_runtime are stable. Everything else is experimental: its interface and output can change between releases.

Compare the routes

RouteInputThe agent's taskRewardLLM at generationRuns onStatus
pr_diffMerged PRs (GitHub, GitLab)Reproduce the change a PR madeDiff similarity, 0–1, with an optional LLM judge at verify timeNoYour machine; no Docker to generateStable
pr_runtimeMerged PRs that add tests (GitHub, GitLab)Make failing tests pass without breaking passing onesGraded F2P × P2P, plus a strict resolved flagYes: bootstrap, cached per repositoryLocal DockerStable
commit_runtimeCommit history (GitHub, GitLab, local path)The same, mined from commits, with the instruction rewritten as a symptomGraded F2P × P2PYes: bootstrap, plus one call per task for the instructionLocal DockerStable
code_instructA Python repository's sourceSolve an LLM-authored problem grounded in the repository's APIsTests pass, 0/1Yes: problem, test and solutionLocal DockerExperimental
equivalence_testsPython functions in a repositoryReimplement a function so it matches a hidden referenceTests pass, 0/1Yes: the equivalence testsLocal DockerExperimental
cve_patchesOSV advisories and their fix commits (GitHub, Python ecosystem)Repair the vulnerabilityGraded F2P × P2PYes: bootstrap, plus a proof-of-concept test when the fix ships noneLocal DockerExperimental
TasksmithA merged PR that changes testable Python sourceImplement the PR's behavior from a human-style requestPrivate behavioral tests, 0/1Yes: a Pi or OpenCode agent, plus review and repairModal or DaytonaExperimental
Research recipesRepositories, PR or commit history, question/answer seeds, Harbor tasks, terminal recordings, problem familiesRepair, reconstruction, terminal, reasoning and optimization tasksTests pass, 0/1; SCALER scores −1/+1; FrontierSmith scores 0–1 continuouslyMost recipes; SCALER is fully programmaticModal or DaytonaExperimental

Graded F2P × P2P is the fraction of fail-to-pass tests the agent fixes times the fraction of pass-to-pass tests it keeps green. Rewards defines every reward kind. The bootstrap is a one-time LLM agent loop that builds a working container for a repository; --max-spend-usd caps it (default $5).

If you want…

  • For a first dataset in minutes, without Docker or an LLM key, use pr_diff. The Quickstart walks through it.
  • For SWE-bench-style tasks graded by the repository's own tests, use pr_runtime.
  • For a repository with few PRs, or one where fixes are committed straight to the main branch, use commit_runtime.
  • For tasks about using a library rather than its history, use code_instruct.
  • For function-level reimplementation tasks, use equivalence_tests locally, or the r2e recipe on remote workers.
  • For security repair tasks, use cve_patches.
  • For one carefully built task per PR, even when the PR's own tests don't cover the change, use Tasksmith. It can write private behavioral tests, and it reviews and repairs each task before export.
  • For many repair tasks from one healthy repository, use swe_smith, which introduces source defects for the agent to fix.
  • For terminal and shell tasks, use seta_seed2synth, tmax, endless_terminals or terminalworld. For broken development environments, use cli_gym.
  • For harder variants of Harbor tasks you already have, use seta_evol or dataarc.
  • For reasoning tasks with no repository at all, use scaler.
  • For open-ended optimization with a graded objective and no known optimum, use frontiersmith.
  • To reproduce a specific published method, use its recipe. Research recipes maps each method to its pipeline name.

Run your choice

Each family starts differently:

FamilyStart with
Native pipelinerepo2rlenv generate --repo org/service --pipeline <name> --out ./tasks, adding --llm provider/model for every pipeline except pr_diff
Research reciperepo2rlenv generate --config <recipe>.yaml, with an execution: block that names a running worker, a runtime wheel and a campaign directory. The repository's examples/owned-*.yaml files are starting points. See Remote execution.
Tasksmithrepo2rlenv tasksmith run <panel>.json --campaign <dir> --output <dir> --runtime-wheel <wheel>, following the one-PR walkthrough

repo2rlenv pipelines list shows every route with its status, and repo2rlenv pipelines describe <pipeline> --recipe <recipe> shows a recipe's input, scope and provenance.

Known limits

  • Status is about the code, not the tasks. A stable pipeline can still emit a weak task. Every exported task starts with the evaluation label unverified; its controls and the review and repair loop decide whether you should train on it.
  • limit caps candidates, not output. Each pipeline discards candidates that fail its filters, and test-based routes also drop candidates whose tests don't flip from failing to passing. A repository whose test suite doesn't run cleanly in a container yields almost nothing on any test-based route. Pipelines lists typical yield per pipeline.
  • Language and host coverage varies. pr_diff works on any language. The bootstrap for pr_runtime and commit_runtime detects Python, Node, Go, Rust, Java and C/C++. code_instruct, equivalence_tests, cve_patches and Tasksmith target Python. cve_patches needs GitHub; generate rejects an unsupported source before doing any work.
  • Remote routes have no local fallback. Research recipes, Tasksmith and the quality loop need a Modal or Daytona worker.
  • SEC-bench is deferred. It appears in pipelines list as planned and isn't implemented.

On this page