Repo2RLEnv

Owned generation recipes

Edit on GitHub

A research recipe is Repo2RLEnv's own implementation of a published task-generation method. It doesn't install or clone the upstream research code at runtime. Each recipe has its own source pin, notices, algorithm, options and RFC. Harbor and ordinary libraries are still dependencies. Target repositories are input data, and they may be cloned on the remote worker.

Pick a recipe by the input you have. Each walkthrough explains the algorithm, numbers every model call, shows where retries loop back, and links to the complete prompts. The prompt guide explains how those templates become the actual requests recorded in a campaign.

Choose a generation route

Walkthroughpipeline.name / recipeInputReference source
SWE-smithrepo_mutate / swe_smithHealthy Python repoOriginal source before mutation
SWE-genpr_to_env / swe_genExplicit merged PR URLsPR-head implementation
SWE-Flowrepo_reconstruct / swe_flowHealthy Python repoOriginal scheduled functions
CodeMidasrepo_reconstruct / codemidasWorking repository or pinned Stack v3 rowOriginal selected implementation
R2Eequivalence_tests / r2eDocumented repo functionsOriginal function in private verifier
SWE-Nextpr_runtime / swe_nextRepository PR historyPost-change source at merge revision
R2E-Gymcommit_runtime / r2e_gymFirst-parent historyPost-change source
CLI-Gymenv_repair / cli_gymHealthy development environmentGenerated, execution-checked recovery
SETA Seed2Synthterminal_synth / seta_seed2synthQuestion/answer recordsGenerated shell reference
SETA Evoltask_evolve / seta_evolOwned Harbor parentsGenerated child reference
DataArcterminal_synth / dataarcHarbor seed tasksGenerated strategy-specific reference
TMaxterminal_synth / tmaxLegacy taxonomy samplerGenerated reference guided by truth/tests
Endless Terminalsterminal_synth / endless_terminalsCategory/complexity/scenario samplerGenerated reference guided by truth/tests
TerminalWorldterminal_reconstruct / terminalworldMetadata and text transcriptExtracted, refined and replayed solution
FrontierSmithoptimization_synth / frontiersmithClosed-ended seed problemsBest sampled solution; not a proven optimum
SCALERreasoning_synth / scalerReleased family JSONAnswer from supplied reference program

pipeline.name names the generation family, and recipe picks the research-inspired implementation within it. Leave out the recipe and you get the existing native behavior. All 16 recipes are experimental. SEC-bench is deferred and not implemented.

repo2rlenv pipelines list
repo2rlenv pipelines describe repo_mutate --recipe swe_smith --json

The catalog marks each method as planned or as an executable experimental implementation. A recipe that runs isn't a claim that its output is good enough to train on. RFC 0011 covers the shared architecture and ownership policy.

Where each stage runs

The diagram shows where each kind of work happens. It isn't one stage order that every recipe follows. The history recipes establish the before-and-after contrast before they write the instruction, TerminalWorld replays before it writes tests, and SCALER has no model stage at all. SWE-smith's fresh Harbor checks are a separate campaign step. Each method's own diagram gives its exact order.

Preflight checks the source, recipe, budget and runtime hash. The model gets stage-specific system and user messages plus an output schema. Execution evidence includes test identities, rewards, logs and artifacts. The cloud worker is a Modal or Daytona sandbox running the owned wheel and Docker.

The controller on your machine handles metadata, parsing, model calls, accounting and file assembly. Image builds, and any run of target or generated code, happen remotely. Repository profiles reuse ensure_bootstrap with an explicit Dockerfile and its cache, but without the bootstrap LLM agent. Terminal builders generate their own task-specific setup and fixtures as part of the recipe. Dependencies are installed before the learner runs, and the learner runs offline.

What a task contains

<task-name>/
  instruction.md             # Learner request
  task.toml                  # Harbor runtime, resources and artifacts
  environment/
    Dockerfile               # Starting state and build-time dependencies
    ...                      # Repository snapshot or fixtures
  solution/
    solve.sh                 # Private reference entrypoint
    ...                      # Original source, recovery script or answer
  tests/
    test.sh                  # Trusted verifier entrypoint
    ...                      # Tests, expected identities and reward code
    Dockerfile               # When using a separate verifier environment

Harbor doesn't copy the whole bundle into the learner's container. It uses the environment, solution and tests, each in its own phase. Model receipts and generation traces stay outside the bundle.

Verification shapeRecipesCurrent reward
Collect allowed source into a separate verifierSWE-smith, SWE-gen, SWE-Flow, R2E, SWE-Next, R2E-GymRequired tests pass with expected nonempty identities; binary 0/1
Inspect terminal state with generated pytest testsSETA, DataArc, TMax, Endless Terminals, TerminalWorldAll required tests pass; binary 0/1. Recorded weights don't make the current grader fractional.
Check environment restorationCLI-GymHealthy tests restored and protected source preserved; binary 0/1
Collect answer file into a separate verifierSCALERNative answer equivalence; −1/+1

Generation checks only show that a task and its reference run. Independent review of instruction quality, and adversarial review, come later.

Accounting and workers

Set the budget once, explicitly. Reinitializing with a different amount is rejected. A model or provider call whose outcome is unknown keeps its reservation. Completed calls are charged from recorded usage estimates and never counted twice.

Point execution.campaign_dir in the generation config at the campaign directory. generate --max-spend-usd only applies to native generation: recipes reject it before dispatch and use the shared campaign ledger instead. --pipeline-opt overrides individual options, even when the pipeline name comes from --config.

For recipes, generate --json prints progress as JSON Lines. Inspection commands such as tasksmith show --json and quality show --json print a single JSON result. CLI failures print {"error": "ExceptionType", "message": "description"} and exit 2. Logs go to stderr, and --verbose adds a traceback there without changing the machine-readable output.

repo2rlenv campaign init workspace/my-campaign --budget-usd 25
repo2rlenv workers start --campaign workspace/my-campaign --provider modal \
  --name my-worker --reserve-usd 3 --timeout-sec 3600
repo2rlenv workers probe workspace/my-campaign/workers/my-worker.json \
  --out workspace/my-campaign/probes/first
repo2rlenv campaign status workspace/my-campaign --json
repo2rlenv workers stop workspace/my-campaign/workers/my-worker.json

Stopping a worker confirms cleanup, but it doesn't make up a bill. Reconcile the reservation with a usage or billing receipt, or with a conservative estimate that's clearly labelled as one:

repo2rlenv campaign settle workspace/my-campaign --operation worker:modal:my-worker \
  --cost-usd 0.30 --evidence workspace/my-campaign/worker-usage.json

The amount above only illustrates the command; it isn't a price quote. Campaign accounting should include failed requests, image builds and runtime. Model receipts record the request, response, schema, usage and cost basis. Keep these private review artifacts out of any task directory the learner can see.

Modal workers run Docker inside a VM. Daytona workers use Daytona's image builds and Docker-in-Docker, and account limits can differ. Modal's timeout is a maximum lifetime. Daytona's configured auto-stop is an idle timeout, so the controller also stops dispatching work once the recorded execution window ends. Always terminate workers explicitly when you're done. There's no local Docker fallback.

Output and acceptance

The release target is 100 generated tasks for each of twelve methods, plus two smaller approved collections: 55 tasks for TMax and 25 for CLI-Gym. The first milestones were 20 tasks each. Tasksmith has its own verified cohort of 50 tasks. The release inventory lists the finished datasets and the remaining counts. Detailed attack and blind-rollout audits come after generation, and they don't hold up work on the next recipe. That ordering doesn't change what a quality-accepted label means when one is given later.

Generation reports how many candidates it attempted and how many it exported. Acceptance is stricter. Every mandatory quality criterion must pass, with evidence for the current bundle hash. A missing check, a failed check or evidence for an older revision isn't a pass. A solver failing on a task doesn't, by itself, make the task invalid. These audits are still being collected, so the tasks the recipes generate are labelled exported.

Harbor 0.22.0 parses the schema 1.3 tasks the recipes emit. Some worker kernels don't support Harbor's nftables-based dynamic firewall. The owned repo2rlenv.execution.harbor_offline:OfflineDockerEnvironment adapter handles Linux Dockerfile tasks that stay offline in every phase, using Docker's network_mode: none. It rejects allowlists, network transitions and extra services defined by the task. Other task shapes need another verified runtime route.

The published research cohort has 14 recipes: SWE-smith, SETA Seed2Synth, SETA Evol, SWE-gen, SWE-Flow, R2E, TMax, Endless Terminals, TerminalWorld, CLI-Gym, DataArc, SWE-Next, R2E-Gym, and SCALER, which adds algorithmic reasoning instances from released families. CodeMidas adds source-driven reconstruction in a separate local campaign using Sol/Luna and Daytona. SEC-bench is still deferred and isn't an implemented recipe.

Contract references

On this page