Repo2RLEnv

Quickstart

Generate, validate and run your first Harbor environment from a real GitHub repository in about five minutes.

Edit on GitHub

Turn recent merged PRs from pallets/click into Harbor tasks, check them, and score them with Harbor. You'll use the pr_diff pipeline, which needs no Docker and no LLM key to generate. Running the tasks in step 5 does need Docker.

You need Python 3.12 or later, Git, and the GitHub CLI (gh).

Install Repo2RLEnv

pip install repo2rlenv

Then sign in to GitHub. Repo2RLEnv uses gh to list merged PRs and fetch their diffs.

gh auth login

If you'd rather not sign in interactively, set GITHUB_TOKEN to a read-scoped token instead. Installation covers extras and other credentials.

Generate tasks from three PRs

repo2rlenv generate \
  --repo pallets/click \
  --pipeline pr_diff \
  --pipeline-opt limit=3 \
  --out ./tasks

This lists the three most recently merged PRs, drops those that make weak tasks (more than five changed files, test-only or docs-only diffs, reverts, diffs under three changed lines), and writes one task per remaining PR to ./tasks. The summary tells you what happened:

candidates  3
emitted     1
skipped     2 ({'too_many_files': 2})

limit counts PRs listed, not tasks emitted, so you often get fewer tasks than the limit. If nothing is emitted, the command exits with status 1. Raise the limit (for example, limit=10) and run it again.

Look at what you made

Each task is a directory named after its source PR:

task.toml
instruction.md
Dockerfile
patch.diff
solve.sh
test.sh
verifier.py
oracle.patch
instruction.md
FileWhat it is
task.tomlHarbor configuration (version = "1.0", the task name default/pallets__click-<pr>, timeouts) plus [metadata.repo2env]: the source PR, base commit, reward kind, difficulty and an evaluation label that starts as unverified.
instruction.mdWhat the agent reads: the PR's title and description rewritten as an issue, with links to the PR and "Closes #N" references removed.
environment/DockerfileA python:3.12-slim image with the repository checked out at the PR's base commit and later history removed.
solution/patch.diffThe merged PR's diff: the reference solution, or oracle.
solution/solve.shApplies patch.diff. Harbor's oracle agent runs this script.
tests/test.shThe verifier entry point: captures the agent's changes as a diff against the base commit and runs the scorer.
tests/verifier.pyThe scorer. It writes the reward to /logs/verifier/reward.txt and a per-component breakdown to reward-details.json.
tests/oracle.patchThe verifier's copy of the reference diff.
tests/instruction.mdThe verifier's copy of the instruction, used by the optional LLM judge.

Harbor uploads tests/ only when it verifies, so the agent never sees the reference diff. The reward is a weighted diff similarity between 0 and 1. It combines file targeting, region overlap, line similarity and patch size, plus an optional LLM judge. pr_diff documents the weights.

Validate the tasks

repo2rlenv validate ./tasks --oracle

This is a static check. It confirms each task.toml parses, the instruction, environment and verifier files exist, and solution/solve.sh and solution/patch.diff are present. It doesn't build the image or run anything. The next step does.

✓ default/pallets__click-<pr>
✓ all 1 tasks valid

Run the tasks with Harbor

Install Harbor and run the oracle agent first. It applies the reference solution, so every task should score 1.0; a lower score means the task is broken.

uv tool install harbor
harbor run -p ./tasks -a oracle --env docker

The nop agent changes nothing, so every task should score 0.0. Together these two runs are the task's controls.

harbor run -p ./tasks -a nop --env docker

Now run a real agent, Claude Code with Sonnet:

export ANTHROPIC_API_KEY=sk-ant-...
harbor run -p ./tasks \
  -a claude-code -m anthropic/claude-sonnet-4-6 \
  --ak max_budget_usd=2.00 \
  --ae ANTHROPIC_API_KEY=$ANTHROPIC_API_KEY \
  --ve ANTHROPIC_API_KEY=$ANTHROPIC_API_KEY \
  --env docker
FlagEffect
-a claude-code -m anthropic/claude-sonnet-4-6The agent harness and the model it uses.
--ak max_budget_usd=2.00An agent option: caps what Claude Code can spend on each task (its --max-budget-usd).
--ae KEY=VALUEPasses an environment variable to the agent.
--ve KEY=VALUEPasses an environment variable to the verifier. Here it turns on the pr_diff LLM judge; without it the judge is skipped and the other components are reweighted.

Harbor writes results to ./jobs/<timestamp>/, with each trial's verifier/reward.txt and verifier/reward-details.json. Browse them with harbor view jobs. To run in a cloud sandbox instead of local Docker, replace --env docker with --env modal, daytona, e2b or runloop.

Push to the Hub (optional)

hf auth login
repo2rlenv push ./tasks <you>/click-pr-diff

push creates the dataset repository (add --private for a private one) and uploads the tasks with a dataset card, a manifest.json and a Harbor-compatible registry.json. pr_diff tasks build from a public base image, so no container registry is involved. Anyone can download the dataset again with repo2rlenv pull <you>/click-pr-diff.

Next steps

pr_diff rewards textual closeness to the merged change. For rewards based on running the repository's tests, switch to pr_runtime. It needs Docker and an LLM key: an LLM agent builds a working container for the repository and caches the image for later runs. That bootstrap loop is capped at $5 by default (--max-spend-usd).

repo2rlenv generate \
  --repo pallets/click \
  --pipeline pr_runtime \
  --pipeline-opt limit=10 \
  --llm anthropic/claude-sonnet-4-6 \
  --out ./tasks-runtime

On this page