Run tasks with Harbor
Download a Repo2RLEnv dataset, check it with the oracle and nop controls, score real agents locally or in cloud sandboxes, and read the rewards.
Repo2RLEnv writes Harbor tasks; Harbor runs them. Use this guide to evaluate an agent on a published dataset or on tasks you generated yourself: check that each task is sound, run an agent, and read the reward of every trial.
Before you start
- Docker, running, for local runs (
--env docker). Cloud sandboxes don't need it. - Repo2RLEnv, to download datasets:
pip install repo2rlenv. - An API key for the model your agent uses, such as
ANTHROPIC_API_KEYfor Claude Code.
Install Harbor
uv tool install harborResearch-recipe and Tasksmith tasks use Harbor's schema 1.3 and were checked with Harbor 0.22.0, the version the repo2rlenv[harbor] extra pins. To run exactly that version, install harbor==0.22.0. For a cloud sandbox, add Harbor's extra for that provider, such as uv tool install 'harbor[modal]' (see Run in a cloud sandbox).
Download a dataset
Pull a published dataset from the Hugging Face Hub. pull flattens the published tasks/ layout into one directory per task, which is what harbor run -p expects.
repo2rlenv pull FineEnvs/repo2rlenv-pr-runtime ./pr-runtimeThe release index lists every published dataset. To try a single task first, add --task encode__httpx-3367. Tasks you generated yourself need no download: point -p at your --out directory.
Datasets published with repo2rlenv release (the research recipes and Tasksmith) also ship tasks.tar.gz. It preserves executable file modes and each task's bundle hash, so their dataset cards download it instead:
hf download FineEnvs/repo2rlenv-swe-smith tasks.tar.gz --repo-type dataset --local-dir ./swe-smith
tar -xzf ./swe-smith/tasks.tar.gz -C ./swe-smithThe tasks land in ./swe-smith/tasks/. Those cards run them on Daytona; see Run in a cloud sandbox.
Check the tasks with the controls
Before you score an agent, run the two controls. The oracle agent applies the reference solution, so every task should score 1.0. The nop agent changes nothing, so every task should score 0.0.
harbor run -p ./pr-runtime -a oracle --env docker
harbor run -p ./pr-runtime -a nop --env dockerA task where the oracle scores below 1.0, or nop scores above 0.0, can't tell a correct solution from an absent one. Leave it out of your evaluation. Troubleshooting covers the usual causes.
The first run of a test-based dataset pulls or builds each task's image, so expect it to take longer than later runs.
Run a real agent
This runs Claude Code with Sonnet, capped at $2 of agent spend per task:
harbor run -p ./pr-runtime \
-a claude-code -m anthropic/claude-sonnet-4-6 \
--ak max_budget_usd=2.00 \
--ae ANTHROPIC_API_KEY=$ANTHROPIC_API_KEY \
--env docker| Flag | What it does |
|---|---|
-p PATH | A dataset directory, or a single task directory. |
-a AGENT | The agent harness. oracle and nop are the controls. |
-m PROVIDER/MODEL | The model the harness uses. |
--ak KEY=VALUE | An agent option. max_budget_usd is Claude Code's spend cap. |
--ae KEY=VALUE | An environment variable for the agent's container. |
--ve KEY=VALUE | An environment variable for the verifier. See Pass API keys. |
--env ENV | Where trials run: docker (default) or a cloud sandbox. |
Every flag that takes KEY=VALUE can be repeated.
Other agent harnesses
Swap -a and -m to run another harness, and pass that harness's own provider key with --ae. Harbor ships, among others: codex, gemini-cli, copilot-cli, cursor-cli, aider, goose, openhands, openhands-sdk, opencode, mini-swe-agent, swe-agent, qwen-coder, kimi-cli, trae-agent and its reference agents terminus-1 and terminus-2. harbor run --help lists every harness in your Harbor version. Agent harnesses explains what each one accepts and how rollout traces leave the sandbox for RL training.
harbor run -p ./pr-runtime \
-a openhands -m openai/gpt-4o \
--ae OPENAI_API_KEY=$OPENAI_API_KEY \
--env dockerPass API keys
Harbor runs the agent and the verifier in separate phases, and each gets only the variables you pass to it.
--ae(agent env) carries the model key the harness needs:ANTHROPIC_API_KEYforclaude-code,OPENAI_API_KEYfor OpenAI models, and so on.--ve(verifier env) matters only forpr_difftasks. Their reward includes an optional LLM judge, which runs at verification time:
harbor run -p ./pr-diff \
-a claude-code -m anthropic/claude-sonnet-4-6 \
--ak max_budget_usd=2.00 \
--ae ANTHROPIC_API_KEY=$ANTHROPIC_API_KEY \
--ve ANTHROPIC_API_KEY=$ANTHROPIC_API_KEY \
--env dockerWithout a judge key, the pr_diff verifier still scores: it records judge_status = "no_api_key" and renormalizes the weights of the other components. To use a self-hosted judge instead of Anthropic, pass --ve R2E_JUDGE_ENDPOINT=… and --ve R2E_JUDGE_MODEL=…. The component weights are also verifier variables (R2E_W_*); see pr_diff reward tuning.
Test-based tasks (pr_runtime, commit_runtime, cve_patches, the research recipes and Tasksmith) grade by running tests and need no verifier keys.
Run a subset
Filters apply to task directory names, and -l is applied after them.
| Flag | Effect |
|---|---|
-l N | Run at most N tasks: the first N after the other filters, not a random sample. |
-i GLOB | Include only matching tasks, such as -i 'encode__httpx-*'. Repeatable. |
-x GLOB | Exclude matching tasks. Repeatable. |
-n N | Run N trials at a time (default 4). |
-k N | Run each task N times, for pass@k and variance. |
harbor run -p ./pr-runtime -i 'pallets__*' -l 10 -a oracle --env dockerTo evaluate only tasks whose quality has been established, list them by evaluation label first: repo2rlenv tasks list ./pr-runtime --status verified.
Run in a cloud sandbox
Replace --env docker with a hosted sandbox provider to run many trials in parallel without local Docker. Install Harbor with that provider's extra and set its credentials in your shell:
--env | Install | Credentials |
|---|---|---|
modal | uv tool install 'harbor[modal]' | modal setup, or MODAL_TOKEN_ID and MODAL_TOKEN_SECRET |
daytona | uv tool install 'harbor[daytona]' | DAYTONA_API_KEY |
e2b | uv tool install 'harbor[e2b]' | E2B_API_KEY |
runloop | uv tool install 'harbor[runloop]' | RUNLOOP_API_KEY |
harbor run -p ./swe-smith/tasks -a oracle --env daytona -n 16These sandboxes run Linux containers. Harbor's own environment docs list further providers.
Read the results
Each harbor run is a job. Harbor writes it to ./jobs/<timestamp>/ (set --job-name to name it and -o to change the parent directory), with one directory per trial:
The job's result.json groups statistics by agent, model and dataset. Its metrics hold the mean reward, and reward_stats maps each reward value to the trials that earned it. For one reward per task, read the trial files:
jq -r '[.task_name, (.verifier_result.rewards.reward // "error")] | @tsv' jobs/my-run/*/result.jsonA trial with no reward failed before verification. Its exception_info says why. To browse trajectories in a web UI, run harbor view jobs.
What a reward means depends on the pipeline:
| Tasks | reward.txt | Details |
|---|---|---|
pr_diff | Weighted diff similarity in [0, 1] | reward-details.json: each component's score, the weights and the judge status |
pr_runtime, commit_runtime, cve_patches | f2p_rate × p2p_rate in [0, 1] | reward-details.json: strict resolved (every fail-to-pass test passes and no pass-to-pass test regresses), pass counts, regressions |
code_instruct, equivalence_tests | 1 or 0 (the test command's exit status) | None |
| Research recipes (except SCALER) and Tasksmith | 1 or 0 | result.json: whether the required tests passed, with per-test statuses. CLI-Gym writes reason.txt when protected source was changed. |
| SCALER recipe | −1 or +1 | result.json: the score and answer accuracy |
Rewards explains each design, and the reward schema documents every field.
Limits
An agent's score means something only if the task is sound, and a freshly exported task has passed nothing beyond its generator's own checks. Run the controls on every dataset you evaluate (the oracle should score 1 and nop 0), and prefer tasks labeled verified when you report a result. Quality explains how a task gets that label.