Repo2RLEnv

Introduction

What Repo2RLEnv builds, why it exists, and where it fits next to Harbor.

Edit on GitHub

Repo2RLEnv turns repositories, pull requests and task seeds into verifiable RL environments for coding agents. Each environment is a Harbor task (an instruction, a container, a hidden verifier and a reference solution) that you can run with any Harbor agent or feed to a trainer.

Why it exists

Reinforcement learning for coding agents needs tasks an agent can attempt and a reward a program can compute. Every task needs a reproducible environment, an instruction that doesn't give the answer away, a verifier that tells a correct change from an incorrect one, and a reference solution that proves the task is solvable.

Building these by hand takes hours per task, and a useful training run needs thousands. Repo2RLEnv generates them from material that already exists, such as merged PRs, commit history, a library's own functions, security advisories and task seeds. It also gives you the checks to decide which ones to keep.

What a Harbor task is

Harbor defines a task as a directory: instruction.md for the agent, task.toml for configuration and metadata, environment/ for the starting container, solution/ for the reference solution (the oracle), and tests/ for the verifier. Harbor runs the agent inside the environment, then runs the verifier, which writes a reward to /logs/verifier/reward.txt. Tasks walks through the layout file by file.

Three ways to generate tasks

  • Native pipelines: six built-in generators that mine a repository's PRs, commits, functions or CVEs. They run on your machine, with Docker for test-based tasks.
  • Tasksmith: an agent investigates one merged PR, builds its environment, designs the task, then reviews and repairs it. It runs on Modal or Daytona workers.
  • Research recipes: 16 in-repo implementations of published methods (SWE-smith, SWE-gen, R2E-Gym, SETA, SCALER, FrontierSmith and more) for repair, terminal and reasoning tasks. These also run on Modal or Daytona workers.

Pipelines compares every route; Choose a pipeline helps you pick one.

Repo2RLEnv and Harbor

Repo2RLEnv writes Harbor's task format directly. Its own provenance and quality labels live under [metadata.repo2env] in task.toml, so any Harbor-compatible runtime or trainer can consume the output unchanged.

Repo2RLEnvHarbor
RoleAuthors, checks and publishes tasksDefines the runtime contract and runs tasks
DoesMines sources; writes instructions, environments, verifiers and reference solutions; validates them; reviews and repairs them; records evaluation labels; pushes datasets to the Hugging Face HubSpecifies the task format; builds sandboxes (Docker, Modal, Daytona, E2B, Runloop and others); runs agent harnesses (Claude Code, Codex, Gemini CLI, OpenHands and others); collects rewards
You runrepo2rlenv generate, validate, quality run, pushharbor run

Generated is not the same as verified

An exported task is a generation result, not a guarantee. You establish its quality in two steps. First, the controls: the reference solution should score 1, and an agent that does nothing (nop) should score 0. Then the review and repair loop checks the instruction, the verifier and leakage against real learner rollouts. The outcome is stored in the task as an evaluation label, and every new task starts as unverified. Quality explains the model.

Published datasets

The Repo2RLEnv collection on the Hugging Face Hub holds six native datasets and 15 Tasksmith and recipe datasets. The 15 hold 1,330 tasks, each carrying its evaluation label; the release inventory records their counts, revisions and validation scope. Download any dataset with repo2rlenv pull and run it with harbor run.

Next steps

On this page