Skip to content

Repo2RLEnv

Turn any GitHub repository into a verifiable RL training and evaluation dataset. End-to-end synthesis → standardize → train + eval, in the Harbor task format.

pip install repo2rlenv

repo2rlenv generate \
    --repo pallets/click \
    --pipeline pr_runtime \
    --pipeline-opt limit=10 \
    --llm anthropic/claude-sonnet-4-6 \
    --out ./datasets/click-pr-runtime
  • Quickstart


    Install, generate your first dataset, and push it to the Hugging Face Hub — in about 10 minutes.

    Start here

  • Pipelines


    Six synthesis pipelines out of the box — from PR mining to LLM-authored coding tasks — each with a graded, verifiable reward.

    Browse the pipelines

  • Reference datasets


    Every pipeline ships with a published HF dataset that you can pull, benchmark against, or use as-is for training.

    Open the collection

  • RFCs


    Design docs for every pipeline. Read the why behind the shape, or draft a new one — the template is in-repo.

    Open the RFC index

What's inside

Three layers — the first is ours, the other two we delegate to Harbor:

Layer Repo2RLEnv ships We rely on
Generation src/repo2rlenv/pipelines/ — six pipelines, gate helpers, quality audits
Spec The [metadata.repo2env] extension to task.toml — provenance so a dataset can be regenerated exactly Harbor's task spec
Consumption HF Hub push bridge (repo2rlenv push), Harbor-compatible registry.json Harbor's full runtime — 5 sandbox backends × 22 agent harnesses

Repo2RLEnv is synthesis-only — we generate the datasets and let Harbor run them. No parallel sandbox runtime; no bespoke evaluation harness.

Pipelines at a glance

Pipeline Task shape Reward Status Reference dataset
pr_diff agent writes a patch matching a real PR's diff 6-component diff-similarity + LLM judge stable repo2rlenv-pr-diff
pr_runtime SWE-bench-style: agent's patch flips F2P tests to green graded F2P × P2P stable repo2rlenv-pr-runtime
commit_runtime commit-level SWE-Gym-style tasks graded F2P × P2P stable repo2rlenv-commit-runtime
code_instruct LLM-authored coding task anchored to a real repo's API binary test_execution experimental repo2rlenv-code-instruct
equivalence_tests R2E-style: agent implements a function equivalent to a frozen reference binary test_execution experimental repo2rlenv-equivalence-tests
cve_patches OSV-driven CVE → fix-commit as a task graded F2P × P2P experimental repo2rlenv-cve-patches

Four more pipelines in the RFC queue: pr_to_env · env_setup · test_synthesis · issue_runtime.

Consuming a dataset

Any dataset from the collection runs end-to-end with harbor run:

uv tool install harbor
repo2rlenv pull AdithyaSK/repo2rlenv-pr-runtime ./workspace/pr-runtime
harbor run --path ./workspace/pr-runtime --agent oracle --env docker
# Every task should score 1.0 with the oracle agent.

harbor run --path ./workspace/pr-runtime --agent claude-code \
    --model anthropic/claude-sonnet-4-6 \
    --sample 10 --backend docker