Input / Output Spec
There are two contracts: what you feed Repo2RLEnv (input) and what comes out (output). The input has the same shape for every pipeline, and pipeline-specific knobs go under pipeline.options (the "kwargs"). The output is pure Harbor with a namespaced [metadata.repo2env] extension.
Input contract: GenerationInput
The single root model is src/repo2rlenv/spec/input.py:GenerationInput. The CLI is a thin shim that builds this object from flags and/or a YAML/TOML config file.
class GenerationInput(BaseModel):
spec_version: Literal["0.1.0"] = "0.1.0"
repo: RepoSpec # source repository
pipeline: PipelineSpec # synthesis method + kwargs
llm: LLMSpec # model driving synthesis
output: OutputSpec # where the dataset lands
qa: QASpec = QASpec() # quality gate (default: diff_parse only for lite)
sandbox: SandboxSpec = SandboxSpec() # execution backend (default: none for lite)
auth: AuthSpec = AuthSpec() # secret references, resolved from envSub-models
| Model | Required fields | Notes |
|---|---|---|
RepoSpec | url | access ∈ {public, private, auto}, optional auth_token_env, ref defaults to HEAD |
PipelineSpec | name, options | name is an enum (see pipelines/); options is validated against the named pipeline's Options model with extra="forbid" |
LLMSpec | provider, model | provider/model resolves to a LiteLLM identifier; endpoint (CLI: --llm-endpoint) targets a self-hosted vLLM/Ollama server, api_key_env (CLI: --llm-key-env) names a non-default key var |
OutputSpec | destination, org, dataset_name | destination is a local path; publish separately via repo2rlenv push |
QASpec | (none) | Defaults to [diff_parse] for the lite path; full pipelines opt into [determinism, oracle_consistency, llm_judge, false_negative] |
SandboxSpec | (none) | See "Sandbox model" below: none for lite, harbor for full pipelines (delegates), local/e2b for lite consumer-side runners |
AuthSpec | (none) | Names of env vars only; values are never stored |
Two equivalent invocation forms
Flag form:
# Generate locally, then push as two explicit steps
repo2rlenv generate \
--repo huggingface/trl \
--pipeline pr_diff \
--pipeline-opt limit=5 \
--llm anthropic/claude-sonnet-4-6 \
--out ./datasets/trl-r2e-v0-1
repo2rlenv push ./datasets/trl-r2e-v0-1 <your-org>/trl-r2e-v0-1Config-file form: --config <path> accepts YAML or TOML and detects the format from the extension. CLI flags override file fields:
spec_version: "0.1.0"
repo:
url: "huggingface/trl"
access: "auto"
pipeline:
name: "pr_diff"
options:
limit: 5
skip_drafts: true
llm:
provider: "anthropic"
model: "claude-sonnet-4-6"
output:
destination: "./datasets/trl-r2e-v0-1"
org: "<your-org>"
dataset_name: "trl-r2e-v0-1"
visibility: "public"Drop this into a file (e.g. repo2rlenv.config.yaml) and run with --config <path>. Publishing is a separate step via repo2rlenv push.
Output contract: Harbor + [metadata.repo2env]
Every pipeline emits standard Harbor task directories. Repo2RLEnv-specific provenance goes into a namespaced subtable inside Harbor's existing [metadata].
Task directory layout
<dataset>/<task_id>/
├── task.toml # Harbor-native + [metadata.repo2env]
├── instruction.md # natural-language prompt
├── solution/patch.diff # oracle (lite) — for diff_similarity scoring
├── environment/Dockerfile # OPTIONAL — emitted by sandbox-required pipelines AND by pr_diff (emit_harbor_env=True)
└── tests/ # OPTIONAL — emitted alongside environment/Dockerfile
├── test.sh # the eval script (Harbor mounts tests/ at /tests in the container)
├── verifier.py # pr_runtime graded F2P/P2P verifier (plain file, read by test.sh)
├── f2p.json # pr_runtime FAIL_TO_PASS test-name list
└── p2p.json # pr_runtime PASS_TO_PASS test-name listpr_runtime ships its verifier + F2P/P2P lists as plain tests/ files
(not base64 inside test.sh); test.sh is a thin orchestrator that reads
them from /tests. The dataset root also carries a manifest.json,
auto-generated by repo2rlenv push, with one machine-checkable row per task
(repo, PR, ref, reward kinds, F2P/P2P counts, difficulty).
pr_diff generates without a sandbox, but by default (emit_harbor_env=True) it still emits an environment/Dockerfile (thin python:3.12-slim) + a tests/test.sh carrying its 6-component diff-similarity verifier, so the task is runnable directly via harbor run. With emit_harbor_env=False only the first three files exist. That's pure text, and the consumer scores the diff themselves.
task.toml example
version = "1.0"
[task]
name = "huggingface__trl-5705"
org = "<your-org>"
description = "..."
[metadata]
difficulty = "medium"
category = "bugfix"
[metadata.repo2env]
spec_version = "0.2.0"
pipeline = "pr_diff"
pipeline_version = "0.1.0"
repo = "huggingface/trl"
ref = "f39373edcd7a..." # base commit SHA
reference = "https://github.com/huggingface/trl/pull/5705"
source_access = "public"
built_at = "2026-05-06T..."
synthesis_llm = "anthropic/claude-sonnet-4-6"
content_hash = "sha256:..."
reward_kinds = ["diff_similarity"]
[metadata.repo2env.pr_diff]
pr_merged_at = "2026-05-05T13:46:07Z"
diff_format = "unified"
context_files = ["trl/trainer/dpo_trainer.py", ...]
# v0.2.0+ only — sandbox-required tasks (pr_runtime / commit_runtime / ...)
# carry this subtable so consumers know exactly what they're getting.
[metadata.repo2env.reproducibility]
mode = "registry" # registry | inline_dockerfile | local_only
image_ref = "ghcr.io/huggingface/r2e-bootstrap-pallets-click@sha256:..."
image_tag = "ghcr.io/huggingface/r2e-bootstrap-pallets-click:a1b2c3d4e5f6-7d8e9f01"
image_visibility = "public" # public | private | unknown
pushed_at = "2026-05-19T11:30:00Z"
pushed_by = "huggingface"
# Inline-mode-only fields (omitted in registry mode):
# inline_recipe_sha256 = "sha256:..."
# inline_recipe_lines = 47
# inline_recipe_source = "agent_replay" # or "user_dockerfile"
# fallback_reason = "no working registry credentials (ghcr.io: L2 auth failed)"
[agent]
timeout_sec = 1800.0
[verifier]
timeout_sec = 300.0Each pipeline writes its own subtable under [metadata.repo2env.<name>] carrying provenance specific to how the task was made. The per-pipeline docs describe each schema.
Dataset-level layout (HF Hub)
When pushed via repo2rlenv push, the dataset on the Hub looks like:
huggingface.co/datasets/<owner>/<name>/
├── README.md # auto-generated dataset card
├── registry.json # Harbor's legacy registry format, pinned to a commit SHA
└── tasks/
└── <task_id>/...registry.json lets any Harbor consumer pull tasks directly:
harbor download <dataset-name> \
--registry-url https://huggingface.co/datasets/<owner>/<name>/resolve/main/registry.jsonImplementation: src/repo2rlenv/hub.py:push_to_hub.
Sandbox model: we don't have one
Repo2RLEnv has no sandbox abstraction of its own. Generation-time execution and consumption-time execution both go through external tools:
| Phase | Pipeline class | What runs the code |
|---|---|---|
| Generation | Lite (text-only generation, e.g. pr_diff) | Nothing: pure text manipulation (no sandbox at gen time) |
| Generation | Full (pr_runtime, commit_runtime, etc.) | Harbor's sandbox layer (harbor invoked under the hood) |
| Consumption | pr_diff (default emit_harbor_env=True) | harbor run against the thin python:3.12-slim env + baked 6-component verifier |
| Consumption | pr_diff (emit_harbor_env=False) or any stored-diff task | from repo2rlenv.reward import calculate_diff_similarity_reward (pure Python, no sandbox) |
| Consumption | Full | harbor run -d <dataset> -e <modal|daytona|e2b|local|runloop> ... |
SandboxSpec exists to describe what the pipeline needs (provider, GPU, network), and at gen-time we lower it onto Harbor's flags. We don't ship a parallel runner. This keeps the surface area small, since Harbor already handles GPU, multi-container, parallelism and provider auth.
GPU
class GPUSpec(BaseModel):
count: int = 1
kind: Literal["any", "a10g", "a100", "h100", "l4", "t4"] = "any"GPU only matters for sandbox-required pipelines on ML repos. For example, mining huggingface/trl with full pr_runtime skips most of the interesting PRs unless the verifier sandbox has a GPU, because the trainer tests require CUDA.
Lite pipelines never use this field. When set on a harbor-provider sandbox, we pass it through to the Harbor backend's GPU config (Modal A100 / H100 / etc.).
Reward kinds
[metadata.repo2env.reward_kinds] is a list naming the reward types this task supports. Two are defined for v0.1:
| Kind | What it is | Where the oracle lives |
|---|---|---|
diff_similarity | Similarity between the predicted and oracle unified diffs (float ∈ [0,1]). pr_diff scores this with a 6-component verifier (format / size / file-targeting / region-overlap / changes-only similarity / LLM-judge) written to /logs/verifier/reward.txt; the simpler stdlib SequenceMatcher path is available for emit_harbor_env=False tasks. | solution/patch.diff |
test_execution | Shell verifier writes a float to /logs/verifier/reward.txt | tests/test.sh |
A task may emit both. The lite pipeline emits only diff_similarity; full sandbox-required pipelines emit test_execution (and may also emit diff_similarity if they capture the oracle as a diff).
The diff-similarity reward function lives in src/repo2rlenv/reward.py:calculate_diff_similarity_reward. It's pure stdlib (difflib.SequenceMatcher) and Apache-2.0, with no SWE-RL CC-BY-NC code vendored.
Image distribution (v0.2.0+)
Sandbox-required tasks (pr_runtime, commit_runtime, …) ship an environment/Dockerfile whose FROM <ref> line points at the bootstrap image, the working Docker environment for the source repo. At generate time the ref is local/r2e-bootstrap/..., which no other machine can pull. repo2rlenv push rewrites it in place to one of two reproducible forms:
| Mode | FROM ref | Reproducibility |
|---|---|---|
registry | ghcr.io/<owner>/r2e-bootstrap-<slug>@sha256:... (or ECR / ACR / GCP AR / Docker Hub equivalent) | Bit-exact: the registry digest is immutable |
inline_dockerfile | full apt-get / pip / ... recipe baked into environment/Dockerfile, no FROM <registry> reference | Recipe-level: assumes mirrors stay stable; rebuilds from scratch on every harbor run |
The mode that was chosen is recorded in [metadata.repo2env.reproducibility] (see the task.toml example above), along with pushed_at, pushed_by and, for inline mode, inline_recipe_source ∈ {user_dockerfile, agent_replay}.
repo2rlenv push decides which mode to use by running the OCI Distribution Spec L1–L4 probe protocol against every registry it finds credentials for in ~/.docker/config.json. The probe never pushes anything to your registry. It confirms reachability, auth, read and write access by starting a blob-upload session and cancelling it straight away. Run repo2rlenv push --check-auth to see the probe output for your machine.
Push flags:
| Flag | Behaviour |
|---|---|
| (default) | Auto-detect a verified registry; fall back to inline-Dockerfile mode with a warning if none |
--image-registry <prefix> | Force a specific registry (e.g. ghcr.io/myorg); probed for write access before push |
--inline-dockerfile | Skip image push; bake the recipe into each task. Recipe-level reproducibility. |
--require-registry | Hard-fail if no verified registry is available (CI / launch mode). No silent fallback. |
--skip-image-push | Rewrite tasks against a remote ref that already exists at the registry. No docker push. |
--image-visibility public|private|inherit | Visibility for the pushed image (GHCR auto-flips via the GitHub API). Default: match dataset. |
--check-auth | Probe every detected registry and exit. --fast skips L3/L4; --json for CI. |
Conformance
A task or dataset is conformant to v0.1 if and only if:
repo2rlenv validate <path>exits 0task.tomlis valid TOML and contains[task].name- The named
[metadata.repo2env.pipeline]matches a registered pipeline solution/patch.diffexists and is non-empty (lite pipelines)- For sandbox-required pipelines:
environment/Dockerfileandtests/test.shexist
v0.2 adds:
- Sandbox-required tasks carry
[metadata.repo2env.reproducibility]withmode ∈ {registry, inline_dockerfile, local_only}.local_onlyis pre-publication, so external consumers should NOT treat these tasks as reproducible.
Deep validation
repo2rlenv validate <path> only checks that each task.toml parses and names its task. --deep also reads each task's assets and metadata; --oracle implies --deep and adds the oracle checks. Errors fail the run (exit 1); warnings are printed but don't. Implementation: src/repo2rlenv/validation.py.
| Check | Applies to | Severity |
|---|---|---|
instruction.md exists and is non-blank | every single-step task | error |
An environment definition exists: environment/Dockerfile, environment/docker-compose.yaml, or [environment].docker_image (Harbor's own rule) | Repo2RLEnv tasks that are runnable (test_execution reward, or an environment/ / tests/ dir); any task with an environment/ dir | error |
tests/test.sh (tests/test.bat when [environment].os = "windows") exists and is non-blank | runnable Repo2RLEnv tasks; any task with a tests/ dir | error |
solution/patch.diff exists and is non-blank | text-only diff_similarity tasks (pr_diff with emit_harbor_env=False), since it's their only oracle | error |
tests/verifier.py parses as Python; tests/f2p.json is a non-empty JSON list of strings; tests/p2p.json is a JSON list of strings | pr_runtime / commit_runtime / cve_patches tasks with a non-empty fail_to_pass (graded reward) | error |
tests/f2p.json matches fail_to_pass in task.toml | same | warning |
tests/<test_filename> exists and is non-blank | code_instruct / equivalence_tests | error |
reproducibility.mode is registry, inline_dockerfile or local_only; typed fields (image_ref, image_visibility, inline_recipe_*, …) have valid values | tasks carrying the subtable | error |
registry mode has a non-empty image_ref; inline_dockerfile mode has environment/Dockerfile | per mode | error |
registry mode's FROM line matches image_ref | per mode | warning |
spec_version >= 0.2.0 task with environment/Dockerfile but no reproducibility subtable | Repo2RLEnv tasks | warning |
Unknown pipeline or reward_kinds entry | Repo2RLEnv tasks | warning |
--oracle: a solve script exists; native pipelines also require solution/patch.diff as a non-blank unified diff (format check skipped for diff_format = "search_replace") | Repo2RLEnv tasks; named recipes can use script-only solutions, but any supplied patch.diff is still checked | error |
What deep validation deliberately does not do:
- Require
solution/without--oracle. Harbor solutions are optional. - Check the executable bit. Harbor
chmod +xes scripts itself. - Reject text-only
pr_diffoutput for lackingenvironment/andtests/. That's a supported shape. - Prove the task runs. It's static: it can't tell whether the Dockerfile builds, whether a shell script's dependencies exist, or whether the patch applies at
base_commit. Pre-v0.2 datasets without a reproducibility subtable, and non-Repo2RLEnv Harbor tasks, only get the checks for assets they actually ship; multi-step[[steps]]tasks skip the layout checks.harbor run --agent oracleremains the ground truth.
Versioning
Before 1.0 the spec is a moving target, and minor bumps may break readers. After 1.0 we honor strict SemVer (additive minors, breaking majors only). Each released spec version freezes its JSON Schema at a stable URL.
v0.2.0 (v0.8.2.post3)
Adds [metadata.repo2env.reproducibility]. It's an additive change: v0.1.0 readers ignore the new subtable; pre-v0.2 datasets that lack it pass validation unchanged but aren't portable across machines.