Repo2RLEnv

Related Work & Provenance

Edit on GitHub

Repo2RLEnv builds on a growing body of research into turning real software repositories into executable, verifiable environments for training and evaluating code agents. This page records two things:

  1. Provenance: exactly what each shipped pipeline draws inspiration from (these mirror the Acknowledgment blocks in the pipeline source files).
  2. Related work & other implementations: adjacent papers, datasets and frameworks worth knowing, including recent model lines from Microsoft and NVIDIA that consume verifiable code-RL environments.

Original Repo2RLEnv code is Apache-2.0. Some owned recipes retain upstream code, prompts or examples under MIT or Apache-2.0; their provenance.md files describe the exact scope and adaptations. The third-party index lists the bundled licenses and notices. SWE-RL is a non-permissive influence (CC BY-NC 4.0); our reward.py independently implements the concept and records that distinction. Referenced research, bundled material and input data have separate licensing boundaries.

Provenance: what each pipeline draws from

PipelineInspirationPaperRepo (license)
pr_diffSWE-RLSWE-RL (Wei et al., NeurIPS '25), arXiv:2502.18449facebookresearch/swe-rl (CC BY-NC 4.0) · SWE-bench/SWE-bench (MIT)
pr_runtimeSWE-bench + SWE-bench-LiveSWE-bench (Jimenez et al., ICLR '24), arXiv:2310.06770; SWE-bench-Live (Zhang et al., NeurIPS '25), arXiv:2505.23419SWE-bench/SWE-bench · microsoft/SWE-bench-Live (both MIT)
commit_runtimeR2E-Gym ("SWE-GEN" commit mining)R2E-Gym (Jain et al., COLM '25), arXiv:2504.07164R2E-Gym/R2E-Gym (MIT)
cve_patchesPatchSeeker + CVE-Bench + OSVPatchSeeker (Le et al.); CVE-Bench (Zhu et al., NAACL '25)hungkien05/PatchSeeker (MIT) · OSV API
mutation_bugsSWE-smithSWE-smith (Yang et al., NeurIPS '25 Spotlight), arXiv:2504.21798SWE-bench/SWE-smith (MIT)
code_instructMagicoder / OSS-InstructMagicoder (Wei et al., ICML '24), arXiv:2312.02120ise-uiuc/magicoder (MIT)
equivalence_testsR2ER2E (Jain et al., ICML '24)r2e-project/r2e (MIT)
refactor_synthesisRefactoringMiner (originally planned, dropped in v0.8)None (the recipe is original)tsantalis/RefactoringMiner (MIT)

Shared components (not pipelines)

ComponentInspirationLink
reward.py (diff-similarity)SWE-RLarXiv:2502.18449 · facebookresearch/swe-rl. Concept only: their reward lib is CC BY-NC 4.0, our reimplementation is Apache-2.0
bootstrap/ (LLM Docker env gen)RepoLaunchmicrosoft/RepoLaunch (MIT) · arXiv:2603.05026
_oss_instruct.pyMagicoderise-uiuc/magicoder (MIT)
_function_extractor.pyR2E (fut_extractor)r2e-project/r2e (MIT)
_mutation_operators.pySWE-smithSWE-bench/SWE-smith (MIT)

Adjacent work that isn't cited above, grouped by the technique it overlaps with. It's useful if you're deciding whether to extend a pipeline, swap a reward or feed our datasets into a trainer.

RL-environment frameworks & training stacks (the "consumption" layer)

  • NeMo Gym (NVIDIA): NVIDIA-NeMo/Gym (Apache-2.0). 100+ RLVR environments for LLMs, including Mini-SWE-Agent / OpenHands SWE agents and a Harbor harness. It's the closest analogue to Repo2RLEnv's goal and shares the Harbor ecosystem.
  • SkyRL (UC Berkeley NovaSky): NovaSky-AI/SkyRL (Apache-2.0) · SkyRL-Agent arXiv:2511.16108. A full-stack RL training framework for long-horizon, SWE-bench-style agentic tasks. It references Harbor.
  • verifiers (Prime Intellect): PrimeIntellect-ai/verifiers. Defines "environment = dataset + harness + rubric/reward", the closest competing spec to our task+verifier emission.
  • OpenEnv (Meta PyTorch + Hugging Face): meta-pytorch/OpenEnv. A Gymnasium-style step()/reset()/state() standard for containerized agentic environments, with a shared Hub. It's a standardization effort alongside our Harbor emission + HF Hub bridge.
  • rLLM (Agentica): rllm-org/rllm. "Run agent → collect traces → reward → update" post-training stack; used to train DeepSWE on R2E-Gym environments.

Automated environment setup / dockerization (the bootstrap/ problem)

  • RepoLaunch (Microsoft): arXiv:2603.05026 · microsoft/SWE-bench-Live (MIT). An LLM agent for polyglot, multi-OS automated build/test env setup, and the canonical citation for our bootstrap layer.
  • Repo2Run: arXiv:2502.13681. Automated building of executable environments for code repositories at scale. Overlaps directly with ensure_bootstrap.
  • EnvBench (JetBrains Research, ICLR '25 DL4Code): arXiv:2503.14443. 994-repo benchmark for automated env configuration with compilation/import success checks.
  • SetupBench (Microsoft): arXiv:2507.09063 · microsoft/SetupBench. 93 tasks isolating the "bootstrap a dev env from a bare OS" skill across 7 ecosystems.

SWE training datasets & RL recipes

  • SWE-Gym (Pan et al., ICML '25): arXiv:2412.21139 · SWE-Gym/SWE-Gym. 2,438 real Python tasks with executable runtimes + tests for training agents/verifiers.
  • SWE-rebench (Nebius): arXiv:2505.20411 · nebius/SWE-rebench. Continuous automated extraction of 21k+ interactive Python tasks from live repos with contamination tracking. It's the closest large-scale mining-to-RL pipeline to ours.
  • Agent-RLVR: arXiv:2506.11425. Training SWE agents via guidance + environment rewards; relevant to our verifier-graded reward design.
  • DeepSWE (Agentica + Together): agentica-org/DeepSWE-Preview. RL-only coding agent (Qwen3-32B) trained on ~4,500 R2E-Gym tasks; a reference end-to-end consumer of mined repo environments.
  • Kimi-Dev (Moonshot AI): arXiv:2509.23045. Agentless skill-prior training + outcome-driven RL on issue resolution.

Synthetic task / test synthesis

  • SWE-Flow (Alibaba Qwen, ICML '25): arXiv:2506.09003 · Hambaobao/SWE-Flow. Builds a Runtime Dependency Graph from unit tests to synthesize verifiable TDD tasks; overlaps equivalence_tests and test-anchored synthesis.
  • SWE-Mirror: arXiv:2509.08724. Mirrors real issues across other repos to produce 60K verifiable tasks in 4 languages, kept only if test-status transitions validate.
  • FEA-Bench (PKU + MSR Asia, ACL '25): arXiv:2503.06680 · microsoft/FEA-Bench. PR-mined, intent-filtered feature-implementation tasks with unit tests. Overlaps our PR mining, structural filters and test verifier.
  • OpenCodeInstruct (NVIDIA): arXiv:2504.04030 · nvidia/OpenCodeInstruct. 5M instruction samples with test cases + execution feedback, close in spirit to code_instruct.
  • Genetic-Instruct (NVIDIA): arXiv:2407.21077. Evolutionary (mutation/crossover) scaling of synthetic coding instructions; comparable to mutation_bugs' procedural generation.

Model lines trained with verifiable code RL (downstream consumers)

This is where datasets like ours end up: model families whose post-training includes RL with verifiable code rewards.


Last refreshed: 2026-06. Spotted something missing or mis-cited? PRs to this page welcome.

On this page