Repo2RLEnv
Pipelines

r2e

Implement a documented function against differential tests.

ExperimentalImplementation and reconstructionR2E recipeRuns on Modal or DaytonaDataset · 100 tasks
Edit on GitHub

The R2E recipe writes differential tests for real functions in a repository: tests that compare a candidate implementation with the original. It repairs those tests using execution and coverage feedback, then uses the observed behavior to refine the task's specification.

Pipeline, step by step

P1, P2, … mark real model calls. Unlabelled stages are code or remote execution.

Extract a real function. The recipe picks supported functions (documented, top-level and synchronous) along with a bounded slice of their module-level dependencies.

Generate and execute tests. The tests compare the function under test with reference_function through fut_module. Remote execution measures the real contrast, and the branch coverage that comes from the generated tests. The repository's existing tests run separately.

Refine the public contract. The specification call only runs once a useful test exists. It sees the function's source, the generated tests and what was observed. The learner gets the refined docstring and instruction, with the function body stubbed out.

Every prompt and its data

A candidate gets between one and max_rounds test-author calls. A successful one then gets one specification call.

CallSystem prompt compositionUser / input materialOutputRetry or branch
P1 · Tests / repairtest_prompt.md + fut_module binding and offline adaptationsfunction_name, dependency context, prior test and execution/coverage feedback.EquivalenceTest: test_codeUp to max_rounds; default three. Default minimum branch coverage is 0.8.
P2 · Specificationspecification_prompt.md + behavioral-only instruction adaptationOriginal function, generated tests, observed executions.RefinedSpecification: docstring, instructionOne call after the test-generation loop succeeds.

The complete r2e prompt reference has every retained template, appended instruction, substitution, example and output schema. The shared prompt guide shows how to inspect the fully resolved request from a real run.

Follow one task

Say the function to reconstruct consumes iterators. The generated tests must compare equivalent fresh inputs and materialize finite iterators, so the verifier measures behavior and not object identity.

What repeats, what is checked

Syntax and schema errors, failed execution, and coverage below the configured threshold all go back to P1 as feedback. The final Harbor check sits outside this loop; if it fails, the task is skipped. The private Python reference runs in the same process as the differential tests, which is a limitation for later adversarial review.

An exported bundle is a generation result. Independent leakage review, shortcut probes and blind solver traces come later, in the quality campaign.

Implementation map

Run and supported profile

Use repo2rlenv generate --config examples/owned-r2e.yaml. The native equivalence_tests pipeline keeps its original options and behavior. Setting recipe: r2e switches to this execution loop, which has its own options. Cloud workers, budget accounting, resume receipts and Rich or JSON progress work as described in the shared interface.

This first profile needs a working public GitHub Python repository with a tests/ directory. It selects documented, module-level synchronous functions and includes a bounded slice of their module-level dependencies. Classes, async functions, nonstandard source roots and reconstructed cross-module import slices need another extraction profile.

Generated tests use the native function / reference_function API. The reference and the test bindings are added only to the separate verifier image. The learner sees a stub in the real repository, plus the refined requirements. Coverage is collected only while the generated tests run. The repository's existing tests run too, but they don't count toward that coverage score. The defaults are three rounds and 80% branch coverage. These are native generation controls, separate from later quality review.

The current verifier loads the private Python reference in the same process as the differential tests. Whether a submission can reach the reference during adversarial grading is still part of the deferred audit, and no quality-accepted claim is made. The first campaign aims for 20 generated tasks with fresh Harbor baseline and reference checks.

On top of the Python source, build and test profile, the options are target, max_candidates, seed, max_rounds and min_branch_coverage. The default build dependencies include pytest and Coverage.py. If you override the dependencies, include both.

Credit: R2E (MIT), commit bcbed156711bb939de14aa46b27eee15073f5272. See RFC 0017 and the packaged recipes/r2e/provenance.md for the source map and adaptations. The execution report uses Coverage.py's branch measurement and JSON reporting.

Cost evidence

See the measured yield and cost and r2e accounting for the pilot/expansion scope, model identities, stage costs, compute resources and validation limits.

On this page