equivalence_tests
Extract real functions from a repository and have an LLM write tests that check a reimplementation against the original.


Each task stubs out a real function from the repository. The agent gets its
signature and docstring and must implement it so it behaves exactly like
reference_<name>, the original kept alongside it. An LLM writes the equivalence
test; the pipeline keeps it only if it fails against the stub and passes against
the original. Because the ground truth is working code rather than a model's
invention, these tasks vary less than code_instruct's.
At a glance
| Input | A Python repository on GitHub, GitLab or a local path |
| Task | Implement a stubbed function in task_module.py so its outputs match reference_<name> |
| Reward | Binary: 1.0 if the hidden equivalence test passes, else 0.0 |
| Needs an LLM | Yes: the one-time bootstrap (cached) and one call per test attempt |
| Needs Docker | Yes, for bootstrap, validation and running tasks |
| Hosts | GitHub, GitLab and local paths |
| Languages | Python only (module-level functions). On GitHub, the primary language is checked first (--force-language skips the check) |
| Status | Experimental |
| Reference dataset | FineEnvs/repo2rlenv-equivalence-tests: 100 tasks from seven utility libraries |
Quickstart
repo2rlenv generate \
--repo pytoolz/toolz \
--pipeline equivalence_tests \
--pipeline-opt limit=5 \
--pipeline-opt seed=42 \
--llm anthropic/claude-sonnet-4-6 \
--out ./tasks/toolz-equivalenceThe example uses pytoolz/toolz because the pipeline only accepts
self-contained, side-effect-free functions: utility libraries have many, while
framework code such as pallets/click has few. limit is the number of tasks to
emit; seed makes the candidate order repeatable. The first run bootstraps the
repository (capped by --max-spend-usd, default 5.0, and cached). Each task lands
in ./tasks/toolz-equivalence/pytoolz__toolz-eqv-<hash>/. Run the oracle, which
should score 1.0:
harbor run -p ./tasks/toolz-equivalence -a oracle --env dockerHow it works
- Bootstrap the repository once and shallow-clone it at
--ref. See Bootstrap. - Extract candidates. Walk files matching
file_globand notexclude_glob, and keep module-level functions that take at least one argument, have a body ofmin_loctomax_loclines, return a value, show no side effects (file, network, process, logging, printing, clock, randomness, framework context), and reference only their own arguments and locals, builtins and a small set of standard-library modules. Async functions and names that start with_ortest_(or aremain,setup,run,init,cli,wrapper) are skipped. Candidates are shuffled. - Write the test. The LLM gets the function's source and writes 5–10
test_*functions, each assertingname(x) == reference_name(x)on one input. - Gate the test. Reject it if it doesn't import and use both names, has fewer than five test functions, has a test function that doesn't call both, contains a constant assert, or duplicates an earlier test suite.
- Verify in the sandbox. Build
task_module.pytwice, with type annotations stripped so the module imports on its own. Withnamestubbed to raiseNotImplementedError, the test must not pass. Withnameset to the original implementation, it must pass. - Retry with feedback. When a gate or sandbox check fails, the next attempt
includes the reason and the last 1,200 characters of the failure log, up to
max_attempts_per_functionattempts. - Emit the task. The stub module is baked into the image; the instruction shows the signature and docstring only.
Options
Pass each option with --pipeline-opt key=value.
| Key | Default | What it does |
|---|---|---|
limit | 50 | Number of tasks to emit |
min_loc / max_loc | 5 / 60 | Function body size, in lines |
file_glob | **/*.py | Files to extract functions from |
exclude_glob | tests, test_*, *_test.py, conftest.py, docs/, examples/, __init__.py, setup.py | Files never used |
seed | none | Random seed for the candidate order |
max_attempts_per_function | 3 | Test-writing attempts per function, each with feedback from the last failure |
llm_temperature | 0.5 | Sampling temperature. Lower than code_instruct's, for stable tests |
max_llm_tokens | 1500 | Output token limit per attempt |
require_test_fails_with_stub | true | Reject tests that pass against the stub |
require_test_passes_with_oracle | true | Reject tests that fail against the original |
validation_timeout_sec | 90 | Timeout for each sandbox test run |
skip_validation | false | Emit without running the sandbox checks. For debugging |
Output
task.toml: Harbor 1.0 task.referencelinks to the function's lines;[metadata.repo2env.equivalence_tests]records the function name, source path and lines, body size, argument names, test file name, bootstrap image andllm_cost_usd. The evaluation label starts asunverified.instruction.md: the function's signature and docstring, where to implement it, and how grading works.environment/Dockerfile:FROMthe bootstrap image, writing/workspace/task_module.pywithreference_<name>and a stub<name>.solution/patch.diff: replaces the stub with the original implementation (the oracle);solve.shapplies it.tests/test_r2e_<hash>.py: the equivalence test, delivered by Harbor at verify time.tests/test.sh: copies the test into/workspaceand runspython -m pyteston it.
The output directory also gets a .debug_skips/<function>/ folder for each
rejected candidate, with its last test and sandbox logs. These aren't tasks.
Reward
tests/test.sh writes 1.0 to /logs/verifier/reward.txt when every equivalence
assertion holds and 0.0 otherwise. There is no partial credit and no
reward-details.json. A no-op agent scores 0, because the stub raises. See
Rewards.
Yield and cost
Completed generation summaries account for at least 200 candidate functions for the 100 reference tasks, including runs on mpmath, setuptools and black that produced none. Other runs lack a final summary, so the overall yield is unknown. Productive runs recorded at least $2.51 in synthesis cost (at least $0.025 per task), a partial floor that excludes the empty runs, bootstrap and compute.
Solver samples on the reference dataset used different tasks per model, so they aren't a leaderboard:
| Model and agent | Tasks | Outcome |
|---|---|---|
| Claude Sonnet 4.6, Claude Code | 5 | 4 scored 1 after setup retries; 1 failed while installing the agent |
| GPT-5.3-Codex, Codex | 5 | 5 scored 1 |
| Qwen3.6-35B-A3B, OpenHands SDK | 5 | 5 scored 1 |
See native results. What moves yield:
repository shape first, since the purity and self-containment filters leave few
candidates in framework code; then the model's choice of inputs the original
handles cleanly, which the feedback loop improves. Cost scales with candidates ×
max_attempts_per_function calls.
Limits
- A generated task isn't a verified environment. Run the controls: the oracle
should score 1.0 and a no-op agent (
-a nop) 0. Record the outcome as an evaluation label; see Quality. - The reference is visible by design.
reference_<name>sits in the agent'stask_module.py, and the instruction says reading it is the intended way to solve the task. The test checks only that outputs match, so a solution that calls or copies the reference also scores 1.0. Treat these as function-reconstruction exercises. - Equality is the only check. Functions whose results don't compare with
==, or that raise on most inputs, rarely make it through; the test covers only the inputs the model chose. - Module-level functions only. Methods, async functions and anything that touches files, network, time or global state are excluded.
- Python and pytest only.
llm_cost_usdis cumulative for the run, not the cost of that task.- Tasks depend on a local image until you publish them with
repo2rlenv push.
Related
- RFC 0005: equivalence_tests
- Reference dataset and its generation evidence
r2e: the research recipe with coverage-guided test repair and a private verifiercode_instruct: tasks invented from a snippet instead of extracted- Tasks, Rewards and Run with Harbor
- Adapted from R2E (Jain et al., 2024): the function filters and the
reference_<name>test pattern. No code is copied.
Implementation notes
Source: pipelines/equivalence_tests.py,
pipelines/_function_extractor.py
and pipelines/_eval_script.py.
The module the agent starts from looks like this:
def reference_<name>(...):
... # the original implementation, annotations stripped
def <name>(<args>):
raise NotImplementedError("implement <name>")- The rename to
reference_<name>rewrites the AST, including recursive calls, so a recursive reference calls itself rather than the stub. - Type annotations are stripped from both functions because annotations such as
def f(x: Argument) -> FCname repository types that don't exist in the standalone module, which would fail at import. - Before any sandbox run, both modules are compiled and their top-level names
resolved against builtins (
stub_module_not_importable,oracle_module_not_importable). - Skip reasons include
llm_parse_failed,test_missing_both_names,too_few_test_functions:<n><5,test_fn_missing_both_names:<test>,trivial_assert_present,duplicate_task,test_passes_with_stubandoracle_does_not_satisfy_test. - The duplicate check hashes the function name with the whitespace-normalized, lower-cased test body.
- The gold patch fills in
<name>in the baked module rather than adding files, so the reference is present for every agent, not only Harbor's oracle agent.