Repo2RLEnv

Reward Schema Reference

Edit on GitHub

Every task emitted by Repo2RLEnv writes its reward to /logs/verifier/ inside the container after the agent's patch is applied and the verifier runs. Harbor reads reward.txt as the primary training signal. Native pipelines with a graded verifier also write reward-details.json, the full breakdown for analysis, filtering, and debugging.

Files written per task

FileAlways present?Content
/logs/verifier/reward.txt✅Single float on one line, e.g. 0.842731
/logs/verifier/reward-details.jsonPipeline-dependent (see below)Full breakdown as JSON (pr_diff, pr_runtime, commit_runtime, cve_patches)
/logs/verifier/result.jsonResearch recipes and TasksmithThe grader's outcome, e.g. {"passed": true, "returncode": 0, "statuses": {…}}; SCALER writes {"score": …, "accuracy": …}. CLI-Gym writes reason.txt instead when protected source was changed.

pr_diff

Reward kind: diff_similarity

reward.txt: weighted combination of 6 components, clamped to [0, 1].

reward-details.json:

{
  "final_reward": 0.74,
  "components": {
    "format_valid":   1.0,
    "size_sanity":    0.85,
    "file_targeting": 0.67,
    "region_overlap": 0.60,
    "similarity":     0.55,
    "llm_judge":      0.80
  },
  "weights": {
    "format_valid":   0.00,
    "size_sanity":    0.08,
    "file_targeting": 0.12,
    "region_overlap": 0.20,
    "similarity":     0.10,
    "llm_judge":      0.50
  },
  "judge_model":  "claude-haiku-4-5-20251001",
  "judge_endpoint": null,
  "judge_status": "ok"
}

Field reference:

FieldTypeDescription
final_rewardfloat [0, 1]Final score, the same value as reward.txt. Hard-capped to ≤ 0.40 when size_sanity < 0.10 (catastrophic-size guard); the breakdown has no separate flag for the cap, so check components.size_sanity.
components.format_valid0 or 1Predicted output parses as a valid unified diff. Weight 0.00: kept as a guard, not a scoring factor.
components.size_sanity[0, 1]min(oracle_loc, pred_loc) / max(oracle_loc, pred_loc). Catches severe over- or under-generation.
components.file_targeting[0, 1]F1 over the sets of changed files (not Jaccard; F1 gives partial credit for TP).
components.region_overlap[0, 1]Predicted hunks overlap oracle hunks within a 5-line slack.
components.similarity[0, 1]difflib.SequenceMatcher ratio over +/- lines only (no credit for context lines).
components.llm_judge[0, 1] or nullLLM semantic judge ("does this address the issue?"). null when disabled or the API key is absent, in which case its weight is redistributed to the remaining components.
weightsobjectEffective per-component weights (overridable via R2E_W_* env vars).
judge_modelstring or nullModel used for llm_judge as the serving API names it (e.g. claude-haiku-4-5-20251001, Qwen/Qwen3.5-4B), or null if the judge did not score.
judge_endpointstring or nullThe OpenAI-compatible server the judge was routed to (R2E_JUDGE_ENDPOINT), or null on the default Anthropic route.
judge_statusstring"ok" | "no_api_key" | "no_judge_model" | "empty_predicted" | "network" | "parse" | "missing_score"

Weight override env vars (set inside the verifier container via --ve): R2E_W_FORMAT, R2E_W_SIZE, R2E_W_FILE, R2E_W_REGION, R2E_W_SIM, R2E_W_JUDGE

Judge routing env vars (also via --ve): R2E_JUDGE_MODEL, R2E_JUDGE_ENDPOINT, R2E_JUDGE_API_KEY. See ENV.md.


pr_runtime · commit_runtime · cve_patches

Reward kind: test_execution (primary) + diff_similarity (fallback)

reward.txt: f2p_rate × p2p_rate, rounded to 6 decimal places.

reward-details.json (normal path: test output parsed successfully):

{
  "reward":                  0.833333,
  "resolved":                false,
  "command_resolved":        false,
  "f2p_total":               3,
  "f2p_passed":              2,
  "f2p_rate":                0.666667,
  "p2p_total":               5,
  "p2p_passed":              5,
  "p2p_rate":                1.0,
  "regressions":             [],
  "untracked_failed_count":  1,
  "untracked_failed":        ["tests/test_other.py::test_legacy"],
  "parse_status":            "ok",
  "runner":                  "pytest",
  "tests_parsed":            12,
  "exit_code":               1
}

reward-details.json (fallback path: test output unrecognised, no per-test status):

{
  "reward":            1.0,
  "resolved":          false,
  "command_resolved":  true,
  "parse_status":      "fallback_exitcode",
  "eval_trustworthy":  false,
  "runner":            "",
  "f2p_total":         3,
  "p2p_total":         5,
  "exit_code":         0
}

Field reference:

FieldTypeDescription
rewardfloat [0, 1]Dense training signal: f2p_rate × p2p_rate. Gold patch → 1.0.
resolvedboolSWE-bench eval signal. At least one F2P test is declared, all F2P tests pass AND all P2P tests pass. Gold patch → true. Use this for benchmark-style scoring.
command_resolvedboolStricter: resolved AND no untracked failures AND exit_code == 0. Filters tasks with pre-existing flaky tests outside F2P/P2P.
f2p_totalintNumber of declared FAIL_TO_PASS tests.
f2p_passedintF2P tests that passed after the agent's patch.
f2p_ratefloat [0, 1]f2p_passed / f2p_total. 0.0 when f2p_total == 0.
p2p_totalintNumber of declared PASS_TO_PASS tests.
p2p_passedintP2P tests still passing after the agent's patch.
p2p_ratefloat [0, 1]p2p_passed / p2p_total. 1.0 when p2p_total == 0 (no regression guard).
regressionslist[str]P2P tests that broke under the agent's patch.
untracked_failed_countintTests that failed but were not in F2P or P2P, often from pre-existing flakiness.
untracked_failedlist[str]Names of untracked failures (capped at 20).
parse_statusstring"ok": per-test status parsed successfully. "fallback_exitcode": runner output unrecognised, so reward is binary and based on the exit code.
eval_trustworthyboolOnly present in fallback_exitcode path. false when an F2P oracle exists but couldn't be verified.
runnerstringDetected test runner: "pytest" | "go" | "cargo" | "jest" | ""
tests_parsedintTotal tests detected in log output (only in parse_status == "ok").
exit_codeintRaw exit code of the test command.

Choosing between reward, resolved, and command_resolved:

  • Training: use reward (dense, graded signal from reward.txt).
  • Benchmark / leaderboard: use resolved (strict SWE-bench tracked resolution, unaffected by untracked flakiness).
  • Strict eval (clean command required): use command_resolved.

code_instruct · equivalence_tests

Reward kind: test_execution

reward.txt: 1.0 (tests pass) or 0.0 (tests fail). Binary.

reward-details.json: not written by these pipelines.

The verifier runs the repo's test suite (or the generated pytest equivalence tests) and maps the exit code directly to reward. No per-component breakdown.

Reading the reward for training:

reward = float(open("/logs/verifier/reward.txt").read().strip())  # 1.0 or 0.0

Per-pipeline summary

Pipelinereward.txtreward-details.jsonTraining signalEval signal
pr_diffweighted [0,1]✅rewardcomponents breakdown
pr_runtimef2p × p2p✅rewardresolved / command_resolved
commit_runtimef2p × p2p✅rewardresolved / command_resolved
cve_patchesf2p × p2p✅rewardresolved / command_resolved
code_instruct1.0 / 0.0❌rewardreward
equivalence_tests1.0 / 0.0❌rewardreward

On this page