Repo2RLEnv

Agent harnesses + how RL traces leave the sandbox

Edit on GitHub

This page covers what happens when someone runs a Harbor task that Repo2RLEnv emitted: which agent harnesses can run it, what inputs each agent accepts (LLM endpoint, model and so on), and how token IDs and logprobs get out of the sandbox and into a trainer. That last path is what makes RL training possible.

All of this is Harbor's responsibility, not Repo2RLEnv's. It's still worth understanding, because it decides which consumers can use our datasets out of the box.

TL;DR

  • Harbor ships 25 agent harnesses: ≈22 distinct coding agents, 2 testing helpers and 1 alias (§1).
  • Every agent extends BaseAgent with shared params (logs_dir, model_name, mcp_servers) plus declarative CliFlag / EnvVar mappings (§2).
  • The LLM never runs inside the task sandbox. The agent makes outbound HTTPS calls to wherever it's hosted, a cloud API or your tunneled vLLM (§3).
  • Token IDs and logprobs are captured inside the agent's container to /logs/agent/, then read back via volume mount or SDK download into trial_result.agent_result.rollout_details (§4).
  • Terminus 2 reports rollout details itself. For other agents, an OpenAI-compatible proxy in front of vLLM does the job instead (§5).

1. The 25 built-in agents

The canonical list is the enum in src/harbor/models/agent/name.py.

Real coding agents (22)

CLI-based proprietary agents:

  • claude-code: Anthropic's official CLI
  • codex: OpenAI Codex CLI
  • gemini-cli: Google
  • copilot-cli: GitHub
  • cursor-cli: Cursor
  • rovodev-cli: Atlassian Rovo
  • kimi-cli: Moonshot
  • qwen-coder: Alibaba (the enum constant is QWEN_CODE, the value is "qwen-coder")

Open-source coding agents:

  • aider
  • goose: Block's Goose
  • openhands and openhands-sdk: OpenHands (formerly OpenDevin)
  • cline-cli: Cline as a CLI
  • swe-agent: Princeton SWE-agent
  • mini-swe-agent: a minimal SWE-agent variant
  • opencode
  • trae-agent: ByteDance Trae
  • hermes
  • pi
  • nemo-agent: NVIDIA NeMo

Harbor-native:

  • terminus-1 and terminus-2: Harbor's reference agents. They use a single tmux tool and capture logprobs and token IDs natively, which makes them RL-friendly.

Testing harnesses (2)

  • oracle: applies the ground-truth solution. Use it to check that a task is internally consistent; the oracle should always score 1.0.
  • nop: does nothing. Use it to debug the harness pipeline itself.

Plus 1 alias

  • terminus: an alias that resolves to one of the terminus variants

Architecture: how they plug in

In src/harbor/agents/factory.py, _AGENT_MAP: dict[AgentName, type[BaseAgent]] is built dynamically from an _AGENTS list. Custom (third-party) agents register with harbor run --agent-import-path my_pkg.agents:MyAgent, so you don't have to patch Harbor.

Two base patterns:

Base classWhen to useExamples
BaseAgentAgent runs as an external process and talks to the container via bash execmost CLIs
BaseInstalledAgentAgent installs into the container with install() and runs there directlyopenhands, terminus-2, claude-code (in some configs)

2. Agent contract

Shared params (every agent)

class BaseAgent(ABC):
    SUPPORTS_ATIF: ClassVar[bool] = False     # Harbor's trajectory format
    SUPPORTS_WINDOWS: ClassVar[bool] = False

    def __init__(
        self,
        logs_dir: Path,                          # written to /logs/agent inside container
        model_name: str | None = None,           # LiteLLM-style "provider/model"
        logger: logging.Logger | None = None,
        mcp_servers: list[MCPServerConfig] | None = None,  # task-declared MCP tools
        skills_dir: str | None = None,           # path to skills config inside container
        *args, **kwargs,
    ): ...

Declarative kwarg → CLI flag / env var mapping

For installed agents (CLI tools that run inside the container), Harbor uses two declarative dataclasses to wire kwargs to the agent's launch:

@dataclass
class CliFlag:
    kwarg: str
    cli: str                                              # e.g. "--max-turns"
    type: Literal["str", "int", "bool", "enum"]
    choices: list[str] | None = None
    default: Any = None
    env_fallback: str | None = None                       # also resolves from this env
    format: str | None = None

@dataclass
class EnvVar:
    kwarg: str
    env: str                                              # e.g. "ANTHROPIC_API_KEY"
    type: Literal["str", "int", "bool", "enum"]
    # ...
    bool_true: str = "true"
    bool_false: str = "false"

A subclass declares CLI_FLAGS: list[CliFlag] and the base class generates the agent's launch command from it. That's how Harbor avoids repeating launch logic across 22+ agents.

Concrete: what claude-code accepts

class ClaudeCode(BaseInstalledAgent):
    SUPPORTS_ATIF: bool = True

    CLI_FLAGS = [
        CliFlag("max_turns",            cli="--max-turns",            type="int",  env_fallback="CLAUDE_CODE_MAX_TURNS"),
        CliFlag("reasoning_effort",     cli="--effort",               type="enum", choices=["low","medium","high","xhigh","max"]),
        CliFlag("thinking",             cli="--thinking",             type="enum", choices=["enabled","adaptive","disabled"]),
        CliFlag("thinking_display",     cli="--thinking-display",     type="enum", choices=["summarized","omitted"]),
        CliFlag("max_thinking_tokens",  cli="--max-thinking-tokens",  type="int",  env_fallback="MAX_THINKING_TOKENS"),
        CliFlag("max_budget_usd",       cli="--max-budget-usd",       type="str"),
        CliFlag("fallback_model",       cli="--fallback-model",       type="str"),
        CliFlag("append_system_prompt", cli="--append-system-prompt", type="str"),
        # ... more
    ]

The Claude Code CLI resolves ANTHROPIC_API_KEY itself; Harbor just passes the environment through.

Concrete: what terminus-2 accepts (RL-targeted)

This is the agent that matters for RL. terminus-2 is built to capture token IDs and logprobs natively.

class Terminus2(BaseInstalledAgent):
    def __init__(
        self,
        logs_dir: Path,
        model_name: str | None = None,
        max_turns: int | None = None,
        parser_name: str = "json",                          # or "xml"
        api_base: str | None = None,                        # ← self-hosted endpoint URL
        temperature: float | None = None,
        reasoning_effort: Literal[
            "none","minimal","low","medium","high","xhigh","max","default"
        ] | None = None,
        collect_rollout_details: bool = False,              # ← FLIP THIS ON FOR RL
        session_id: str | None = None,
        enable_summarize: bool = True,
        proactive_summarization_threshold: int = 8000,
        max_thinking_tokens: int | None = None,
        model_info: dict | None = None,                     # max_input_tokens, cost-per-token, ...
        trajectory_config: TrajectoryConfig | None = None,  # raw_content, linear_history (for SFT export)
        tmux_pane_width: int = 160,
        tmux_pane_height: int = 40,
        store_all_messages: bool = False,
        record_terminal_session: bool = True,
        interleaved_thinking: bool = False,
        suppress_max_turns_warning: bool = False,
        use_responses_api: bool = False,
        llm_backend: LLMBackend | str = LLMBackend.LITELLM,  # litellm | openai-direct | hf-router | ...
        llm_kwargs: dict | None = None,                      # passed to backend constructor
        llm_call_kwargs: dict[str, Any] | None = None,       # per-call (top_p, top_logprobs, ...)
        extra_env: dict[str, str] | None = None,
    ): ...

Two of these matter most when you host the LLM yourself:

  • api_base: points at vLLM / SGLang / Ollama / your cloudflared tunnel
  • llm_backend: picks the Python client driver (LiteLLM by default, with adapters for direct providers)

3. LLM hosting

The LLM never runs inside the task sandbox. The agent process inside the sandbox makes outbound HTTPS calls to wherever the LLM is hosted.

Where the LLM runsHow the agent reaches it
Cloud API (Anthropic / OpenAI / etc.)model_name="anthropic/claude-sonnet-4-6" + ANTHROPIC_API_KEY env. No api_base.
Self-hosted vLLM / SGLang on your laptop or tunnelmodel_name="vllm/Qwen3.5-4B" + api_base="https://your-tunnel/v1" + OPENAI_API_KEY="dummy"
HF Inference Router (Together / Nscale / Scaleway)model_name="huggingface/Qwen/...:together" + HF_TOKEN; LiteLLM auto-points at the router
Bedrock / Vertex / etc.LiteLLM provider strings + provider-specific env auth

Implications

  • Sandbox network policy must allow outbound (or run with network: open for build phase)
  • The LLM endpoint sees only the agent's prompts, never the verifier's reward
  • A self-hosted LLM behind your tunnel keeps both code AND prompts internal
  • Cold-API agents (Claude Code, Codex CLI) need their provider env var passed through task.toml's [environment].env map

4. RL traces + logprobs: how data leaves the sandbox

This is the pipe that makes RL training possible. It has four layers.

Layer 1: agent stores rollout per turn

Inside terminus_2.py, on each LLM response:

if response.prompt_token_ids is not None:
    rollout_detail["prompt_token_ids"] = [response.prompt_token_ids]
if response.completion_token_ids is not None:
    rollout_detail["completion_token_ids"] = [response.completion_token_ids]
if response.logprobs is not None:
    rollout_detail["logprobs"] = [response.logprobs]

These accumulate in self._rollout_details: list[RolloutDetail]. They're only collected when collect_rollout_details=True. It's off by default because the data is large (~500B/token, see the HF inference doc).

Layer 2: written to /logs/agent/ inside the container

By Harbor convention, the agent's logs_dir is /logs/agent. At the end of each turn the rollout dicts are serialized to JSONL there, next to:

  • trajectory.jsonl: per-step ATIF format (chat history, tool calls, observations)
  • tmux_session.cast: a full asciinema recording of the terminal (Terminus 2 specifically)

Layer 3: Harbor pulls them out post-trial

There are two paths, depending on the sandbox provider's EnvironmentCapabilities.mounted:

ProvidermountedHow rollout data exits
Local Docker✅ TrueBind mount: Harbor reads /logs/agent/* directly from the sandbox host
Apple Container✅ TrueBind mount
Daytona✅ TrueVolume mount at /harbor/logs/agent on the sandbox
Islo✅ TrueBind mount + CA bundle for TLS
Modal / E2B / Runloop❌ FalseHarbor calls the provider SDK's download_file() to pull each file

Both paths populate the same end state: trial_result.agent_result.rollout_details is a list of RolloutDetail Pydantic objects, each with prompt_token_ids: list[list[int]], completion_token_ids: list[list[int]], logprobs: list[list[float]].

Layer 4: trainer consumes

trial = job.run(...)
for trial_result in trial.results:
    reward = trial_result.verifier_result.rewards.get("reward", 0)   # float ∈ [0, 1]
    for rd in trial_result.agent_result.rollout_details:             # one per turn
        prompt_ids = rd.prompt_token_ids                             # list[list[int]]
        completion_ids = rd.completion_token_ids                     # list[list[int]]
        logprobs = rd.logprobs                                       # list[list[float]]
        # Feed into TRL / SkyRL / Prime-RL policy gradient computation

Every Harbor-compatible trainer uses this shape today: harbor-cookbook/sky-rl, harbor-cookbook/prime-rl, harbor-cookbook/tinker-rl and harbor-cookbook/harbor-rl all consume this interface.

5. Two paths for token-ID capture

Token IDs and logprobs can be captured at two different layers, depending on how cooperative your agent is.

Path A: agent self-reports (Terminus 2)

The LLM API response carries token IDs and logprobs. vLLM and SGLang return them natively; HF Router returns them depending on the provider (see references/hf_inference.md). The agent unpacks them directly. It's the cleanest path, and it works whenever the agent and the LLM endpoint cooperate.

Requires:

  • Agent supports rollout capture (Terminus 2 does; most CLI agents don't)
  • LLM endpoint returns logprobs (cloud APIs that don't return logprobs are unusable here)

Path B: vLLM proxy intercept

For agents that don't support collect_rollout_details natively (most CLI agents, such as Claude Code and Codex), Harbor can run a vLLM-compatible proxy in front of your hosted model. The proxy logs every (prompt, completion, logprobs) triple keyed by trial ID, regardless of which agent makes the call.

Requires:

  • Self-hosted LLM (or any OpenAI-compatible endpoint you can put a proxy in front of)
  • Agent talks to the proxy URL (set OPENAI_BASE_URL in the agent's env)

Comparison

Path A (self-report)Path B (proxy intercept)
CleanlinessCleanest: data flows through one channelTwo data sources to reconcile
Agent supportOnly Terminus 2 todayAny agent that talks OpenAI-compatible HTTPS
LLM supportEndpoint must return logprobsEndpoint must return logprobs (proxy logs them)
Cost overheadZeroOne extra hop per LLM call
Best whenYou control the agent and use Terminus 2You want to use Claude Code / Codex / etc. and still do RL

Path A is cheaper and cleaner; Path B works for any harness that talks to an OpenAI-compatible endpoint.

6. What this means for Repo2RLEnv

When we ship full pipelines (v0.2+), the loop looks like:

  1. Generate dataset with repo2rlenv generate ... --pipeline pr_runtime
  2. Push to HF Hub
  3. User trains with their RL framework of choice:
    from harbor import Job, JobConfig, TaskConfig
    tasks = TaskConfig(dataset="myorg/django-r2e@1.0", ...)
    job = Job(JobConfig(...), tasks)
    results = job.run()
    for r in results:
        reward = r.verifier_result.rewards.get("reward", 0)
        rollout = r.agent_result.rollout_details
        # ... policy gradient update
  4. Standard policy-gradient update from there

We don't need to touch the rollout-capture pipe at all, because Harbor already plumbs it end to end. Repo2RLEnv just produces tasks. Everything from "agent runs in sandbox" to "trainer gets rollout tensor" is Harbor's responsibility.

The one RL concern on the Repo2RLEnv side is emitting tasks tagged with the right reward_kinds (see SPEC.md):

Pipeline classReward kinds emittedRollout pipe used?
Lite (pr_diff)diff_similarity onlyNo (the trainer compares text directly)
Full (pr_runtime, commit_runtime, etc.)test_execution (and optionally diff_similarity)Yes (the full Harbor rollout pipe)

There's no official integration with TRL (HF's RL library) yet. A GRPO training loop (ORS reward server + TRL trainer) is planned for v0.9. It will support pr_diff datasets through the diff-similarity reward first; execution-verified pipelines (pr_runtime, commit_runtime, etc.) follow once a Harbor-wrapping reward server is in place.

7. References

On this page