Claude Code Under the Hood: A Recovery-Oriented Agent Harness

Cover Image for Claude Code Under the Hood: A Recovery-Oriented Agent Harness

Listen with Article TTS Reader

Checking for Article TTS Reader…

Most coding-agent explanations begin with a tiny loop:

while true:
    response = llm(messages)
    if response has no tool calls: break
    messages += run_tools(response)

That loop is real, yet it leaves out the work that makes an agent usable in a terminal for hours: interrupted streams, malformed provider signals, rejected commands, oversized context, resumable sessions, hooks that block completion, and delegated work.

This is the first post in the Agent Harness series. The next installment examines Codex.

Scope and provenance: I reviewed an unaffiliated claude-code repository snapshot at revision 6a2590911df2. Its publisher describes it as a leaked snapshot, but Anthropic has not verified its origin, completeness, or correspondence to any official release. Every implementation observation in this article applies only to that unverified snapshot. It should not be read as a claim about current Claude Code.

Within that boundary, the design is strikingly pragmatic. The code reads like a system that has accumulated defenses around each place the simple loop can fail.

The loop records why it continues

The entry point is an async generator, query(), which delegates to queryLoop() in src/query.ts. The central while (true) is familiar. The surrounding state handling is not.

Each continuation carries structured state and a transition value that records why the prior iteration continued. query/transitions.ts names reasons such as tool_use, reactive_compact_retry, stop_hook_blocking, and token_budget_continuation. The type intentionally leaves room for other strings, so this is an open diagnostic field rather than a closed enum.

That small choice changes the maintenance story. A bare continue says that the loop moved on. A transition reason says what happened before it moved on, which gives logs, tests, and recovery code a common vocabulary.

The snapshot also treats provider completion metadata as advisory. A comment near src/query.ts:554-558 says that stop_reason === 'tool_use' is unreliable and that streamed tool-use blocks provide the loop's exit signal. Production agent loops need this kind of redundant interpretation because a provider can report a stop while the message itself still contains work to perform.

Several recovery paths sit alongside that decision:

  • Output-token exhaustion has a bounded recovery limit.
  • A Stop hook can block completion and inject a follow-up message, with a guard against hook-and-retry loops.
  • A prompt-too-long error triggers reactive compaction before retrying; the snapshot recognizes different status codes on different API paths.

The result is a state machine built around a loop, rather than a loop patched with scattered exception handling.

Tools begin conservative and can start during streaming

The tool interface in src/Tool.ts makes the safe defaults explicit. Tools built without declarations are neither concurrency-safe nor read-only. In practical terms, an undeclared tool is serialized and subject to permission checks. This avoids silently granting a new tool the properties needed to bypass either control.

Permissions are part of the tool contract through checkPermissions(input, context). Deny rules can remove a tool from the list shown to the model in filterToolsByDenyRules, so a forbidden tool can be unavailable before the model proposes a call.

For a permitted call, src/services/tools/toolExecution.ts follows a recognizably defensive pipeline:

  1. Validate the input with Zod.
  2. Run PreToolUse hooks, which can allow, deny, request confirmation, or rewrite the input.
  3. Resolve permissions. A hook allowing a call does not necessarily remove an interactive confirmation requirement.
  4. Execute the tool.
  5. Run PostToolUse hooks.

Failures are converted to <tool_use_error> text for the model. That keeps a recoverable tool failure inside the conversation rather than automatically turning it into a process failure.

The performance-oriented detail is StreamingToolExecutor in src/services/tools/StreamingToolExecutor.ts. When a streamed response yields a complete tool-use block, execution can begin before the rest of the assistant message arrives. It coordinates safe and exclusive queues. The non-streaming fallback uses a separate partition-and-run path, batching consecutive concurrency-safe calls with a default limit of ten. These are two distinct implementations, though both try to overlap tool latency with model generation.

Session history is a chain that can branch

The snapshot persists transcripts as JSONL. Each entry carries a parentUuid; append operations create a chain, and resume walks backward from a selected leaf to reconstruct the conversation. That representation naturally supports branches: a resume can start from a chosen historical leaf rather than only from the latest message.

Compaction boundaries use a null parentUuid together with a logicalParentUuid, preserving the logical connection where a physical chain break is needed. Subagents use separate JSONL side chains and metadata, keeping their transcripts out of the main conversation chain.

The ordering principle matters as much as the format. QueryEngine.ts writes a user message before requesting the API response, so an interrupted session has a durable user action to resume from. It is a mundane durability rule with a large effect on whether resume feels trustworthy.

Context management is a set of independent defenses

The most elaborate part of the reviewed snapshot is context management. It does not run a fixed sequence of five shrinking stages on every iteration. The stages are independently gated by feature flags, model capabilities, thread type, and runtime thresholds.

  • A compact boundary limits which historical region remains active.
  • Tool-result byte budgets constrain unusually large results.
  • HISTORY_SNIP can replace old large history blocks with markers.
  • microcompact targets old results from particular tools, including Read, Bash, and Grep, when its feature and cache-editing conditions are met.
  • CONTEXT_COLLAPSE can project REPL history into a shorter granular view backed by a separate collapse store.
  • autocompact operates near the effective context-window limit, reserving a 13,000-token buffer in this snapshot and opening a circuit after three consecutive failures.
  • Reactive compaction remains the final recovery path after a prompt-too-long response.

This separation lets a small, local intervention happen before an expensive full-summary operation. TOKEN_BUDGET adds a related control: at roughly 90 percent of a task budget it can nudge the run to wrap up, then stop early after several rounds of insufficient progress.

The important lesson is operational rather than architectural purity. Context pressure has several causes, so the snapshot uses different mitigations for stale tool output, long histories, provider rejection, and tasks that keep consuming budget.

A subagent is a constrained nested run

runAgent() in src/tools/AgentTool/runAgent.ts calls the same query() mechanism with a subagent context. The implementation does not introduce a second, independent execution engine for delegation. Built-in Explore and Plan agents carry explicit read-only instructions, and multiple AgentTool calls can participate in the main loop's concurrency behavior.

The snapshot's optional backgrounding behavior is also conditional. The 120-second automatic background path requires the relevant setting or feature gate and can be disabled. That condition is worth retaining because it prevents a configuration-dependent behavior from becoming a universal product claim.

What this snapshot suggests

The design priority here is recovery depth. The simple agent loop remains visible, while state transitions, permissions, streaming execution, transcript structure, context controls, and subagents surround it with explicit failure handling.

That comes with a readability cost. In the reviewed snapshot, main.tsx is nearly 4,700 lines and query.ts is about 1,730 lines. The optional operating-system Bash sandbox is another useful example of why scope matters: it is disabled by default in this snapshot and can fall back to unsandboxed execution when unavailable unless failIfUnavailable is enabled.

Those details support a narrow conclusion: this unverified snapshot is optimized around keeping a long-running terminal agent moving through known failure modes. They do not establish what any official or current Claude Code build does.

Source index

  • Main loop and transitions: src/query.ts:265-307
  • Streamed tool-use completion handling: src/query.ts:554-558
  • Conservative tool defaults: src/Tool.ts:757-769
  • Tool hook pipeline: src/services/tools/toolExecution.ts:800-861
  • Streaming executor: src/services/tools/StreamingToolExecutor.ts:55
  • Transcript chain: src/utils/sessionStorage.ts:1001-1067
  • Autocompact threshold: src/services/compact/autoCompact.ts:62
  • Nested subagent invocation: src/tools/AgentTool/runAgent.ts:748-757