Can DeepSeek Harness Reproduce Six Other Agent Harnesses?

Cover Image for Can DeepSeek Harness Reproduce Six Other Agent Harnesses?

Listen with Article TTS Reader

Checking for Article TTS Reader…

This is the final article in the Agent Harness series. Read the individual breakdowns of Claude Code, Codex, opencode, Pi, OpenClaw, Hermes, and DeepSeek Harness first.

Every product in this series begins with the same small program:

while the agent should continue:
    response = call_model(messages)
    execute_requested_tools(response)
    append_results_to_the_next_request()

The systems differ because that loop runs inside a different product contract. A coding assistant needs reliable edits, command execution, and recovery. A personal gateway needs channel routing and durable conversations. A self-improving agent needs a way to turn past trajectories into future behavior. The loop is common; the operational system around it is the product.

This comparison uses the following repository snapshots: Claude Code 6a2590911df2 from an unofficial source that Anthropic has not confirmed; Codex 13fe2bcb7a3b; opencode b72b50006b24; Pi 8fa7eebd2353; OpenClaw 888c8422013e; Hermes 5b82658b3cca; and DeepSeek Harness b150a551b8d4. Third-party DSH plugin availability was checked on 2026-08-26. A repository snapshot supports an implementation observation, not a claim about every later release.

Seven answers to the same operational problems

The harnesses share several production requirements. They must decide whether a turn really ended, preserve or compact history, contain side effects, route new user input during work, and restore useful state after interruption. Their answers reflect different priorities.

Harness Primary shape Loop and interaction model Durable-history direction
Claude Code Official coding agent Generator-style state transitions with pending input and stop hooks JSONL parent chain with several compaction guards
Codex Official coding agent Sampling loop combines tool handling, pending input, and end_turn JSONL rollout with compatibility and model-aware compaction
opencode Terminal coding agent and server Single loop reloads durable state and processes queued work Durable events with SQLite projections
Pi Small, extensible coding agent Nested loop with steering and follow-up queues Append-only JSONL tree with compaction entries
OpenClaw Multi-channel personal-agent gateway Nested loop with steering able to preempt unstarted tool calls Per-agent SQLite transcript tree
Hermes Self-improving personal agent Budgeted loop with stop and redirect controls SQLite and FTS5, with soft archive and optional compaction
DSH Plugin-composed agent runtime Turn and step loops with two classes of inbox Append-only event log with replacement projections

Two patterns appear across the implementations. First, production systems do not trust a provider's finish field as the only termination signal. Tool state, queued input, and harness-level control signals also decide whether to continue. Second, recoverable tool failures usually return an error result to the model rather than crashing the entire loop. Those are mundane defenses, but they separate an interactive product from the pseudocode above.

History design is more varied. Some systems preserve original records and alter the model's view through compaction; others also retain explicit cleanup or retention paths. It would be wrong to describe any of them as universally non-destructive. DSH is unusually strict about one narrower property: messages visible to the model must be derivable from its event log, including streamed chunks. That helps replay and audit, while still allowing the prompt projection to hide old material.

The sandbox comparison needs the same precision. Codex has deep OS-level isolation and command-policy work. Claude Code offers a Bash sandbox but does not require it by default. opencode emphasizes approvals without an OS-level sandbox on its main path; Pi runs under the user's permissions; OpenClaw and Hermes have different optional backends. DSH's fail-closed behavior applies to its constrained subprocess runner in standard and headless compositions, not to every preset, network action, or plugin side effect. A single permission-mode label is not enough to rank these systems by security.

Treating DSH as a substrate

DSH is the unusual entry in the comparison because much of its product surface is assembled through bundles, presets, services, providers, and tool waterfalls. That raises a tempting question: can it reproduce the other six harnesses?

The word "reproduce" must be qualified. Having a place to add a feature differs from having a feature that behaves like the target system under real workloads. This article uses three explicit evidence tiers:

Evidence tier Meaning
shipped The behavior is present in DSH's distributed implementation and can be exercised through its documented composition.
partial plugin evidence A plugin or integration demonstrates a limited part of the behavior. It does not establish full product semantics, production quality, or parity.
architectural seam DSH has a plausible extension point for the work. No complete implementation or parity test establishes the target behavior.

These tiers are deliberately asymmetric. A seam is useful design information, but it is not a completed feature. A published plugin proves that code exists and can be obtained; it does not prove that its reliability, security, or interaction contract matches a mature harness.

What the evidence supports

Target capability DSH evidence tier What the evidence supports What remains unproven
Pi's minimal two-tool agent shipped The minimal preset provides a persistent PTY shell and str_replace_editor, with no compaction. Equivalence beyond this intentionally small tool surface.
opencode-style provider depth partial plugin evidence Provider details fit naturally in LLM adapters; dsh-coding-plan@0.1.0 and dsh-delegate-router@0.4.0 demonstrate provider and routing work. opencode's accumulated per-provider compatibility behavior, retries, and edge-case coverage.
Delegation to Claude Code, Codex, or ACP agents shipped Optional Claude Code and Codex profile bundles, outbound ACP subagents, and an inbound ACP server establish integration directions. Reproduction of each target's own tools, permission policy, history model, and compaction semantics.
Codex-style execpolicy architectural seam tools/pre-execute can decide whether to allow, prompt, or deny a tool call; existing escalation approval proves the decision point can block execution. Shell parsing, prefix rules, persistent amendments, sandbox exceptions, and parity tests.
OpenClaw-style multi-channel gateway partial plugin evidence Host, session, inbox, and approval services are candidate locations. @dsh-suite/plugin-notify@0.2.0, dsh-interconnect@0.10.0, and dsh-mobile@0.2.1 demonstrate notifications, handoff, and an additional surface. A full bidirectional, multi-channel gateway with identity, thread routing, attachments, outbound replies, and approval forwarding.
Hermes-style learning memory loop partial plugin evidence Tool and hook surfaces can host retrieval and experience capture. dsh-agent-memory@0.8.4, @openviking/dsh-memory-plugin@0.2.1, and dsh-continual-evolve@0.5.0 show active work in this direction. An auditable, reversible, quality-gated loop that turns trajectories into sustained product improvements.

The Pi case is the cleanest result. DSH ships a minimal preset that already resembles Pi's deliberately narrow coding-agent surface. Even there, "resembles" is the appropriate word: matching two tools does not inherit Pi's simplicity, extension conventions, or operational ergonomics.

Provider compatibility is a different class of work. DSH can place vendor-specific behavior behind an adapter, and community packages show that routing and provider delegation are practical. opencode's value, however, includes a large body of provider-specific adjustments: unusual reasoning parameters, retry behavior, rate-limit interpretation, and backoff details. Porting those adjustments remains engineering work one provider at a time.

The Claude Code, Codex, and ACP integrations deserve another distinction. DSH can delegate work to an external agent through optional profile bundles or ACP. That is an actual integration capability. It does not make DSH a reimplementation of the delegated agent. The child still brings its own tool, approval, sandbox, and context-management semantics.

Why the missing parts are hard

Codex's command policy illustrates why an architectural seam cannot be promoted to parity by optimism. The pre-execute waterfall is a credible home for an Allow, Prompt, or Deny decision. Complete execpolicy behavior also needs correctly parsed shell commands, prefix matching, persistent user amendments, and carefully defined interactions with sandbox escape conditions. Each boundary needs tests. The DSH seam reduces where the implementation must attach; it does not eliminate the implementation.

OpenClaw's channel gateway has the same shape. A notification plugin proves that messages can leave DSH. A mobile interface proves that a second client surface can exist. A multi-channel gateway also needs ingress adapters, sender identity, conversation-to-session routing, attachment handling, outbound replies, and approval forwarding. The available plugin sample is partial plugin evidence, not evidence that this complete bidirectional contract already exists.

Hermes-style improvement loops require an even stronger contract. Retrieval tools and memory stores can make past information available. A continual-evolution plugin can demonstrate an attempt to update behavior. A dependable learning loop must additionally define when learning runs, what it may change, how changes are reviewed and rolled back, what quality gates apply, and how conflicts are resolved. Without those properties, "self-improvement" can be an unbounded source of configuration drift rather than an operational capability.

High ceiling, incomplete coverage

DSH has a broad set of explicit extension points: tool waterfalls, providers, session services, host surfaces, bundles, and presets. That architecture gives engineers a named location for many kinds of work. The event-sourcing invariant and data-driven composition are especially useful when an installation needs to explain what the model saw and why a capability was present.

Architecture does not transfer the labor already embedded in the other products. Channel adapters, provider quirks, security policies, data migrations, and replay compatibility still need design, implementation, and maintenance. Some differences are structural rather than missing plugins: DSH's multi-package runtime and vendored Cordis dependency are less compact than Pi's model, and its 0.1.x version line signals active change.

The conclusion is therefore bounded. DSH has shipped evidence for a minimal preset and for several external-agent integration directions. It has partial plugin evidence for provider routing, memory, notifications, and alternate user surfaces. It has an architectural seam for more ambitious features such as Codex-style command policy. None of this establishes complete cosplay of Claude Code, Codex, opencode, Pi, OpenClaw, or Hermes.

Choosing a harness starts with whether a team needs a finished product or a configurable substrate. The established harnesses each offer a more opinionated answer for a particular workload. DSH is compelling when composition, replayable model-visible history, and the willingness to build and verify the remaining layers matter more than immediate feature completeness.

This comparison is a snapshot, not a final verdict. I will revisit it as the DSH kernel stabilizes and the community produces more mature plugins, reproducible evaluations, and parity evidence.