DeepSeek Harness: An Auditable, Composable Agent Runtime

Listen with Article TTS Reader
Checking for Article TTS Reader…
This is part seven of the Agent Harness series. Earlier articles examine Claude Code, Codex, opencode, Pi, OpenClaw, and Hermes. The series closes with a cross-harness comparison.
DeepSeek Harness, or DSH, is easiest to misunderstand when viewed as a command-line wrapper around an LLM. Its core loop is familiar:
while the agent should continue:
response = call_model(messages)
execute_requested_tools(response)
append_results_to_the_next_request()
The interesting work begins around that loop. DSH treats the agent, its tools, its history policy, and even its product surfaces as a composition of plugins. That makes it a useful case study in what an agent harness must add once a one-shot tool caller becomes a long-running application.
The implementation details below describe the upstream DeepSeek Harness repository at revision b150a551b8d4, with the standard web and headless compositions as the primary subject. They do not describe every possible third-party plugin or custom profile. For the runtime model beneath the plugin system, see Cordis Runtime From First Principles. For the composition mechanics behind the assembled product, read How DeepSeek Harness Uses Cordis for Data-Driven Assembly.
The product is assembled, not hard-coded
DSH has a plugin engine, a base bundle, and product-specific bundles such as the web application. Bundles contribute ordered patch manifests. Presets then select a capability set for a particular session, including standard, code, and minimal configurations. The resulting agent is therefore a declared composition rather than a single fixed binary behavior.
This changes where to look for an answer. If a tool is absent, the question is rarely only whether the loop supports it. The relevant answer may be in a bundle, a preset, a profile overlay, or a plugin that was never installed. Flexibility is the benefit; tracing behavior across those layers is part of the operational cost.
The composition mechanism also explains why DSH can offer distinct product shapes. A web profile layers a web bundle over the shared base, while headless execution layers a smaller headless bundle over the same base. The shared loop remains available, but the tools, host services, and user interface can differ without duplicating the whole runtime.
A turn contains several steps
The first pressure on a naive loop arrives when the user sends input while tools are running. DSH separates a turn from a step. The outer loop advances turns; the inner loop builds a model request, streams a response, and handles its tool calls one step at a time.
That split gives incoming input two useful destinations. send, steer, inject, and followup are ultimately routed into either a next-turn queue or a next-step queue. A follow-up can wait for the current turn to finish. A steering message can be injected at the next step boundary without interrupting an in-flight tool call. In the web surface, this distinction appears directly as queued input versus steering input, while cancellation stops the current turn without discarding queued work.
The loop also does not delegate termination entirely to a provider finish reason. Pending tool calls and queued next-step input are considered before a turn ends. This is a recurring production-harness pattern: continuation is a property of the harness state, not merely a field returned by a model API.
Tools are a pipeline with a permission escalation point
DSH tools carry more than a function name and a JSON-shaped argument object. A tool can declare its output schema, its rendering behavior, and whether it is safe to run concurrently. Tool calls pass through four waterfall extension points:
pre-execute -> execute -> post-execute -> result
This gives policy, auditing, retry behavior, and telemetry a place to attach without changing the agent loop. Concurrent tool calls are still committed to the model in the original order supplied by the model, not in completion order. That preserves a replayable conversational sequence even when individual calls finish at different times.
The standard preset provides the expected coding-agent surface: shell and background-job tools, file reads and image reads, globbing and grep, edits and writes, plans, subagents, user questions, and web search. It also includes DSH-specific mechanisms such as skills, workflows, and a bounded self-driving ralph loop. A different preset can make a deliberately smaller tool surface.
Approval asks about leaving the boundary
The default standard and headless configuration is workspace-write with an ask approval policy. That wording can suggest that every shell command is presented to a human. It does not work that way.
Commands and file changes that remain within the active sandbox mode can run without an approval prompt. An approval request occurs when the model asks to escalate permissions:
read-only -> workspace-write or danger-full-access
workspace-write -> danger-full-access
The escalation request must carry a justification. If the host declines it, or if no approval handler is available, the escalation is rejected. The web interface presents that request as a focused allow-once or refuse decision, and the request can be reconstructed after a page refresh because it is part of session history.
This is a more precise policy than approving every individual command. It reduces prompt volume while preserving a human decision at the point where the model requests a broader file-effect boundary.
The sandbox claim has a specific scope
DSH has a meaningful fail-closed property, but it is narrower than a general claim that every agent action is sandboxed.
For the constrained subprocess runner used by the standard and headless configurations, DSH supports read-only, workspace-write, and danger-full-access modes. Its platform implementations include Linux confinement through Bubblewrap and Landlock, macOS Seatbelt, and a Windows restricted-token ACL path. If that constrained runner cannot obtain usable confinement, it rejects execution rather than silently running the original command unsandboxed. The Windows enforcement path is marked partial.
The policy is about file effects on the same host. It does not by itself bound network access or process visibility. A tool with its own network behavior or another host-side side effect does not automatically enter this file-effect policy. The minimal preset can use the unsandboxed local filesystem backend, which ignores sandboxPolicy. DSH also has an experimental remote E2B provider, but moving a provider to a remote sandbox is a separate execution model.
The accurate claim is therefore:
In the standard and headless constrained-subprocess path, unavailable confinement fails closed, and permission escalation is explicitly approved and logged.
It would be inaccurate to extend that claim to all presets, all plugins, all network activity, or every possible host effect. The comparison with Codex is similarly conditional: both have meaningful OS-level isolation work, while DSH's default mode alone does not prove a stricter overall security boundary.
Event sourcing makes history an operational artifact
A growing messages array has two failure modes: it consumes the context window, and it disappears when the process ends. DSH records an append-only event stream for user messages, assistant chunks, tool calls, tool results, and lifecycle transitions. Events receive sequence numbers, including streamed model output.
The central invariant is simple: model-visible means logged. DSH checks that the messages sent to the model can be derived from the event log. This lets the runtime distinguish the durable record from the current model view.
Compaction changes that view by appending a replacement event that hides an older section from the active prompt. It does not need to destructively rewrite the original events. The standard configuration first trims oversized tool results, then begins summary compaction at 80% of the context window while retaining a recent tail. Its summarizer reuses the session's prompt prefix, tools, and earlier messages to improve the chance of provider KV-cache reuse.
This event model has practical consequences:
- A session can resume by replaying its log.
- A fork can seed a child session from the last complete turn.
- The interface can reconstruct approvals and streamed output.
- Auditing and compaction can use the same durable source without pretending that the live prompt contains every historical byte.
Sessions carry composition as well as conversation
DSH creates web sessions lazily. Once a session is created, its selected preset is recorded and mounted as part of the agent setup. If setup fails, creation rolls back rather than leaving a half-configured session. On resume, preset membership is recovered from the log rather than guessed again from a mutable header.
Preset composition is shared across sessions in a generation, while session state remains isolated. When a preset file changes, new sessions can use a later generation while existing sessions remain associated with the generation that created them. This makes configuration changes visible without rewriting the history of already-running work, although it also creates a lifecycle-management problem for old generations.
Subagents follow the same model. A child starts with a fresh conversation context, inherits selected execution configuration such as its route and working context, and retains persistent lineage. It does not inherit the parent transcript. Delegation depth is bounded, and child agents are configured not to request human approval themselves. That prevents a delegated tree from producing an uncontrolled approval surface, but it also means the parent must choose delegation boundaries carefully.
What DSH optimizes for
DSH's distinguishing idea is not a novel call to an LLM. It is the discipline that every model-visible fact has a durable account, while capabilities can be assembled at the product, profile, preset, and session levels.
That offers a high ceiling for teams that need an agent runtime to evolve. The price is a multi-package system with a vendored plugin runtime, several composition layers, and an early 0.1.x maturity signal. A missing capability may be intentional, configured elsewhere, or genuinely absent. Engineers evaluating DSH should inspect the assembled profile and test the security boundary of the exact preset and tool path they intend to deploy.
The final article, Can DeepSeek Harness Reproduce Six Other Agent Harnesses?, puts that composability claim under a stricter evidence model.