How DeepSeek Harness Uses Cordis for Data-Driven Assembly

Cover Image for How DeepSeek Harness Uses Cordis for Data-Driven Assembly

Listen with Article TTS Reader

Checking for Article TTS Reader…

The DeepSeek Harness breakdown describes the assembled product. This companion looks at the mechanism that assembles it: how DeepSeek Harness (dsh) uses Cordis to make capabilities data, then composes that data into different runtime shapes.

The underlying Cordis model matters here. If Context, services, plugins, lifecycle ownership, and inject are unfamiliar, start with Cordis Runtime From First Principles. This post focuses on the next layer: how a runtime can compose a large plugin graph without encoding every product variation in bootstrap code.

The implementation details describe the upstream DeepSeek Harness repository at revision b150a551b8d4. They are bounded observations of that revision rather than claims about every later release or third-party composition.

From a bootstrap function to composition data

Most plugin applications begin with code resembling this:

const app = createApp()
app.use(timer())
app.use(llm({ model: 'deepseek-v4-flash' }))
app.use(tools([BashTool, ReadTool]))
app.run()

That is a reasonable implementation when one product shape has one stable set of components. It becomes awkward when the same runtime needs a web application, a headless tool, several capability tiers, and user-specific customization. A model choice or a disabled module becomes a code change, and alternative shapes tend to copy most of the original bootstrap.

dsh moves the component list into patch data. Each entry has an explicit id, which becomes the stable address for later changes:

- id: llm
  config:
    model: deepseek-v4-flash
- id: tools
  config:
    list: [bash, read, edit]

The decisive idea is modest: composition has data semantics. Components are named; layers modify named entries; the final data drives a Cordis application. The runtime does not need a separate mechanism for every product variation.

Patch merging gives layers a concrete meaning

Once two layers want to change the same component, the system needs predictable merge rules. dsh keeps the patch vocabulary deliberately small:

# Add a new component.
- insert:
    - id: webserver
      config: {}

# Find an existing component by id and replace the fields named here.
- id: hmr
  disabled: true

The rules are intentionally constrained.

  • insert adds rows, including rows that later patches in the same ordered sequence can target.
  • A normal patch addresses an existing row by id.
  • Assigned fields replace the corresponding field as a whole. They are not recursively deep-merged.
  • Patches apply in order, so later layers win where they touch the same field.
  • A patch whose target is absent is skipped with a warning; an optional name guard can prevent a patch from landing on an unexpected entry.

Whole-field replacement deserves attention. If a layer changes one value inside a component's config, it may need to restate the entire config object it owns. That creates more repetition than a deep merge. In return, the resulting configuration is easier to calculate and review: each layer says exactly which complete value it supplies, rather than relying on inherited fragments that can be difficult to reconstruct.

The same patch semantics should serve both runtime composition and inspection tools. That shared evaluator is a practical reliability feature. A dumped effective configuration is useful only if the process uses the identical merge behavior at startup.

Bundles produce product shapes without copying the core

A bundle is a named layer of component patches. A profile chooses an ordered bundle list, such as a shared base followed by a web-specific or headless-specific layer:

const profileTemplates = {
  web: ['base', 'web-app'],
  headless: ['base', 'headless'],
}

The order carries composition meaning. First apply the base, then apply the selected product layer. A web bundle can disable a set of core entries and insert its host-specific equivalents without duplicating the full base definition. A headless bundle can make a smaller set of changes for its own environment.

This is distinct from activation order. Layer order decides the final component graph. It does not say that the first row in a resulting file must start first. Keeping those meanings separate avoids a common failure mode in configuration systems: a file gradually becomes an accidental startup script because authors rely on positional effects it was never designed to guarantee.

Presets reuse a composition generation

Bundles choose a product shape. Presets choose a capability tier within that shape. A coding-focused session may need tools and agents that a minimal session should not mount.

dsh's preset model is more than a templating convenience. A preset is mounted as a standing composition generation and can be shared by sessions that select the same preset. The first relevant session causes the preset graph to mount; later sessions bind into that already-mounted subtree. When the preset source changes, newly created sessions can use a new generation while existing sessions continue on the prior one until it is safe to retire it.

The operational distinction is useful: composition and its expensive shared services can be reused, while session-specific state remains isolated by scope. A preset therefore behaves like a versioned, shared capability graph rather than a configuration file re-executed independently for every session.

This design also needs an explicit failure boundary. If mounting the selected preset fails during session setup, session creation should roll back. Producing a half-configured session is worse than rejecting a bad preset because the resulting behavior is harder to diagnose and replay.

Activation follows service availability

Patch order determines how dsh calculates the component graph. Activation follows a different rule: service availability.

An entry declares the services it needs through Cordis dependency injection. It activates when those services are available. A tool that depends on a provider can wait for that provider rather than relying on a particular line number in a configuration file. The component list can be organized for review instead of being fragile boot choreography.

The tradeoff is that an unavailable component can otherwise hide in a successfully parsed configuration. A robust boot process must settle the graph and then check for entries that failed to activate, reporting the missing injected services. This is a fail-loud requirement: configuration syntax being valid does not establish that the requested runtime exists.

Together, ordered patch composition and availability-driven activation give each concern one job:

Concern Rule
What the final graph contains Ordered patch merging; later layers override earlier ones
When a component can start Its required services are available
What happens after boot Unresolved requested components are reported as failures

User overlays use the same language as the product

The cleanest part of the design is that user customization does not introduce an unrelated extension API. It is another set of patch layers over the factory bundles:

factory bundles, in declared order
  -> profile patch, watched and recomposed on change
  -> user-home patch, watched and higher priority
  -> command-line patch, one startup-only overlay

The two persistent user layers can trigger recomposition when their files change. The command-line overlay is intentionally different: it participates only in the startup composition and is not a watched source. This separation prevents a temporary invocation override from becoming a surprising long-lived configuration input.

Recomposition must begin from a clean representation of the original layers. Patch processing can add inserted entries to the composition tree, so reusing a mutated in-memory structure risks accumulating or duplicating entries after successive reloads. Cloning or rebuilding before each merge is a small implementation detail with large operational consequences.

Third-party packages can join the same flow when a profile includes them as bundles. The application gains a consistent model for built-in capability, product-specific changes, optional plugins, and user policy. The model remains understandable only if its layering and activation errors are visible to the operator.

Failure modes are part of the architecture

Data-driven composition makes variation cheap. It also moves failures into declarative inputs that may be edited independently of the code that evaluates them. The design needs clear responses to predictable faults:

Failure mode Consequence Useful response
A patch targets a missing id Intended customization does not take effect Warn with the target and layer; review the effective configuration
A whole-field replacement omits a value The layer removes part of the prior field value Restate owned objects completely and test the resolved graph
Layer order is misunderstood A lower-priority change is silently overridden Make precedence documented and inspectable
A dependency never becomes available The component does not activate Fail loudly after graph settlement and name the missing service
A preset mount throws A session could otherwise be only partly configured Roll back session creation
Reload reuses mutated patch input Entries can accumulate across updates Rebuild or clone source layers before every composition
Configuration evaluates executable expressions A configuration edit crosses from data into code Treat authorship, review, and permissions as code-level security concerns

The last item is particularly easy to underestimate. Any configuration facility that permits executable expressions is no longer purely declarative. It can be appropriate for controlled deployment configuration, but it requires the same provenance, review, permission boundaries, and audit expectations as a code change.

Compatibility: mapping outward or adapting inward

There are two useful, complementary approaches to the growing ecosystem of agent configuration files.

OpenPackage's supported-platforms mapping maintains a universal file and directory mapping for many coding-agent platforms. Its practical value is that a package manager can place rules, commands, agents, skills, and MCP configuration into each platform's expected locations and apply platform-specific conversions.

dsh can shift some compatibility work inward. Rather than exporting every external asset into its original platform layout, an importer or bridge can adapt external agent assets to dsh's own runtime: sessions can become dsh sessions, skills can become dsh skills, and instructions can be registered through dsh's composition model. The destination runtime then owns the resulting capability graph and lifecycle.

These approaches solve different portions of the problem. OpenPackage is a portability layer across many file conventions. A dsh bridge is an adaptation layer into one runtime. This is not automatic universal compatibility. Asset semantics, tool permissions, instruction precedence, session fidelity, security boundaries, and evaluation behavior still need validation for every source and destination pair.

Demonstrated community bridge candidates

The following public repositories demonstrate narrowly scoped import or centralization directions. Their documentation is the source for these descriptions; they should be evaluated against the installed dsh version and local security policy before use.

Those are bridge and import candidates, not proof that every external configuration preserves its behavior after conversion. A migration plan should test representative workloads, inspect the permission model, compare the resolved prompt and tool surfaces, and retain a reversible path.

The "BIOS of agentic OS" idea needs a control plane

There is an attractive forward-looking interpretation of this architecture: a canonical body of composition knowledge could help a future trained model recognize runtime patterns and modify an extensible agent environment faster. It is reasonable to imagine such knowledge functioning like a BIOS-level map of an agentic operating system: services, dependencies, composition layers, scopes, and lifecycle transitions would have stable names and explicit relationships.

Training data alone does not deliver safe self-improvement. A model that can propose changes still needs explicit policies that constrain authority, tests and evaluations that establish behavioral evidence, replay and audit records that explain a change, sandboxing that limits blast radius, and human approval for consequential actions. Composition data makes a runtime more legible and more editable; it does not remove the need for governance.

A small mechanism with a large assembly surface

dsh illustrates a disciplined composition stack:

Need Composition mechanism
Addressable components Explicit entry IDs
Product variants Ordered bundle layers
Capability tiers Standing preset generations with scoped sessions
Correct startup Service-availability activation plus fail-loud checks
User customization Watched profile and user patches, plus a startup-only overlay

The mechanism remains small because the same patch language appears repeatedly. Bundles describe factory composition, presets refine capability, and users apply overlays without learning a parallel configuration system. The hard engineering work lives in the semantics around that language: deterministic precedence, lifecycle-aware activation, failure visibility, safe reloads, and security review.

That is the lasting lesson for agent runtimes. A large product can have a compact composition core when its rules are few, explicit, and shared by both the runtime and the tools used to inspect it.