Distributing Agentic Capabilities

Cover Image for Distributing Agentic Capabilities

Listen with Article TTS Reader

Checking for Article TTS Reader…

Lately, I have been helping build tools and plugins that let our development teams use coding agents—or whatever agent harness they happen to prefer—to do useful work: analyse data, investigate logs, work with Jira and internal wikis, interact with CI/CD pipelines, create reports, or simply write better code.

The engineering challenge is making those tasks available to a broad and increasingly mixed set of runtimes: local agents on developers' machines, hosted agents, coding agents, chat assistants, CI workers, and harnesses that have not been invented yet.

That has made one thing very clear to me: we need a common model for distributing agentic capabilities—one that keeps them portable and compatible while respecting the firm's permission and access model.

The conversation about AI agents is still mostly framed around models and harnesses. Which model reasons best? Which coding agent has the nicest terminal UI? Which framework can coordinate tools, memory, subagents, permissions, and a few heroic retries?

Those questions matter, but they describe the platform more than the business application.

A useful mental model is that the model and harness together form an agent operating system. The model provides general-purpose judgement. The harness provides or integrates the process lifecycle, context, tool discovery, sessions, permissions, sandboxing, and user interaction. The combination increasingly resembles iOS or Android: a general environment in which many specialised things can run.

This is a distribution and lifecycle analogy, not a security boundary. The model does not confer authority, and the harness often delegates isolation and enforcement to the host operating system, identity provider, gateway, or target service.

The distributable capability is the app on top.

That framing also makes portability more realistic. iOS and Android are not interchangeable, and an app may need different builds or adapters for each. Its identity, purpose, business rules, support contract, and expected behaviour can still remain stable across both. An organisation will probably use several agent operating systems at once; its business capabilities should not be trapped inside any one of them.

The harder question is therefore:

How do we make an organisation's useful, safe, business-specific actions available to any of those agent operating systems?

My answer is to treat those actions as distributable capability packages.

The app layer between agents and business systems

Most companies already run on stable systems.

There are order-management platforms, risk engines, payment workflows, data warehouses, internal APIs, approval queues, reconciliation services, operational runbooks, and a small archaeological dig of spreadsheets. These systems may not be fashionable, but they are where the business actually happens.

An agent becomes useful when it can work with that reality. It needs a controlled way to:

  • retrieve the right information;
  • understand the domain constraints;
  • execute a bounded action;
  • request approval where needed;
  • produce evidence of what it did;
  • fail safely.

That bridge is the capabilities layer.

Agent operating system
  model + harness
        ↓ installs and invokes
Capability package
  instructions + code + interfaces + policy + evals
        ↓
Stable business systems, data, workflows, and controls

The model supplies judgement and the harness supplies the runtime. The capability package supplies the organisation-specific thing that the agent is allowed and equipped to do.

Like an app with separate iOS and Android builds, one capability may expose an MCP adapter, a local CLI, and a harness-specific plugin. Portability does not require those adapters to contain identical bytes. The package defines shared business and policy invariants, while each adapter profile must declare its principal, credential path, approval experience, and enforceable operations, then pass the relevant contract and policy tests.

That distinction matters because models and harnesses will change faster than the business systems underneath them.

MCP was an important first move

Model Context Protocol gave the industry a useful common language for exposing tools, resources, and prompts. Its data and transport layers standardise connection, discovery, invocation, and version negotiation without dictating how an application orchestrates its model.

That is a real improvement over every agent framework inventing its own connector format.

An MCP server can say, in effect:

  • here are the tools you may call;
  • here are the resources you may read;
  • here are the prompts or workflows I expose;
  • here is the schema for interacting with them.

The protocol also places important responsibilities on implementations. Servers must validate inputs and enforce access controls; clients should confirm sensitive operations, validate results, apply timeouts, and log usage.

MCP is intentionally a protocol rather than a complete enterprise packaging model. It does not answer enough of the questions that an internal app platform must answer:

  • What business capability does this represent?
  • Which team owns and supports it?
  • Which users, agents, or environments may use it?
  • Is it read-only, approval-gated, or able to make irreversible changes?
  • What runbook, policy, and evaluation suite come with it?
  • How is it released, deprecated, rolled back, observed, and audited?
  • Can its business contract survive a move from MCP to a local CLI, hosted worker, or another harness?

MCP defines an important platform API. The packaging and support contract sits above it.

Skills pointed in a useful direction

The recent “skill” pattern is more interesting than it first appears.

The open Agent Skills specification defines a small folder anchored by SKILL.md, with optional scripts, reference material, and assets. The agent loads the instructions when relevant and can run the bundled code or use the bundled templates. Claude and Codex consume this shape, while OpenCode and other agents support similar filesystem or plugin mechanisms.

This starts to look like an application bundle:

reconcile-cash-breaks/
├── SKILL.md
├── scripts/
├── references/
├── assets/
├── tests/
└── manifest.json

The minimal standard does not require tests/ or a richer manifest.json; those are deliberate organisational additions in this example. The gap is revealing. A portable folder format solves loading and progressive disclosure, while a production package also needs trust, policy, quality, and support metadata.

A capability is more than a function signature. A function can tell an agent how to call create_payment_instruction. It cannot by itself teach the agent what a payment instruction means, when it should be rejected, how to validate downstream state, which approvals are required, or how to handle a partial failure at 17:25 on month-end.

The operational knowledge belongs with the action.

Skills package that knowledge more naturally than a bare API schema. CLI tools do too. So do well-designed plugins and npm packages. Each distributes some combination of code, instructions, context, assets, and conventions.

The industry now has several plausible application bundles. What it lacks is a common contract for what a production capability must contain.

From tools to capability packages

I think an organisation should treat a capability package as the versioned unit that it builds, reviews, promotes, distributes, and supports.

A package might expose an MCP server, a CLI, a REST adapter, or a harness-specific plugin. Those are implementation choices. The package's business contract should remain durable across them.

At a minimum, it should declare:

  • Identity and ownership — stable name, purpose, owning team, domain, support path, and risk tier.
  • Versions and compatibility — package version, adapter versions, supported models and harnesses, and required runtime dependencies.
  • Interfaces — tools, CLI commands, APIs, input/output schemas, and supported adapters.
  • Instructions — how an agent should use it, including domain language and decision boundaries.
  • Permissions — requested scopes, data classification, environment restrictions, approval requirements, and credential flow.
  • Operational semantics — idempotency, retries, timeouts, rate limits, compensating actions, and rollback behaviour.
  • Evidence — audit events, correlation IDs, receipts, provenance, and human-readable summaries.
  • Quality gates — deterministic tests, scenario evals, expected failure modes, and policy checks.
  • Supply-chain integrity — immutable artifact digest, source and build provenance, dependency inventory, licences, and attestations.
  • Lifecycle and support — release channels, support level, compatibility guarantees, deprecation dates, migration guidance, and emergency revocation.

The package can request entitlements; it cannot award them to itself. Permission declarations are inputs to policy enforced by the trusted harness, gateway, or target business system.

The folder in Git is source material. A production release should resolve a stable package name and version to an immutable artifact that the runtime can verify before loading. Existing software-supply-chain practices already give us useful machinery for this: provenance, attestations, signed artifacts, dependency review, and admission policy. Agent instructions deserve that treatment because a malicious or compromised package can direct an agent to execute code, invoke tools, or move data.

This is the structure that makes reuse safe. Without it, every agent team rebuilds fragile adapters: a tool definition here, a prompt pasted into a system message there, perhaps a shell script living in somebody's home directory. The result works until the original author goes on holiday or a new agent runtime interprets the prompt differently.

That is a collection of demos wearing a trench coat.

An app needs a support lifecycle

Calling something a package creates a release-management obligation. An app becomes dependable because somebody owns each version, declares where it runs, supports it for a period, updates it, and can remove it when it becomes unsafe. Capability packages need the same discipline.

The reproducible unit is larger than the package version. Observed behaviour comes from a combination:

capability version
  × adapter version
  × harness version
  × model version
  × policy version

A package that passes with one combination can regress when a model changes its tool selection, a harness changes context compaction, or an adapter changes approval handling. Production deployments should pin these versions where the platform allows it, record them on every trace, and re-evaluate supported combinations before promotion. “Latest” is a convenient development channel, not a support policy.

Testing the full Cartesian product would quickly become impossible. Package owners should publish a small set of supported profiles, certify the combinations that matter for their consumers, and label or reject everything else as unsupported. A support contract for a high-impact package should also name service levels, downstream dependencies, incident severity and escalation, the rollback owner, and the end-of-support date—not merely an email alias.

A simple package lifecycle could look like this:

State Required evidence and control
Draft Named owner, intended users, risk tier, initial threat model, fixtures, and no production distribution.
Candidate Immutable build, provenance and dependency review, passing contract tests and evals against the intended model-harness combinations.
Approved Signed release, scoped entitlements, support level, staged rollout, observability, rollback path, and a defined end-of-support policy.
Deprecated Named replacement or migration path, end date, usage reporting, and no silent expansion to new consumers.
Revoked Execution denied within a declared maximum revocation window, including for cached copies; credentials and grants withdrawn where necessary, owners notified, and audit evidence retained.

Governance should scale with blast radius. A read-only documentation helper does not need the same approval path as a capability that can release a payment. Both still need an owner, a compatibility statement, a support path, and an end-of-life mechanism. Otherwise an internal catalogue slowly fills with ownerless packages that everyone is afraid to update and nobody is authorised to remove.

Self-improvement should look like release engineering

The phrase “self-improving capability” is easy to misunderstand. Letting a production agent rewrite its own instructions or scripts after a bad run collapses observation, authoring, approval, and deployment into one opaque loop. It also gives noisy or adversarial production input a path to change future behaviour.

A safer improvement loop is explicit:

Production traces, incidents, and user feedback
        ↓ curate, redact, and label
Versioned regression and evaluation set
        ↓
Candidate change to package, adapter, model, or policy
        ↓
Offline replay + deterministic tests + security checks
        ↓
Shadow or canary rollout + required approval
        ↓
Promote, monitor, or roll back

This loop needs several kinds of evidence because a single score hides too much:

  • Contract tests check schemas, scripts, deterministic transformations, and adapter behaviour.
  • Scenario and trajectory evals check capability selection, tool choice, action sequence, source grounding, recovery, and final outcome.
  • Policy evals check that prohibited actions remain prohibited and approval boundaries cannot be bypassed.
  • Operational evals measure latency, cost, retries, rate-limit behaviour, and downstream side effects.
  • Human review resolves ambiguous or high-impact cases that automated graders cannot safely judge.

The practical pattern is already emerging. OpenAI's guide to evaluating skills reduces an eval to a prompt, a captured trace and its artifacts, a set of checks, and a score that can be compared over time. Real failures become new regression cases. AWS now exposes both continuous production evaluation and targeted evaluation of selected traces. The useful idea is independent of either platform: operational evidence should feed a controlled release pipeline.

A model can help cluster failures, propose an instruction change, generate a new fixture, or draft a patch. Promotion authority remains outside the execution loop. The candidate must pass independent gates, and the evidence used to improve it must be access-controlled and reviewed for sensitive data, poisoning, and misleading proxy metrics.

That turns “self-improvement” from configuration drift into ordinary, inspectable release engineering—with an unusually capable contributor inside the loop.

The registry grows into a control plane

Once an organisation has more than a handful of packages, a registry cannot remain a directory listing. In the operating-system analogy, it becomes a combination of an enterprise app store, mobile-device management, and release control.

The control plane should know:

  • which package names exist, who owns them, and which immutable artifacts they resolve to;
  • which versions are draft, candidate, approved, deprecated, quarantined, or revoked;
  • which model, harness, adapter, and policy combinations have passing evaluation evidence;
  • which users, agents, environments, and data classifications may install or invoke them;
  • which approvals, credentials, budgets, and rate limits apply at runtime;
  • which release channel and rollout cohort should receive each version;
  • how the capability is performing, who supports it, and when its support ends;
  • how to stop it quickly without waiting for every cached installation to update.

The control plane defines desired and permitted state. The execution plane—a local agent, hosted worker, or chat assistant—loads an approved package, obtains appropriately scoped credentials, invokes business systems, and emits receipts. The control plane does not have to proxy every low-risk call, but its policy must remain authoritative through an online check or a signed, short-lived policy snapshot. Offline snapshots and short-lived credentials create an explicit maximum revocation delay; capabilities that require immediate shutdown need an online authorisation check at each consequential action. A revoked package sitting in a developer's cache cannot be allowed to act beyond its declared revocation window because the laptop missed an update.

The industry is assembling these pieces from several directions. The preview MCP Registry verifies publisher namespaces and points to artifacts in package registries, while deliberately delegating security scanning and curation to package registries and downstream aggregators; it recommends private registries for private servers. OpenAI's Skills API supports versioned bundles and explicit version pinning, while its plugin format bundles skills and MCP configuration into a larger installable unit. Microsoft describes Agent 365 as a unified registry and control plane for cross-platform agent inventory, governance, and security. AWS groups its preview agent registry with identity, gateway policy, observability, evaluations, and optimisation around model- and framework-independent agents.

These products govern different units and do not yet form one capability-package standard. The shared direction is more important than the current product names: discovery is being surrounded by identity, admission, policy, telemetry, evaluation, rollout, and revocation.

A simple example

Imagine a capability called cash-break-investigation.

Its job is broader than “query database,” which is a primitive rather than a business capability. The package could provide:

  • a manifest naming the operations team, risk tier, support hours, supported agent platforms, and end-of-support policy;
  • read-only tools for retrieving break details, positions, ledger entries, and prior investigation notes;
  • a deterministic script that assembles an evidence pack;
  • instructions explaining break categories, cut-off times, and escalation rules;
  • a policy stating that the agent may propose a resolution but cannot book an adjustment;
  • an approval-gated command for creating an investigation case;
  • eval scenarios covering stale data, missing identifiers, conflicting records, attempted policy bypass, and end-of-day cut-offs;
  • an audit receipt linking every conclusion to the retrieved source records;
  • a signed release plus MCP, CLI, and chat-workflow adapters that preserve those semantics.

A local coding agent could use this through the CLI adapter. A hosted agent could use it through MCP. A chat assistant could use a narrow UI workflow. The control plane could approve one version for an initial read-only cohort, observe its traces, promote it more broadly after the eval gate, and revoke it if a downstream schema change made its conclusions unsafe.

The business capability remains the same even when its platform-specific builds differ. That is the portability worth optimising for: the same governed capability wherever the work happens.

The organisation model matters as much as the file format

This changes how teams should divide responsibility.

Platform teams should provide the package conventions, registry and control plane, identity integration, policy enforcement, secrets handling, observability, evaluation infrastructure, and common adapter tooling.

Domain teams should own and support capabilities close to the systems and operational knowledge they already maintain. They decide what a correct cash-break investigation means and which failures demand escalation.

Agent-product teams should compose those packages into end-to-end experiences and verify the combinations of model, harness, and capability that they release.

Security, risk, and compliance teams should define cross-cutting admission policy and review the high-impact packages without becoming the operational owner of every capability.

That division is healthier than centralising all agent work in one AI team. A central team cannot become the expert on every payment exception, reporting process, onboarding workflow, or production support path. Shared platform services also spare each domain team from building its own agent operating system.

The capability package becomes the handshake between them, and the control plane records the terms of that handshake.

Standardise the app contract, not the operating system

I do not expect one grand standard to replace MCP, skills, CLIs, plugins, and framework-specific extensions. That would be both unrealistic and slightly suspicious in the way all grand standards are suspicious.

The goal is narrower: standardise the seam between an agent operating system and an organisational capability.

Platform-specific builds are expected. The stable contract is the capability's identity, business semantics, policy model, evaluation suite, evidence format, release history, and support commitment. MCP can be one adapter. A skill folder can supply operating guidance. An npm or OCI package can be a distribution mechanism. A DSH, OpenCode, or other plugin can integrate the capability deeply into a particular runtime.

Those mechanisms are builds of the app, not the app's durable organisational contract.

The durable asset is the capability ecosystem

The most defensible asset in enterprise agent adoption will be the catalogue of trusted capabilities that safely connects intelligent systems to the organisation's actual work.

Models will improve. Harnesses will proliferate. Some vendors will disappear with impressive velocity.

A well-owned capability package can survive those changes when it captures domain logic, policy, interfaces, evaluation evidence, support commitments, and operational history. It can be rebuilt for the next agent operating system without rediscovering what the organisation permits, how success is measured, or who answers the pager when it fails.

That is the layer worth building deliberately.

The future of enterprise agents depends on organisations learning to package, govern, support, and improve what they know how to do.


References