An Execution Contract for Skill-Bundled Scripts

Listen with Article TTS Reader
Checking for Article TTS Reader…
Two days ago, I wrote about distributing agentic capabilities: treating the useful thing above a model and harness as a package with instructions, code, interfaces, policy, tests, and an owner.
In that article, a capability meant an owned, governed unit of useful work: its business semantics, interfaces, permissions, evaluation evidence, and support lifecycle. A skill can contribute instructions and scripts to that unit, but it does not supply the whole contract. A capability might combine several skills, services, and tools, or expose an API or CLI without a skill at all.
This follow-up is a smaller experiment. I built @lunarmoon26/agent-skill-runtime to give scripts bundled with skills a shared, bounded execution contract and reusable host adapters. My immediate use case is exposing those scripts as tools through npm-backed plugins in OpenCode and DeepSeek Harness, with an MCP adapter for consumers that need it. This explores one integration problem inside the larger capability-package idea; it is not an implementation of the whole model.
Introducing Agent Skills
An Agent Skill is a folder anchored by SKILL.md: instructions explaining when and how an agent should approach a task. It can also contain scripts, references, and assets that support the workflow. The agent interprets the instructions; a bundled script can handle the deterministic part, such as resizing an image. That does not make the entire skill deterministic.
Codex and Claude Code can read those instructions and invoke bundled scripts through their existing shell or execution tools, using the available interpreter and dependencies under the host's permissions. They do not need this npm runtime or a custom npm plugin merely to execute a skill's script. The script is being called through an existing tool, not registered as a new native tool. OpenCode and DSH can use a direct execution path too; my adapter is useful when I want the script exposed through their plugin tool interfaces.
Where execution glue starts to repeat
The trouble begins when that same script needs to appear as a native tool in several hosts. It is tempting to write a small MCP wrapper, an OpenCode tool, and a DeepSeek plugin around it. Each wrapper then has to rediscover the same facts:
- the tool name and description;
- its input and output schemas;
- the fixed executable and its arguments;
- required capabilities, limits, and preparation;
- cancellation and cleanup behavior.
When those declarations live in parallel wrappers, they drift. One host can expose a permission that another forgot, or validate input differently from the tool that ultimately runs. The runtime makes the skill, rather than an adapter, the owner of that contract.
The design in one picture
SKILL.md explains the workflow to the agent
skill-runtime.json describes the executable tool contract
scripts/ contains the deterministic implementation
│
▼
shared runtime discovers, validates, prepares, and executes tools
│
├── MCP exposes the contract through a stdio server
├── OpenCode translates it into native tool definitions
└── DeepSeek registers it with Cordis
This diagram shows the optional adapter path. A host can also read SKILL.md and invoke a bundled script directly through its existing execution tools, without passing through this runtime.
The manifest sits beside SKILL.md because it describes the script interface supplied by that skill. It names a tool, defines a closed JSON input shape, selects an author-controlled entrypoint, and records its requested access and execution limits. The manifest's capabilities field means low-level access categories such as filesystem writes or network use, distinct from the broader business capability discussed in the previous article. A host can use the description to expose a tool interface while the runtime validates and executes calls to it.
The hosts still remain different systems. OpenCode needs its own tool shape and permission flow. DeepSeek needs a Cordis plugin that follows its lifecycle. MCP needs a server and transport. The shared part is deliberately smaller: discovery, validation, preparation, execution, and result semantics.
A tool call, from discovery to result
The first step is discovery. Given a plugin root, the runtime finds skill manifests, rejects duplicate tool names, and resolves the manifest and entrypoint to their real filesystem paths. A declared script must remain within its own skill directory after symbolic links are resolved. That makes the script path part of the reviewed package rather than another piece of model input.
Next comes validation. The runtime checks a tool's input before it starts a process, and can check the structured result before it returns it to a host. The executable receives one JSON request on standard input and returns one JSON result on standard output. Diagnostics belong on standard error, so they do not corrupt the result channel.
Then the runtime applies the operational boundary. It starts the fixed entrypoint with an argument array rather than a shell command, gives it a constrained environment, enforces declared input and output limits, and forwards cancellation. Tool authors can use Node, Python, Python managed by uv, or Wasmtime; dependency preparation is explicit so ordinary tool calls do not quietly become package-install events.
This sequence is intentionally mundane. A reliable tool boundary comes from making the normal path predictable: known input, known program, known permissions, known output.
A concrete skill stays small
The photo-processing skill was the first useful example. Its manifest exposes a bounded set of image operations and accepts a command plus an array of argument tokens. The manifest declares that it reads and writes files; a host can therefore ask for approval before the script starts.
The skill's Python bridge does very little. It reads the structured request, rejects unknown operations or malformed argument arrays, passes the values to the existing photo CLI without invoking a shell, and turns the CLI's output into a structured result. The image implementation stays where it belongs: inside the skill. The runtime and adapters only provide a consistent way to call it.
This was a useful test of the ownership boundary. I could add a machine-readable interface to an existing skill without moving its domain logic into a framework or teaching every host about image processing.
Adapters should translate, not reinterpret
The host adapters stay thin because they have one job: translate the portable contract into the host's native extension surface, then delegate the call back to the shared runtime.
- The MCP adapter registers each discovered tool on a stdio server and returns the validated result as structured content where the protocol supports it.
- The OpenCode adapter converts the portable schema into its native schema type, forwards the workspace and cancellation signal, and uses OpenCode's permission UI for privileged capabilities.
- The DeepSeek adapter registers the same schema with Cordis. Cordis owns that registration, so unloading the plugin also removes its tools.
Harness Alchemist uses the same split in its generated plugin repositories. Product skills own their instructions, manifests, and scripts. Its OpenCode and DeepSeek entrypoints name the skill manifest and delegate to the runtime package. The generated project retains native install and lifecycle behavior without maintaining another execution engine.
That distinction is important. A portable contract does not make hosts interchangeable. It gives them one stable place to agree on the parts that should be identical, while leaving host policy and lifecycle in the code that understands each host.
Interoperability travels in two directions
Harness Alchemist approaches portability from the package side: one project carries shared skills and the integration surfaces needed by different harnesses. Within that project, this runtime centralizes the execution contract for scripts exposed as plugin tools. Together they explore part of the universal capability-package direction, while leaving the broader business and governance contract to the package owner.
The other direction belongs to the harness. A host can import useful package formats from other ecosystems and map them into its own model, rather than asking authors to rebuild their whole library for every new runtime. OpenCode already searches .claude/skills and .agents/skills alongside its own directories, accepts additional local or remote catalogs, defines precedence for duplicates, and applies permission rules when a skill is loaded.
DeepSeek Harness has a similar extension seam. Its filesystem skill provider searches DSH and shared-agent roots by default, and can be given extra roots through configuration. That makes it possible to mount a compatible skill library from another location, including a Claude-style skill directory, while still publishing the result through DSH's own catalog and tool model.
The boundary matters. Importing a SKILL.md bundle is not the same as executing an arbitrary Claude Code plugin. A plugin can include agents, hooks, MCP configuration, and host lifecycle behavior beyond the shared skill shape. Claude Code documents those as separate plugin surfaces. A compatibility layer should preserve the standard subset it understands, make unsupported extensions visible, and require an explicit adapter before foreign host behavior runs.
The ecosystem can converge from both sides: package authors can adapt one capability outward to several hosts, and harnesses can import several compatible package forms inward. Neither route requires one agent runtime to replace the others.
Permission policy and discovery guardrails remain unresolved
The runtime validates access declarations and checks approval before launching a tool. It also discovers manifests under a configured root. Those mechanisms do not establish which business capabilities a person or agent is entitled to use, or which packages should enter a session's catalog. Its adapters register the discovered tools; the runtime supplies no organization-wide admission policy or task-aware selection layer. Those decisions need a trusted host or control plane.
I now separate three of them:
- Admission decides which artifacts may be installed or imported, with a known origin, owner, support status, and risk classification.
- Discovery decides which approved skills and tools are visible to a particular agent and principal. Catalog ordering, duplicate names, source labels, and visibility rules need to be deliberate because they affect what the model can select.
- Invocation decides whether one concrete action may proceed. It includes approval, credentials, rate limits, and authorization at the target system.
MCP makes the distinction concrete. Its tool specification permits a server's visible tool list to vary with the authorization presented on each request. It also says that tool annotations are untrusted unless the client already trusts the server, and that servers must enforce access controls. Describing a tool as read-only or destructive can inform a policy decision; it cannot replace enforcement.
The emerging registry layer has the same boundary. The MCP Registry verifies publisher namespaces and hosts metadata, while security scanning and curation remain responsibilities of package registries and downstream aggregators. Discovery is becoming richer, but it has not yet become a complete admission and authorization system.
An execution contract is not a sandbox
The runtime enforces the contract it owns. A capability declaration can require explicit approval. Process groups, timeouts, output limits, and cancellation make failures bounded and observable. It also avoids passing ambient credentials through to child processes by default.
It does not make a native script safe by declaration alone. The operating system, container, identity provider, and target service remain the actual enforcement boundaries. The runtime cannot grant itself filesystem, network, or credential access, and its process cleanup guarantees remain platform-dependent. Keeping those limits explicit is more useful than calling the design a sandbox.
CLI-first, adapters where they help
For a local coding agent with shell access, a skill plus a CLI is often the shortest path. The skill supplies workflow knowledge when it is needed, and the CLI performs the deterministic operation. That path can use the script's native environment without this runtime. When the shared contract is useful, the runtime also offers a CLI that discovers, prepares, and runs declared tools without starting an MCP server.
MCP is useful when a consumer already expects MCP tools. OpenCode and DeepSeek plugins benefit from their native adapters for the same reason. Changing transports alone does not reduce an agent's context cost, though. A large always-visible tool catalog can be expensive regardless of whether it arrived through MCP or a native API.
A prose-only skill also does not need an invented executable wrapper. The runtime is for deterministic operations with a stable tool boundary. Instructions without a bounded program should remain instructions.
The layer this establishes
The experiment gives skill-bundled scripts one optional execution contract across host adapters. It keeps script logic in the skill and reduces duplicated plugin glue. It does not turn a skill into a governed capability package, and it does not require every agent to route script execution through npm.
The broader capability-package goal will need industry and community agreement on the boundaries between these layers: portable instructions and interfaces, trustworthy package identity, scoped discovery, and authorization enforceable by the host and target service. Compatibility profiles can state which parts a package exports and which parts a harness imports. Vendors can keep different execution environments while agreeing on what compatibility preserves and what must fail when a required policy cannot be enforced.
This runtime deliberately leaves provenance, organization-wide policy distribution, evaluation history, revocation, and operating-system isolation to other layers. Those are consequential problems, and pretending a small runtime has solved them would make the boundary less trustworthy.
For now, I have a practical way to expose the same bundled script through several tool interfaces without rewriting its execution glue. The next question is how that small contract can participate in the governed distribution model from the previous article, where ownership, discovery, and authority have to survive a change of harness too.