Pi's Small Agent Harness: Queues, Session Trees, and an Honest Security Boundary

Listen with Article TTS Reader
Checking for Article TTS Reader…
Pi takes a usefully different position from the agent harnesses that accumulate layers of policy, modes, and subsystems. Its core runtime is deliberately small, and its project documentation is clear about features it leaves out, including subagents and plan mode. The point is to keep the harness understandable enough that users can extend it for their own work.
This is the fourth article in the Agent Harness series. I reviewed the upstream earendil-works/pi repository at revision 8fa7eebd2353. The next article looks at how a related loop evolved into the channel-oriented architecture of OpenClaw.
A loop with two kinds of interruption
The main engine is a pair of nested while loops. The inner loop handles tool calls and pending messages; the outer loop returns for queued follow-up work after the agent would otherwise stop.
That structure supports two queues with different timing semantics:
- Steering messages are injected between steps so a user can redirect work already in progress.
- Follow-up messages wait until the agent reaches a natural stopping point, which suits work the user wants to queue for later.
The distinction is small in code and large in interaction design. A single inbox either delays every correction until the end of a turn or makes every incoming message disruptive. Pi exposes both choices and lets configuration determine whether a queue is drained all at once or one message at a time.
Tool execution favors replayable history
Pi defines tools with TypeBox schemas. A tool exception is caught and represented as an error result, so one failing operation does not break the conversation protocol.
Two implementation choices are especially conservative:
- If a model response ends because it hit its length limit, Pi fails every tool call in that response without executing any of them. A truncated call may contain incomplete arguments, and executing a partially specified operation is a worse failure mode than returning an error.
- Tool calls may run concurrently, while their result messages are written in the order in which the assistant requested them. Execution can be concurrent without making the saved transcript nondeterministic.
That second point matters whenever a session must be replayed or inspected after the fact. Runtime scheduling should not decide the apparent order of the conversation.
A session is an append-only tree
An ordinary chat history is an array. Pi stores its session as JSONL entries with id and parentId fields, plus a leaf pointer identifying the current position. Adding an entry creates a child of the current leaf. Moving the leaf to an earlier entry creates a new branch without rewriting the old one.
This makes an "undo and try another route" workflow cheap and auditable. The abandoned branch remains in the same session file, and Pi can create a summary for it. A cross-session fork is also represented explicitly through a new session file that records its parent session.
Compaction uses the same representation. A CompactionEntry records a summary, the first retained entry, and the prior token count. When Pi reconstructs the model-visible history, the latest compaction entry replaces the earlier portion of the transcript for that view. The original records remain in the file.
Pi compacts when the current context exceeds the model window after reserving response space:
context tokens > context window - reserved tokens
Its default reserve keeps the most recent 20,000 tokens. The compactor avoids splitting a tool call from its result, though it can split a larger conversational turn. That is a specific consistency guarantee, not a promise that every turn remains whole.
Pi deliberately does not provide a sandbox
This is Pi's most consequential boundary. It has no built-in permission system for filesystem, process, network, or credential access. By default, tools run with the permissions of the user and process that launched Pi.
The project treats that as an honest security statement. An incomplete in-process permission layer can look like containment while failing to provide it. Pi directs users who need real isolation toward operating-system, virtualization, or container boundaries.
There are still protections with narrower purposes. Project-local extensions pass through a trust approval gate before they are loaded, which helps avoid automatically loading code from an untrusted repository. Extensions can also block tool calls through a beforeToolCall hook. The repository includes plan-mode and sandbox-oriented extensions as examples.
Those mechanisms can enforce workflow policy. They do not turn the core process into a security boundary. Pi should therefore not be treated as safe for unattended operation merely because an extension asks for approval.
Extensions are the product surface
Pi makes lifecycle hooks first-class: extensions can register tools and commands, observe trust and tool events, participate in compaction, and prepare the next turn. Subagents and plan mode are examples implemented through that extension model rather than permanent parts of the core runtime.
This is the trade Pi is making. A small core is easier to understand, inspect, and adapt, while more of the out-of-the-box experience depends on extensions and their surrounding ecosystem. A single-file JSONL session model is also less attractive than a database once sessions become extremely long or require shared, multi-user coordination.
Pi is not trying to fully govern an agent. It is trying to supply a harness small enough to understand completely, then make the right places available for users to grow it.
Source scope
The implementation details in this article refer to upstream Pi revision 8fa7eebd2353, including its agent loop, session manager, compaction logic, security documentation, and extension types. Continue with OpenClaw's persistent agent gateway.