Testing a Recursive Self-Improvement Pipeline on a CrunchDAO Competition

Listen with Article TTS Reader
Checking for Article TTS Reader…
I wanted a place to test a recursive self-improvement idea where the score was real, the constraints mattered, and a bad change could not be explained away by a toy benchmark.
I found one in the CrunchDAO Structural Break Challenge. The task is to watch a time series one observation at a time and estimate whether it has permanently changed. The model must work causally, produce deterministic output, and run as a competition package outside my laptop. Its score is Time-Stratified AUC, a measure of whether the model ranks changed and unchanged series correctly at each point in time.
My immediate goal is simple: build a better structural-break detector. The longer experiment is whether the research process that builds it can improve safely over time.
The idea: two loops, two jobs
I am building the system as two nested loops.
Outer loop: human + OpenCode (or another harness)
choose the objective, budget, rules, and evaluator
review evidence, approve changes, and decide when to submit
Inner loop: DSH researcher + DAL evidence layer
run bounded experiments on the competition model
record outcomes and failures
propose one improvement to the research workflow
evaluate the proposal and return the evidence to the human
DeepSeek Harness, or DSH, is the agent runtime I want to use as the researcher. DSH Adaptive Loop, or DAL, is my local layer for recording evidence, staging candidate changes, and enforcing review boundaries.
OpenCode is useful in the outer loop because it gives me a practical control surface for directing work, inspecting results, and deciding when the inner loop should continue or stop. Any other harness could play that role. The important choice is that the researcher and the supervisor are separate jobs.
What I mean by recursive self-improvement
People use this phrase for everything from an LLM revising an answer to an agent rewriting the machinery that produces its future behavior. I care about the second version, with a very important limit: the system has to prove each proposed improvement against an evaluator it does not control.
The Darwin Godel Machine is a useful reference: it creates modified coding agents, evaluates them, and keeps an archive of useful variants. AlphaEvolve follows a related pattern for algorithms: language models propose candidates, and an automated evaluator selects promising ones.
Earlier work such as STOP and Automated Design of Agentic Systems makes the same direction concrete: a language-model scaffold or agent design becomes an object of search. These systems differ in what they change and how they score it. The common lesson is that candidate generation is only one part of an improvement loop. The evaluator and the search boundary decide whether its conclusions mean anything.
My version is much narrower. DAL is not an autonomous research scientist. It is a way to make an agentic workflow more inspectable: preserve useful evidence, propose a scoped change, test it, and let a human decide whether it becomes the new baseline.
Why DSH became my substrate
I arrived here through the Agent Harness series. After examining Claude Code, Codex, OpenCode, Pi, OpenClaw, Hermes, and DSH, I saw DSH as unusually interesting for this experiment. Its model provider, prompt assembly, tools, permissions, session history, and product surfaces can be composed from explicit plugins, bundles, profiles, and patches. Its append-only event model also gives a durable account of what the model saw and what the runtime did. I covered those mechanics in my DSH architecture post and the follow-up on Cordis assembly.
That does not mean DSH already ships a self-improving agent. A place to attach a plugin is not evidence that the resulting behavior is complete, safe, or comparable to another harness. In my cross-harness comparison, I separated shipped behavior, partial plugin evidence, and architectural seams for exactly this reason.
DSH gives DAL a runtime substrate with named modification points and useful session evidence. DAL supplies the missing control plane: independent evaluation, an explicit editable surface, a holdout boundary, candidate lineage, human promotion, and a path back to a known baseline. The agent can create a candidate broadly; the path that makes it live stays narrow.
DAL before CrunchDAO
I started DAL before I had the CrunchDAO task. The original question was: after a repeated agent workflow succeeds or fails, how can I preserve enough evidence to improve it later without treating a raw chat transcript as memory or letting the system rewrite itself freely?
The first answers were conservative. Records are privacy-scanned and immutable. Evaluation is deterministic where possible. A human owns promotion. I took design ideas from GEPA, SkillOpt, and PenguinHarness, while keeping them as design references rather than DAL runtime dependencies.
By the end of that first phase, DAL had structured feedback and approval records, run records, deterministic checks, a holdout boundary, a small workflow benchmark, and DSH-facing recorder and improvement-plugin prototypes. I separated harness health from business success because an agent can run perfectly while still failing the actual task. Receipts, manifests, and provenance make later comparisons reproducible rather than dependent on somebody's memory of a session.
The runtime work also gave me an early lesson in restraint. I explored staging runtime changes and looked at live reload as a possible route to faster iteration. Source inspection and runtime probes did not establish that a live DSH generation could be safely replaced. The safe response was to quarantine live application and keep candidate staging inactive. An extension point is not proof that an automated deployment path is safe.
When the CrunchDAO workspace began, DAL was therefore an evidence and staging system looking for a real workload. The competition supplied one. I did not build DAL to make a structural-break detector. I use the detector to find out whether the DAL workflow holds up when there is a real metric, real resource limits, and a package that has to run somewhere else.
How DAL turns work into a candidate
DAL has two separate state machines. One describes what happened in the work. The other describes a proposed change to the worker.
Evidence path
task feedback + run records
-> validation and immutable local records
-> failure clustering and compatible-batch measurement
-> human review
Candidate path
observed -> proposed -> sandbox evaluated
-> approved or rejected -> applied -> measured
That separation is more than bookkeeping. A measurement can show that a failure rate looks worse across comparable runs, including its uncertainty. It cannot promote a new skill. A sandbox can show that a candidate passes a fixed evaluator. It cannot apply the candidate. Each transition has a distinct owner and a distinct form of evidence.
In a normal workspace, I start DAL with dal init. It creates repository-local evidence stores and installs the local workflow instructions without touching a shared DSH profile. During work, the agent records structured task feedback and, when appropriate, a privacy-safe run record. The record carries bounded facts such as outcome, failure category, tool and harness identities, digests, and elapsed time. It does not preserve prompt text, tool arguments, results, credentials, or hidden reasoning as optimization material.
The smallest local setup is intentionally ordinary:
npm install -g @lunarmoon26/dal
dal init
# After a batch of work:
dal feedback summary
dal cluster run
Those commands create and inspect evidence; they do not automatically modify a skill, runtime, or model.
At the end of a batch, the human runs the reconciliation work: summarize feedback, cluster failures, and inspect the resulting evidence. A pile of unrelated agent traces is not a benchmark.
Only then does a proposal make sense. A valid proposal names one editable surface, such as a skill, tool description, routing rule, retry policy, or bounded harness patch. It has to state the failure evidence it addresses and a prediction that can fail. DAL can keep alternative candidate branches, run deterministic evaluation gates, and preserve rejected branches as negative evidence. Applying a change remains a separate, exact human approval.
DSH is optional for this basic workflow. DAL works as a local CLI and repository convention first. DSH adds a useful runtime for collecting agent-work evidence and experimenting with skills or plugins, but its optional plugin modes are not mounted automatically and candidate application is still unavailable. That is intentional: the control plane should become more capable only after the evaluator and deployment boundary can support it.
Three planes keep the loop honest
The same system looks different depending on where you stand. I separate its responsibilities into a run plane, an improvement plane, and a governance plane.
The run plane does normal work. Improvement code does not sit in its hot path deciding how a live task should behave. The improvement plane studies completed evidence and may produce candidate skills, prompts, tool descriptions, routing rules, or bounded harness patches. The governance plane decides whether any candidate has earned the right to affect a future generation.
This separation blocks the most tempting self-confirming loop: a proposer edits its prompt, changes the evaluator, and lets the new evaluator declare success. In DAL, the proposer never controls the evaluator, sealed holdout, permission policy, maximum budget, promotion policy, audit record, or rollback procedure. Those are immutable anchors, not additional prompts the agent is asked politely to respect.
Today the governance plane is intentionally human-operated. DAL can validate an exact approval and record a proposal transition, while promotion and rollback remain ordinary reviewed deployment or version-control procedures. Automated canaries, generation switching, and executable rollback are future work, not capabilities I am quietly assuming into existence.
Feedback is not yet learning data
"This response was too verbose" is useful feedback, but it does not identify the cause. The prompt might be too broad, a skill could have been selected incorrectly, a tool may have returned too much data, the context policy may have failed, or the user may simply have had a one-off preference.
DAL therefore treats a useful experience as a compact packet of evidence rather than a raw transcript. It binds a task class, pinned model and harness identity, expected and observed outcomes, failure category, side-effect receipts where available, evaluator version, and privacy classification. The package says enough to compare runs without treating prompts, tool arguments, results, credentials, or hidden reasoning as an optimizer's training corpus.
That distinction lets the later evaluator separate several important cases: the harness completed its steps while the business task failed; the harness itself failed; the agent correctly refused a prohibited action; or a tool had a side effect while the final response still failed. Conflating all of those into one thumbs-up signal would teach the system the wrong lesson.
Measurement before control
Repeated workflows invite a useful control-systems analogy. One harness generation runs a batch of comparable tasks, produces measurements, and then a later generation may receive one bounded change.
Generation Gk -> task batch -> measurements -> candidate edit -> Generation Gk+1
That is why DAL has a run-to-run observation layer. Given a human-authored policy and a compatible batch, it estimates configured success proportions with explicit exclusions and 95% Wilson intervals. Mixed task sets, models, evaluator versions, context policies, or runtime identities fail closed rather than being averaged into a comforting number.
The important boundary is in the name: it is an observation layer, not a PI controller or Model Predictive Controller. A future governor could use deadbands, hysteresis, bounded search budgets, and eventually an explicit model of how a proposed edit changes future runs. Until there is a verified observe-predict-optimize-act-replan loop with a real deployment seam, calling it MPC would be marketing rather than engineering.
What can evolve, and at what scope
The competition model and the research harness are different objects. Training a better detector is normal research. Changing the process that researches the detector is the RSI experiment. Keeping them separate prevents a lucky model run from becoming accidental evidence that the harness improved itself.
I use this as an engineering taxonomy rather than a claim about an industry standard:
| Scope | What changes | Does it persist? | Example |
|---|---|---|---|
| In-run adaptation | Current context, plan, or retry | No | Reflection after a failed tool call |
| Workspace evolution | Prompt, skill, memory, or routing | Yes | A skill updated from repeated failure evidence |
| Harness evolution | Plugin, workflow, tool, or agent code | Yes | A bounded guard or verifier adapter |
| Model evolution | Weights, training data, or training process | Yes | A separately evaluated fine-tuning experiment |
| Open-ended RSI | Several layers and the improvement machinery | Yes | A system that improves its own ability to improve |
DAL currently concentrates on workspace evolution and bounded harness candidates. Model evolution is a separate future plane. Open-ended RSI is a research question, not a product claim.
Skills and research methods
The most practical place to start is skills: instructions that tell the researcher how to inspect the metric, study relevant methods, propose causal features, run an experiment, and report an invalid result honestly. Skills are easy to diff, review, and test against a fixed task.
Runtime and capability packages
The next surface is the runtime: tool descriptions, routing rules, project packages, and removable plugins. I am deliberately avoiding an upstream fork of DSH for this work. A repository-owned package with a clear review path is easier to inspect and remove than a pile of private changes to somebody else's runtime.
The model itself
Model-level evolution is a future control plane. The competition detector is trained as part of the task, but that differs from changing the model that powers the researcher. A future experiment might compare a new proposer model, a small local model for trace analysis, or an isolated weight update. Each needs its own training data, fixed benchmark, budget, and rollback plan.
The order matters. A skill edit is visible. A runtime change is more consequential. A weight update is harder to inspect. The evidence needed to keep a change should grow with its blast radius.
The competition keeps the loop honest
CrunchDAO gives this experiment an external score that I cannot rewrite after seeing it. The evaluator, train/development/holdout splits, causal inference contract, dependency checks, and export rules are fixed. A candidate has to cross that boundary as an exported competition package, rather than merely look promising inside my local research loop.
That boundary has already changed my view of a candidate. An earlier CUSUM model improved development TS-AUC from 0.558713 to 0.560579, then scored 0.552166 on holdout, below the previous model's 0.554467. The package was valid and the experiment was useful. Its holdout result ruled it out as an overall improvement.
The current best candidate is a fixed, equal-weight ensemble of three AR-feature models. It improved the local development result over the prior canonical model and returned an out-of-sample AUC of 58.97% on its fourth CrunchDAO submission. The components, random seeds, and weights were fixed before evaluation, and each member was trained from scratch.
I ran matching controls, determinism and fresh-process checks, and CPU-inference-cost checks alongside the local evaluation. They documented the local result's limits and made the development gain worth submitting. The competition then evaluated the exported artifact on data I did not use to tune it, giving the local observation an external score.
This score is evidence about the structural-break detector. A claim that DSH or DAL improved the research process needs a different comparison: a frozen researcher configuration versus an approved methodology change, under the same task set, budget, model, and evaluator. The competition result constrains detector evaluation; it does not supply that researcher-method comparison.
How I run it in CrunchDAO
The two modes stay separate in the competition workspace.
Getting DAL into the workspace
Today the setup is deliberately manual. I clone the DAL repository beside the research workspace as ~/Workspace/dsh-adaptive-loop, install it locally, and initialize DAL inside the competition repository. That creates the local CLI surface, workspace instructions and skill, and .dal evidence stores. OpenCode can use those workspace-level pieces directly; DAL is an evidence layer around the research work, not a plugin that silently changes the agent.
DSH has a second integration step. For session recording, I link a selected source package from ~/Workspace/dsh-adaptive-loop/plugins/ into a dedicated profile under ~/.dsh/profiles/, then mount it through that profile's cordis.patch.yml. My crunch-codex-research and crunch-codex-acp profiles currently mount @lunarmoon26/dal-run-record and write privacy-safe session records to .dal/runs. Profile mounting changes a capability boundary, so I treat it as a separate reviewed decision rather than an automatic installation step.
This is an early developer workflow. I am working toward versioned npm packages and a smoother onboarding path, so a fresh research project will not need a local clone, manual profile links, and Cordis edits. Until then, the manual setup makes the installed source and profile configuration easy to inspect.
In research mode, DSH can read the competition brief, inspect the canonical model, study relevant methods, propose causal features, and run approved experiments against the development evaluator. It cannot submit to the competition, inspect the holdout to tune a candidate, or recursively hide another researcher inside one trial. The output is an experiment result and a candidate package, not a decision to promote it.
In optimization mode, DAL receives reviewed evidence from completed research sessions. The output is a draft change to the researcher: perhaps a better skill, a reporting rule, a feature-research procedure, a routing choice, or a bounded runtime package. It does not train the competition model, apply itself, or modify the active profile. A separate, matched research batch must test whether that methodology change helped.
The human keeps control of the evaluator, holdout, permissions, maximum budget, promotion policy, and rollback. That is less dramatic than an agent that rewrites every layer of itself after each run. It is also more useful for real work.
A practical way to adopt DAL
DAL is most useful when the work has a repeatable shape and an external definition of success. A support workflow with a state/effect grader, a regression-tested engineering task, an operations routine, or a competition pipeline can work. Open-ended creative work can still produce useful notes, but it does not support the same improvement claim.
- Pick a workflow with an evaluator you are willing to keep outside the agent's control.
- Start with one editable surface, usually a skill or a workflow instruction.
- Initialize DAL in the repository and collect structured evidence from real tasks rather than inventing a synthetic history.
- Compare one candidate against a fixed baseline with the same tasks, model, budget, and evaluation conditions.
- Review the evidence, then promote or reject the candidate through an ordinary version-control change.
The DAL repository contains the CLI, schemas, fixtures, and operator documentation. The important part is the discipline around it. If the evaluator can move, the batch is not comparable, or the candidate can deploy itself, the loop is no longer telling you whether it improved the workflow.
If this arrangement improves the quality or efficiency of my competition research, I will have evidence for a small and practical form of recursive self-improvement. If it only adds ceremony, the same records should make that visible. Either result is better than declaring a system self-evolving because it can edit its own files.