Native Evolution Protocols
Native evolution protocols make iterative experiments durable without creating a privileged Evolver agent or an automatic optimization daemon. They are disabled by default:
MARINA_EVOLUTION_PROTOCOLS=trueThe existing evolve coach and experiment commands are unchanged when the feature is disabled.
When enabled, a protocol attaches to an existing experiment and records its objective, immutable
policy, proposal lineage, evidence, and attributed decisions.
Constitutional boundaries
Section titled “Constitutional boundaries”- The controller validates and records; it never prompts participants or continues a loop.
- Acceptance is a record, not activation or promotion.
- Metrics are advisory evidence and never alter standing, competence, permissions, or memory.
- Evolver, evaluator, and reviewer are voluntary activities, not entity classes.
- Humans, in-system agents, and external agents use the same world commands.
- Protocol state recovers from SQLite rather than an LLM-generated summary.
Workflow
Section titled “Workflow”Create the underlying experiment first, then attach a protocol:
experiment create PromptTrial arms baseline,candidate metric accuracy goal higherexperiment start PromptTrialevolve create PromptTrial | Improve accuracy without latency regression | max-runs=12 | min-trials=3 | independent-review=true | guardrail=latency:lowerevolve start PromptTrialParticipants explicitly propose, evaluate, and decide:
evolve propose PromptTrial | A shorter prompt reduces distraction | prompt:candidate-2experiment record PromptTrial baseline accuracy 0.74experiment record PromptTrial candidate accuracy 0.81evolve analyze PromptTrialevolve evaluate PromptTrial 1 | benchmark:prompt-trial-2026-08-04evolve decide PromptTrial 1 acceptAn accepted run remains inactive. Promotion or activation must happen separately through the ordinary command and review path for that candidate type.
Trials: a candidate against the incumbent, side by side
Section titled “Trials: a candidate against the incumbent, side by side”evolve trial <experiment> <run> incumbent:<role> [benchmark:smoke] [limit:N] [seed:N] [model:<m>] [timeout:30m] measures a role:<name> candidate without adopting it. It runs only in a child or
parallel world (World Collective; or MARINA_EVOLVE_TRIALS=here for a dedicated parallel world),
where that world’s daily spend cap bounds the cost. It spawns a temporary agent on each role (same
model, same budget), joins each to its own model channel in code, runs the same benchmark on both,
waits with a deadline, then tears agents and channels down. Nothing is activated. Rank 4 plus the
agent.spawn gate; one trial at a time.
# from the parent — the trial runs inside the child worldworld seed-role trial2 scout # incumbentworld seed-role trial2 scout-v2 # candidateworld run trial2 evolve propose ScoutTrial | a precise tone answers better | role:scout-v2world run trial2 evolve trial ScoutTrial 1 incumbent:scout benchmark:smoke timeout:15mworld run trial2 evolve trial ScoutTrial 1 result# candidate scout-v2 completed 100.0% (15/15) benchmark:br_…# incumbent scout completed 93.3% (14/15) benchmark:br_…# candidate − incumbent: +6.7 pointsThe result is kept (a trial started over world run outlives the bridge’s connection) and names both
runs, which an evaluator other than the proposer cites with evolve evaluate.
Held out by default. Every benchmark’s items are split once, by a hash of the item id, into a
holdout fifth and a tune rest (benchmarks/partition.ts) — disjoint and stable, unlike reseeding.
A trial judges on 100 items of ARC-Challenge’s holdout split (227 items) when that dataset is cached,
else on smoke with a warning; iterate on benchmark run <name> --partition tune, never on holdout.
The trial reports a 95% interval on the difference (Agresti–Caffo — honest at small n and at 0% /
100%): on smoke, 15/15 vs 14/15 is +6.7 points with an interval that includes zero, i.e. noise.
evolve replicate requires the interval to sit above zero as well as the fishing margin.
Each arm also shows its answer rate and its accuracy on answered items, with a split: line — the
overall score counts an unanswered item as wrong, so it mixes answer quality with reliability.
The first real held-out trial read +6.0 overall: +0.2 points of quality, +6 answered items. A 100-item
trial runs ~200 agent answers — roughly $10 on Claude Sonnet 5, inside a child’s $50 day.
Earned replication: a winner spawns copies of itself
Section titled “Earned replication: a winner spawns copies of itself”evolve replicate <experiment> <run> [n:1] [budget:200] [model:<m>] lets a candidate that WON spawn
copies of itself — in the same child or parallel world, never the parent. It refuses unless:
- the run is accepted (recorded by
evolve decide, with the protocol’s independent review); - its trial completed on both arms and both cited benchmark runs still resolve;
- the difference’s 95% interval is above zero — not distinguishable from noise is not a win;
- the candidate beat the incumbent by the fishing margin —
0.02 + 0.01·log₂(1 + candidates already trialed in this session), the same rule arena signal discovery uses, so every try raises the bar; - there is room: at most 5 copies per run ever, the caller’s standing-scaled spawn budget,
MAX_AGENTS, and the world’s daily spend cap.
Each copy runs the candidate role with its own call budget, is spawned_by the caller, and gets a
lineage record (evolve-trials / evolve_replica: run, role, agent, parent, delta, margin). Nothing in
the parent world changes; bringing a role back is a separate, reviewed step.
world run trial2 evolve replicate ScoutTrial 1# Not earned: the candidate beat the incumbent by 0.0 points; this session needs 2.0world run trial2 evolve replicate ScoutTrial 2 n:2# earned: +100.0 points over hasty (bar 3.0) · spawned scoutv2r2n1, scoutv2r2n2Adopting a winner into the parent world is a separate, reviewed step — see World Collective → Bringing a winner home.
Evidence you can check
Section titled “Evidence you can check”Cite a benchmark run by its id and evolve evaluate checks it: the run must exist, be completed and
have answered something, or the evaluation is refused. Its verified score is stored with the
evidence, together with what it measured — for a marina:<name> target, the agents serving it,
their role and the hash of the exact system prompt in force:
benchmark run smoke --model marina:scout-v2 # → br_…evolve evaluate ScoutTrial 1 | sharper answers on the smoke set benchmark:br_da6d8043_muhobbx8# verified benchmark:br_da6d8043_muhobbx8 smoke 93.3% (15/15) · marina:scout-v2 · Scout2 role scout-v2 prompt ab12cd34ef56Free-text labels (benchmark:prompt-trial-2026-08-04) are still accepted as notes; only run ids
(br_…) are resolved.
Use parent=<run-id> to preserve experimental lineage:
evolve propose PromptTrial | Refine the successful shorter prompt | prompt:candidate-3 | parent=1Available status controls are explicit:
evolve sessionsevolve status PromptTrialevolve pause PromptTrialevolve resume PromptTrialevolve complete PromptTrialRun and time budgets refuse new proposals after exhaustion but do not silently pause, complete, or otherwise change the session. The creator decides what happens next.
Qualification and transport equivalence
Section titled “Qualification and transport equivalence”After participants complete a controlled trial, verify its durable evidence outside the agent loop
(participants can preview the same verdict in the world with evolve qualify; the gate is the
external probe, which reads the evidence over HTTP rather than through the loop being judged):
bun run qualify:evolutionbun run qualify:evolution http://marina.example:3300The bounded probe waits for a started session, proposal, attributed evidence, and decision. It also fails if independent-review attribution collapses or passive constitutional defaults were changed. It is read-only: it cannot propose, decide, continue, activate, or promote.
In-system Pi agents receive the scoped marina_evolve tool only while participating in an active
session. MCP clients receive an evolve tool, SDK clients use client.command("evolve ..."), and
humans use the same command directly. All surfaces converge on the same command handler,
participant checks, persistence, and audit trail.
Bounded soak testing
Section titled “Bounded soak testing”Exercise socket churn and command latency against a disposable or staging world:
bun run soak --connections=25 --duration=60 --rate=2 --max-errors=0 --max-p95=2000bun run soak:churn --url=ws://localhost:3300 --clients=12 --cycles=20bun run soak:churn:localThe driver exits non-zero when clients fail to connect, errors exceed the threshold, no responses arrive, or round-trip p95 exceeds the bound. Its final line is machine-readable JSON for CI and demo qualification. Do not aim a high-rate run at a production world without considering its normal participants and provider limits.
For a disposable three-role real-model trial, run bun run trial:evolution:local. It starts and
stops an empty temporary world, asks three agents to join voluntarily, stages proposer → evaluator →
reviewer work, and exits non-zero unless durable evidence passes qualify:evolution. Use
MARINA_TRIAL_MODEL and MARINA_TRIAL_TIMEOUT_MS to select the provider-neutral model identifier
and bound. A failed run is evidence: the harness never fills in an agent’s proposal, evaluation, or
decision.
The Experiments dashboard expands each evolution session into lineage, candidate, evidence, attribution, decision, guardrail, activity, communication, tool-use, and latency detail. Token and cost fields remain explicitly unavailable until provider-neutral per-session attribution is durable; Marina does not infer them from activity.
