Skip to content

Native Evolution Protocols

Native evolution protocols make iterative experiments durable without creating a privileged Evolver agent or an automatic optimization daemon. They are disabled by default:

Terminal window
MARINA_EVOLUTION_PROTOCOLS=true

The existing evolve coach and experiment commands are unchanged when the feature is disabled. When enabled, a protocol attaches to an existing experiment and records its objective, immutable policy, proposal lineage, evidence, and attributed decisions.

  • The controller validates and records; it never prompts participants or continues a loop.
  • Acceptance is a record, not activation or promotion.
  • Metrics are advisory evidence and never alter standing, competence, permissions, or memory.
  • Evolver, evaluator, and reviewer are voluntary activities, not entity classes.
  • Humans, in-system agents, and external agents use the same world commands.
  • Protocol state recovers from SQLite rather than an LLM-generated summary.

Create the underlying experiment first, then attach a protocol:

experiment create PromptTrial arms baseline,candidate metric accuracy goal higher
experiment start PromptTrial
evolve create PromptTrial | Improve accuracy without latency regression | max-runs=12 | min-trials=3 | independent-review=true | guardrail=latency:lower
evolve start PromptTrial

Participants explicitly propose, evaluate, and decide:

evolve propose PromptTrial | A shorter prompt reduces distraction | prompt:candidate-2
experiment record PromptTrial baseline accuracy 0.74
experiment record PromptTrial candidate accuracy 0.81
evolve analyze PromptTrial
evolve evaluate PromptTrial 1 | benchmark:prompt-trial-2026-08-04
evolve decide PromptTrial 1 accept

An accepted run remains inactive. Promotion or activation must happen separately through the ordinary command and review path for that candidate type.

Trials: a candidate against the incumbent, side by side

Section titled “Trials: a candidate against the incumbent, side by side”

evolve trial <experiment> <run> incumbent:<role> [benchmark:smoke] [limit:N] [seed:N] [model:<m>] [timeout:30m] measures a role:<name> candidate without adopting it. It runs only in a child or parallel world (World Collective; or MARINA_EVOLVE_TRIALS=here for a dedicated parallel world), where that world’s daily spend cap bounds the cost. It spawns a temporary agent on each role (same model, same budget), joins each to its own model channel in code, runs the same benchmark on both, waits with a deadline, then tears agents and channels down. Nothing is activated. Rank 4 plus the agent.spawn gate; one trial at a time.

# from the parent — the trial runs inside the child world
world seed-role trial2 scout # incumbent
world seed-role trial2 scout-v2 # candidate
world run trial2 evolve propose ScoutTrial | a precise tone answers better | role:scout-v2
world run trial2 evolve trial ScoutTrial 1 incumbent:scout benchmark:smoke timeout:15m
world run trial2 evolve trial ScoutTrial 1 result
# candidate scout-v2 completed 100.0% (15/15) benchmark:br_…
# incumbent scout completed 93.3% (14/15) benchmark:br_…
# candidate − incumbent: +6.7 points

The result is kept (a trial started over world run outlives the bridge’s connection) and names both runs, which an evaluator other than the proposer cites with evolve evaluate.

Held out by default. Every benchmark’s items are split once, by a hash of the item id, into a holdout fifth and a tune rest (benchmarks/partition.ts) — disjoint and stable, unlike reseeding. A trial judges on 100 items of ARC-Challenge’s holdout split (227 items) when that dataset is cached, else on smoke with a warning; iterate on benchmark run <name> --partition tune, never on holdout. The trial reports a 95% interval on the difference (Agresti–Caffo — honest at small n and at 0% / 100%): on smoke, 15/15 vs 14/15 is +6.7 points with an interval that includes zero, i.e. noise. evolve replicate requires the interval to sit above zero as well as the fishing margin. Each arm also shows its answer rate and its accuracy on answered items, with a split: line — the overall score counts an unanswered item as wrong, so it mixes answer quality with reliability. The first real held-out trial read +6.0 overall: +0.2 points of quality, +6 answered items. A 100-item trial runs ~200 agent answers — roughly $10 on Claude Sonnet 5, inside a child’s $50 day.

Earned replication: a winner spawns copies of itself

Section titled “Earned replication: a winner spawns copies of itself”

evolve replicate <experiment> <run> [n:1] [budget:200] [model:<m>] lets a candidate that WON spawn copies of itself — in the same child or parallel world, never the parent. It refuses unless:

  • the run is accepted (recorded by evolve decide, with the protocol’s independent review);
  • its trial completed on both arms and both cited benchmark runs still resolve;
  • the difference’s 95% interval is above zero — not distinguishable from noise is not a win;
  • the candidate beat the incumbent by the fishing margin — 0.02 + 0.01·log₂(1 + candidates already trialed in this session), the same rule arena signal discovery uses, so every try raises the bar;
  • there is room: at most 5 copies per run ever, the caller’s standing-scaled spawn budget, MAX_AGENTS, and the world’s daily spend cap.

Each copy runs the candidate role with its own call budget, is spawned_by the caller, and gets a lineage record (evolve-trials / evolve_replica: run, role, agent, parent, delta, margin). Nothing in the parent world changes; bringing a role back is a separate, reviewed step.

world run trial2 evolve replicate ScoutTrial 1
# Not earned: the candidate beat the incumbent by 0.0 points; this session needs 2.0
world run trial2 evolve replicate ScoutTrial 2 n:2
# earned: +100.0 points over hasty (bar 3.0) · spawned scoutv2r2n1, scoutv2r2n2

Adopting a winner into the parent world is a separate, reviewed step — see World Collective → Bringing a winner home.

Cite a benchmark run by its id and evolve evaluate checks it: the run must exist, be completed and have answered something, or the evaluation is refused. Its verified score is stored with the evidence, together with what it measured — for a marina:<name> target, the agents serving it, their role and the hash of the exact system prompt in force:

benchmark run smoke --model marina:scout-v2 # → br_…
evolve evaluate ScoutTrial 1 | sharper answers on the smoke set benchmark:br_da6d8043_muhobbx8
# verified benchmark:br_da6d8043_muhobbx8 smoke 93.3% (15/15) · marina:scout-v2 · Scout2 role scout-v2 prompt ab12cd34ef56

Free-text labels (benchmark:prompt-trial-2026-08-04) are still accepted as notes; only run ids (br_…) are resolved.

Use parent=<run-id> to preserve experimental lineage:

evolve propose PromptTrial | Refine the successful shorter prompt | prompt:candidate-3 | parent=1

Available status controls are explicit:

evolve sessions
evolve status PromptTrial
evolve pause PromptTrial
evolve resume PromptTrial
evolve complete PromptTrial

Run and time budgets refuse new proposals after exhaustion but do not silently pause, complete, or otherwise change the session. The creator decides what happens next.

After participants complete a controlled trial, verify its durable evidence outside the agent loop (participants can preview the same verdict in the world with evolve qualify; the gate is the external probe, which reads the evidence over HTTP rather than through the loop being judged):

bun run qualify:evolution
bun run qualify:evolution http://marina.example:3300

The bounded probe waits for a started session, proposal, attributed evidence, and decision. It also fails if independent-review attribution collapses or passive constitutional defaults were changed. It is read-only: it cannot propose, decide, continue, activate, or promote.

In-system Pi agents receive the scoped marina_evolve tool only while participating in an active session. MCP clients receive an evolve tool, SDK clients use client.command("evolve ..."), and humans use the same command directly. All surfaces converge on the same command handler, participant checks, persistence, and audit trail.

Exercise socket churn and command latency against a disposable or staging world:

bun run soak --connections=25 --duration=60 --rate=2 --max-errors=0 --max-p95=2000
bun run soak:churn --url=ws://localhost:3300 --clients=12 --cycles=20
bun run soak:churn:local

The driver exits non-zero when clients fail to connect, errors exceed the threshold, no responses arrive, or round-trip p95 exceeds the bound. Its final line is machine-readable JSON for CI and demo qualification. Do not aim a high-rate run at a production world without considering its normal participants and provider limits.

For a disposable three-role real-model trial, run bun run trial:evolution:local. It starts and stops an empty temporary world, asks three agents to join voluntarily, stages proposer → evaluator → reviewer work, and exits non-zero unless durable evidence passes qualify:evolution. Use MARINA_TRIAL_MODEL and MARINA_TRIAL_TIMEOUT_MS to select the provider-neutral model identifier and bound. A failed run is evidence: the harness never fills in an agent’s proposal, evaluation, or decision.

The Experiments dashboard expands each evolution session into lineage, candidate, evidence, attribution, decision, guardrail, activity, communication, tool-use, and latency detail. Token and cost fields remain explicitly unavailable until provider-neutral per-session attribution is durable; Marina does not infer them from activity.