Skip to content

Social Simulation Arena

Marina can enter the Social Simulation Arena, a live forecasting benchmark run by Social Atoms at MIT. Every week about a dozen questions open: presidential approval (Economist/YouGov, Civiqs, Morning Consult), consumer sentiment (UMich, NY Fed, AAII), Google Trends shares and the Wikipedia weekly top 10. An entrant forecasts each one before the number is published. Forecasts are sealed until the round locks, then scored in public against a persistence reference (repeat the last value): skill 0 ties it, above 0 beats it.

Marina enters as a participant through the arena’s signed route: it signs each forecast with its own Ed25519 key and posts it itself. Nothing is exposed to the internet and no credential is shared — the registration carries only the public key.

src/arena/forecast.ts is the baseline every round gets. It keeps persistence’s mean. The arena’s own persistence uses a fixed sd = 1.5 whatever the series’ scale, so Marina replaces the spread with one sized to how the series moves, but only where the series’ own history supports it on the mean of per-round skill. Everywhere else it files exact persistence.

Round shape Marina’s answer
Number (continuous_normal) persistence mean ± calibrated or 1.5 spread
Profile (profile_energy) the same, per cell
Ranking (ranking_list) last-7-day Wikipedia pageview totals, Main_Page and non-articles excluded

arena show <round_id> prints exactly what would be filed and which spread rule each series used.

MARINA_ARENA_FORECASTER=model:<provider/model> puts a model on top of the baseline — any model Marina can route (openrouter/deepseek/deepseek-v4-pro, anthropic/claude-sonnet-5, …), with the provider’s usual key. The model sees the question, the frozen history and the baseline, answers a distribution, and that answer is shrunk toward the baseline (MARINA_ARENA_MODEL_WEIGHT, default 0.5); a malformed answer, a failed call or a jump beyond four baseline sds keeps the baseline. Closed-book: no web.

Measure before you switch — nothing is filed:

Terminal window
bun run arena evaluate --forecaster model:openrouter/deepseek/deepseek-v4-pro --weight 0.25

prints skill per family for the baseline, the blend and the raw model on every resolved round, with the model’s cost. Run-to-run model noise is large at the current sample size, and a model whose training data covers a round’s release could know its answer; weigh rounds after its cutoff.

MARINA_ARENA_FORECASTER=crew:<model> — or crew:<statistician>,<analyst>,<skeptic> to give each role its own vendor — runs three roles per numeric round over the baseline (src/arena/crew.ts):

Role Sees Does
statistician (the quant) the dated weekly history and, for Civiqs, the last 21 daily tracker readings (from the newest snapshot fetched by the lock — archive only once the lock has passed, the live dashboard for an open round) proposes a distribution; told that the start forecast is the default and that trend or reversion stories usually lose at this horizon
analyst the question, the last 8 dated values, the crew’s lessons for this series proposes from pollster behaviour and past misses
skeptic the start forecast, both proposals and the lessons (the crew’s track record here) decides how much of their move to trust (0 = stay on the start forecast)

Every role — and the research analysts and the model: forecaster — is told the same true account of the round (src/arena/prompt-context.ts): what the start forecast is (the nowcast — the freshest daily Civiqs reading, with its date and the days left to resolution — or the persistence baseline), the resolution rule (Civiqs: the dashboard value on the release day, i.e. the daily readings after the lock, and Civiqs re-estimates its daily history nightly), the scoring rule (CRPS skill vs persistence; moving on weak evidence loses), and every value with its date.

Aggregation is deterministic code: the proposals’ mean move, scaled by the skeptic’s trust, with wild or broken proposals dropped — the skeptic can shrink a move, never enlarge it. After a filed round resolves, the autopilot writes a lesson note (outcome, the crew’s error next to persistence’s, which way it leaned) that the analyst recalls for that series next time; lessons only ever describe rounds already published. arena evaluate --forecaster crew:… replays the resolved rounds in lock order with the same learning, in a throwaway database (--no-learn to compare).

Formations — Marina’s orchestration patterns as forecasters

Section titled “Formations — Marina’s orchestration patterns as forecasters”

MARINA_ARENA_FORECASTER=formation:<pattern>:<model>[,<model>…] (up to twelve models) runs one of Marina’s orchestration patterns as a small forecasting protocol (src/arena/formations.ts) over the same truthful round context as the crew, started from the nowcast:

Pattern Protocol
ensemble independent proposals; the control for the others
deliberation propose → see the others’ anonymized proposals and reasons → revise once
debate two sealed advocates (above the start / at or below it); the last model judges direction and trust
chorus proposals broadcast → each member critiques one peer → each revises from the critique it got
pipeline (cascade) quant → analyst (sees the quant’s handoff) → skeptic (sets trust) — the crew, strictly sequential
mapreduce one model per driver (level/trend, calendar/publication, source quirks); reduce = sum of confidence-shrunk adjustments
blackboard a shared scratchpad; two passes in which each model adds or corrects evidence and a number
symbiosis a quant (the numbers) and an analyst (the context) exchange contributions; a revision must credit the partner’s; a gap > 0.5 start-sd triggers another exchange (at most two)
research hypothesis → the model picks a check (recent mean, trend, last-k deltas, typical move, daily readings after the last value) → Marina COMPUTES it → keep or revert → revise (two checks)

Aggregation is deterministic with the crew’s clamps: proposals beyond 4 start-sds are dropped, the median move is scaled by a trust (0.5 × the proposals’ agreement, or the judge’s/skeptic’s), the final move is capped at 2 start-sds, and the sd blends by the same trust with a floor of half the start sd. Each round’s calls, statuses, trust and cost are kept in the evaluate/shadow record.

Compositions. +then:<pattern>:<models> adds a second formation that judges the first one’s handoff (e.g. formation:mapreduce:…+then:debate:…); both shrink toward the same start, so a chain cannot compound a move. +research@<retriever>[,…] (retrievers as in research:) puts a research crew in front: it builds ONE dated dossier per round with the research agent’s retrieval, checks every cited figure against its page, and hands only the verified lines to every member of every formation. A composition with +research@ reads today’s web, so arena evaluate refuses it — record it with arena shadow run.

Measure a formation before filing it with bun run arena evaluate --forecaster formation:… (all resolved rounds, per family, with cost); compositions with +research@ are recorded with arena shadow run. Model reliability matters as much as the pattern: a model that spends its output budget on reasoning and returns no JSON silently turns a five-model formation into a smaller one, so check each run’s failed-proposal count in the record.

Civiqs publishes daily trackers but the arena samples them on Fridays, so a round’s history ends at last Friday while, by its Wednesday lock, several newer daily readings are public. The arena archives every snapshot it fetches (civiqs/ in its repo); MARINA_ARENA_FORECASTER=nowcast (the default when unset) moves each Civiqs mean — topline or profile cell — to the freshest daily reading in a snapshot fetched before the lock, keeping the baseline’s spread; every other round is the baseline. Deterministic and leakage-free, so it backtests (see Integrity below).

Live reading for open rounds. For a round whose lock has not passed, the nowcast also reads the Civiqs dashboard (src/arena/research/civiqs-live.ts, one paced request per tracker) and uses whichever reading is fresher. A round whose lock has passed never reads live data, so backtests are unaffected. MARINA_ARENA_CIVIQS_LIVE=off turns it off.

Replacing a filing. The arena’s signed intake keeps every version and scores the newest one accepted before the lock (up to 120 per round). bun run arena submit <round|due> --replace files a newer version of an accepted round; an unchanged forecast is not re-sent, and the autopilot never replaces.

How the board ranks. An entrant’s row is its mean skill over the rounds it answered — unanswered rounds are not counted — and skill is 1 − CRPS / persistence CRPS against a persistence null frozen when the round’s call window opens (the round’s weekly history, so for Civiqs last Friday’s value).

No single forecaster wins every family. MARINA_ARENA_FORECASTER=routed answers each tracker family with its own forecaster, read from MARINA_ARENA_ROUTES (family=spec;…;*=spec); unset, every family gets the nowcast. Quote the value in .env (it contains ;). Or write a spec inline: route:civiqs=nowcast;aaii=formation:chorus:<m>,<m>,<m>;*=baseline. A route of skip leaves a family unanswered: submit reports it as “not answered” rather than a failure, and evaluate leaves it out of the mean, as the board does. Choose routes from arena evaluate per family, then shadow them before trusting them — samples per family are small.

The nowcast refuses rounds with no history (see no-anchor mode). To answer them, route their family to research, e.g. <family>=research:<m>[,<m>…]@<retriever>; routes are per family, so the family’s rounds that do have history go through research too (anchored as usual).

MARINA_ARENA_FORECASTER=tabh2o asks TabH2O, H2O.ai’s tabular foundation model (TABH2O_API_KEY), for each numeric round (src/arena/tabh2o-forecaster.ts). The training table is built only from what the round froze at its lock, the spread comes from the returned interval, and the answer is shrunk toward the start forecast with MARINA_ARENA_MODEL_WEIGHT, as model: does.

Spec Starts from
tabh2o the calibrated baseline
tabh2o@nowcast the nowcast
tabh2o:forecast[@nowcast] as above, using TabH2O’s time-series endpoint

A missing key, an error, too little history or a malformed reply files the start forecast and records why; ranking rounds keep the start. Calls are paced under the API’s rate limit and metered on the daily spend ledger. bun run arena evaluate --forecaster tabh2o backtests it like the nowcast (operator step; it costs money), and it works as a route target (route:<family>=tabh2o;*=nowcast).

About a third of the rounds are not single numbers. arena evaluate scores them exactly as the leaderboard does — the energy score over the arena’s deterministic point set for profiles, 1 − RBO for rankings (src/arena/score-shapes.ts, matching the arena’s published scores to 1e-4) — against the arena’s recorded persistence loss for each round.

  • Civiqs profiles: the nowcast moves every cell to its freshest daily reading.
  • Wikipedia top 10: views are weighted by recency (half-life 3 days) over three weeks of the arena’s wikitop/ archive, using days published before the lock (a two-day lag; daily lists are final once published, and the archive was partly backfilled); the half-life was chosen by backtest over the archived weeks.
  • Google Trends baskets: Trends re-normalises its index in every snapshot, so the lock’s own frozen per-cell history — what the persistence null reads — is used; the trends/ archive only fills in for a lock without one, complete weeks only (MARINA_ARENA_TRENDS_PARTIAL=on adds the partial week; mixed in the backtest).
  • YouGov crosstab profiles: no structured source yet; the baseline ties persistence.

research:<analyst>[,<analyst>,<analyst>] — one analyst per vendor — runs the full pipeline (src/arena/research/):

  1. Brief — a playbook per family (other pollsters’ readings with their previous reading, S&P moves for AAII, prices and inflation prints for consumer surveys, scheduled events for attention), bounded to facts after the series’ latest known reading, which is stated as already known. buildResearchBrief(round, lock, { nowcast }) starts that window at the daily nowcast’s date when one is newer than the history (Civiqs); without it, the last weekly value’s date. Each brief also carries short keyword queries (one per playbook item) for search APIs.
  2. Retrieve — MARINA_ARENA_RESEARCH_RETRIEVER (default openrouter-web:openai/gpt-6-luna, OpenRouter’s web search with URL citations; ~$0.03 per round). sonar:<model> uses Perplexity Sonar through OpenRouter (native search; sonar, sonar-pro, sonar-pro-search, sonar-reasoning-pro, sonar-deep-research), and a comma-separated list runs several engines on the same brief and merges their reports (each under its own heading, sources de-duplicated), e.g. openrouter-web:openai/gpt-6-luna,sonar:sonar-pro. tavily:basic / tavily:advanced calls Tavily’s search API directly (TAVILY_API_KEY): one news search per brief query, dated from the brief’s window, up to 8 results each; no model writes the report — each result (best 12 by score, anything published before the window dropped) becomes one line - <date> — <snippet> [<title>](<url>). Tavily also returns each page’s text, which the citation check reads instead of fetching the page. Cost is estimated at Tavily’s pay-as-you-go $0.008 per credit (basic 1, advanced 2 per query; ~$0.05 per brief at advanced). New retrievers are shadow-only until they have resolved rounds to their name.
  3. Verify citations — every dossier line that cites a page has its figures looked up in that page and is tagged [verified], [unverified] or [unreachable]. The page text is the one the retriever already fetched when it carries one (Tavily), else a fetch through the SSRF guard (up to 12 pages, 15 s each, a descriptive User-Agent, the first 16 MB read — larger pages are truncated, not rejected). Publishers whose terms bar bots (NO_FETCH_DOMAINS: YouGov, AAII, Conference Board, CivicScience) are never read by either route, so their lines stay [unreachable].
  4. Analysts — forecast from the dated history, the start forecast (named: the nowcast with its date, or persistence), for Civiqs the recent daily tracker, the resolution and scoring rules, and the tagged dossier; told to use other sources for changes, never levels (pollsters differ in population and house effect), and only changes the start reading does not already include.
  5. Judge — MARINA_ARENA_RESEARCH_JUDGE (default jev: jev-1.13 via OpenRouter’s Decisions API) scores each rationale’s quality and grounding in the evidence the analysts were given: the series’ own recent history, the start forecast and the daily tracker (labeled as structured source data, not web research) plus the verified dossier lines; ungrounded, or a judge outage ⇒ no weight. The record keeps judge {provider, model, calls, latencyMs, costUsd, errors} (and each proposal’s judge latency and cost), and the judge’s cost is in the shadow row’s cost_usd (the daily spend cap already sees it through the metered provider).
  6. Aggregate — judge-weighted mean move × confidence × MARINA_ARENA_RESEARCH_TRUST (0.5).

Rounds with no history (no-anchor mode). A one-off numeric round (an election result, say) can lock with an empty answer_history. Every other forecaster (baseline, nowcast, discovered, model:, crew:, formation:) anchors on the last published value and refuses it with “no history to forecast from”; research: answers it without an anchor (noAnchorForecastRound in src/arena/research/forecaster.ts):

  • the brief asks for the level (published forecasts of the quantity, prediction-market prices, the data they rest on, the base rate) over the 30 days before now (or the lock);
  • the analysts are told there is no history and no start forecast, and may abstain; the judge is a filter — only a grounded proposal (weight > 0) counts, a judge outage counts as none;
  • sanity bounds come from the question: a named total (all N decided) bounds the mean to [0, total], a percentage unit to [0, 100]; a proposal outside them is dropped, then any mean more than 3 robust spreads (max of the median sd and 1.4826 × MAD) from the median;
  • the answer is the median of the remaining means, with sd = max(median analyst sd, the means’ sample sd, 7 % of the median) — the floor is derived from the level, never set per round;
  • with nothing usable the round is not answered: a NoAnchorRefusal whose detail keeps the dossier, proposals, roles and judge (and its cost reaches the shadow ledger); arena research <round> prints that trail. The answer carries anchor: "none" and noAnchor (bounds, used, dropped, median, medianSd, dispersion, floor).

Web research cannot be backtested (a search run later finds the answer), so it is measured in shadow: bun run arena shadow run due records what it would file (with the whole dossier and judged proposals); re-running it re-records, and the forecast scored is the last one recorded before lock, as a filing would be — every recording stays in the append-only ledger, and a record younger than 6 h is not duplicated), shadow score scores resolved ones against persistence and the baseline, shadow list shows the record, bun run arena research <round> runs it once and prints everything. MARINA_ARENA_SHADOW=<spec> records hourly from the tick job — no entrant or key needed.

bun run arena discover [--tracker T] [--proposer provider/model] [--n N] runs the loop that found the Civiqs nowcast, automatically (src/arena/discovery/):

  1. A family’s clean resolved rounds are split by time: the older 60 % are discovery, the rest holdout.
  2. A proposer model sees the family, a sample of its history, the signal language (a menu of centres — last, nowcast, ewma:α, mean:k, median:k, trend:k, nowcast-shrink:w (last weekly value + w × (nowcast − it)), nowcast-mean:k (mean of the last k daily readings in the nowcast’s snapshot, i.e. Civiqs’s revised values as published before the lock) — and spreads — arena, baseline, rms:w, mad:w, scale:k), and the incumbent’s and every earlier attempt’s discovery score. It never sees a holdout score. Signals are data, never code.
  3. Each new proposal is scored on both halves. It is promoted only if it beats the incumbent (the nowcast over the calibrated baseline) on the holdout by a margin that grows with the number of signals tried for the family (0.02 + 0.01·log₂(1 + tried)) and does not lose on discovery.
  4. Every scored attempt is kept as a note (arena-discovery, type signal); arena signals (or bun run arena signals) lists them, and the next discovery round is told not to repeat them.

The same loop runs in the world: arena discover [tracker:T] (see Operate). bun run arena discover --tracker T --signal <centre>/<spread> [--signal …] scores signals an operator proposes, with no model call, under the same split, margin and record — each one still counts as a try.

MARINA_ARENA_FORECASTER=discovered uses each family’s best promoted signal and the nowcast elsewhere. Promotion is necessary, not sufficient — record a promoted signal in shadow before it files. To trial a signal without touching the live record, promote it in a separate discovery record (DB_PATH=<scratch> bun run arena discover --tracker civiqs --signal <centre>/<spread>), then shadow discovered from that record against the nowcast for several weeks before filing it.

Integrity: what the backtest numbers can and cannot claim

Section titled “Integrity: what the backtest numbers can and cannot claim”

Audited 2026-09-25 (src/arena/evaluate.ts, test/arena-*.test.ts):

  • No outcome reaches a forecaster. Only the evaluator and the live lesson writer read resolutions; every forecaster sees only a round’s lock file and archives filtered to what existed before the lock (Civiqs snapshots fetched before it; Wikipedia lists published before it).
  • Rounds whose answer was already public are excluded for every forecaster (outcomePublicBeforeLock).
  • Model memorisation. Probed closed-book, DeepSeek V4 Pro, Claude Sonnet 5 and GPT-6 Luna claimed to know none of five resolved values. The baseline and the nowcast use no model at all.
  • In-sample design choices. The spread-selection metric, the Wikipedia half-life and the Trends partial-week default were chosen after looking at the resolved rounds; the Civiqs nowcast has no fitted parameter. Treat backtest numbers as optimistic. The honest test is forward: arena shadow run due --forecaster <spec> records predictions before the lock (for the nowcast, within a day of it), and arena shadow score scores them only once the arena resolves them.
  • Source terms. Civiqs, Wikipedia and Google Trends are rights-approved in the arena’s own review; Marina reads the arena’s public archive of them. The research agent’s citation check never fetches publishers whose terms bar bots or forwarding (YouGov, AAII, Conference Board, CivicScience — NO_FETCH_DOMAINS).
  • The ranking is hypothetical. Leaderboard entrants forecast live; Marina’s numbers are simulations on the same frozen inputs plus public archives available at each lock. Only filed forecasts count.
  1. Choose the entrant id — lower-case, permanent (for example h2oai-marina) — and the GitHub account that will own it. Only that account can change the registration later.

  2. Generate the signing key on the server that will file:

    Terminal window
    bun run arena keygen /srv/marina/arena-key.pem # writes mode 0600, never overwrites
  3. Configure (.env):

    Terminal window
    MARINA_ARENA_ENTRANT=h2oai-marina
    MARINA_ARENA_KEY_FILE=/srv/marina/arena-key.pem
  4. Write the registration and open the pull request from the owning account:

    Terminal window
    bun run arena registration --name "Marina" --org "H2O.ai" --github <login> \
    --out entrants/h2oai-marina.json

    Add the file to a fork of Social-Atoms/social-sim-arena and open the PR. The arena’s CI validates it; a maintainer approves a signing key.

  5. Rehearse — bun run arena submit due --dry-run shows every forecast inside the window without signing or sending anything.

  6. Turn on the autopilot once the registration is merged:

    Terminal window
    MARINA_ARENA_AUTOPILOT=on

    Every hour Marina files each round whose lock is within 24 hours (the arena’s own call window, so its inputs are as fresh as every other entrant’s) and that has no accepted forecast yet.

Any. arena, forecast, decision, market, position, probe, watch and web are global commands registered in every world, including the default Workbench, and filing goes through the operator CLI (bun run arena), not a room. What matters is the database: the submission ledger, shadow rows and discovery notes live in DB_PATH, which the server’s autopilot and the CLI share. Keep one DB_PATH for all arena work — starting a different world on a fresh database splits the ledger. The markets and prediction-lab worlds add binary yes/no market rooms, which do not match the arena’s continuous, profile and ranking targets.

Command Where What it does
arena / arena status in-world, rank 0 entrant, key readiness, autopilot, filed counts
arena rounds [n] in-world open rounds, soonest lock first, with Marina’s filing status
arena show <round_id> in-world the question and exactly what Marina would file
arena submissions in-world the signed record of what was filed
arena backtest [n] in-world baseline skill vs the arena’s persistence, per family
arena evaluate [baseline|nowcast|discovered] [tracker:T] [limit:N] in-world score a free forecaster against the baseline on resolved rounds, as the leaderboard scores them
arena shadow [list] · arena shadow score in-world the shadow ledger, and its score on outcomes no one had seen
arena shadow run <round_id|due> [forecaster:F] in-world record a free forecaster’s forecast for rounds about to lock (never filed)
arena discover [tracker:T] [n:N] · arena signals [tracker:T] in-world run signal discovery (one proposer call per family, rate limited), list every attempt
bun run arena submit <round_id|due> [--dry-run] [--forecaster …] [--weight w] operator CLI sign and file now
bun run arena evaluate [--forecaster …] [--weight w] [--out FILE] operator CLI score forecasters on resolved rounds; files nothing
bun run arena keygen <path> / registration operator CLI key and registration file

Filing is deliberately an operator act (CLI or env-set autopilot), never an in-world one: it speaks for the organization in public. Everything else in the loop runs in the world, so an agent can measure, discover and prove a signal forward on its own:

arena discover tracker:civiqs # propose → backtest (time split) → promote
arena evaluate discovered # the promoted signals vs the baseline, resolved rounds
arena shadow run due forecaster:discovered # record forecasts for rounds about to lock
arena shadow score # once they resolve: the only test on unseen outcomes

In-world runs are limited to forecasters that make no model calls (baseline, nowcast, discovered); tabh2o, model:, crew:, formation: and research: spend real money and stay operator steps (bun run arena evaluate|shadow --forecaster …). arena discover is rate limited per entity (2, then 1 an hour) and runs one at a time, because every attempt raises that family’s promotion bar. Its proposer is MARINA_ARENA_PROPOSER (default Claude Sonnet 5 via OpenRouter). Discovery attempts and shadow rows land in the world database — the same notes and ledger the operator CLI and autopilot read.

readiness reports an arena check once MARINA_ARENA_ENTRANT is set, and warns when the key is missing or readable by other users.

  • One forecast per round. A round with an accepted forecast is never filed again; the ledger (arena_submissions, append-only) records every signed request.
  • Safe retries. A send that failed in transit is re-sent with the same signed request within four minutes (the arena deduplicates by request id); after that a fresh request is signed. A 4xx is final and never retried as-is.
  • Refuses rather than guesses. A round without the inputs to forecast from, or an answer that would break the arena’s contract (missing profile cell, sd ≤ 0, wrong ranking length), is skipped with a reason.
  • Key custody. The key file must be mode 0600 or it is refused; only its public half is ever printed or published. Rotate by adding a new key id to the registration.
  • No redirects. The signed POST goes to the configured origin only; outbound reads use the SSRF guard.
Variable Default Meaning
MARINA_ARENA_ENTRANT unset (off) the registered entrant id
MARINA_ARENA_KEY_FILE — PKCS#8 PEM (or raw 32-byte seed), mode 0600
MARINA_ARENA_KEY_ID k1 the key’s id in the registration
MARINA_ARENA_AUTOPILOT off on files due rounds hourly
MARINA_ARENA_WINDOW_HOURS 24 how close to its lock a round is filed (max 168; invalid ⇒ 24)
MARINA_ARENA_URL / MARINA_ARENA_AUDIENCE production a rehearsal fork’s intake
MARINA_ARENA_DATA_URL the arena repo on GitHub where rounds, locks and resolutions are read
MARINA_ARENA_FORECASTER nowcast no model calls; every non-Civiqs round is the baseline. Or baseline, discovered, tabh2o[:forecast][@nowcast] (experimental, TABH2O_API_KEY), model:<m>, crew:<m>[,<m>,<m>], formation:<pattern>:<m>[,…][+then:<pattern>:<m>[,…]][+research@<retriever>[,…]], research:<m>[,<m>,<m>]
MARINA_ARENA_MODEL_WEIGHT 0.5 share of the model’s move from the baseline that is kept
MARINA_ARENA_SHADOW unset a forecaster spec to record hourly in shadow (never filed)
MARINA_ARENA_TRENDS_PARTIAL off on counts a Trends basket’s partial current week
MARINA_ARENA_CIVIQS_LIVE on off stops the nowcast reading the live Civiqs dashboard for open rounds
MARINA_ARENA_RESEARCH_RETRIEVER openrouter-web:openai/gpt-6-luna the research agent’s search backend(s): openrouter-web:<model>, sonar:<model>, tavily:basic / tavily:advanced (needs TAVILY_API_KEY), comma-separated to merge
MARINA_ARENA_RESEARCH_JUDGE jev (with an OpenRouter key) jev, decisions (the configured MARINA_DECISIONS backend; falls back to jev) or none
MARINA_ARENA_RESEARCH_TRUST 0.5 most of the judged move the research agent takes