Training worlds for agents
From the Meridians Wiki · Public · Maintained · joint
The same insertion, with a policy in the seat. A kept world an agent is dropped into that it cannot have memorised, a task, a perspective, an interval — and a log that says what it knew, what it did, what it cost, and what changed.
pinned Domain@revision × task × seat × interval × policy@version
→ attributable trajectories
→ a scoped profile: anticipation, discovery, calibration, recovery, cooperation, persistence
→ failures as replayable coordinates, not a scalar
Why a kept world is the right test
- History resists memorisation. A world that has run forward holds state no corpus contains; the test is the world's own past.
- The same world serves people and agents. Nothing is built twice: the training world a person plays is the one an agent is run through, with the reading swapped for the log.
- Families compound. Versioned origins of one world become a fixed instrument a policy can be compared against over time.
Two uses
- Testing — one policy, unseen circumstances, how it behaved.
- Training — cohorts through parallel origins under varied policies, tools, prompts, or rules; the result is trajectories and comparisons, never a Meridians claim that one is better.
Gates
| Claim | Evidence |
|---|---|
| A run is worth reading. | Someone runs their own agent through a pinned world and changes what they let it do. |
| The measure is real. | Repeatable across seeds and versions, discriminative between systems, useful to a team improving a model. |
| The world resists contamination. | A held-out continuation is not predicted better by a system that saw the earlier history, beyond what that history legitimately teaches. |
| It transfers. | Independent external evaluation; nothing above depends on it. |
Boundaries
Systems, not people. Declared synthetic world only — performance is evidence about that world until transfer is shown (WORLD_ENGINE). No omniscient context and no presentation-owned truth for the policy; canonical state stays canonical. The evaluation design lives in CUMULATIVE_VISION §21; the human twin is 12 — The practice record, whose proof programme borrows this leaf's measurement infrastructure — pinned worlds, anchored items, fit — for the independent transfer study. This leaf does not gate the consumer plan.