Remember The Truman Show? They built an entire town so real that Truman never knew it was fake.
Agent evals and reinforcement learning need the same thing: worlds so realistic that the agent doesn't know it's being tested. Production and staging can't do that. They're shared, and you can't reset them. Hand-written mocks reset fine, but they aren't real enough to fool an agent.
So we built Seahaven, an open source Python framework for building synthetic worlds for AI agents. It's named after Truman's town.
Why we built it
We've been optimizing long-running agents at Kiln: rewriting their prompts, tools, skills and subagents, and keeping the changes that score better. We've been working on some version of this problem for over a decade (at Apple, our own startup, now Kiln).
One-shot tasks never needed a world. You send a prompt and grade the answer. Agents that run for dozens of turns across lots of tools are different. Their evals need three things:
- A known starting state, the same for every run.
- Stateful tools. A write on turn 3 changes a read on turn 30.
- A state diff. Did the agent make the right change in the world? Grade that, not the transcript.
And they need it thousands of times, in parallel, with every run isolated from the others.
We built a few of these environments by hand. Doable, but hard, and every one rebuilt the same layer: a database per run, frozen starting states, parallel instances, clock control, a log of every change. That layer deserved a great framework, so we built one. You write just the logic specific to your world. Seahaven handles the rest.
What's a synthetic world?
A Seahaven world is a working copy of your agent's tools and data. It's a schema (SQLite tables) plus tools (plain Python functions). The agent calls the tools exactly like it would call the real thing. It can read, write and even break things, without touching anyone else's run.
How realistic can it get? Our Stripe World mocks Stripe's Billing and Payments core: 24 tables and 155 API operations, behind the same tools as Stripe's own MCP server. It also serves Stripe's REST API, so the official Stripe SDKs work against it unchanged.
Seahaven vs. production vs. mocks
| Seahaven | Production or staging | Hand-written mocks | |
|---|---|---|---|
| Realistic tools and data | |||
| Stateful across arbitrary tool calls | |||
| A private instance for every run | |||
| Hundreds of parallel instances | |||
| Every run starts from a known state | |||
| Reproducible | |||
| Every change logged for grading | |||
| Safe for the agent to break things |
Production is realistic but shared. Mocks are private but fake. Seahaven gets you the realism of production with the isolation of a mock.
What Seahaven handles
- A private world per run. Each instance is its own SQLite database, copied from a fixture in milliseconds and thrown away at the end. The agent can break anything.
- Fixtures. Freeze known starting states like
small_startup,agencyorbig_co, and reuse them across every run. - Hundreds of parallel instances. Serve hundreds of worlds per process, at thousands of requests per second. Each new connection gets its own world.
- Grade on state, not transcripts. Every row the agent changed is logged, before and after, in one JSON document. Did the agent refund the right charge? Check the row.
- Reproducible. Same fixture, same clock, same random seed:
the same run, every time. The clock and seed are wired into SQLite itself,
so even
CURRENT_TIMESTAMPandrandom()in SQL replay exactly. - Composable. Build a world once and reuse it everywhere. Your company world can add Stripe World and a chat world, and the agent sees one tool list.
Build a world in minutes
Start with the scaffold:
uvx seahaven new crm_world
cd crm_world && uv syncThat writes a complete project: schema, tools, tests, a fixture generator, and
an AGENTS.md.
A tool is just a Python function:
@world.tool
def create_lead(ctx: seahaven.Ctx, email: str) -> dict[str, str]:
"""Add a contact to the pipeline as a new lead."""
lead = {"id": ctx.ids.uuid(), "email": email, "stage": "lead", "updated_at": ctx.clock.iso()}
ctx.db.execute("INSERT INTO contacts VALUES (?, ?, ?, ?)", *lead.values())
return leadSeahaven builds the tool's JSON schema from the signature, and the docstring is what the agent reads.
Then run your agent against a fixture, and grade what it changed:
for rollout in range(100):
with world.instance("big_co", seed=rollout) as inst:
run_agent(inst) # your agent, your harness
reward = grade(inst.state()) # every row the agent changedOr let your coding agent build it. Seahaven is designed to be
built by coding agents. seahaven new writes an AGENTS.md that points your agent at the docs for the version you
have installed, not stale ones from the web. seahaven check catches
the mistakes that are easy to make and hard to notice (a table without a primary
key, a timestamp from the wall clock) and tells the agent the exact fix. Describe
the product you want to mock, and let it work.
Plugs into what you already use
- OpenEnv. Every world is an OpenEnv environment, the open standard for RL environments. Use it with TRL's OpenEnv support or any OpenEnv client, in any language, or publish it to Hugging Face for anyone to use.
- MCP.
seahaven mcpconnects a world to Claude, Cursor or any MCP client. Explore a world by hand, or try a task yourself before you give it to an agent. - Web console.
seahaven serveincludes a web UI: open instances, call tools and inspect state in your browser.
Evaluate and optimize agents with Kiln
Kiln connects to any Seahaven world. Write scenarios against a fixture, evaluate your agent on the state it leaves behind, then use the Kiln Harness Optimizer to improve the harness: prompts, tools, skills, subagents and model choice, all scored against those evals.
Seahaven is standalone, and you don't need Kiln to use it. But if your goal is a better agent, Kiln is built for exactly that loop.
Get started
Seahaven is open source (MIT) and on PyPI.
Which world should we build next? Tell us on Discord.