BLOG

Introducing Seahaven: Synthetic Worlds for Agent Evals and RL

Seahaven is an open source Python framework for building synthetic worlds for AI agents: stateful, resettable, parallel environments for evals and RL.

Remember The Truman Show? They built an entire town so real that Truman never knew it was fake.

Agent evals and reinforcement learning need the same thing: worlds so realistic that the agent doesn't know it's being tested. Production and staging can't do that. They're shared, and you can't reset them. Hand-written mocks reset fine, but they aren't real enough to fool an agent.

So we built Seahaven, an open source Python framework for building synthetic worlds for AI agents. It's named after Truman's town.

Why we built it

We've been optimizing long-running agents at Kiln: rewriting their prompts, tools, skills and subagents, and keeping the changes that score better. We've been working on some version of this problem for over a decade (at Apple, our own startup, now Kiln).

One-shot tasks never needed a world. You send a prompt and grade the answer. Agents that run for dozens of turns across lots of tools are different. Their evals need three things:

  • A known starting state, the same for every run.
  • Stateful tools. A write on turn 3 changes a read on turn 30.
  • A state diff. Did the agent make the right change in the world? Grade that, not the transcript.

And they need it thousands of times, in parallel, with every run isolated from the others.

We built a few of these environments by hand. Doable, but hard, and every one rebuilt the same layer: a database per run, frozen starting states, parallel instances, clock control, a log of every change. That layer deserved a great framework, so we built one. You write just the logic specific to your world. Seahaven handles the rest.

What's a synthetic world?

A Seahaven world is a working copy of your agent's tools and data. It's a schema (SQLite tables) plus tools (plain Python functions). The agent calls the tools exactly like it would call the real thing. It can read, write and even break things, without touching anyone else's run.

How realistic can it get? Our Stripe World mocks Stripe's Billing and Payments core: 24 tables and 155 API operations, behind the same tools as Stripe's own MCP server. It also serves Stripe's REST API, so the official Stripe SDKs work against it unchanged.

Seahaven vs. production vs. mocks

SeahavenProduction or stagingHand-written mocks
Realistic tools and data
Stateful across arbitrary tool calls
A private instance for every run
Hundreds of parallel instances
Every run starts from a known state
Reproducible
Every change logged for grading
Safe for the agent to break things

Production is realistic but shared. Mocks are private but fake. Seahaven gets you the realism of production with the isolation of a mock.

What Seahaven handles

  • A private world per run. Each instance is its own SQLite database, copied from a fixture in milliseconds and thrown away at the end. The agent can break anything.
  • Fixtures. Freeze known starting states like small_startup, agency or big_co, and reuse them across every run.
  • Hundreds of parallel instances. Serve hundreds of worlds per process, at thousands of requests per second. Each new connection gets its own world.
  • Grade on state, not transcripts. Every row the agent changed is logged, before and after, in one JSON document. Did the agent refund the right charge? Check the row.
  • Reproducible. Same fixture, same clock, same random seed: the same run, every time. The clock and seed are wired into SQLite itself, so even CURRENT_TIMESTAMP and random() in SQL replay exactly.
  • Composable. Build a world once and reuse it everywhere. Your company world can add Stripe World and a chat world, and the agent sees one tool list.

Build a world in minutes

Start with the scaffold:

terminal sh
uvx seahaven new crm_world
cd crm_world && uv sync

That writes a complete project: schema, tools, tests, a fixture generator, and an AGENTS.md.

A tool is just a Python function:

crm_world/tools.py python
@world.tool
def create_lead(ctx: seahaven.Ctx, email: str) -> dict[str, str]:
    """Add a contact to the pipeline as a new lead."""
    lead = {"id": ctx.ids.uuid(), "email": email, "stage": "lead", "updated_at": ctx.clock.iso()}
    ctx.db.execute("INSERT INTO contacts VALUES (?, ?, ?, ?)", *lead.values())
    return lead

Seahaven builds the tool's JSON schema from the signature, and the docstring is what the agent reads.

Then run your agent against a fixture, and grade what it changed:

rollouts.py python
for rollout in range(100):
    with world.instance("big_co", seed=rollout) as inst:
        run_agent(inst)  # your agent, your harness
        reward = grade(inst.state())  # every row the agent changed

Or let your coding agent build it. Seahaven is designed to be built by coding agents. seahaven new writes an AGENTS.md that points your agent at the docs for the version you have installed, not stale ones from the web. seahaven check catches the mistakes that are easy to make and hard to notice (a table without a primary key, a timestamp from the wall clock) and tells the agent the exact fix. Describe the product you want to mock, and let it work.

Plugs into what you already use

  • OpenEnv. Every world is an OpenEnv environment, the open standard for RL environments. Use it with TRL's OpenEnv support or any OpenEnv client, in any language, or publish it to Hugging Face for anyone to use.
  • MCP. seahaven mcp connects a world to Claude, Cursor or any MCP client. Explore a world by hand, or try a task yourself before you give it to an agent.
  • Web console. seahaven serve includes a web UI: open instances, call tools and inspect state in your browser.

Evaluate and optimize agents with Kiln

Kiln connects to any Seahaven world. Write scenarios against a fixture, evaluate your agent on the state it leaves behind, then use the Kiln Harness Optimizer to improve the harness: prompts, tools, skills, subagents and model choice, all scored against those evals.

Seahaven is standalone, and you don't need Kiln to use it. But if your goal is a better agent, Kiln is built for exactly that loop.

Get started

Seahaven is open source (MIT) and on PyPI.

Which world should we build next? Tell us on Discord.

Jump to section
Newsletter

New posts in your inbox.

Build AI that actually works.

Ship custom AI products with evals, fine-tuning, and prompt optimization built in.

macOS, Windows, and Linux