Skip to main content
Obversa lets you build teams of agents around the way people work: writing, reviewing, making decisions and asking for help. One agent writes a draft. Another reads it and returns notes. A judge decides whether another revision is worth doing, and the workflow sends the feedback back to the writer. A person makes the calls you leave to them. These patterns fit an article, a research brief or a support reply as well as a software change. Put that process in a TypeScript file. Use Claude Code, Codex, Grok or OpenCode with the CLIs you already have signed in, or give a role an API engine. Tests run as commands. The workflow sends their failures and the review notes back to the writer, and keeps the run in a plain-file record. No server or database is required.

Familiar patterns

Start with a part of the work you want to organise:
  • Get a second opinion. A different model reviews the work and the writer revises it with the notes.
  • Know when to stop. A judge reads the review history and decides whether another revision is worth doing.
  • Ask a panel. Several reviewers read the work, with a threshold for agreement and each opinion kept in the result.
  • Ask a person. Put a decision in front of a person and pause until they answer.
  • Automate the routine work. Let a command’s result decide whether a model needs to revise its work.
  • Pick up unfinished work. Resume a recorded workflow() without repeating its finished stages.
Browse the patterns for more ways to work together. The complete workflows combine them around real use cases, including research, editorial work, support and software delivery. The feature team gives Jev the decision about another revision. judge() from @obversa/runtime puts that decision inside the workflow:
examples/teams/feature-delivery.ts (excerpt)
The judge weighs every finding, a blocking one included, and decides whether another round is worth doing. The run ends when it stops the rounds or the review passes. Know when to stop shows the same pattern in a writing workflow. Add a person’s decision where your process needs one.

Motivation

Agents aren’t deterministic. Left to themselves, there’s no guarantee that every step of a process is followed, in the right order, under the right conditions. Obversa puts a deterministic process around the inference: many small jobs, each with its own review. The agent’s work stays narrow, so each job spends fewer tokens. The workflows generated in most coding agent harnesses are internal, and can’t be saved, shared or repeated. Obversa workflows are simple, harness-agnostic TypeScript files. You save one, share it and run it again, and the process runs the same way each time. It models process at every layer: the organisation, the team and the individual.

Inference at the leaves

Inference belongs mostly at the leaves. Around each inference step you can build as much process as you like: deterministic scripts, tools and checks, so the job an agent has to do is clearer, simpler, better bounded and cheaper. Squeeze the determinism out of the agents and what’s left is a set of very specific, focused agents doing narrow tasks. That improves quality. It also means a small open-weight model tuned to a narrow, well-bounded task can do it as well as a frontier model would, at a fraction of the cost. Process engineers call this value stream mapping: understand the process, then make the steps that don’t add direct value cheaper or faster, or cut them out. Here that means lifting every deterministic step out of the agents, so their context shrinks to the parts that truly need inference.

Key concepts

  • Workflows. The team as a TypeScript file: the steps, who does each one, where the reviews are, and where a person decides. workflow() is the short form; dag() is where your own code, a branch, a tournament or a team inside a team go.
  • Feedback loops. A review or a test fails a step, and the step runs again with the findings, up to a limit you set.
  • The record. An append-only event log of the run, the single source of truth a stopped run carries on from.
  • Running. The runtime runs one bounded step at a time. The runner restarts a killed run from its record.
  • Surfaces. One question on a local page in front of a person. The run waits for the answer, then the page is gone.
  • Memory. Files a step can open again later, behind one port with an adapter for each place they can live.
  • Workspace. A Git worktree per writer, captured and verified, so two writers never collide.

Install

Install the bundle:
Terminal
Installation covers the tools and sign-in needed by each engine. First run walks through a complete file and its output.

Next steps

Installation

Add the packages and prepare your project.

First Run

Have Claude write, run the tests, and ask Codex to review.

Examples

Complete teams by field, with what each run printed.

Patterns

Familiar ways to review work, make decisions and hand over a task.

Evals and Gates

The checks that decide the next step, and the person at the gate.

Recording Runs

What the event log holds and how a run carries on from it.

Driving Runs

Supervised local runs that restart after a crash.

Surfaces and Hosts

Watch a run in the browser, or drive it from cmux.

Writing Workflow Shapes

Build a shape of your own over the public contract.

Memory

One port, three adapters, and the reasoning behind each change.

Workspace

Capture, verify and fork a Git workspace.