Familiar patterns
Start with a part of the work you want to organise:- Get a second opinion. A different model reviews the work and the writer revises it with the notes.
- Know when to stop. A judge reads the review history and decides whether another revision is worth doing.
- Ask a panel. Several reviewers read the work, with a threshold for agreement and each opinion kept in the result.
- Ask a person. Put a decision in front of a person and pause until they answer.
- Automate the routine work. Let a command’s result decide whether a model needs to revise its work.
- Pick up unfinished work. Resume a recorded
workflow()without repeating its finished stages.
judge() from @obversa/runtime puts that decision
inside the workflow:
examples/teams/feature-delivery.ts (excerpt)
Motivation
Agents aren’t deterministic. Left to themselves, there’s no guarantee that every step of a process is followed, in the right order, under the right conditions. Obversa puts a deterministic process around the inference: many small jobs, each with its own review. The agent’s work stays narrow, so each job spends fewer tokens. The workflows generated in most coding agent harnesses are internal, and can’t be saved, shared or repeated. Obversa workflows are simple, harness-agnostic TypeScript files. You save one, share it and run it again, and the process runs the same way each time. It models process at every layer: the organisation, the team and the individual.Inference at the leaves
Inference belongs mostly at the leaves. Around each inference step you can build as much process as you like: deterministic scripts, tools and checks, so the job an agent has to do is clearer, simpler, better bounded and cheaper. Squeeze the determinism out of the agents and what’s left is a set of very specific, focused agents doing narrow tasks. That improves quality. It also means a small open-weight model tuned to a narrow, well-bounded task can do it as well as a frontier model would, at a fraction of the cost. Process engineers call this value stream mapping: understand the process, then make the steps that don’t add direct value cheaper or faster, or cut them out. Here that means lifting every deterministic step out of the agents, so their context shrinks to the parts that truly need inference.Key concepts
- Workflows. The team as a TypeScript file: the
steps, who does each one, where the reviews are, and where a person
decides.
workflow()is the short form;dag()is where your own code, a branch, a tournament or a team inside a team go. - Feedback loops. A review or a test fails a step, and the step runs again with the findings, up to a limit you set.
- The record. An append-only event log of the run, the single source of truth a stopped run carries on from.
- Running. The runtime runs one bounded step at a time. The runner restarts a killed run from its record.
- Surfaces. One question on a local page in front of a person. The run waits for the answer, then the page is gone.
- Memory. Files a step can open again later, behind one port with an adapter for each place they can live.
- Workspace. A Git worktree per writer, captured and verified, so two writers never collide.
Install
Install the bundle:Terminal
Next steps
Installation
Add the packages and prepare your project.
First Run
Have Claude write, run the tests, and ask Codex to review.
Examples
Complete teams by field, with what each run printed.
Patterns
Familiar ways to review work, make decisions and hand over a task.
Evals and Gates
The checks that decide the next step, and the person at the gate.
Recording Runs
What the event log holds and how a run carries on from it.
Driving Runs
Supervised local runs that restart after a crash.
Surfaces and Hosts
Watch a run in the browser, or drive it from cmux.
Writing Workflow Shapes
Build a shape of your own over the public contract.
Memory
One port, three adapters, and the reasoning behind each change.
Workspace
Capture, verify and fork a Git workspace.