> ## Documentation Index
> Fetch the complete documentation index at: https://obversa.ai/docs/llms.txt
> Use this file to discover all available pages before exploring further.

# Runtime

> @obversa/runtime: declare a team or a graph, review inside it, ask a person, run it and record every step.

`@obversa/runtime` is the package you import to write and run a workflow:
the shapes (`workflow`, `pipeline`, `dag`, `loop`), the steps, the reviews,
the person's question, `run()` and the record it writes. It imports no
engine or memory adapter; your program supplies the instances.

## Install

<CodeGroup>
  ```bash npm theme={null}
  npm install @obversa/runtime
  ```

  ```bash pnpm theme={null}
  pnpm add @obversa/runtime
  ```
</CodeGroup>

Included in `@obversa/obversa`. The package also installs the
`obversa-record` command, which prints a run's record as a page a person
scans: [Read a record](/docs/recording/read-a-record).

## Quickstart

One agent, one job, one run. The engine is a stand-in from the testing
module, so the file runs offline:

```ts examples/one-agent-job.ts {6-10,12} theme={null}
import { agentJob, run } from '@obversa/runtime';
import { MockEngine } from '@obversa/runtime/testing';

const engine = new MockEngine(() => 'ready');

const job = agentJob({
  label: 'prepare-item',
  engine: 'offline',
  prompt: 'Prepare the item.',
});

const result = await run(job, {
  engine: 'offline',
  engines: { offline: engine },
});

console.log(result.outcome.status);
```

`agentJob` makes a step from a prompt and an engine name. `run` runs it
with the engines you pass and returns the outcome. Run it with
`npx tsx one-agent-job.ts`:

```text Output theme={null}
pass
```

## Declare a team

`workflow(name, { brief, roles, stages })` is the team, in a file: the
brief, the roles named once, the stages in order. Each stage says who does
it, what it writes, who reads it and where a red result goes:

```ts examples/teams/writer-reviewer-pair.ts (excerpt) {2,5-8,11,13,18,22} theme={null}
  return workflow('writer-reviewer-pair', {
    brief: briefFromFile('briefs/add.md'),
    options: { timeout: '10m' },

    roles: {
      write: engines.claude('claude-sonnet-4-5'),
      review: [engines.codex('gpt-5.6-luna')],
    },

    stages: [
      stage('write', {
        agent: 'write',
        writes: ['src/add.mjs', 'test/add.test.mjs'],
        desc: 'Write the function and its test from the brief.',
        gate: 'The files named in the brief exist in the workspace.',
        refine: 1,
      }),
      stage('test', {
        run: ['node', '--test', 'test/add.test.mjs'],
        desc: 'Run the test command against the written files.',
        gate: 'The test command exits 0.',
        sendsBackTo: 'write',
      }),
```

`stage(name, config)` is one step; exactly one of `agent`, `run`, `panel`
or `input` says what kind. Seats come from the engine plugins:
`claude('claude-sonnet-4-5')` from `@obversa/engine-claude-cli`,
`codex('gpt-5.6-luna')` from `@obversa/engine-codex-cli`,
`opencode(model, { executable })` from `@obversa/engine-opencode-cli`.
`person(question)` is a role a person fills. `briefFromFile(path)` reads a
Markdown file as the brief; its front matter may carry `files`, the paths
the brief expects written. `formatEvent(event, totals?)` turns one run
event into the line a person reads, and `RunResult.usage` is what the run
spent in the shape it takes. Every seat of a `reviewedBy` panel must report
a different model family from the stage's writer, or the team is refused
before any model runs. The whole file and a real run are on
[A writer and a reviewer](/docs/patterns/writer-and-reviewer).

## Build a graph

Under `workflow()` sit the shapes: `pipeline` for stages in order, `dag`
for a dependency graph, `sequence` and `parallel` for the two plain cases,
`loop` for a body that runs until a condition holds. Each returns a `Job`
you can run or nest. The steps are `fnJob` for a function, `agentJob` for a
prompt to an engine, `commandJob` for a command whose exit code decides:

```ts examples/feature-team.ts (excerpt) {7,10} theme={null}
export const featureDelivery = pipeline(
  'feature-delivery',
  [
    { name: 'analyse', job: analyse },
    { name: 'implement', job: implement },
    { name: 'test', job: testStage },
    { name: 'review', job: review, acceptsKickbackTo: ['implement'] },
    { name: 'approve', job: approve },
  ],
  { maxKickbacks: 2 },
);
```

A step that sends work back lists its targets in `acceptsKickbackTo`, and
each must be a step it depends on. A send-back to a step it does not list
fails the sender with an error that names both steps. `maxKickbacks` sets
the limit: a number is the send-backs allowed in the whole graph, and a map
gives each target its own number of
[refinements](/docs/concepts/feedback-loops#counting-rounds).
A target's number can be a [`judge()`](/docs/patterns/judge-stops-the-loop)
instead. The judge reads the dag's `brief` (the brief or ticket text the work
answers) and `useCase` (what the work is for), the target node's `desc` and
`gate`, and the file the target node names in `file`. `brief` and
`useCase` are optional strings on the dag's config.

The target's next run reads the send-back as `ctx.lastReview`, once. A
writing `agentJob` that runs again after a send-back and leaves its node's
`file` as it was fails: it returned the work unchanged, whatever else it
wrote. With no `file`, a change to any file in the workspace counts. When a
graph is a `loop()`'s body, each step's first run in a round reads the
loop's last review the same way. A step that fails with an error a retry
cannot fix fails the graph with that error, so a loop around the graph
stops, as it does around the step alone.

A `fnJob` function returns a full outcome, a one-line summary (a pass) or
nothing (a pass with the label as its summary); a throw is a fail carrying
the error. `commandJob` passes on exit 0 and otherwise fails with the
output as evidence, and `target` names the step that owns the fix.
`passed(name)` and `failed(name)` are conditions for a node's `when` that
read a named dependency's outcome, so an exit code chooses the branch; the
node `failed` reads must be `optional: true`. A node that ran out of
rounds meets `failed`. `isolated(job)` runs a step
in its own worktree and branch and lands the change on pass. `jobMeta`,
`copyJobMeta` and `renderPlan` read a job's declared shape and render it.

## Review inside it

`reviewPanel` runs several reviewers on one attempt and passes on a
threshold; a failing panel returns the findings to `target`, which runs
again:

```ts examples/feature-team.ts (excerpt) {8-9} theme={null}
const review = reviewPanel({
  label: 'review',
  reviewers: [
    { name: 'correctness', job: checks.correctness },
    { name: 'safety', job: checks.safety },
    { name: 'scope', job: checks.scope },
  ],
  pass: 2, // two of three agree and the step passes
  target: 'implement', // a failing panel returns the work here
});
```

`tournament` runs `n` candidates, each in its own worktree, and a `judge`
scores the finished ones; the highest lands:

```ts examples/tournament.ts (excerpt) {2-3,10} theme={null}
  const result = await run(
    tournament({
      name: 'retry-implementation',
      n: ANGLES.length,
      candidate: (i) => fnJob(`candidate-${i}`, async (ctx) => {
        await writeFile(join(ctx.workspace.dir, 'src/retry.ts'), TASK[1] + ANGLES[i]!);
        await writeFile(join(ctx.workspace.dir, 'candidate.test.ts'), CANDIDATE_TEST);
        await runNodeTest(ctx);
        return { status: 'pass' as const, data: { candidate: i } };
      }),
      judge: score,
    }),
    { cwd: repo },
  );
```

`kickback(to, reason)` and `revisionRequest(input)` build the outcome a
reviewer returns to fail a step with findings aimed at a named stage.
`quorum(k, ...conditions)` is met when `k` of its inputs hold. `agentCheck`
asks an engine a scored question as a condition; `confidenceCondition` and
`minConfidence` read a confidence and gate on it. `RECORDED_ENGINE_USAGE`
is the state key under which a run records which model answered each
engine call, so the review gate compares writers with reviewers by family.

## Ask a person

`humanReview(name, { question, input, interaction })` makes a review step
that keeps what you changed, what you wrote and whether you approved.
Use it as a loop's `review`. Its response
contains `feedback`, a composed `prompt`, and an explicit `decision`:
`changes-requested` sends work back; only `approved` passes. The writer's
`consumeFeedback: true` includes the full response on its next turn.
[Surfaces](/docs/concepts/surfaces#refine-together) shows an interactive example.

`interaction` names the handler with `id`, declares `responseSchema`, and
optionally supplies an async `answer(request, signal)` function. That
function can open a surface and return its payload. Return `undefined` if
the person cancels. Without a handler, answer through the run's stored
callbacks client.

An `agentJob` can also declare `interaction`. To ask a question, that agent
returns JSON with an `interaction` object containing `question` and `input`.
The answer's `feedback` and `prompt` return to the same model in a fresh
turn with the saved task context. This asks for input; it does not require
an approval decision. `person(question, { interaction })` uses the same
rich response in a workflow input stage.

`approval(label, { question, target?, answer? })` is a person's decision as
a step. Yes passes; no goes to `target` with the note as the finding, or
fails the step when there's no target; no answer pauses the run:

```ts examples/approval.ts (excerpt) {2-3} theme={null}
const approve = approval('approve', {
  question: 'Ship this change?',
  target: 'implement',
  answer: (request: CallbackRequest) => {
    const approved = JSON.stringify(request.input).includes('header');
    decisions.push(approved ? 'yes' : 'no');
    return approved
      ? { approved: true }
      : { approved: false, note: 'the ticket asked for a header row' };
  },
});
```

Under it sit the callback gates: `createCallbackGate` makes a question with
an identity, `createCallbackClient` holds it for one run,
`createStoredCallbackClient` keeps it across a process exit,
`replayCallbackClient` rebuilds pending questions from a history, and
`directRouter` is the smallest router. `createApprovalCallbackGate`,
`approvalSubjectDigest` and `resolveApproval` bind an approval to the exact
bytes it approves. [Callback gates](/docs/reviewing/callback-gates) is the guide.

## Run it

`run(job, options)` runs a job from a fresh record and returns a
`RunResult`: the `outcome`, the `usage`, and `monitor` when the page is on.
Turn the page on and read its address from the record:

```ts examples/monitor.ts (excerpt) {3,5} theme={null}
let address: string | undefined;
const result = await run(digest, {
  monitor: true,
  onEvent: (event: LoopEvent) => {
    if (event.kind === 'monitor') address = event.url;
  },
});
```

Every run emits one root `run:start` before it dispatches work and one
`run:end` before it returns, both as `LoopEvent`s to `onEvent`. Cancelling a
waiting workflow reports the run as aborted and leaves the question in the
record, so a later resume can still answer it. An attempt that runs past
its time budget comes back as a typed outcome, not a hang.

## Keep the evidence

`writeProofArtifact` stores one proof packet addressed by its digest, and
`createAcceptedResultRecord`, `resolveAcceptedResult` and
`acceptedResultMatches` bind a result to the inputs, proof, graph and
workspace it was judged on:

```ts examples/proof-bound-approval.ts (excerpt) {1} theme={null}
  const proofArtifact = await writeProofArtifact(
    storage.artifactStore,
    { namespace: storage.record.namespace, runId },
    {
      inputs: { proposal: digest('1') },
      result: { passed: true, tests: 21 },
    },
  );
```

`createProofCache` shares declared read-only evidence between workers.
[Proof-bound acceptance and approval](/docs/reviewing/proof-acceptance) is the
guide.

## Store and execute

For hosts and for shapes of your own: `compileGraph` validates a definition
against a graph type, `resolveGraphPlan` freezes what the run may use,
`persistRunDefinition` stores it, and `createGraphExecutor` runs it one
recorded decision at a time:

```ts examples/custom-graph.ts (excerpt) {1,3} theme={null}
const compiled = compileGraph(graphType, definition);
const state = events.reduce(compiled.reduce, compiled.initialState());
const plan = resolveGraphPlan(compiled.describe(), resolution);
```

`dagGraphType`, `convergence` and `teamGraphType` are the forms that ship.
`loadRunDefinition` reads a stored run back; `readRunPreflight` and
`interruptRunPreflight` read and close a run's engine checks;
`validateDomainEventId` checks an event id before a resume.
`createLocalRunStorage` from `@obversa/runtime/storage/local` is the
on-disk event and artifact store the examples use, and
`createGitWorktreeProvider` the workspace provider. The guides are
[Outside graph types](/docs/graphs/contract) and [Graph executor](/docs/graphs/executor).

## Options

### `recordTo` and `resume`

`recordTo` appends every event as one JSON line to the file you name, or
under `.obversa/records` with `'auto'`. `resume: true` reads that record
first: a `workflow()` or a `dag()` carries on from it and repeats no
finished step; a `loop()` or a plain job appends to it and runs again. A
graph that runs more than once, such as a loop's body, skips finished steps
only the first time. To resume, keep the graph's name, workspace and
shape the same (for a workflow, its brief, stage list and role seats too);
change any and the run starts again. [The record](/docs/concepts/record) says
what a resumed run repeats.

### `callbacks` and `onCallback`

`callbacks` is the client the run's questions go through, in memory (the
default, fresh per run) or stored. `onCallback: 'wait'` keeps the process
up until an answer arrives; the default `'exit'` records the question and
returns, so a schedule can start the same file again until the answer is
there. A record written in one mode resumes in the other.

### `engine` and `engines`

`engines` is the map of ready-made engine instances by name, and `engine`
the one a job or condition uses when it names none. The runtime imports no
adapter, so both are yours to supply. `withEnv(job, env)` pins environment
variables for a job and every job beneath it.

### The rest

| Field | Type | Default | Description |
| - | - | - | - |
| `signal` | `AbortSignal` | none | Abort the run from outside. |
| `cwd` | `string` | `process.cwd()` | The run's root working directory. |
| `environment` | `Environment` | none | An environment brought up for the run's root workspace. |
| `onEvent` | `(event: LoopEvent) => void` | none | Where every run event goes. |
| `params` | `RunBrief` | none | A frozen JSON object every job reads as `ctx.params`. |
| `state` | `Record<string, unknown>` | none | Seed for the shared mutable run state. |
| `memory` | `Memory` | none | The memory adapter every job sees as `ctx.memory`. |
| `monitor` | `boolean` | off; on under `supervise` | Serve the run's own page on a loopback port. |
| `budget` | `number` or `BudgetConfig` | none | Cap on total tokens; `{ limit, headroom?, soft? }`. |
| `supervise` | `boolean` | `false` | Register under `~/.obversa/runs/<runId>` for outside inspection. |
| `runId` | `string` | generated | The registry id, `[a-z0-9][a-z0-9-]*`. |
| `onLimit` | `LimitPolicy` | `'auto'` | What a rate limit, quota or budget does: `auto`, `wait`, `exit` or `fail`. |
| `maxWaitMs` | `number` | `300000` | Ceiling on one interruptible limit wait. |
| `judgeContextLimit` | `number` | `50000` | How many characters, counted as JSON, a [judge](/docs/patterns/judge-stops-the-loop#the-size-limit) reads. Long parts are cut to fit. |
| `cost` | `{ prices, baselineModel? }` | none | Price the measured usage; `baselineModel` adds the counterfactual. |

### Stage keys

| Key | What it means |
| - | - |
| `agent` | The role that does the work. The stage must say what it `writes`. |
| `effort` | For an agent stage, the reasoning effort for this stage only, over the seat's own. See [Engines](/docs/packages/engines#reasoning-effort). |
| `run` | A command, one string with no quoting or an array. Exit 0 passes. A red run goes to `sendsBackTo` with the output as the finding. It may change the files it `writes` and no other declared file. |
| `panel` | A role that is a list of seats. `agree` is how many must accept; the default is all of them. |
| `input` | A person role. The stage asks the question and waits. |
| `fn` | A plain function job: regular code, not another agent or command. The same `writes` rule as `run`; a fail with no target of its own goes to `sendsBackTo`. |
| `writes` | The files this stage writes, relative to the workspace. For an agent stage, a missing or empty one fails the stage by name, and only these files count when the stage checks for work left unchanged after feedback. |
| `desc`, `gate` | What the stage does and what must be true for it to count. Both go into the writing seat's prompt, the plan and the record. |
| `reviewedBy` | A panel or a person role that reads this stage's work. A rejection runs the stage again with the findings; a writer that leaves the work unchanged after feedback fails the stage. A run that dies mid-review runs the round in flight again. The earlier rounds stay used, and the next build reads the findings of the last rejected round. |
| `agree` | On a stage a panel reviews: how many reviewers must accept; the default is all of them. |
| `sendsBackTo` | The earlier stage a rejection or a red run goes to. A stage takes `reviewedBy` or `sendsBackTo`, not both. It declares the target when the graph is built, so a send-back to any other stage fails this stage with an error that names both. |
| `synthesise` | On a stage whose `reviewedBy` or `panel` role has more than one seat: a seat, or `true` for the first reviewer's seat, that merges the findings naming the same problem. Each reviewer then votes once on the merged findings it did not raise, and the votes drop, replace or dispute findings by rule. Off unless set. [Ask a panel](/docs/patterns/review-panel) has the steps. |
| `goal` | On a stage a panel reviews: a seat, ideally from another model family than the builder's, that checks the brief was met before the reviews. Each round, it reads the brief, `desc`, `gate` and the work, and marks each requirement met or unmet with evidence. Any unmet requirement sends the round back to the builder with its evidence: the reviewers do not run that round, and a judge does not decide it. Each round adds a `goal:check` event to the record. [Check the brief was met](/docs/patterns/check-the-brief) has the steps. |
| `refine` | How many [refinements](/docs/concepts/feedback-loops#counting-rounds) this stage gets: rounds of rework after the first build. A count, default 1, or `judge(seat)`. `refine: N` allows N+1 builds in all, so `refine: 0` is one build and no second. This holds on a stage a panel or a person reviews and on a stage a later stage's `sendsBackTo` names. When the refinements run out before the review passes, the run fails with the last review's findings. With `judge()`, a seat such as `jev()` from [`@obversa/engine-jev-api`](/docs/packages/engine-jev-api) decides between a review's verdict, a person's refusal included, or a send-back and the next round. The judge decides each finding, a block included, `act` or `skip`, with a one-line reason; a block's question sets a higher bar to skip it. The builder gets only the findings it acts on, each with the judge's reason; when the judge gives no reason of its own, as with `jev()`, the reason is what the chosen option means. The next round's reviewers are told which findings it skipped and why. When it acts on any finding, another round runs. When it skips every finding, the work stands as a pass, whatever its round answers say. Its `product_decision` asks a person. With your own `questions`, the whole round goes back or stands on the judge's answers to your questions, unless you add `perFinding: true`. With no `cap`, the rounds end when the judge stops them or the review passes. `judge(seat, { cap: N })` adds an optional backstop that takes the place of the count: at most N refinements. The judge reads the review of the last build too, and is told it is the last round: `holds` or `over_polishing` (or skipping every finding) passes, with the open findings on the outcome as `openFindings`; `product_decision` asks a person and records the answer, but no further round runs, so the stage fails; anything else fails with the cap named in the reason. A block in that review goes to the judge too. It belongs to the stage that repeats. |
| `retrySafe` | `true` when the stage may run again after an interruption without a person reconciling it. |
| `when` | A condition read before the stage runs: `passed(name)`, `failed(name)`, or a predicate over `ctx`. A stage whose `when` is not met is recorded as skipped and counts as passed. |
| `optional` | A red run of this stage doesn't fail the workflow. `failed(name)` on a later stage needs the named stage optional, and is met when that stage fails or runs out of rounds. |
| `needs` | Earlier stage names this stage also depends on, so a condition can read further back. |

### Workflow keys

| Key | What it means |
| - | - |
| `brief` | The work, as text or `briefFromFile`. Every model in the team reads it. |
| `roles` | Every seat and person, named once. |
| `stages` | The steps, in order. Each needs the one before it. |
| `options.timeout` | A duration such as `'10m'`, applied to each agent and review stage. A command stage keeps its own ten-minute limit. |
| `post.always` | Runs after the graph settles, with the record. |

## Errors

* **`LoopError`** carries a `code`: `ENGINE`, `TIMEOUT`, `ABORTED`,
  `VALIDATION`, `WRITE_BOUNDARY`, `CONFIG`, `BUDGET`, `RATE_LIMIT`, `QUOTA`,
  `BODY` or `UNKNOWN`, and a `phase`: `start`, `body`, `until`, `stopOn`,
  `review` or `engine`.
* **`WorkspacePolicyError`**: `INVALID_WORKSPACE_POLICY`,
  `NOT_GIT_REPOSITORY`, `GIT_INSPECTION_FAILED`, `ABORTED`.
* **`EngineIdentityUnresolvedError`**: a recorded auth failure names no
  provider and the lane can't supply one.
* **Re-exported from `@obversa/api`**: `EngineError`,
  `EngineIncompleteResultError`, `GraphValidationError`,
  `GraphExecutionError`, `StorageError`, `ApprovalSubjectError`,
  `JsonValueError`. Their codes are on [API](/docs/packages/api#errors).

## API

**Declare a team.** `workflow`, `stage`, `person`, `briefFromFile`,
`formatEvent`, `defineJob` (returns the job it's given, so a file's default
export keeps the exact `Job` type), `defineAgent`, `defineAgentFromMarkdown`,
`defineSkill`, `fromFile`.

**Build a graph.** `dag`, `pipeline`, `sequence`, `parallel`, `loop`,
`fnJob`, `agentJob`, `commandJob`, `commandSucceeds`, `gateJob`, `prove`,
`isolated`, `writeScope`, `team`, `jobMeta`, `copyJobMeta`, `renderPlan`,
`assertGraph`, `withEnv`.

**Conditions.** `passed`, `failed`, `all`, `any`, `not`, `always`, `never`,
`predicate`, `toCondition`, `describeConditions`, `bodyPassed`, `quorum`,
`sampled`, `ratchet`, `agentCheck`, `confidenceCondition`, `minConfidence`,
`confidenceFromText`, `lastDecisionLine`, `lastGateBrief`, `reviewContext`.

**Review.** `reviewPanel`, `goalCheck`, `tournament`, `kickback`, `revisionRequest`, `judge`, `isJudge`, `stopQuestions`,
`RECORDED_ENGINE_USAGE`, `RESUME_STAGE_OUTCOMES`.

**Ask a person.** `approval`, `createCallbackGate`, `createCallbackClient`,
`createStoredCallbackClient`, `replayCallbackClient`, `directRouter`,
`callbackRequestDigest`, `createApprovalCallbackGate`,
`approvalSubjectDigest`, `resolveApproval`.

**Run.** `run`, `EXIT_PAUSED`, `exitCodeFor`, `costReport`,
`formatCostReport`, `renderRecord`, `summarizeRecord`,
`preflight`, `preflightEngine`, `formatPreflight`,
`fallbackEngine`, `classifyEngineFailure`, `LANE_DEAD_FAILURES`,
`finalResultPart`, `finalResultText`, `validateAgentResult`.

**Evidence.** `writeProofArtifact`, `createAcceptedResultRecord`,
`resolveAcceptedResult`, `acceptedResultMatches`, `createProofCache`.

**Graph layer.** `compileGraph`, `resolveGraphPlan`, `validateGraphDescription`,
`persistRunDefinition`, `loadRunDefinition`, `createGraphExecutor`,
`readRunPreflight`, `interruptRunPreflight`, `dagGraphType`, `convergence`,
`teamGraphType`, `projectTeamRooms`, `createGitWorktreeProvider`, and the
`validate*` functions for stored records, re-exported from `@obversa/api`.

**Subpaths.** `@obversa/runtime/testing`: `MockEngine`, `MockEnvironment`,
`mockVerdict`, `recordedJudge` (replays a judge's recorded answers from a
JSON file, in place of `jev()`), `defineGraphDefinition`,
`createGraphEventTrace`, and the conformance kits. `@obversa/runtime/memory`: `ground`, `curate`,
`consolidate`. `@obversa/runtime/workflow-support`: `outcomeFromAgentText`,
`INVALID_TEAM_DECISION`, `seatIdentity`, `assertDistinctSeats`,
`requireNonEmptyFiles`, `requireNoFiles`, `teamAgent`, `panelReviewers`.
`@obversa/runtime/env/command`: `commandEnvironment`.
`@obversa/runtime/storage/local`: `createLocalRunStorage`.

## Next steps

* [Workflows](/docs/concepts/workflows): the idea behind `workflow()` and the
  stages.
* [Feature delivery](/docs/workflows/feature-team): a complete workflow with
  model reviews and human approval.
* [Running](/docs/concepts/running): what `run()` records and what a resumed run
  repeats.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.