Skip to main content
@obversa/runtime is the package you import to write and run a workflow: the shapes (workflow, pipeline, dag, loop), the steps, the reviews, the person’s question, run() and the record it writes. It imports no engine or memory adapter; your program supplies the instances.

Install

Included in @obversa/obversa. The package also installs the obversa-record command, which prints a run’s record as a page a person scans: Read a record.

Quickstart

One agent, one job, one run. The engine is a stand-in from the testing module, so the file runs offline:
examples/one-agent-job.ts
agentJob makes a step from a prompt and an engine name. run runs it with the engines you pass and returns the outcome. Run it with npx tsx one-agent-job.ts:
Output

Declare a team

workflow(name, { brief, roles, stages }) is the team, in a file: the brief, the roles named once, the stages in order. Each stage says who does it, what it writes, who reads it and where a red result goes:
examples/teams/writer-reviewer-pair.ts (excerpt)
stage(name, config) is one step; exactly one of agent, run, panel or input says what kind. Seats come from the engine plugins: claude('claude-sonnet-4-5') from @obversa/engine-claude-cli, codex('gpt-5.6-luna') from @obversa/engine-codex-cli, opencode(model, { executable }) from @obversa/engine-opencode-cli. person(question) is a role a person fills. briefFromFile(path) reads a Markdown file as the brief; its front matter may carry files, the paths the brief expects written. formatEvent(event, totals?) turns one run event into the line a person reads, and RunResult.usage is what the run spent in the shape it takes. Every seat of a reviewedBy panel must report a different model family from the stage’s writer, or the team is refused before any model runs. The whole file and a real run are on A writer and a reviewer.

Build a graph

Under workflow() sit the shapes: pipeline for stages in order, dag for a dependency graph, sequence and parallel for the two plain cases, loop for a body that runs until a condition holds. Each returns a Job you can run or nest. The steps are fnJob for a function, agentJob for a prompt to an engine, commandJob for a command whose exit code decides:
examples/feature-team.ts (excerpt)
A step that sends work back lists its targets in acceptsKickbackTo, and each must be a step it depends on. A send-back to a step it does not list fails the sender with an error that names both steps. maxKickbacks sets the limit: a number is the send-backs allowed in the whole graph, and a map gives each target its own number of refinements. A target’s number can be a judge() instead. The judge reads the dag’s brief (the brief or ticket text the work answers) and useCase (what the work is for), the target node’s desc and gate, and the file the target node names in file. brief and useCase are optional strings on the dag’s config. The target’s next run reads the send-back as ctx.lastReview, once. A writing agentJob that runs again after a send-back and leaves its node’s file as it was fails: it returned the work unchanged, whatever else it wrote. With no file, a change to any file in the workspace counts. When a graph is a loop()’s body, each step’s first run in a round reads the loop’s last review the same way. A step that fails with an error a retry cannot fix fails the graph with that error, so a loop around the graph stops, as it does around the step alone. A fnJob function returns a full outcome, a one-line summary (a pass) or nothing (a pass with the label as its summary); a throw is a fail carrying the error. commandJob passes on exit 0 and otherwise fails with the output as evidence, and target names the step that owns the fix. passed(name) and failed(name) are conditions for a node’s when that read a named dependency’s outcome, so an exit code chooses the branch; the node failed reads must be optional: true. A node that ran out of rounds meets failed. isolated(job) runs a step in its own worktree and branch and lands the change on pass. jobMeta, copyJobMeta and renderPlan read a job’s declared shape and render it.

Review inside it

reviewPanel runs several reviewers on one attempt and passes on a threshold; a failing panel returns the findings to target, which runs again:
examples/feature-team.ts (excerpt)
tournament runs n candidates, each in its own worktree, and a judge scores the finished ones; the highest lands:
examples/tournament.ts (excerpt)
kickback(to, reason) and revisionRequest(input) build the outcome a reviewer returns to fail a step with findings aimed at a named stage. quorum(k, ...conditions) is met when k of its inputs hold. agentCheck asks an engine a scored question as a condition; confidenceCondition and minConfidence read a confidence and gate on it. RECORDED_ENGINE_USAGE is the state key under which a run records which model answered each engine call, so the review gate compares writers with reviewers by family.

Ask a person

humanReview(name, { question, input, interaction }) makes a review step that keeps what you changed, what you wrote and whether you approved. Use it as a loop’s review. Its response contains feedback, a composed prompt, and an explicit decision: changes-requested sends work back; only approved passes. The writer’s consumeFeedback: true includes the full response on its next turn. Surfaces shows an interactive example. interaction names the handler with id, declares responseSchema, and optionally supplies an async answer(request, signal) function. That function can open a surface and return its payload. Return undefined if the person cancels. Without a handler, answer through the run’s stored callbacks client. An agentJob can also declare interaction. To ask a question, that agent returns JSON with an interaction object containing question and input. The answer’s feedback and prompt return to the same model in a fresh turn with the saved task context. This asks for input; it does not require an approval decision. person(question, { interaction }) uses the same rich response in a workflow input stage. approval(label, { question, target?, answer? }) is a person’s decision as a step. Yes passes; no goes to target with the note as the finding, or fails the step when there’s no target; no answer pauses the run:
examples/approval.ts (excerpt)
Under it sit the callback gates: createCallbackGate makes a question with an identity, createCallbackClient holds it for one run, createStoredCallbackClient keeps it across a process exit, replayCallbackClient rebuilds pending questions from a history, and directRouter is the smallest router. createApprovalCallbackGate, approvalSubjectDigest and resolveApproval bind an approval to the exact bytes it approves. Callback gates is the guide.

Run it

run(job, options) runs a job from a fresh record and returns a RunResult: the outcome, the usage, and monitor when the page is on. Turn the page on and read its address from the record:
examples/monitor.ts (excerpt)
Every run emits one root run:start before it dispatches work and one run:end before it returns, both as LoopEvents to onEvent. Cancelling a waiting workflow reports the run as aborted and leaves the question in the record, so a later resume can still answer it. An attempt that runs past its time budget comes back as a typed outcome, not a hang.

Keep the evidence

writeProofArtifact stores one proof packet addressed by its digest, and createAcceptedResultRecord, resolveAcceptedResult and acceptedResultMatches bind a result to the inputs, proof, graph and workspace it was judged on:
examples/proof-bound-approval.ts (excerpt)
createProofCache shares declared read-only evidence between workers. Proof-bound acceptance and approval is the guide.

Store and execute

For hosts and for shapes of your own: compileGraph validates a definition against a graph type, resolveGraphPlan freezes what the run may use, persistRunDefinition stores it, and createGraphExecutor runs it one recorded decision at a time:
examples/custom-graph.ts (excerpt)
dagGraphType, convergence and teamGraphType are the forms that ship. loadRunDefinition reads a stored run back; readRunPreflight and interruptRunPreflight read and close a run’s engine checks; validateDomainEventId checks an event id before a resume. createLocalRunStorage from @obversa/runtime/storage/local is the on-disk event and artifact store the examples use, and createGitWorktreeProvider the workspace provider. The guides are Outside graph types and Graph executor.

Options

recordTo and resume

recordTo appends every event as one JSON line to the file you name, or under .obversa/records with 'auto'. resume: true reads that record first: a workflow() or a dag() carries on from it and repeats no finished step; a loop() or a plain job appends to it and runs again. A graph that runs more than once, such as a loop’s body, skips finished steps only the first time. To resume, keep the graph’s name, workspace and shape the same (for a workflow, its brief, stage list and role seats too); change any and the run starts again. The record says what a resumed run repeats.

callbacks and onCallback

callbacks is the client the run’s questions go through, in memory (the default, fresh per run) or stored. onCallback: 'wait' keeps the process up until an answer arrives; the default 'exit' records the question and returns, so a schedule can start the same file again until the answer is there. A record written in one mode resumes in the other.

engine and engines

engines is the map of ready-made engine instances by name, and engine the one a job or condition uses when it names none. The runtime imports no adapter, so both are yours to supply. withEnv(job, env) pins environment variables for a job and every job beneath it.

The rest

Stage keys

Workflow keys

Errors

  • LoopError carries a code: ENGINE, TIMEOUT, ABORTED, VALIDATION, WRITE_BOUNDARY, CONFIG, BUDGET, RATE_LIMIT, QUOTA, BODY or UNKNOWN, and a phase: start, body, until, stopOn, review or engine.
  • WorkspacePolicyError: INVALID_WORKSPACE_POLICY, NOT_GIT_REPOSITORY, GIT_INSPECTION_FAILED, ABORTED.
  • EngineIdentityUnresolvedError: a recorded auth failure names no provider and the lane can’t supply one.
  • Re-exported from @obversa/api: EngineError, EngineIncompleteResultError, GraphValidationError, GraphExecutionError, StorageError, ApprovalSubjectError, JsonValueError. Their codes are on API.

API

Declare a team. workflow, stage, person, briefFromFile, formatEvent, defineJob (returns the job it’s given, so a file’s default export keeps the exact Job type), defineAgent, defineAgentFromMarkdown, defineSkill, fromFile. Build a graph. dag, pipeline, sequence, parallel, loop, fnJob, agentJob, commandJob, commandSucceeds, gateJob, prove, isolated, writeScope, team, jobMeta, copyJobMeta, renderPlan, assertGraph, withEnv. Conditions. passed, failed, all, any, not, always, never, predicate, toCondition, describeConditions, bodyPassed, quorum, sampled, ratchet, agentCheck, confidenceCondition, minConfidence, confidenceFromText, lastDecisionLine, lastGateBrief, reviewContext. Review. reviewPanel, goalCheck, tournament, kickback, revisionRequest, judge, isJudge, stopQuestions, RECORDED_ENGINE_USAGE, RESUME_STAGE_OUTCOMES. Ask a person. approval, createCallbackGate, createCallbackClient, createStoredCallbackClient, replayCallbackClient, directRouter, callbackRequestDigest, createApprovalCallbackGate, approvalSubjectDigest, resolveApproval. Run. run, EXIT_PAUSED, exitCodeFor, costReport, formatCostReport, renderRecord, summarizeRecord, preflight, preflightEngine, formatPreflight, fallbackEngine, classifyEngineFailure, LANE_DEAD_FAILURES, finalResultPart, finalResultText, validateAgentResult. Evidence. writeProofArtifact, createAcceptedResultRecord, resolveAcceptedResult, acceptedResultMatches, createProofCache. Graph layer. compileGraph, resolveGraphPlan, validateGraphDescription, persistRunDefinition, loadRunDefinition, createGraphExecutor, readRunPreflight, interruptRunPreflight, dagGraphType, convergence, teamGraphType, projectTeamRooms, createGitWorktreeProvider, and the validate* functions for stored records, re-exported from @obversa/api. Subpaths. @obversa/runtime/testing: MockEngine, MockEnvironment, mockVerdict, recordedJudge (replays a judge’s recorded answers from a JSON file, in place of jev()), defineGraphDefinition, createGraphEventTrace, and the conformance kits. @obversa/runtime/memory: ground, curate, consolidate. @obversa/runtime/workflow-support: outcomeFromAgentText, INVALID_TEAM_DECISION, seatIdentity, assertDistinctSeats, requireNonEmptyFiles, requireNoFiles, teamAgent, panelReviewers. @obversa/runtime/env/command: commandEnvironment. @obversa/runtime/storage/local: createLocalRunStorage.

Next steps

  • Workflows: the idea behind workflow() and the stages.
  • Feature delivery: a complete workflow with model reviews and human approval.
  • Running: what run() records and what a resumed run repeats.