Skip to main content
If the worker dies, a new one reads the record and a finished step does not run again. Start a graph in a worker and let the calling process be its watchdog: it checks each agent you named before the first step, waits for the worker, stops its child processes, and restarts it within the time and restart limits you set. Use it for a run that must survive a crash or an engine that isn’t ready yet. For a run you start and watch yourself, run() from the runtime is enough. The run is durable: every step is written to the run’s record, so a killed run carries on where it stopped. The smallest supervised run starts one graph under the watchdog:
examples/preflight-supervised-run.ts (excerpt)
startSupervisedRun returns a handle with done, status() and stop(). done resolves to the graph’s result, complete, pause or fail, and infrastructure errors reject it. A supervised run needs:
  • A runner directory and a worker root, with the host module as a relative specifier inside that root.
  • The storage settings and the workspace provider.
  • The stored definition: the run id, the graph, the resolved plan and the resolved inputs.
  • Limits, a restart policy and a cleanup grace.

Bind the host module

The runner loads host modules and supervises processes; the runtime never loads a module. The module exports bindRun, which receives the stored definition and a scratch directory and returns the compiled graph, the node bindings and the engine bindings. examples/preflight-host.mjs is one; the example copies it into the disposable repository it creates and the runner loads that copy. The module must stay inside the worker root after symlinks resolve; a path outside it is refused with HOST_MODULE. The worker root must be the top level of the repository the workspace provider captures; a provider for another repository is refused with WORKSPACE_ROOT before any run is stored. The runner stores the module’s path and byte digest and checks the digest before and after each import. Changed bytes fail with HOST_MODULE_CHANGED. The digest covers the module file alone, not what it imports. The worker inherits only PATH, HOME, TMPDIR, TMP, TEMP, SystemRoot, USERPROFILE and PATHEXT from the watchdog. To pass an engine credential, list its variable name in environmentVariables on the start or resume call. The runner copies the value when it’s set and stores neither the name nor the value in the run. resolvedInputs is durable storage, so keep secrets out of it.

Check the engines

When the stored plan carries a preflight policy, the worker checks every declared engine seat before it dispatches anything.
  • A static check asks the engine to admit the seat, with the node’s real configuration and no prompt. The engine reports the identity it will run under (adapter, provider, model family, model, executable), and the runtime compares it with the saved identity. An engine that doesn’t support the check is unsupported, and the lane’s policy says whether that blocks the seat or is allowed.
  • A live check, when the lane requires one, is one tool-free engine call with no workspace. It runs once per eligible seat, one at a time, and stops at the first seat that answers. It proves the seat can answer, not only that it’s configured.
  • A failed check that retires nothing pauses the run on a PREFLIGHT_PAUSED record before any dispatch, carrying the pause event id. Nothing has run, so there’s nothing to reconcile.
  • A lane with no admissible seat ends the run. When every target of a lane is excluded or blocked, the worker returns PREFLIGHT_FAILED and done resolves to fail with that code. Nothing was dispatched and there’s no pause to resume from.
  • A check the worker died in the middle of becomes a pause. Once the watchdog has verified the worker’s processes are gone, it closes the check as interrupted with interruptRunPreflight from @obversa/runtime, and done resolves to a PREFLIGHT_PAUSED pause you resume like any other. If the watchdog can’t verify cleanup, the check stays open.
Command-line adapters record the path and version they were admitted with, not a hash of the executable. An API-key adapter is admitted locally with no executable and no capabilities, and its retry layers are off for the check only. Environment variable names are validated and their values captured before any asynchronous work; no value is written to run storage.

Start, inspect, stop

Call stop() while work is running to request a stop and await cleanup. A clean stop without a recorded worker result returns fail with STOPPED; after completion it returns the settled result. When the worker has recorded a result and cleanup and lease release are verified, the watchdog returns that result even if a stop or timeout arrives before the worker exits. Cleanup and lease failures take precedence over a recorded result. readSupervisedRunStatus reads a run’s recorded progress from another process with the same storage settings. It doesn’t take ownership of the run or stop another process’s watchdog. Completion output is stored once as an artifact under the storage policy’s limits, and the watchdog verifies the stored bytes before returning done.output. Engine responses are stored the same way, as runner-engine-parts artifacts. A response that fits the node’s output limit can still exceed a storage limit: a size or quota refusal raises StorageError with STORAGE_LIMIT_EXCEEDED, and a match on a known secret raises KNOWN_SECRET. Inside the worker the refused write becomes a failed node with EFFECT_FAILED, and with the built-in dag form a failed required node makes done resolve to fail with DAG_NODE_FAILED.

Recover after a crash

The watchdog waits for worker exit and cleanup, then starts a replacement that reads the record and decides what remains. A saved start without a saved result doesn’t prove whether an outward effect happened. Set retrySafe: true only when repeating the node is acceptable after an unknown outcome. Otherwise the run pauses with a reconcile-attempt request for a person to answer. The watchdog holds the workspace lease while the worker runs and while cleanup executes. Cleanup uses the public @obversa/core/command API and its owner markers; on Linux it also finds processes carrying the worker’s inherited marker, and elsewhere it follows the observed process tree. Incomplete cleanup or a failed lease release retains the process lock and starts no further worker. The lock is scoped to the storage directory, namespace and run id: a second watchdog for the same run gets PROCESS_LOCKED, and if a watchdog dies its lock and lease remain, with no forced takeover.

Resume a paused run

resumeSupervisedRun reopens one recorded pause. For a preflight pause, pass the run id and the preflightEventId; for a node pause, the run id and the exact position from its graph:node-paused event:
examples/preflight-supervised-run.ts (excerpt)
The definition, host module and limits come from the stored run. Resume writes no second start event and repeats no finished occurrence. For a preflight pause the static checks run again, the live receipts that still apply are reused, and a stale id is refused with RESUME_EVENT_MISMATCH naming both ids. Before releasing ownership after a node pause, the watchdog saves a snapshot of the workspace. Resume checks the workspace against it before starting a worker and never adopts edits made while paused. A changed workspace pauses again with WORKSPACE_DRIFT, a missing snapshot with WORKSPACE_ANCHOR_MISSING, an unreadable one with WORKSPACE_ANCHOR_INVALID. If the snapshot can’t be captured at pause time, the run fails with WORKSPACE_ANCHOR_WRITE and no resumable pause is recorded. The host’s action policy runs again on resume; record approval where that policy reads it, because calling resume isn’t an approval. Run the example with npx tsx preflight-supervised-run.ts. It starts a one-node graph under the watchdog with a scripted engine that isn’t ready, pauses before any work is dispatched, makes the engine ready, stops the old watchdog, and resumes from the recorded pause:
Output
Five engine calls, in order: a static check of the seat, a live check that found it not ready, and after resume a second static check, a live check that found it ready, and the one ordinary call the graph asked for. No node was dispatched before the resume, and the resume used the exact pause event the first run returned.

Read progress

Status carries the run phase, worker liveness, inspected processes, the restart count, backoff, elapsed and remaining time, and pause reasons. cleanupVerified is null before a terminal result, true when cleanup was verified within the platform’s capability, and false when it couldn’t be. leaseRetained reports a lease still held after a terminal failure. A dead worker or an empty process list isn’t proof of cleanup: read these two fields before treating the workspace as released. Usage is reported per engine call. Missing usage isn’t zero: a call without a receipt counts as unknown, and a node’s usage is partial when some of its calls have receipts and some don’t. The check calls before the run report their own usage, separate from the nodes’.

Failure

A check that fails tells the run what not to try again, and the scope follows what the failure proves:
  • Bad credentials retire every model on that adapter and provider, because the credential is theirs.
  • A missing model, exhausted credit or an exhausted quota retire that provider and model, and also the provider and model the engine reported, if they differ. A quota is an allowance gone for hours or longer.
  • A missing command-line tool or an invalid configuration retire the adapter.
  • A rate limit or a transport error retire nothing, because they clear in seconds or minutes.
  • An old auth record with no provider recovers one from its lane. When that’s ambiguous, the run refuses before any work with ENGINE_IDENTITY_UNRESOLVED.
If the pause event itself can’t be written, the run returns fail with RUN_STORAGE and no resumable pause exists. If the pause was written and only the watchdog’s final record failed, the result is a pause with RUN_STORAGE, and a later resume can reopen it. Neither path retries the write on its own.

Limits

  • Elapsed time. limits.timeoutMs covers worker execution and restart backoff across pauses and resumes. It counts from the stored run timestamp, freezes at each settled pause, and resumes at the next worker launch. Resume and preflight checks before launch don’t spend it; engine checks inside the worker do.
  • Dispatches. limits.maxDispatches counts recorded graph dispatches across workers.
  • Restarts. restart.maxRestarts caps replacement workers, with backoff from initialBackoffMs to maxBackoffMs.
  • Cleanup grace. teardownGraceMs is the time processes get to stop before they are killed.
  • Worker output. Each worker launch may write 1,000,000 bytes to its combined stdout and stderr. Beyond that, the watchdog stops the worker without restarting it and done resolves to fail with OUTPUT_LIMIT.
examples/preflight-supervised-run.ts
The engine and its bindings live in examples/preflight-host.mjs, the host module the example copies into its disposable repository.

Next steps