run() from the
runtime is enough.
The run is durable: every step is written to the run’s record, so a killed
run carries on where it stopped. The smallest supervised run starts one
graph under the watchdog:
examples/preflight-supervised-run.ts (excerpt)
startSupervisedRun returns a handle with done, status() and stop().
done resolves to the graph’s result, complete, pause or fail, and
infrastructure errors reject it. A supervised run needs:
- A runner directory and a worker root, with the host module as a relative specifier inside that root.
- The storage settings and the workspace provider.
- The stored definition: the run id, the graph, the resolved plan and the resolved inputs.
- Limits, a restart policy and a cleanup grace.
Bind the host module
The runner loads host modules and supervises processes; the runtime never loads a module. The module exportsbindRun, which receives the stored
definition and a scratch directory and returns the compiled graph, the node
bindings and the engine bindings. examples/preflight-host.mjs is one; the
example copies it into the disposable repository it creates and the runner
loads that copy.
The module must stay inside the worker root after symlinks resolve; a path
outside it is refused with HOST_MODULE. The worker root must be the top
level of the repository the workspace provider captures; a provider for
another repository is refused with WORKSPACE_ROOT before any run is
stored. The runner stores the module’s
path and byte digest and checks the digest before and after each import.
Changed bytes fail with HOST_MODULE_CHANGED. The digest covers the module
file alone, not what it imports.
The worker inherits only PATH, HOME, TMPDIR, TMP, TEMP,
SystemRoot, USERPROFILE and PATHEXT from the watchdog. To pass an
engine credential, list its variable name in environmentVariables on the
start or resume call. The runner copies the value when it’s set and stores
neither the name nor the value in the run. resolvedInputs is durable
storage, so keep secrets out of it.
Check the engines
When the stored plan carries a preflight policy, the worker checks every declared engine seat before it dispatches anything.- A static check asks the engine to admit the seat, with the node’s
real configuration and no prompt. The engine reports the identity it will
run under (adapter, provider, model family, model, executable), and the
runtime compares it with the saved identity. An engine that doesn’t
support the check is
unsupported, and the lane’s policy says whether that blocks the seat or is allowed. - A live check, when the lane requires one, is one tool-free engine call with no workspace. It runs once per eligible seat, one at a time, and stops at the first seat that answers. It proves the seat can answer, not only that it’s configured.
- A failed check that retires nothing pauses the run on a
PREFLIGHT_PAUSEDrecord before any dispatch, carrying the pause event id. Nothing has run, so there’s nothing to reconcile. - A lane with no admissible seat ends the run. When every target of a
lane is excluded or blocked, the worker returns
PREFLIGHT_FAILEDanddoneresolves tofailwith that code. Nothing was dispatched and there’s no pause to resume from. - A check the worker died in the middle of becomes a pause. Once the
watchdog has verified the worker’s processes are gone, it closes the
check as interrupted with
interruptRunPreflightfrom@obversa/runtime, anddoneresolves to aPREFLIGHT_PAUSEDpause you resume like any other. If the watchdog can’t verify cleanup, the check stays open.
Start, inspect, stop
Callstop() while work is running to request a stop and await cleanup. A
clean stop without a recorded worker result returns fail with STOPPED;
after completion it returns the settled result. When the worker has
recorded a result and cleanup and lease release are verified, the watchdog
returns that result even if a stop or timeout arrives before the worker
exits. Cleanup and lease failures take precedence over a recorded result.
readSupervisedRunStatus reads a run’s recorded progress from another
process with the same storage settings. It doesn’t take ownership of the
run or stop another process’s watchdog.
Completion output is stored once as an artifact under the storage policy’s
limits, and the watchdog verifies the stored bytes before returning
done.output. Engine responses are stored the same way, as
runner-engine-parts artifacts. A response that fits the node’s output
limit can still exceed a storage limit: a size or quota refusal raises
StorageError with STORAGE_LIMIT_EXCEEDED, and a match on a known secret
raises KNOWN_SECRET. Inside the worker the refused write becomes a failed
node with EFFECT_FAILED, and with the built-in dag form a failed required
node makes done resolve to fail with DAG_NODE_FAILED.
Recover after a crash
The watchdog waits for worker exit and cleanup, then starts a replacement that reads the record and decides what remains.
A saved start without a saved result doesn’t prove whether an outward effect
happened. Set
retrySafe: true only when repeating the node is acceptable
after an unknown outcome. Otherwise the run pauses with a
reconcile-attempt request for a person to answer.
The watchdog holds the workspace lease while the worker runs and while
cleanup executes. Cleanup uses the public @obversa/core/command API and
its owner markers; on Linux it also finds processes carrying the worker’s
inherited marker, and elsewhere it follows the observed process tree.
Incomplete cleanup or a failed lease release retains the process lock and
starts no further worker. The lock is scoped to the storage directory,
namespace and run id: a second watchdog for the same run gets
PROCESS_LOCKED, and if a watchdog dies its lock and lease remain, with no
forced takeover.
Resume a paused run
resumeSupervisedRun reopens one recorded pause. For a preflight pause,
pass the run id and the preflightEventId; for a node pause, the run id
and the exact position from its graph:node-paused event:
examples/preflight-supervised-run.ts (excerpt)
RESUME_EVENT_MISMATCH
naming both ids.
Before releasing ownership after a node pause, the watchdog saves a snapshot
of the workspace. Resume checks the workspace against it before starting a
worker and never adopts edits made while paused. A changed workspace pauses
again with WORKSPACE_DRIFT, a missing snapshot with
WORKSPACE_ANCHOR_MISSING, an unreadable one with
WORKSPACE_ANCHOR_INVALID. If the snapshot can’t be captured at pause time,
the run fails with WORKSPACE_ANCHOR_WRITE and no resumable pause is
recorded. The host’s action policy runs again on resume; record approval
where that policy reads it, because calling resume isn’t an approval.
Run the example with npx tsx preflight-supervised-run.ts. It starts a
one-node graph under the watchdog with a scripted engine that isn’t ready,
pauses before any work is dispatched, makes the engine ready, stops the old
watchdog, and resumes from the recorded pause:
Output
Read progress
Status carries the run phase, worker liveness, inspected processes, the restart count, backoff, elapsed and remaining time, and pause reasons.cleanupVerified is null before a terminal result, true when cleanup
was verified within the platform’s capability, and false when it
couldn’t be. leaseRetained reports a lease still held after a terminal
failure. A dead worker or an empty process list isn’t proof of cleanup:
read these two fields before treating the workspace as released.
Usage is reported per engine call. Missing usage isn’t zero: a call without
a receipt counts as unknown, and a node’s usage is partial when some of
its calls have receipts and some don’t. The check calls before the run
report their own usage, separate from the nodes’.
Failure
A check that fails tells the run what not to try again, and the scope follows what the failure proves:- Bad credentials retire every model on that adapter and provider, because the credential is theirs.
- A missing model, exhausted credit or an exhausted quota retire that provider and model, and also the provider and model the engine reported, if they differ. A quota is an allowance gone for hours or longer.
- A missing command-line tool or an invalid configuration retire the adapter.
- A rate limit or a transport error retire nothing, because they clear in seconds or minutes.
- An old auth record with no provider recovers one from its lane. When
that’s ambiguous, the run refuses before any work with
ENGINE_IDENTITY_UNRESOLVED.
fail with
RUN_STORAGE and no resumable pause exists. If the pause was written and
only the watchdog’s final record failed, the result is a pause with
RUN_STORAGE, and a later resume can reopen it. Neither path retries the
write on its own.
Limits
- Elapsed time.
limits.timeoutMscovers worker execution and restart backoff across pauses and resumes. It counts from the stored run timestamp, freezes at each settled pause, and resumes at the next worker launch. Resume and preflight checks before launch don’t spend it; engine checks inside the worker do. - Dispatches.
limits.maxDispatchescounts recorded graph dispatches across workers. - Restarts.
restart.maxRestartscaps replacement workers, with backoff frominitialBackoffMstomaxBackoffMs. - Cleanup grace.
teardownGraceMsis the time processes get to stop before they are killed. - Worker output. Each worker launch may write 1,000,000 bytes to its
combined stdout and stderr. Beyond that, the watchdog stops the worker
without restarting it and
doneresolves tofailwithOUTPUT_LIMIT.
Full file
Full file
examples/preflight-supervised-run.ts
examples/preflight-host.mjs, the host
module the example copies into its disposable repository.
Next steps
- Runner: the package’s public entry points.
- The record: what a resumed run skips, repeats and asks a person to reconcile.
- Watch a run in the browser: the page a supervised run serves while it works.