examples/preflight-executor.ts (excerpt)
run() returns a GraphExecutorResult: complete, pause, fail or
waiting. An executor needs:
- The run id and the storage the run definition and plan were persisted to.
- The compiled graph, node bindings and engine bindings, here from a
host module’s
bindRun. Node behaviour is keyed by node id: a data-only node gets a function, an engine-backed node a prompt builder that receives that dispatch’s typed input. The node uses the lane in the stored plan, and the engine binding joins that target to the full engine identity. preflightScratchDirectorywhen the plan carries a preflight policy. That path and each node binding’sscratchDirectorymust be absolute, real and normalised.
Run a dispatch
- The executor loads the run definition and the frozen plan from storage and checks that the supplied graph matches them.
- When the plan carries a preflight policy, it checks every engine seat: a
static check of its identity and, where the lane requires it, one
tool-free live call with no workspace. The live call proves the seat
answers and does no work. A failed check pauses the run with
PREFLIGHT_PAUSEDor retires the engine, as below. Nothing is dispatched before the checks finish. - It folds the run’s recorded events into graph state and asks the graph for its next command.
- It records the whole dispatch decision, then
node-attempt-startedimmediately before each node starts. - It runs each node through the safe attempt lifecycle, records
node-completed,node-pausedornode-failed, and asks the graph again.
unsupportedStatic policy to say whether an engine without
admit blocks the run or is allowed; the static check of such an engine
comes back unsupported, never a success receipt. Usage from the checks is
kept separate from usage from nodes, and missing usage counts as unknown,
not zero.
Fall back and retire
The host chooses the lane order once. The executor tries the effective target and at most one live fallback. Only a failure that means the lane is dead moves to the fallback; a rate limit, a timeout, a policy pause or a denied action doesn’t run again elsewhere. Each dead-lane event records the selected identity and, when available, the identity the engine reported. What the failure retires follows what it proves: bad credentials retire every model on that adapter and provider; a missing model, exhausted credit or an exhausted quota retire that provider and model, and the reported ones if they differ; a missing command-line tool or an invalid configuration retire the adapter; a rate limit or a transport error retire nothing. When a recorded auth failure names no provider, the executor recovers one from the lane, and if that’s ambiguous it refuses before any work withENGINE_IDENTITY_UNRESOLVED. Later
dispatches skip whatever was retired.
During a run, if every declared target of a lane is dead, the executor
records ENGINE_UNAVAILABLE as a normal node failure, so the graph can
retry, fail or recover in its usual way. Before the first step, when every
target of a lane is excluded or blocked by its checks, the run ends with
PREFLIGHT_FAILED instead: terminal, with preflight:failed on the record,
run() returning { kind: 'fail', code: 'PREFLIGHT_FAILED' }, nothing
dispatched and no pause to resume from.
Read engine identity records
Every engine call records agraph:engine-attempt-recorded event before
fallback routing or the node’s result, with the node id, position,
sequence, requested identity and reported effective identity. The sequence
starts at 1 for each position and continues across fallback calls and
resume. When resume recovers an engine node that started and never
finished, it adds a record with both identities null: a mark of
uncertainty, not a count of calls. If a receipt can’t be saved, run()
rejects and the affected attempt stops before fallback or completion, while
other attempts in the batch can continue. The records describe what the
bound engine reported; they don’t see inside an adapter.
Resume after a stop
A fresh executor returnswaiting, with the open positions, when a dispatch
has no result. Call resume with one exact position, and only after the
earlier process has stopped, because the executor doesn’t lock the run
between processes. If two dispatches were open, the first resume returns
waiting again with the other position. What resume does depends on how
far the attempt got:
For a preflight pause, call
resume with the { preflightEventId } from
the pause the first run returned:
examples/preflight-executor.ts (excerpt)
RESUME_EVENT_MISMATCH, which names both ids.
The saved retrySafe value and the saved plan govern recovery; a change in
the live binding doesn’t rewrite an earlier record. Run the file with
npx tsx preflight-executor.ts. It stores a one-node graph, runs the
executor against a scripted engine that isn’t ready, pauses before any
dispatch, makes the engine ready and resumes with the exact pause event:
Output
readRunPreflight reads a stored run’s
preflight record without touching an engine.
Failure
INVALID_PREFLIGHT_CONFIG:preflightScratchDirectoryis optional in the type, so TypeScript won’t catch a missing path.await createGraphExecutor(...)rejects with aGraphExecutionErrorof this code when the plan needs the path and it’s missing or malformed, and a relative or symlinked binding path rejects with the same code insiderun()andresume()when its lane is checked. The error doesn’t say which path failed, so handle the code at both call sites.- A denied or aborted attempt records
node-failedwithDENIEDorABORTED. An action-policy wait recordsnode-pausedwith its reason and request. A result too large for the event stream records a smallnode-failedwithRESULT_TOO_LARGE, so the dispatch never stays in flight. - An interrupted check. A live check is an engine call, and the process
can die in the middle of it, leaving an open probe on the record. A fresh
executor’s
run()then fails withPROTOCOL, names the probe and tells you what to do: check that the old worker and its engine process are gone, then callinterruptRunPreflight(storage, runId). That call closes the probe asinterruptedand writes a preflight pause in one write that fails if anything else wrote first, returns the pause for the next resume, returnsnullwhen no probe is open, and returns the existing pause when one is already recorded. It never retries and never calls an engine. Call it only when you own the run and have checked the worker is gone: if the old call is still alive, the record says interrupted while the call keeps spending.
Limits
- The executor doesn’t lock the run. Two processes over the same store are yours to keep apart; the supervised runner holds a process lock for you.
- One live fallback per attempt. The lane order is the host’s, chosen once.
Full file
Full file
examples/preflight-executor.ts
examples/preflight-host.mjs. The types
are GraphExecutorResult for what run() and resume() return,
RunPreflightPolicy for the plan’s checks, RunPreflightState for what
readRunPreflight returns, and PreflightPauseResult and
PreflightFailureResult for the two ways the checks stop a run early.
Next steps
- Supervised local runs: the watchdog that runs this executor in a worker and restarts it.
- Safe node attempts: what one dispatch records.
- Plan admission: the frozen plan the executor checks the graph against.