Skip to main content
Turn a stored run into work, one recorded decision at a time: your graph decides, the executor records each decision before any node runs, runs the node inside the safe attempt lifecycle, and records what came back. Use it directly when you write a host. For a watchdog that starts the executor in a worker, restarts it after a crash and checks engines for you, read Supervised local runs; this page keeps only what is the executor’s own. The smallest executor binds a stored run to its nodes and engines and runs it:
examples/preflight-executor.ts (excerpt)
run() returns a GraphExecutorResult: complete, pause, fail or waiting. An executor needs:
  • The run id and the storage the run definition and plan were persisted to.
  • The compiled graph, node bindings and engine bindings, here from a host module’s bindRun. Node behaviour is keyed by node id: a data-only node gets a function, an engine-backed node a prompt builder that receives that dispatch’s typed input. The node uses the lane in the stored plan, and the engine binding joins that target to the full engine identity.
  • preflightScratchDirectory when the plan carries a preflight policy. That path and each node binding’s scratchDirectory must be absolute, real and normalised.

Run a dispatch

  1. The executor loads the run definition and the frozen plan from storage and checks that the supplied graph matches them.
  2. When the plan carries a preflight policy, it checks every engine seat: a static check of its identity and, where the lane requires it, one tool-free live call with no workspace. The live call proves the seat answers and does no work. A failed check pauses the run with PREFLIGHT_PAUSED or retires the engine, as below. Nothing is dispatched before the checks finish.
  3. It folds the run’s recorded events into graph state and asks the graph for its next command.
  4. It records the whole dispatch decision, then node-attempt-started immediately before each node starts.
  5. It runs each node through the safe attempt lifecycle, records node-completed, node-paused or node-failed, and asks the graph again.
Set the lane’s unsupportedStatic policy to say whether an engine without admit blocks the run or is allowed; the static check of such an engine comes back unsupported, never a success receipt. Usage from the checks is kept separate from usage from nodes, and missing usage counts as unknown, not zero.

Fall back and retire

The host chooses the lane order once. The executor tries the effective target and at most one live fallback. Only a failure that means the lane is dead moves to the fallback; a rate limit, a timeout, a policy pause or a denied action doesn’t run again elsewhere. Each dead-lane event records the selected identity and, when available, the identity the engine reported. What the failure retires follows what it proves: bad credentials retire every model on that adapter and provider; a missing model, exhausted credit or an exhausted quota retire that provider and model, and the reported ones if they differ; a missing command-line tool or an invalid configuration retire the adapter; a rate limit or a transport error retire nothing. When a recorded auth failure names no provider, the executor recovers one from the lane, and if that’s ambiguous it refuses before any work with ENGINE_IDENTITY_UNRESOLVED. Later dispatches skip whatever was retired. During a run, if every declared target of a lane is dead, the executor records ENGINE_UNAVAILABLE as a normal node failure, so the graph can retry, fail or recover in its usual way. Before the first step, when every target of a lane is excluded or blocked by its checks, the run ends with PREFLIGHT_FAILED instead: terminal, with preflight:failed on the record, run() returning { kind: 'fail', code: 'PREFLIGHT_FAILED' }, nothing dispatched and no pause to resume from.

Read engine identity records

Every engine call records a graph:engine-attempt-recorded event before fallback routing or the node’s result, with the node id, position, sequence, requested identity and reported effective identity. The sequence starts at 1 for each position and continues across fallback calls and resume. When resume recovers an engine node that started and never finished, it adds a record with both identities null: a mark of uncertainty, not a count of calls. If a receipt can’t be saved, run() rejects and the affected attempt stops before fallback or completion, while other attempts in the batch can continue. The records describe what the bound engine reported; they don’t see inside an adapter.

Resume after a stop

A fresh executor returns waiting, with the open positions, when a dispatch has no result. Call resume with one exact position, and only after the earlier process has stopped, because the executor doesn’t lock the run between processes. If two dispatches were open, the first resume returns waiting again with the other position. What resume does depends on how far the attempt got: For a preflight pause, call resume with the { preflightEventId } from the pause the first run returned:
examples/preflight-executor.ts (excerpt)
The static checks run again and live receipts that still apply are reused. A stale id is refused with RESUME_EVENT_MISMATCH, which names both ids. The saved retrySafe value and the saved plan govern recovery; a change in the live binding doesn’t rewrite an earlier record. Run the file with npx tsx preflight-executor.ts. It stores a one-node graph, runs the executor against a scripted engine that isn’t ready, pauses before any dispatch, makes the engine ready and resumes with the exact pause event:
Output
No node was dispatched before the resume, and the resume used the pause event the first run returned. readRunPreflight reads a stored run’s preflight record without touching an engine.

Failure

  • INVALID_PREFLIGHT_CONFIG: preflightScratchDirectory is optional in the type, so TypeScript won’t catch a missing path. await createGraphExecutor(...) rejects with a GraphExecutionError of this code when the plan needs the path and it’s missing or malformed, and a relative or symlinked binding path rejects with the same code inside run() and resume() when its lane is checked. The error doesn’t say which path failed, so handle the code at both call sites.
  • A denied or aborted attempt records node-failed with DENIED or ABORTED. An action-policy wait records node-paused with its reason and request. A result too large for the event stream records a small node-failed with RESULT_TOO_LARGE, so the dispatch never stays in flight.
  • An interrupted check. A live check is an engine call, and the process can die in the middle of it, leaving an open probe on the record. A fresh executor’s run() then fails with PROTOCOL, names the probe and tells you what to do: check that the old worker and its engine process are gone, then call interruptRunPreflight(storage, runId). That call closes the probe as interrupted and writes a preflight pause in one write that fails if anything else wrote first, returns the pause for the next resume, returns null when no probe is open, and returns the existing pause when one is already recorded. It never retries and never calls an engine. Call it only when you own the run and have checked the worker is gone: if the old call is still alive, the record says interrupted while the call keeps spending.

Limits

  • The executor doesn’t lock the run. Two processes over the same store are yours to keep apart; the supervised runner holds a process lock for you.
  • One live fallback per attempt. The lane order is the host’s, chosen once.
examples/preflight-executor.ts
The engine and its bindings are in examples/preflight-host.mjs. The types are GraphExecutorResult for what run() and resume() return, RunPreflightPolicy for the plan’s checks, RunPreflightState for what readRunPreflight returns, and PreflightPauseResult and PreflightFailureResult for the two ways the checks stop a run early.

Next steps