Skip to main content
Turn one option on and the run serves a page about itself while it works: each declared step with its live state, the findings that made a step run again, the questions waiting for a person, and the tail of the record. Use it to watch a long run without tailing a log. To answer questions with a page of your own, use a surface; this page’s one control answers a waiting question through the same callbacks client. The smallest use turns the page on and reads the address from the record:
examples/monitor.ts (excerpt)
monitor: true binds a free port on the loopback address and folds the run’s own events into the state of each declared step, served as a page and as JSON. The address is written once as a monitor event, so it’s in the record and in any onEvent sink; the run never prints it. result.monitor carries the same url and a close(). The page needs:
  • monitor: true on a plain run; the default is off. Under supervise it’s on unless you pass monitor: false.
  • A process that stays up for as long as you want to watch. The server doesn’t keep the process alive: a script that finishes exits as it would without the page, and the page goes with it.

Read the state

The example runs two plain steps offline, reads the page’s state as JSON once the run is done, and closes the page:
examples/monitor.ts (excerpt)
Run it with npx tsx monitor.ts. Open the printed address while a longer run is going to watch it move:
Output
The page said done after the run and both steps were done. The page shows:
  • Each declared step, in order, with what it needs, its desc, and its phase: declared, running, done with its outcome, or skipped. A step that ran more than once says how many times.
  • Findings that made a step run again: which step returned them, to which step, why, and how much of that step’s budget is used.
  • Questions waiting for a person, each with what it is about and a form for the answer that the question asks for.
  • The tail of the record, the last forty of the two hundred events the page keeps in memory. Read the record file for all of them.
The page polls once a second and stops when the run is done. Only the steps of the outermost graph appear; a graph nested inside a step shows through that step’s outcome.

The record panel

The panel prints the record the way the console does, one line per event, through the same formatEvent the console uses, so the browser and the terminal never disagree about a run. A line starts with the time and the step’s path, then says what happened:
  • A tool call names the tool, the phase and what it acted on: the file for a read or an edit, the first two words of a command, the URL of a fetch. Claude Code, the Claude Agent SDK, Grok and OpenCode report the target. Codex and the Anthropic API report no tool calls at all, so a step on one of those seats shows none.
  • A step’s end carries its summary, so you read what the step concluded without opening the record.
  • A step that runs again gets one line: which step returned the findings, which step runs again, why, and how much of that step’s limit is used.
  • Streamed text and thinking never appear as rows. They arrive as fragments, and the step’s summary says what the model concluded.
This is the panel during a run of the writer, reader and judge team over a docs page, at the point the route returned the draft to the writer:
Example record panel
The route’s line ends with its summary, the judge’s reason. The kickback line says the draft went from route back to write, the first of six rounds the writer is allowed, and the next line is the writer starting again.

The approval card

When a step asks a person a question, the card shows what they are approving before it asks for the answer:
  • The question, in the words the step wrote it.
  • The input, the thing the question is about. For the approval steps in the examples, that is the file and the sha of its bytes, so the yes covers exactly the bytes the reviewer saw. A changed file is a new question.
  • A link, when the input carries a url, so the reviewer opens the diff or the page the question is about.
Below those sit a field for a note and the Yes and No buttons. A no with a note sends the note to the step the question names as its target, and that step runs again with it. A question with no target fails its step, with the note as the reason.

Other questions

The card builds each answer from the shape the question asks for. That shape is the question’s responseSchema, and /state sends it with each waiting question.
  • A judge’s product decision gets a box for your written decision and a Send button. Write the decision the work should follow. The page sends it as the answer’s prompt. The page fills in feedback when the question accepts an empty one or allows only one value. Otherwise the card asks you for it too. The judge page says what the run does with it.
  • Any other question gets one labelled field for each field the answer needs, and a Send button.
  • A question that accepts more than one shape of answer also gets an Answer with list. Pick a shape, fill in its fields, and send. The page sends only the fields of the shape you picked.
When the run refuses an answer, the page says why, and what you typed stays in the form. The reason stays on the page until you send again or the question is answered.

Answer a waiting question

Nothing on the page starts, stops or edits the run. Any other route is 404, any other method 405, and a body that isn’t JSON 400. A request whose Host isn’t the bound address is refused, and /answer takes only application/json, so a page from another origin can’t post to it.
Keep the page off on a shared host. It carries no token, so anyone on the machine the run is on can answer a waiting question.

Failure

  • A run that never started still binds the port. Call result.monitor.close() anyway.
  • An environment that fails to come up fails the run, and the page shows that final state. A process that stays up keeps serving the final state until close() is called.
examples/monitor.ts

Next steps