Run it
Set the project up as Installation describes, putbriefs/backlog.md, backlog/raw.md and judge.json from
examples/use-cases/engineering/ beside the file, sign in to Claude Code and
Codex, and run it. Without JUDGE=jev the judge replays the answers in
judge.json; with it, and a TypeSafe endpoint and key in the environment,
the judge is Jev:
Terminal
Output, from the offline proof
The file
One model grooms, a model from another family reviews, and the owner is a person:examples/use-cases/engineering/backlog-groom-then-rank.ts (excerpt)
clarify follows the same shape and writes backlog/questions.md: for each
story, the questions that must be answered before anyone writes code, with a
proposed answer for each. Then the run waits for the owner:
examples/use-cases/engineering/backlog-groom-then-rank.ts (excerpt)
backlog/ for the owner to read. A no
goes to split with the owner’s note as the finding. Priority is the
owner’s call: the brief tells the models not to decide it, so any story
order in stories.md would only be a model’s guess.
The “done when” sentences are each stage’s gate. The writing seat reads
them in its prompt, but the package doesn’t check them. It checks what each
stage declares: a file in writes that’s missing or empty fails the stage,
a reviewed stage passes only when its reviewer accepts, and a person stage
waits until the person answers. A raw ticket with no story fails the split,
because the reviewer reads the raw tickets as well as the stories; that’s
what happens in backlog-groom-then-rank.proof.ts. A split sent findings
must change: if the same bytes come back, the stage ends instead of running
a second identical review.
Full file
Full file
examples/use-cases/engineering/backlog-groom-then-rank.ts
The team’s shape
When the loop stops
After each round the reviewer did not accept the stories, the judge reads the use case from the brief, the reviewer’s findings and every round so far, and answers whether the stories hold, whether the findings are worth acting on, and whether another round is worth it. The loop goes round again on continue and stops when the judge chooses a reason to stop: the stories hold, the rest is polish, or the same findings keep coming back. The judge decides a finding tagged block too. With no cap, a run ends when the judge stops it or the review passes. To bound the rounds as well, passjudge(judgeSeat, { cap: 3 }): it allows three
refinements, the judge reads
the review of the last draft too, and the run fails unless it lets the
stories stand. The questions stage keeps a plain count of three
refinements, so the file shows both shapes. The judge’s recorded answers for the offline run
are in judge.json, one set per round.
A judge stops the loop has the questions
and the rule.
Next steps
- Feature delivery: what happens to a story once it’s ranked and picked up.
- A judge stops the loop: the judge’s questions, and the same loop as plain graph nodes.
- A person decides: the owner’s question as a step, and how the answer reaches a paused run.