Skip to main content
The week leaves a backlog behind: support threads, a sales ask, a one-line wish from the founder. Before anyone codes, someone has to turn each into a story a developer could pick up, with acceptance checks and the questions to settle first. On most teams that’s an afternoon nobody wants. You want stories you can trust: every raw ticket covered, every story checked against the ticket it came from by someone who wasn’t the writer, and the open questions written down instead of discovered mid-sprint. You want to keep the one decision that is yours, the order. Obversa gives the grooming to two models and keeps the ranking for you. One model turns the tickets into stories and lists the questions; a model from another family reads them back against the raw tickets and fails the set when a ticket is missed or a story is vague; a judge, Jev, says when another round on the stories is worth it; then the run waits for the product owner. This file is a three-stage team: split, clarify, rank. Nothing in it writes code. The work is deciding what’s worth writing.

Run it

Set the project up as Installation describes, put briefs/backlog.md, backlog/raw.md and judge.json from examples/use-cases/engineering/ beside the file, sign in to Claude Code and Codex, and run it. Without JUDGE=jev the judge replays the answers in judge.json; with it, and a TypeSafe endpoint and key in the environment, the judge is Jev:
Terminal
When the run reaches the owner, it waits and prints a page to answer on. The owner answers there, and the run finishes. Keep the process running until then: if it stops, the next run starts again from the first stage. Every event prints as one line as it happens, usage lines included, and the outcome prints last as JSON. The output below is the proof’s offline run, with scripted seats standing in for the models, a recorded judge, and the proof answering as the owner on the page, so the words are the script’s and the shape is the run’s:
Output, from the offline proof
Two rounds of stories, two answers from the judge, one set of questions. The first set of stories left the founder’s one-line wish uncovered; the reviewer named the ticket, and the judge said another round was worth it. The second set covered every ticket and drew one taste note, and the judge said the stories hold. The questions passed their review first time, and the run waited for the owner. The owner’s yes on the page finished it. A grooming loop that stops with the work unfinished is doing its job. One that runs out of rounds and passes the stories on anyway isn’t. Splitting a backlog of vague tickets has many defensible answers, so no count says how many rounds it needs. Here the judge says, round by round, and the rounds end when the judge stops them or the reviewer accepts the stories.

The file

One model grooms, a model from another family reviews, and the owner is a person:
examples/use-cases/engineering/backlog-groom-then-rank.ts (excerpt)
clarify follows the same shape and writes backlog/questions.md: for each story, the questions that must be answered before anyone writes code, with a proposed answer for each. Then the run waits for the owner:
examples/use-cases/engineering/backlog-groom-then-rank.ts (excerpt)
The stories and the questions are in backlog/ for the owner to read. A no goes to split with the owner’s note as the finding. Priority is the owner’s call: the brief tells the models not to decide it, so any story order in stories.md would only be a model’s guess. The “done when” sentences are each stage’s gate. The writing seat reads them in its prompt, but the package doesn’t check them. It checks what each stage declares: a file in writes that’s missing or empty fails the stage, a reviewed stage passes only when its reviewer accepts, and a person stage waits until the person answers. A raw ticket with no story fails the split, because the reviewer reads the raw tickets as well as the stories; that’s what happens in backlog-groom-then-rank.proof.ts. A split sent findings must change: if the same bytes come back, the stage ends instead of running a second identical review.
examples/use-cases/engineering/backlog-groom-then-rank.ts

The team’s shape

When the loop stops

After each round the reviewer did not accept the stories, the judge reads the use case from the brief, the reviewer’s findings and every round so far, and answers whether the stories hold, whether the findings are worth acting on, and whether another round is worth it. The loop goes round again on continue and stops when the judge chooses a reason to stop: the stories hold, the rest is polish, or the same findings keep coming back. The judge decides a finding tagged block too. With no cap, a run ends when the judge stops it or the review passes. To bound the rounds as well, pass judge(judgeSeat, { cap: 3 }): it allows three refinements, the judge reads the review of the last draft too, and the run fails unless it lets the stories stand. The questions stage keeps a plain count of three refinements, so the file shows both shapes. The judge’s recorded answers for the offline run are in judge.json, one set per round. A judge stops the loop has the questions and the rule.

Next steps