Skip to main content
A post has to go out on the company page this week. It has a brief, a house style with six rules, and one editor who says what is published. Most first drafts break a rule or two: a superlative slips in, a number appears that nobody gave. Catching those is a grader’s job, and it is tedious for the editor to do by hand every time. You want the draft written and then read against the rules by someone who didn’t write it, with each finding naming the line and the rule. You want the writer to go again with those findings, but not for ever: someone has to say when the draft is good enough for a company page and another round has stopped being worth it. And you want the editor’s decision to be the only thing that publishes. Obversa makes the grader a stage and the stopping rule a judge. A Claude seat writes the draft. A Codex seat, a different model family, reads it against the style guide and returns findings; the writer runs again with them. After each round, Jev, a small decision model, reads the findings and the rounds so far and says whether another round is worth it. When the draft holds, the run stops for the editor’s yes. This file is two stages, a draft with its grader and the editor’s decision, and every draft, every finding and every answer from the judge is on the record.

Run it

Set the project up as Installation describes. Copy the file with briefs/, style/ and judge.json beside it, sign in to Claude Code and Codex, and run it from that directory. Without JUDGE=jev the judge replays the answers in judge.json; with it, and a TypeSafe endpoint and key in the environment, the judge is Jev:
Terminal
The output below is the proof’s offline run, with scripted seats standing in for the models, so the words are the script’s and the shape is the run’s.
Output, from the offline proof
Three drafts, three rounds of findings, three answers from the judge. The first draft called the review “the best in the business”, and the grader named the line and the rule against superlatives; the judge said another round was worth it. The second draft explained the loop and added a number the brief never gave, and the grader named that; the judge said the same again. The third draft drew one taste note, and the judge said the draft holds, so the run stopped for the editor. Nothing was published.

The file

The brief names the post, the style guide, the grader’s job and the editor’s:
briefs/post.md
The style guide is what the grader reads against, one rule per line:
style/house.md
The judge’s answers for the offline run, one set per round. A live run asks Jev the same questions instead:
judge.json
The whole team is one workflow(): two roles from different model families, an editor, a judge on the draft stage, and two stages:
examples/use-cases/editorial/writer-grader-cap.ts (excerpt)
reviewedBy: 'grade' puts the Codex seat on every draft. When the grader does not pass a draft, refine: judge(judgeSeat) asks the judge before the writer runs again. The rounds end when the judge stops them or the grader passes the draft. There is no round limit unless you set one. workflow() refuses the team before any model runs if the writer and the grader share a family. The publish stage is the editor’s question; a no goes to draft with the editor’s note as the finding.
examples/use-cases/editorial/writer-grader-cap.ts

The team’s shape

What the run did

The proof runs the file against scripted seats and a recorded judge, and checks the loop the page describes: three writer turns, three grader reads, the judge answering continue, continue, holds, the superlative and the invented number gone from the draft on disk, and the run stopped at the editor. Obversa recorded the run as one event log, so it shows each draft, each set of findings and each answer from the judge in order, with the reasons on the record for the editor to read.

When the loop stops

A count says how many rounds a draft may take. A judge says whether the next one is worth taking. After each round the grader did not pass, Jev reads the use case from the brief, the grader’s findings and every round so far, and answers whether the draft holds, whether the findings are worth acting on, and whether another round is worth it. The loop goes round again on continue, and stops when the judge chooses a reason to stop: the draft holds, the rest is polish, or the same findings keep coming back. The judge decides a finding tagged block too. To bound the rounds as well, pass a cap as a backstop: judge(judgeSeat, { cap: 3 }). It allows three refinements, four drafts in all. When the fourth draft does not pass, the judge is told it is the last round. If it lets the draft stand, the draft goes to the editor, with the open findings on the record. Any other answer fails the run, and the reason names the cap, so the editor is never asked. A judge stops the loop has the questions and the rule.

Next steps