Skip to main content
An article has to go out in French next week. There’s a glossary the company has settled on, a tone the readers expect, and one person who knows those readers. A translation from one model call reads fine to anyone who isn’t one of them. You want the glossary honoured even where another word reads better, and you want to know where that cost something. You want an editor’s pass by someone who didn’t write it, and the last call on nuance made by the person who knows the audience, not by a model. Obversa makes that a refinement loop with a person at the end. One model translates with the glossary open. A model from another family reads the translation the way an editor would and fails it with what to change, not a score. The translator runs again with those findings, and after each round a judge, Jev, reads the findings and the rounds so far and says whether another round is worth it. Then the translator writes down how every glossary term was rendered and where the glossary and natural French pulled apart, and the run waits for the person who knows the readers. This file is a three-stage team: translate, terms, nuance.

Run it

Set the project up as Installation describes, put briefs/translation.md, source/article.md, glossary/en-fr.md and judge.json from examples/use-cases/other/ beside the file, sign in to Claude Code and Codex, and run it. Without JUDGE=jev the judge replays the answers in judge.json; with it, and a TypeSafe endpoint and key in the environment, the judge is Jev:
Terminal
When the run reaches the editor, it waits and prints a page to answer on. The editor answers there, and the run finishes. Keep the process running until then: if it stops, the next run starts again from the first stage. Every event prints as one line as it happens, and the outcome prints last as JSON. The output below is the proof’s offline run, with scripted seats standing in for the models, a recorded judge, and the proof answering as the editor on the page, so the words are the script’s and the shape is the run’s:
Output, from the offline proof
Two translations, two reads, two answers from the judge. The first translation rendered two glossary terms with the words the glossary rules out; the reviewer named them, and the judge said another round was worth it. The second followed the glossary and drew one taste note, and the judge said the translation holds. The run waited at nuance with the translation in fr/article.md and the terms note in fr/terms.md, and the editor’s yes on the page finished it. A no goes to translate with the editor’s note as the finding.

The file

The brief’s front matter names the source and the glossary as workspace files, so every seat knows they exist and the translating seat can’t write them. The brief holds the translator to the glossary even where another word reads better, and asks for each such place in the terms note, so the person decides with the trade-off in front of them.
examples/use-cases/other/translate-reflect.ts (excerpt)
The translation is reviewed by the Codex seat, reading as an editor. When the reviewer does not accept it, refine: judge(judgeSeat) asks the judge whether another round is worth it. With no cap, the loop ends when the judge stops it or the reviewer accepts the translation. The terms note has no reviewer, so only its own seat judges whether every glossary term is on the list. Then the run waits for the editor:
examples/use-cases/other/translate-reflect.ts (excerpt)
The “done when” sentences are each stage’s gate. The package doesn’t check them itself; it checks what each stage declares. A file in writes that’s missing or empty fails the stage, a reviewed stage passes only when its reviewer accepts, and a person stage waits until the person answers.
examples/use-cases/other/translate-reflect.ts

The team’s shape

When the loop stops

A translation that honours a glossary and still reads naturally has many defensible answers, so the reviewer and the translator need room to meet, and no count says how much. After each round the reviewer did not accept, the judge reads the use case from the brief, the reviewer’s findings and every round so far, and answers whether the translation holds, whether the findings are worth acting on, and whether another round is worth it. The loop goes round again on continue and stops when the judge chooses a reason to stop. The judge decides a finding tagged block too. With no cap, a run ends when the judge stops it or the review passes. To bound the rounds as well, pass judge(judgeSeat, { cap: 3 }): it allows three refinements, the judge reads the review of the last draft too, and the run fails unless it lets the translation stand. The judge’s recorded answers for the offline run are in judge.json, one set per round. A judge stops the loop has the questions and the rule.

Next steps