Skip to main content
Ask several models from different families to review the same work. Their reviews become one list before anyone acts on it, and a judge decides each finding on that list: act on it or skip it. Use it for a page, a plan or a change where one review isn’t enough. Three things go wrong when several reviews are simply added together. Two reviewers raise the same problem, and the writer gets it twice. One reviewer raises a point the others would reject, and it goes back anyway. A reviewer who passes the work has its notes thrown away. With synthesise set, the panel fixes all three before the writer or the judge reads a single finding. For a review by a single model, get a second opinion.

Shape

The team

One writer and three reviewers, each from a different model family:
examples/teams/review-battery.ts (excerpt)
The writer’s stage names the panel in reviewedBy, turns on synthesise, and puts a judge on refine:
examples/teams/review-battery.ts (excerpt)

The steps

Each round runs these steps in order. When the stage sets goal, a goal check runs before them, every round. When it finds a requirement of the brief unmet, the round goes back to the writer, and none of these steps run. The judge never decides an unmet requirement. Check the brief was met shows it.
  1. Review. The three reviewers read the page at the same time. Each one passes it or asks for a revision, and tags each finding block, should-fix or nice-to-have. A reviewer that passes can still list findings; the panel keeps them. A finding from a reviewer that passed, with no tag, counts as nice-to-have.
  2. Merge. One seat reads every finding and groups the ones that name the same problem, even when the reviewers word it differently. Each group becomes one finding. It takes the strongest severity in the group, credits every reviewer who raised it in raisedBy, and keeps the evidence and the fix the merging seat picked as clearest. synthesise: true uses the first reviewer’s seat for this step. Name a seat instead, synthesise: claude('claude-sonnet-4-5'), to use another.
  3. Cross-review. Each reviewer sees the merged findings it did not raise. It answers each one agree, disagree or better fix, with a one-line reason, and for better fix the fix it would make. This is one round: each reviewer replies once. Each reviewer gets the brief and the files under review again, because it answers in a fresh turn. A reviewer whose seat declares tools can read the work to check a claim.
  4. Votes. Plain code applies the votes. It counts only the reviewers who voted on that finding.
    • Most disagree: the finding is dropped, but only when the reviewers who disagree also outnumber the ones who raised it, agreed or offered a better fix. Their reasons stay in the record. Otherwise the finding stays, marked disputed, for the judge. So when two of three reviewers raise the same problem, one disagree from the third cannot drop it.
    • Exactly half disagree: the finding stays, marked disputed, for the judge.
    • Most prefer a better fix: the finding carries the first fix offered, in place of its own.
    • Otherwise the finding stays as it is.
    Votes never drop a block. When most voters disagree with one, it stays, marked disputed. Votes also never drop a finding shown to a reviewer whose seat failed or sent no readable reply in this step. It stays, marked disputed.
  5. Judge. When a judge is set, it decides every finding that is left, a block included. It reads each one with who raised it and how the others voted. It decides each one, act or skip, with a one-line reason. The writer gets only the findings the judge acts on, each with the judge’s reason. The next round’s reviewers are told which findings it skipped and why. When the judge acts on any finding, another round runs. When it skips every finding, the page stands as a pass. With no cap on the judge, the rounds end when it skips every finding or every reviewer passes. A block goes straight back to the writer without a judge only when there is no judge. Know when to stop has the judge’s questions.
The panel passes or fails on how many reviewers passed the work, with one exception. When the votes drop every finding a failing panel had, nothing is left to send back, so the panel passes. A panel of one reviewer has nothing to merge or vote on, so synthesise leaves it as it is. synthesise takes the same value on a panel: stage and on reviewPanel(). On reviewPanel(), every reviewer names its seat, which answers the cross-review round, and context gives the voters the brief and the files.

What the run did

The proof runs the file offline. A stand-in plays the Claude, Codex and OpenCode command line tools, and the judge replays the answers recorded in judge.json. The file printed:
Output
The reviewers are named after the stage: write-1 is the Claude seat, write-2 the Codex seat and write-3 the Gemini seat. In round one, write-1 and write-2 both said the page uses “backoff” without saying what it means. The merge made that one finding, crediting both. It kept the evidence from write-1 and the fix from write-2. write-3 asked for a table of every client setting. The other two voted against it, so it was dropped and never reached the writer. write-3 also raised a small wording point. write-1 agreed and write-2 disagreed, so it stayed, marked disputed. The judge then decided the three findings left. It acted on the merged “backoff” finding and on the missing last retry, and skipped the disputed wording point as taste. The writer got only those two findings, each with the judge’s reason. In round two, every reviewer passed the page, so there was nothing to merge, vote on or judge, and the run passed. Each panel’s outcome carries synthesis, one PanelSynthesisEntry per finding. Its result is kept, better fix, disputed or dropped. Its finding carries raisedBy, and votes, a list of FindingVote answers, one per voter. The record also holds a review:synthesis event for each round with findings, with the same entries. The event says when the merger or a voter failed or sent no readable reply: mergeFailed means the findings went on unmerged, and noVotesFrom names the reviewers whose votes are missing. The judge’s decision on each finding is in the refine:judge event, under findings. Run the same file from the directory the work belongs in, with the three command line tools signed in, to use the real seats; add JUDGE=jev to ask Jev.
examples/teams/review-battery.ts

A simpler panel: count the acceptances

When you only need to know how many reviewers accept, leave synthesise off. The panel counts the acceptances and sends every finding from the reviewers that rejected the work back to the writer.

Shape

The team

One implementer and a panel of two reviewers from two other model families. agree is how many must accept:
examples/teams/threshold-panel.ts (excerpt)
The reviewers run at the same time, one job each, once the test passes. agree is a whole number from one to the number of reviewers, and the default is all of them. Set it to the number of reviewers when one dissent must be enough to fail the step; below that, the change can pass despite a dissent, and the dissent is still on the record. A panel below the threshold carries its findings to implement, which runs again with them, as many times as that stage’s refine allows. When those rounds run out, the run fails with the last findings. workflow() refuses the team before any model runs if a reviewer shares a model family with the seat it reviews, or two reviewers share one: two reviewers from one family would be one opinion counted twice. The third seat runs through the OpenCode command line tool, whose adapter needs an absolute command path; resolveCommandExecutable('opencode') finds it on your PATH when the file runs.

What the run did

Run the file from the directory the work belongs in, with the Claude Code, Codex and OpenCode command line tools signed in. One real run printed:
Output
The Codex reviewer accepted. The OpenCode seat returned no decision: a reviewer’s decision is its reply, the panel reads the first JSON object in it, and a reply with none is asked for once more, then counted as an engine error, not a rejection. Nothing went to the implementer on its account. With agree: 1 the change passed on the one acceptance. Set agree to two and the same run would have paused at the panel instead, because an error isn’t an acceptance. The implementer wrote src/double.mjs:
src/double.mjs
and test/double.test.mjs:
test/double.test.mjs
examples/teams/threshold-panel.ts

Next steps

  • Feature delivery: a synthesised panel of three as one of seven steps.
  • Review loop: the same idea as a graph form, with a quorum and requireDiversity you set yourself.
  • Runtime: workflow, stage, panel and agree.