Motivation
Left alone, a model is the only judge of its own work. A loop that never stops is the other failure, and both waste a budget. A feedback loop puts a check between the attempt and the next step, names who fixes what the check found, and caps the rounds. The writer’s job stays narrow: fix these findings, not start again. The record keeps every round, its findings and the count against the limit, so you can read why a run stopped.Parts
- The check. A reviewer on another model, a panel of them, a test command, or a person. Each returns a pass or a fail with findings.
- The target. The step that owns the fix. A fail carries the findings to it, and it runs again with them, not from the start of the run.
- The stopping rule. A number of refinements, or a judge that reads the findings and the rounds so far and says whether another round is worth it. A judge needs no number: the run ends when it stops the rounds or the review passes. When the refinements run out, the run fails with the last findings instead of looping. A judge can take an optional cap as a backstop; the judge’s answer about the last round the cap allows decides how the run ends.
Counting rounds
A round is one build and the check of it. The first build is round 1.refine: Nallows N refinements. A refinement is one more build after the first, with the findings. Sorefine: Nallows N+1 builds in all, andrefine: 0allows one build and no second. The default is 1.- The same number means the same in every form.
refine: Non a stage a panel or a person reviews,refine: Non a stage another stage sends work back to,maxKickbacks: { target: N }in adag(), and a judge’scap: Neach allow N refinements. - A reviewed stage that is also sent work back to has one count. Its
own reviews and the send-backs from later stages share its refinements
and its judge’s rounds in one run. With
refine: 1, a review that rejects the first build uses the one refinement, so a later send-back fails the run. - A judge with no cap sets no limit. The judge or a passing review ends the rounds.
- A judge with a cap reads the last round too. With
cap: N, the judge is asked about the review of build N+1. Only an answer that lets the work stand passes; any other answer fails the step. - An unmet requirement from a goal check counts as a round. It goes back to the builder without the reviews or the judge.
- A
loop()counts its body runs. Itsmaxis the total number of runs, the first included, somax: N + 1allows the same rounds asrefine: N.
Shapes
- One reviewer. A model from a different family reads the writer’s result and passes or fails it. Anthropic calls the pattern evaluator-optimizer. A writer and a reviewer is that case.
- A panel with a threshold. Several reviewers read at once and the step passes on the count you set. Set the pass count to the number of reviewers when one voice must be enough. Some people call this an eval loop. A review panel with a threshold is that case.
- A panel whose reviews become one. With
synthesiseset, a seat merges the findings that name the same problem, each reviewer votes once on the findings it did not raise, and a finding most of them reject is dropped, unless it is a block. A judge then decides each finding that is left: act on it or skip it. Ask a panel shows the whole process. - A tournament with a judge. Several candidates run in their own worktrees, a function scores each, and only the winner lands. The rest leave nothing behind. Three candidates, one winner.
- A command whose exit code decides. A test runs; red carries its output to the target, or picks the branch that runs next. No model in between. A command decides the path.
- A person’s decision. Yes passes the step. No carries the person’s note to the target as the finding. Silence pauses the run. A person decides.
- A judge that says when to stop. A small model reads what the
reviewer found and how the rounds have gone, and decides each finding:
act on it or skip it. The builder gets only the findings it acts on, each
with the judge’s reason. When the judge gives no reason of its own, as
with
jev(), the reason is what the option it chose means. Every review that can send work back to that step is told what the judge skipped and why: an agent reviewer reads the list in its prompt, and your own review code reads it asctx.skippedFindings. When the judge skips every finding, the work stands. The judge decides a block finding too. A block goes back without a judge only when there is no judge. A judge stops the loop. - A draft refined over rounds. Translate, reflect, revise, then a person reads the result. That page calls it a refinement loop. Translate and reflect.
Example
The offline feature team names the target and the threshold on the panel:examples/feature-team.ts (excerpt)
implement, which runs again. The second attempt passes on two of three:
Example record
Limits
- A panel needs a target. Without one, a failing panel fails the run instead of running a step again.
- A step declares where it sends work back. In a
dag()or apipeline(), the step that sends work back lists its targets inacceptsKickbackTo, and each must be a step it depends on. A send-back to a step it does not list fails the sender with an error that names both. Inworkflow(), a stage’ssendsBackTodeclares it. maxKickbackssets the limit. Without it, a graph or pipeline never runs a step again on a fail.workflow()sets it for you from each target’srefine, which defaults to 1; a target withrefine: judge(seat)asks the judge before each round and sends back only the findings it acts on. With no cap, its rounds end when the judge stops them or the review passes. A number is one budget for the whole graph: N send-backs in all, to any target. A map,{ plan: 3, 'tests-first': 3 }, gives each target its own N refinements, so a review that spends its budget on one step can’t starve another. A target the map doesn’t name gets none.- Every round is recorded. Each request to run again is in the record with its count, and with its limit when the budget has one.