Shape
The stopping rule
The judge’s seat isjev() from
@obversa/engine-jev-api, one line like any
other seat. Offline, recordedJudge from @obversa/runtime/testing stands
in for it and replays the answers in judge.json:
examples/judge-stops-the-loop.ts (excerpt)
write stage. The reader reviews the page, and
refine names the judge:
examples/judge-stops-the-loop.ts (excerpt)
block, should-fix or
nice-to-have. A block is a false claim or a sentence a person would not
understand.
When a judge is set, it decides every finding on its own, a block included:
act or skip, with a one-line reason. A taste note does not cost a round,
and a real finding is not waved through because the rest of the list was
trivial.
- The judge acts on at least one finding. The page goes back. The writer gets only the findings the judge acts on, each with the judge’s reason.
- The judge skips every finding. The page stands as a pass, with the judge’s reason.
- The reader’s next round. The reader is told which findings the judge skipped and why, so it does not raise them again.
refine or maxKickbacks.
When the stage also sets goal, a goal check runs before the reader,
every round. A requirement of the brief that it finds unmet goes straight
back to the writer. The reader does not run that round, and the judge is
not asked, because a requirement in the brief is never skipped as polish.
Check the brief was met shows it.
The approval is a plain function. It reads the page, puts its sha256 in the
question and sends a person’s no back to the writer.
With several reviewers, set synthesise on the stage as well. The panel
then merges the findings that name the same problem, lets each reviewer
vote once on the findings it did not raise, and drops a finding when the
reviewers who reject it outnumber those who raised or backed it, unless it
is a block. The judge decides each finding on that one
list, with who raised it and how the others voted.
Ask a panel shows it.
A cap as a backstop
To bound the rounds as well, pass a cap:judge(judgeSeat, { cap: 3 }).
A cap of 3 allows three refinements:
the first draft and three more, four builds in all. It means the same on a
workflow() stage’s refine and on a dag()’s maxKickbacks.
When the review of the last build does not pass, the judge is asked about
it too and told that this is the last round. Its answer decides the outcome:
holdsorover_polishing. The work stands as a pass with the judge’s reason. The findings still open are on the outcome asopenFindings, and in the judge’s entry in the record.product_decision. A person is asked, and their answer is recorded. No build round follows to apply it, so the run fails, and the reason quotes their answer.continue,not_convergingor no clear answer. The run fails, and the reason names the cap.
What the judge sees
The judge reads no file itself. The runtime puts together what it reads, in code, with no extra model call. For the same rounds, aworkflow() stage
and a dag() give the judge the same input. The judge reads it in four
labelled parts, always in this order, and then the questions. The four
parts, and cut when the size limit cut something, are the state of the
judge’s { state, questions } prompt.
why: what the work is for. The brief, the use case, and the target’sdescandgate. Aworkflow()passes itsbrief, and takes the use case from the brief’sUse case:paragraph. Adag()passes its ownbriefanduseCase. A person’s answers to the judge’s earlier product decisions are here too.what: the change. In a git workspace, each file changed since the run of the workflow or dag began, with the lines added and removed. For each changed file and line a finding cites, such assrc/app.ts:42, the parts of the diff within ten lines of that line. A finding that cites a file with no line gets that file’s whole diff. When the target writes a file, the judge also gets its content: the first file a stagewrites, or a dag node’sfile.how: this round’s evidence. The latest result of each check step in the graph, arun:stage or acommandJob: its command, its status and, when it did not pass, its output. The verdicts of the target’s own goal check, when it has one. Each finding, with the id its question uses, who raised it and its severity. When the panel synthesised, each finding also carries the other reviewers’ votes and their reasons.when: where the rounds stand. The round number, the cap, and whether this is the last round. Every earlier round, with its findings, the judge’s decision on each and its reason, and how many lines that round changed. The findings the judge skipped, with its reasons. In a git workspace, the files changed since the last round.
The size limit
What the judge reads can be large, so it is kept to 50000 characters, counted as JSON. Set another limit with thejudgeContextLimit option of
run(), for example run(job, { judgeContextLimit: 20000 }).
When the four parts are over the limit, the runtime cuts every long part
down to one shared length, set so that the whole fits. A short part is cut
only when cutting the long ones is not enough. These parts can be cut:
- Texts. The brief, the use case, the
desc, thegate, each diff, the file’s content and each check’s output. A text keeps its start. - Lists. A person’s earlier answers, the changed files, the goal check’s verdicts, the earlier rounds, the findings skipped before and the files changed since the last round. A list keeps whole items. The earlier rounds, the skipped findings and the answers keep the latest. The other lists keep their first items.
cut, after the four parts, names each
cut, for example
the content of page.md: kept the first 20000 of 31250 characters.
Each judge entry in the record names the target and the round. It also
gives the size of what the judge read, in characters, and the same list of
cuts.
The questions
For each finding in the round, the judge gets onechoice question, keyed
finding-1, finding-2 and so on in the order of the round’s findings. It
chooses act or skip. A block keeps its severity in what the judge sees,
and its question sets a higher bar: skip it only when the case it names is
outside how the work is really used, or the same class of finding keeps
returning after it was answered. Otherwise the judge acts on it. When an
answer carries no reason of its own, the
reason is the text of the chosen option. The judge
answers these in the same call as the questions about the whole round. The
default set of those is stopQuestions() from
@obversa/runtime. It asks four questions:
holds. For this use case, does the draft hold as it stands?worth_doing. For this use case, are the latest findings worth acting on?worth_another_round. Given the whole history, is another round worth it?stop_reason. Are we at the point of diminishing returns? The judge choosesholds,over_polishing,not_converging,continueorproduct_decision.
act, another round runs. When every finding is skip, the
work stands as a pass. This holds whatever the round answers say, so a
chosen not_converging neither drops a finding the judge acts on nor fails
a round whose findings it skips. The one exception is product_decision:
the judge asks a person as described below. A finding the judge does not
answer follows the round answer: it goes back when that answer is
continue, and it is skipped when that answer is holds or
over_polishing.
A dag() takes the same judge() value on maxKickbacks. The target that
the work goes back to gets only the findings the judge acts on, and a stop
means the same thing at both levels. The judge runs as part of the node that
sent the work back, inside that node’s time limit. The unmet requirements a
goalCheck() node sends back go to the target without the judge.
A stage a person reviews takes the same judge. Each refusal goes to the
judge with the person’s note as a finding, as a reviewer’s findings do.
Pass your own set as judge(seat, { questions }) when this wording
does not fit your use case. With your own set, the judge is not asked about
each finding, and the whole round goes back or stands on one answer. When
the judge chooses holds or over_polishing, the work stands as a pass
carrying the judge’s reason. Every other chosen stop leaves the review’s
failure in place. A custom question set that wants a reason of its own to
let the work stand names that choice holds or over_polishing. Add
perFinding: true to ask about each finding with your own set as well. The
judge engine takes { state, questions } as its prompt and returns one
answer per question; the Jev API engine page has
the shape.
The built-in stopQuestions() also offers product_decision: the work
needs a person’s judgement before review can continue. Pass an interaction
binding to judge() to open your review surface. Its request includes the
work and the review history. After the person answers, the work goes back
to the writer. The writer gets the person’s structured feedback, their
composed prompt and the findings they answered. With a cap, that round
counts against it like any other round. The review then runs on the new
draft, and only after that does the judge decide again. After the last
review a cap allows, a person is still asked and their answer is recorded,
but no round follows to apply it, so the run fails. Without an interaction handler,
answer through the run’s stored callbacks client, or on the
run page with a written decision. This gets an answer while
the work is being revised. A later human review can still require you to
approve the finished work.
What the run did
Run offline, with the judge’s answers replayed fromjudge.json beside the
brief and the person’s answer recorded in approve.json, the example
printed:
Output
should-fix (the page never
says where to run the command) and a nice-to-have. The judge acted on the
first and skipped the second, so the writer got only the first. The third
pass found one nice-to-have, and the judge skipped it, so the page stood.
The record holds each round’s findings and the judge’s answers. Each answer
also says whether the page went back or stopped, which answer decided that,
and for a stop, whether the step passed or failed. It lists each finding’s
id with the judge’s decision and reason, for example:
JUDGE=jev, and TYPESAFE_ENDPOINT and TYPESAFE_API_KEY in the
environment, the same file asks Jev instead of replaying.
Next steps
- Feedback loops: the check, the target and the stopping rule, and the other shapes a loop takes.
- Jev API engine: the judge’s question types and how it reports its identity.
- A person decides: the approval step this loop ends on, bound to the exact bytes.