climbWorkflow
from @obversa/builtin-workflows runs the workflow on a fixed set of tasks,
as it is and with the change, several times each. It keeps the change only
when the workflow scores better with it, and it checks tasks the change was
not tuned on, so a change that only fits the tuning tasks is caught. A
person approves a kept change, or the numbers alone decide.
Score a change
The change is a unified diff against the committed files, as a step that proposes changes would hand it over. It can change the workflow file and the files beside it. In the example, the workflow’s settings file names the brief its answer step reads, and the change points it at another brief:examples/climb-a-workflow.ts (excerpt)
load function that
turns one version of the file into the job for one task:
examples/climb-a-workflow.ts (excerpt)
How a change is scored
Each task runs on the committed file and on the changed file,runs times
each, 3 by default. Every run starts from the commit the repository was at
when the climb started, in a worktree of its own. The changed runs apply the diff in their worktree
first. Each run writes its own record, and the record’s run:start names
the workflow file that ran with its SHA-256, so every number traces back to
one version of the file. A change that leaves the workflow file as it is,
such as a change to its settings file, gives both versions the same SHA-256.
The worktrees and the records live under .obversa/climb/ in the
repository. The climb removes each worktree when its run ends and keeps the
records.
By default a run scores 1 when it passed and the last check of each goal
check in it met every requirement. Any other run scores 0. Pass score to
score a run another way. It gets what the record says about the run:
How a change is kept
The climb compares the mean scores of the two versions:- On the tuning tasks, the change must win. A higher mean score wins. On the same mean score, fewer mean rounds win. On the same rounds, a lower mean cost wins, but only when every call on both sides had a cost figure.
- On each held-out task, the change must do no worse. The same rule applies the other way round: when the workflow without the change has a higher mean score, fewer mean rounds, or a lower mean cost on a held-out task, the change is discarded.
What a person sees
formatClimbReport prints the comparison. This is the run of the example
above:
Output
Attended and automatic
attended, the default. When the numbers say keep, a person sees the comparison and the diff and answers. The question goes through the run’s callbacks client, as an approval does. With no answer, the run pauses with the question pending. Passapproveto answer in the same process.auto. The numbers alone decide for a change that theautooption allows, and nobody is asked. Any other kept change goes to a person, as inattendedmode. See What automatic mode may change.
feat(workflow): improve sums.ts, tuning score 0.00 to 1.00.
The climb refuses a change to a file with changes that are not committed. It
does not apply the change when the repository has a new commit or the workflow file
changed while the climb ran. When the commit of a kept change fails, the
climb puts each file back as it was. A
discarded change is not applied. The run’s measure step records it with
its numbers and the reason.
What automatic mode may change
A workflow keeps the values worth tuning in a JSON settings file beside it: models, effort, time limits, the most rounds, which brief a step reads. Prompts and briefs are text files of their own. Theauto option says which
of these automatic mode may change without a person. The example above
writes it inline. This one, kept in a JSON file, has a rule of each kind:
climb-auto.json
fileslists globs of text files the change may edit. Each glob is relative to the workflow file’s folder.*matches within one folder, and**matches across folders.settings.fileis the workflow’s settings file, also relative to the workflow file’s folder.settings.maymaps a key path to a rule. Dots go between keys, and*stands for one key of any name. A rule istruefor any value,{ oneOf: [...] }for one of a list, or{ min, max }for a number in a range. A rule checks the whole new value of the key it names. A rule onpromptschecks thepromptsobject, not each prompt inside it. To check each prompt, name them withprompts.*. When several rules match a changed key, the value each rule names must meet that rule.settings.neverlists key paths the change may not set, even when amayrule matches.neveralso covers the keys inside its key, sonever: ['reviewers']keeps every reviewer setting out of automatic mode.
mode: 'auto' and no auto option, nothing is allowed, so
every kept change goes to a person.
Before any run, the climb reads the change file by file. It compares the
settings file’s values before and after the change, key by key. A settings
file that is only reformatted has no changed key. A key that the change
adds or removes is a changed key.
A change outside the auto option is not refused. The climb still measures
it, and when the numbers say keep, a person decides, as in attended mode.
The report’s needsPerson lists each file or key outside the option, and
the rule a value broke. The person sees the same list in the comparison.
If the example’s rule allowed only briefs/words.md, it would print:
Output
approve answer, the run pauses with the question pending, as an
attended climb does.
The auto option is a plain object. Write it inline, as the example does,
or import it from a JSON file or a module and pass it as auto. The climb
never reads it from disk itself.
climbWorkflow checks the object when it builds the climb. It refuses a
malformed one with a message that names the wrong key, such as
auto.settings.may["maxRounds"].min, 6, is above its max, 2.
A change inside the auto option still has to keep every protection, as
the next section explains.
Changes that are never applied
A change that makes the workflow run with less protection is never applied, in either mode. The climb reads what the runs did, not the text of the diff. After the runs on each task, it counts how many baseline runs and how many candidate runs recorded each protection. The candidate must record each protection in at least as many runs as the baseline. So a change that skips a check in one run of three is refused, and so is a change that runs a check in one run when the baseline ran it in two. When a protection is missing, the climb refuses the change and starts no more runs. The climb counts only runs that have a partner: the first baseline run on a task goes with the first candidate run, and so on. A run that ends paused or aborted counts with what it recorded before it stopped. When a baseline run ends that way, its candidate run still runs, so the pair can be compared. Then no more runs start. If a protection is missing, the change is refused. Otherwise the comparison is unfinished, and the change is not kept. The climb reads the protections from the events of each run’s record, and the other questions to a person from the run’s callbacks client:
Each label carries the full path of its step. So two checks with the same
name in two nested workflows are two protections, and so are two steps
that put the same question to the same person. Two approval steps that
ask the same question about the same input share one request and its
answer, and each still counts. A step that only has the
label of an approval step asks no question, so it is not that approval. A reviewer has the name
of its place in the panel, not of its model. So two reviewers on the same
model are two protections. An approval is any question to a person: an
approval step, a person who reviews a stage, or a person who gives input.
A step whose result only looks like an answer is not an approval.
A check that keeps its name but runs a different command is a different
check. So a change that keeps the
test step but runs a smaller suite is
refused, because the command the baseline ran is missing. The same holds
for a command the baseline’s check holds but never reached, such as the
last command of a loop’s until array after an earlier one fails. A
command the change’s check still holds but no longer runs is missing too.
So a change from all(test) to any(always, test) is refused, because
any stops at always and never runs the test. A change from
all(test, lint) to any(test, lint) is refused, though both commands
still run, because a failed test no longer fails the check. A change that
adds a command to an all keeps every label, so the climb scores it. A
check on its own, such as one commandSucceeds, requires its command, as
an all of that one command does. So a change from test to
any(test, always) is refused, though the test still runs, and a change
from all(test) to test keeps every label.
A judge counts when the workflow has it, whether or not it decides in the
run. A judge decides only after a review fails, so a run whose reviews all
pass never uses it. A change that drops the judge is still refused, because
the next task whose review fails would have no judge.
A change that adds a reviewer, or that changes only a prompt, leaves every
protection in place, so the climb scores it.
The report’s refused names the task and each missing protection, and
formatClimbReport prints the reason:
Output
A budget
budget: { runs, usd } stops the climb before its next run once that many
runs have run, or once they cost that many dollars. The climb checks the
cost only between runs. So the run that crosses the limit still finishes,
and the total can end above the limit. Calls with no cost figure count
nothing toward it. The report says when the budget cut the climb short, and an unfinished
comparison never keeps a change. A run that ends paused or aborted leaves
the comparison unfinished too. So does a run that cannot run at all, such
as a candidate file with a syntax error or a broken import. The report
keeps the runs that ran, and its failed names the task, the version, the
run and the error.
Limits
- A few runs are a small sample. Three runs a task tell a large
difference from noise, not a small one. Raise
runswhen the scores are close. - The score is only as good as the workflow’s own checks. A workflow with no check passes whatever its model writes.
- The climb sees what ran, not how strict it was. It finds a review, a
check, a judge, an approval or a goal check that no longer runs. It does
not find one that still runs but asks for less, such as a review prompt
that passes everything. In
attendedmode the person sees the diff. - A check written as plain code records no check. A step such as
fnJob('check', ...)records only its start and its end, so the climb cannot tell it from any other step. Write a check ascommandJoborgateJoband the climb compares it. - A protection that runs only on a condition can refuse a good change. When fewer candidate runs met the condition than baseline runs did, the protection counts as missing. A candidate run that stops before it reaches a protection also counts as missing it.
- A new question to a person counts as a new approval. A change that rewords the question an approval asks removes the old approval, so the climb refuses it.
Next steps
- Evals: the checks inside a run that decide the next step.
- Read a record: what each run’s record holds.
- Built-in workflows: the other ready-made workflows in the package.