the long view
every merged commit, one row each. every number is a deterministic count, so any machine extends the series and none can smear it. ci measures the commit under test, appends it as the last point, and redraws this chart on every pull request; what merges is what you see here.
the welfare column is written under one formula. when the formula moves, every row's score is recomputed on the counters that row already holds and the column is rewritten in place, so any two points on the line are answers to the same question. it is also the one line with a rule attached: no change to the compiler may lower it, and a change that would gets optimized until it holds. every other number here is a softer gate — a counter may worsen only if something else improved, and each worsening needs a written sentence naming why.
an older row holds fewer counters than the formula reads, because the counter set grew with the project. each term is scored on the counters the row carries; a term with none of them is dropped and the remaining weights are renormalized, so a row that scores well on what it has reads like a full row that scores well. a row carrying none of the counters any term reads gets no welfare column at all, and the line has a gap where it sits. every row records the three things that say what scored it: scored_by names the formula, scored_weight gives the share of that formula's weight the row's counters cover, and scored_base names the epoch its baseline was measured in. the compile counters have been re-measured three times on smaller workloads, and a row is scored against the baseline as it stood before the next of those changes ahead of it; a row past all three reads as measured.
read scored_weight before reading a step in the line. the formula reads five counters and an older row carries fewer of them. compile allocations and compile memory go back to the first row, 2026-08-09, minus the stretch between 2026-08-13 and 2026-08-24 that the miscompilation below cost; compile instructions joins on 2026-09-03; the run program's retired instructions and its peak bytes join on 2026-09-06, when the objective's run side became one consolidated program. a row holding two of the five is scored at coverage 0.28, four at 0.74, all five at 1.00, and a row holding none of them has no score at all. the chart marks every one of those changes with a dashed rule and draws the score faded wherever the coverage is below 1.00, because a step across one of them is a change in what was recorded rather than in what the compiler costs. there are two, and these are their figures on 2026-09-09. compile instructions joining on 2026-09-03 takes 0.28 to 0.74 and the score from 88.28 to 62.67. the run peak joining on 2026-09-06 takes 0.74 to 1.00 and the score from 64.95 to 56.73, because the compile terms carry the advantage they have accumulated since august while the run terms start at parity against a baseline measured on the day they joined. the forty-eight rows at 0.74 hold four of the five rather than three because their run instruction count is REBUILT from the eight per-phase counters that were the measurement before runbench existed, not taken on the day; the rebuild is what the second dashed rule sits beside. neither is the compiler. every other fall in the file is a counter that got worse.
a change to the scoring function is a change to what the project says it wants, and it is the one thing allowed to lower the number. every such change is in the history of bench/welfare_floor.json with the reason written beside it.
each line is scaled to its own range, because an instruction count and a byte peak cannot share an axis and stay readable. what the chart shows is shape — which way a number went and when it stepped. the horizontal axis is commit order, because commits are what the ratchet acts on. a line starts where its counters start, and the schema grew with the project, so compile allocations and compile memory reach back to the first recorded row, compile instructions reaches back to 2026-09-03, and the two run lines to 2026-09-06.
two of these lines have a gap in them. the run-memory and compile-memory series were blank for a stretch because the program that writes these rows was miscompiled: an interpolation shared one buffer across a loop, so five groups of counters left it with their key names run together into a single key each. the rows were written, the job was green, and the chart simply could not find the numbers. the compiler bug is fixed, and so is the reason nothing caught it — the row is now checked against the same lists on the way out as on the way in, and the check runs on every pull request rather than only after a merge. a completeness check on the inputs is not one on the outputs.
loading the chart…
a dashed vertical rule marks each change in scoring coverage, labelled with the share of the formula the rows to its right are scored on; the welfare line is drawn faded wherever that share is below 1.00. how each line is derived, from the code that draws it.
the first five are the counters the welfare formula reads, one line each;
bench/objective_sources.txt is that list and a spec replays it
against the code below. run instructions is
run_instructions, retired instructions for the one consolidated
run program, counted under callgrind. run memory is
run_peak_bytes, that program's arena, held and permanent peaks
summed. compile instructions, compile
allocations and compile memory are
compile_instructions, compile_allocs and
compile_peak_bytes — what deciding and emitting cost, in retired
instructions, allocator calls, and peak bytes held. binary size
is the decoder's .text section in bytes, from
bench/text_golden.txt — its own only, since every binary links the
same runtime and summing them would count it four times. it is drawn dashed
because it is the one line the score does not read: a machine-code-size term
was ruled out of welfare on 2026-09-05, and .text kept its own
exact vein instead.
welfare is the row's welfare column, written by
scripts/welfare_rescore under the formula
scripts/welfare emits: for every term the row can be scored on,
the mean over its present counters of r / (r + satiation) where
r is baseline over the row's reading, weighted, summed, and
divided by the weight those terms carry, times a hundred.
every one of these is an exact integer a rerun reproduces, which is what lets them be diffed rather than eyeballed. the memory and size lines reproduce on any machine; the instruction line reproduces on one compiler, because instruction count is what an optimiser's inlining choices move most, and a toolchain change steps every row of it at once.
retired instructions are the line to read for speed, and allocation counts answer a different question: how often the allocator was called. the two come apart, because a guard that runs per element allocates nothing. a decoder measured at 2,545,249,871 instructions and later at 2,762,364,162 — eight and a half per cent more work, tracking wall clock — reported byte-identical allocation counters across the same span. counting the work directly is what makes that difference visible.
counterswhat every commit costs
the chart shows shape; these are the numbers themselves, as ci recorded them on the last push to main. each is deterministic, so a change here is somebody's deliberate edit rather than weather. the panel opens with the seven lines the chart draws, in the chart's order, so a line you cannot pick out up there can be read off down here.
the numbers above were measured by hand on a named day. these are measured by ci on every push to main, and they are the deterministic ones — allocation counts, arena blocks, rewind iterations, kernel presence, and what the compiler spent deciding and emitting. a noisy runner cannot move any of them, so a change in this panel is somebody's deliberate edit rather than weather. select a row to read what the counter measures and why it is worth watching.
loading the history…
the counts rise across the span because the corpus does: the basket gained a hundred-thousand-element accumulator and a fifty-kilobyte text workload, the encode vein gained lib/json, and each widening was re-baselined so the score read the same either side. these are the cost of the work as it stands each day, not of a frozen file.