guardrails, not lifetimes
here is a json decoder reading a 188 kb file—nested objects, arrays, escapes, unicode. it is written in a pure language with no pointers, no manual memory, and not one lifetime annotation. beside it, on the same machine and the same file, are the decoder a rust team hand-tunes and ships, a rust parser written the way most people would write it, and go's standard library. the clock is milliseconds of cpu per decode and the four lanes ran interleaved. shorter is faster.
- 01constraints buy freedom
- 02the journey, honestly
- 03the strict tier: rewinding arenas
- 04the last stretch to serde
- 05the rest of the license
- 06the deep end
- 07the lazy bet
- techniquesthe ledger
- 08the receipts
- 09the surface a program needs
- codathe mushroom test
| decoder | ms/decode | peak mem | |
|---|---|---|---|
| kanso + arenas | 0.87 | 4.2 mb | |
| serde_json (rust, hand-tuned) | 0.90 | 6.8 mb | |
| reasonably-written rust | 1.04 | 6.9 mb | |
| go encoding/json | 2.05 | 11.8 mb |
sat 2026-08-07, seven rounds, four lanes interleaved so contention leans on all of them at once. milliseconds of cpu per decode, read as a slope from the same decoder built to run 150 times and 450 times so process startup and the file read cancel. cpu time bills every thread, so go's collector counts against it. this board is re-sat by hand when a release goes out; nothing rewrites it on a merge.
the person who wrote the kanso decoder never typed the word lifetime, never named a buffer, never freed anything, never chose a memory strategy. a reasonably-written rust parser running the same algorithm trails by a fifth, and go takes better than twice as long with its collector working. the recipe at the bottom of this page runs the whole field on your own machine in a few minutes.
a pure, annotation-free language keeping pace with hand-tuned rust is a claim worth doubting. so the rest of this page is the evidence: the papers we read, the experiments that failed (twice, in one case), the one idea that did most of the work, and the recipe. no lifetimes, no garbage collector, no annotations — guardrails instead.
and the json table is now the older of the page's two scoreboards. kanso is a lazy language—a value is computed when it's demanded, or never—and on workloads where demand is conditional, the second scoreboard (§07) shows the clear kanso form fifteen times ahead of rust as naturally written and within forty percent of rust hand-restructured to skip the same work, with counters proving 95% of the work was never born. the strict-tier story below and the lazy story in §07 are two halves of one design.
1constraints buy freedom
picture a fast car on a mountain road with a cliff on one side and no railing. even a great driver takes the corners slowly there, because a single mistake goes over the edge, so caution has to live in every turn.
now bolt a guardrail along the drop. the same driver takes the same corners faster, because the railing has made going over the edge impossible, and once the worst outcome is off the table there is nothing left to be careful about. the guardrail lets a good driver off the brakes.
kanso is built out of guardrails. functions are pure. nothing mutates. every program has one canonical way it's written — a parenthesis pair that groups nothing is rejected the same way a stray blank line is, so 1 + (2 * 3) does not sit beside 1 + 2 * 3 as a second spelling of one thing. the only branch in the language is dispatch—choosing a function by the shape of its arguments—so there is no free-floating if on values for a compiler to squint at. to the person writing kanso, these are just the rules of a plain, readable language, and following them costs nothing.
to the compiler, they are railings. an ordinary compiler has to assume the worst at every turn: this value might be shared, this call might have printed something, this branch might do anything. so it drives slowly. kanso's compiler knows the worst can't happen, because the language already forbade it, and so it can take optimizations that would be reckless, even unsound, in a language without the railings. it treats each constraint as something to spend rather than work around.
2the journey, honestly
the goal was straightforward: run faster than an ordinary rust program without asking the programmer to do rust's work by hand. there are two known ways to manage memory safely, and both charge the programmer for it. rust makes you prove in the source exactly when each value dies; those proofs are lifetimes, and writing them is real work. garbage collection skips the proofs and instead pauses the program every so often to find and free dead memory; the cost moves from your keyboard to your latency. we wanted a third option, so we spent a while in the literature.
the card we expected to play: perceus
the most promising idea in the literature is reference counting with reuse. give every object a small tally of how many things currently point at it. when the tally hits zero, the object is dead and can be freed on the spot—no janitor, no pause. and if an object is about to die and the program is right then building a new one of the same shape, don't free-and-reallocate, just overwrite the old one in place. this is called perceus; koka invented it, and lean 4 and roc ship it in production. it is real, and it is good.
we expected to build it. then we did the arithmetic on kanso's own budget, and it failed twice over.
first, the tally has to live somewhere: a small counter attached to every object, paid for on every object, whether or not the reuse ever pays off. second, and worse, is what happens when a big structure dies. dropping the last pointer to the root of a large tree walks the whole tree, decrementing each child and freeing as it goes. that walk is a pause — a small garbage collection by another name, and it happens exactly when a big result is thrown away, which in a parser is constantly.
a cascade is me, but for memory, and it doesn't announce itself. one pointer drops and ten thousand children get freed on a hot path. at least i ride the return values in the open, where you can see me coming.
the confession
there is an embarrassing part of this story, and leaving it out would make the page worse. for a while our own notes listed in-place reuse as built and working. then a review actually read the code instead of the notes. the analysis that was supposed to find reusable objects ran correctly and handed its answer to—nothing. no part of the compiler consumed it. it was dead code that had been dead for weeks, and meanwhile the per-object tally it was meant to justify was still sitting on every object, costing us the overhead with none of the benefit.
that's the kind of thing a differential test corpus and an append-only log exist to catch, and it's why the honesty tags on this page are load-bearing rather than decorative. the one reuse that does fire today is the narrow, provable case—appending to a list you uniquely own, overwriting in place. the general machinery got pulled back to the bench.
where perceus landed
we didn't delete the idea. we demoted it. reference counting survives in the design as a fallback — something we can drop in for a specific pattern if the static approach ever turns out to copy too much — rather than the main mechanism. the mechanism we actually use came almost for free from the shape of the program itself, which is the next section.
a few other roads we walked to the end and turned back from, so no one has to walk them again:
interaction nets (the hvm/bend line) are spectacular on a narrow class of problems and roughly ten times slower on ordinary numeric code—the retrospective saying so is by the person who invented them. adopting them means replacing the whole runtime, not adding one pass. they stay on a watch list, not the road.
an eight-byte ascii skip in the utf-8 validator looked like free speed: gulp eight bytes of text at once, and if they're all plain ascii, skip the per-byte checking. we tried it. it made the benchmark slower by 3%, because the strings in real json are mostly shorter than eight bytes, so the setup cost was paid every time and the skip almost never fired. then—months later, having forgotten—we tried the exact same thing again, and it lost again by the same margin. the log now says, in as many words: do not try a third time without a workload full of long strings.
3the strict tier: rewinding arenasshipped
this section teaches the shipped allocator for strict values — one of the two tiers named in the intro, and the one behind the json scoreboard. the lazy tier and the counted destination are §07's story; read them as two halves.
the idea that does most of the work needs no background beyond "programs use memory." each step below is small.
the problem everyone has
a running program constantly asks for little pieces of memory—a list here, a string there, thousands of times a second. asking is easy. giving the pieces back is the hard part, because handing back a piece something still needs corrupts the program, and never handing anything back fills the machine until it dies. the three classic answers are the three we've already met: do it by hand (fast, and famously how programs get hacked), collect the garbage at runtime (safe, with pauses), or prove to the compiler when each piece dies the way rust does (fast and safe, but you write the proof).
programs are loops
kanso's answer starts from a plain observation about the shape of real programs: they repeat. serve one request. handle one keystroke. decode one file. take one step of a loop. work arrives, work happens, an answer leaves—and then the whole thing repeats. call each cycle a beat. almost everything a program makes, it makes inside a beat, purely in service of that beat's answer, and the instant the answer leaves, all of that scaffolding is trash at once.
two kinds of memory
so kanso keeps two kinds of memory, matching the two kinds of thing a beat touches.
the shelf holds long-lived state—the thing an imperative program would keep in a mutable global. in kanso that's one value with exactly one owner, handed from beat to beat, so the compiler just updates it in place and never has to wonder about it.
the workbench holds everything else: the intermediate lists, the parsed fragments, the half-built strings, all the scaffolding a beat erects on the way to its answer.
the sweep
at the end of a beat, kanso does not go looking through the workbench asking which scraps are garbage. they all are—the answer already left. so it sweeps the whole bench with a single move: one pointer slides back to where the bench started, and everything past it is gone. a thousand things made or one thing made, the sweep costs the same, because it never touches the things individually. it's about as expensive as setting a variable to zero.
the strategy, precisely
so the story reads honestly end to end, here is the whole memory strategy in four sentences. counting is the model: kanso's values are immutable and the data graph is acyclic, so reference counting is complete—every cell freeable at its last use, no tracing, no collector, no soundness gap. the compiler then proves most counts away: a batch it can show dies together carries no counts at all and lives on a rewinding arena—one pointer reset, no walk, which is this section and the json receipts in §08; a value it can show is used linearly is reused in place, no count either. real counts survive only where lifetime is a runtime fact: thunks—a rewindable arena structurally cannot hold a pending computation (§07 has the receipts from jhc, ghc, and the ml kit)—and structures whose sharing escapes proof. the objections to counting-everywhere are real—a tally on every object, a cascading walk when a big tree dies—and they are exactly why the arena is the proved fast path inside the counted model, not a casualty of it.
the experiments behind that ruling are in §07; the engines walk there in the open, ratchet by ratchet.
a garbage collector spends its time working out which objects are still alive. the sweep never does that work, because at the end of a beat everything on the workbench is already dead.
and now sharing is free
the entire reason reference counting exists is to answer one question: when two things point at the same piece of memory, who frees it last? counting answers by tallying pointers—and paying, a little on every share and occasionally a lot on that cascade.
inside a beat, the question has no meaning. point at something twice, point at it ten times, let two half-built results share a chunk—nobody tallies anything, because it all dies at the same instant no matter who was pointing where. "who frees it last?" has one answer everywhere: the beat does.
the two loose ends
survivors. once in a while a value made on the workbench genuinely needs to outlive its beat. the compiler's escape analysis—built and tested—spots these at compile time. a small survivor is copied onto the shelf, and the copy is placed before the program ever runs. a big survivor isn't copied at all: its bench is kept whole, a fresh one is started, and the old bench carries a single count of how many survivors still point into it—one integer per bench, never per object—swept the ordinary way when that count hits zero. no walk, no cascade, no per-object tally.
long beats. a beat that piles up a mountain of scrap before it finishes can be handed intermediate sweep points mid-beat. because kanso's function bodies are flat and the whole program is visible to the compiler, it knows exactly which values are still alive at any line, so those extra sweeps are placed statically too.
what the programmer does
nothing. there is no keyword in this section. no lifetime, no annotation, no region name, no tuning flag. you write the loop, the parser, the handler; the cycle is already in the shape of the program, and the compiler reads it and places every memory decision at compile time. nothing does this work at runtime, because by the time the program runs there is nothing left to decide.
the receipt
kanso's allocator was once the worst case of this design: one enormous workbench that never got swept, so a long-running program's memory just climbed. we made the bench resettable and let the compiler drop a beat boundary between decodes of the json benchmark—the same 188 kb file, over and over. the difference is not a matter of percentages. grow-only, the memory climbs without bound; swept every beat, it stays flat.
grow-only, the peak tracks the work: 0.76 gb at 150 decodes, 15 gb at 3000, because the program keeps every scrap it ever touched. swept every beat, the peak is 5.8 mb at 150 decodes and 5.8 mb at 3000 — flat, because a decode's scraps die with the decode. and the swept version ran faster too: instead of asking the operating system for fresh, cold pages, it reused the same warm ones, so the cache stayed hot for the whole run.
the honesty tier. the compiler now emits the beat boundaries itself: an analysis proves that a loop's iterations keep nothing across the line. with the per-object header gone from the arena, the emitted version measures 5.8 mb peak—under serde's 6.7 and the reasonably-written rust's 6.8. the copy-or-pin split for survivors was the last thing here still called planned; it is now declined on its own numbers, which the ledger carries — pinning the survivor's page died to a measurement before it was built, and retaining the region instead was built and measured and lost. what remains planned is the static sweep points for long beats.
"almost everything a program makes, it makes inside a beat" is the claim this whole section rests on, and it now has a number on the four boards. at a beat boundary the live set is exactly what the carry stages, and the runtime already sizes it to allocate a buffer, so counting it costs nothing:
shelf boundaries live crossing (max) arena held then
decode 1 112 bytes 1,048,576
encode 401 80 bytes 4,194,304
one-shot 3 995,504 3,145,728
basket 10,000 0 1,048,576
encode crosses a hundred and twenty-eight bytes in total across four hundred and one boundaries, and basket crosses nothing at all across ten thousand. "rare" is rarer than the design assumed. one-shot is the exception at 31.6%, which is the same board the evacuation counter singles out, and the two agree about where the remaining cost lives.
what this measures is liveness at the rewind, so it bounds how much of what the arena holds is already dead — at encode's rewinds, essentially all four megabytes of it. it does not say what a reference-counted runtime's peak would be, because such a runtime frees as values die rather than at boundaries, and its peak is the largest set live at one instant. that quantity is still unmeasured, and it is the last unknown in the comparison. the honest ceiling on the gap is smaller than the table looks in any case: the arena's floor is one block, so eighty bytes is not a footprint any policy here could reach.
this is the corpus, not the verdict. the number that will ultimately judge the design is still the one this section has always named—write the kanso compiler in kanso and count survivors on a program nobody tuned for the boards.
4the last stretch to serdeshipped
the arenas closed most of the gap on their own. rust pays a separate malloc and free for every object unless its programmer hand-builds an arena to avoid it—which is exactly the kind of expert move most code never gets around to. kanso's shape makes that arena automatic, sound, and invisible. but the arena alone left a representational gap against serde—tagged values and per-call dispatch in the hot loop—and three changes, in this order, closed most of it. (laziness later spent a slice of the margin back; the board above is the current, honest ledger.)
rung one: the inline twins −10.7%
the runtime has a handful of tiny helpers—is this value truthy? did that call fail?—each one line long. we had marked them to be folded directly into the code that calls them, and asked the linker's whole-program optimizer (the -flto pass) to do the folding at the very end of the build. it quietly declined. deep in the release binary, twenty-seven of those tiny calls were still there, each one a little detour off the hot path. finding them cost an evening of staring at disassembly, waiting for the missing pieces to explain themselves.
the fix was to stop asking the linker later and have the compiler do the inlining itself, as it emits the code. we moved the decision into the ir. it was worth more than the other two changes combined, and it left a lesson in the log in block capitals: never trust -flto to inline across that boundary—verify it in the disassembly.
rung two: look at sixteen bytes at once −1.4%
the string scanner is hunting for two bytes: the closing quote and the backslash. the plain way checks one byte, then the next, then the next. the fast way loads sixteen bytes into a single wide register and asks "is either target anywhere in here?" in one instruction—the vector units every modern chip has (neon on apple silicon, sse2 on intel). it's serde's own trick, borrowed in the open.
rung three: the integer fast path −2.2%
kanso integers are arbitrary precision—they never overflow. but almost every integer in a real json file is small enough to fit in a single machine word. so the common case runs a tight loop that just accumulates digits, and only the rare genuinely-enormous number falls through to the general, bignum-capable path. floats are left exactly as they were, because their shortest-round-trip rendering is sacred and not worth risking for a few percent.
all three produce byte-identical output to the versions they replaced, checked on both intel and apple-silicon ci. and they're held in place by cost goldens—the receipts section explains that ratchet—so none of the three can quietly regress.
the write path runs the same playbook
everything above is about reading json. writing it out uses the same ideas, and each piece is there because the obvious alternative loses.
the encoder threads one byte builder through the whole tree: an accumulator that owns a capacity-headed buffer and claims its own frontier, the same discipline the arena gives list push. a fold of appends is amortized linear, every intermediate stays an ordinary value, and each byte is written once. the obvious alternative—build a string per node and join them upward—recopies every byte at each nesting level on its way to the root, about six copies per byte on a flat document.
escaping dispatches on bytes, so the compiler gives it the same switchboard the decoder's scanner gets—dispatching on single-character strings would put a memcmp probe on every character written. and it keeps the decoder's favorite bet: one simd pass proves a string holds no quote, backslash, or control byte—almost every string in real json—and a proven-clean string is appended whole, in one copy.
numbers get the treatment their distributions deserve. almost no double survives shortest-round-trip rendering in under fifteen digits, so the precision probe starts at fifteen instead of one—probes below that are dtoa calls that cannot win. integers render through a dedicated digit loop rather than the general formatter.
the receipt lives where a user can feel it: kq's pretty-printer rides the same builder and wins every board of its interleaved race against jq—hardest on the biggest documents, where the recopying alternative hurts most. bench/kq_race.sh in that repo reproduces it, byte-identity checked before any timing starts.
5the rest of the license
the arenas and the ladder are where the headline number comes from. but the guardrails pay out in other places too, and these are the ones worth seeing before the deep end.
dispatch becomes a switchboard
a recursive-descent parser is one long question asked over and over: what character am i looking at? in kanso you answer it by writing one small function per case—one for a quote, one for a bracket, one for the letter t, and so on—and the language picks the right one by the argument. that's dispatch, and it's the only kind of branch kanso has.
the naive way to run seven cases is to check them one after another: is it a quote? no. a bracket? no. a t?—like walking into a room and asking every person in turn whether the call is for them. kanso's cases are a fixed, known set of literal bytes, so the compiler instead builds a switchboard: take the byte, use it as an index, jump straight to the one case that matches, in a single hop. seven functions in the source; one indexed jump in the machine.
and because the compiler knows the exact set of value shapes that can reach each call site, it stamps out a specialized copy of the code for each one, with the type-checking removed—the checks aren't predicted well or branch-hinted, they're simply gone, because on that path the answer was already known. the word for making one specialized copy per concrete shape is monomorphization, and it's exactly what you just watched happen to the switchboard. the receipts, with the emitted machine code, are in chapter 07.
recursion that never overflows
recursion is kanso's only loop, so a call in tail position—the last thing a function does—cannot be allowed to grow the stack, or every loop would eventually crash. it doesn't. the native backend emits musttail, the strongest promise llvm offers and one the code generator is required to honor; the browser backend emits wasm's return_call. the receipts: ten million frames of mutual recursion in constant stack natively, one million inside a browser tab. deep recursion carries no risk to manage, because in kanso it is simply how you write a loop.
arbitrary precision, paid for only when used
an int in kanso never overflows—that's a guarantee that holds on every engine and isn't for sale. what the compiler gets to decide is how rarely you pay for it. it emits each function in two versions. the fast one bets the number fits in a machine word and runs the plain, fast code, with a cheap check that bails out of the whole version if the bet is wrong. the bail restarts the call in a second version compiled against heap-backed bignums.
restarting a half-finished function would be a disaster in most languages—it might already have printed something, or mutated something, and you can't un-print. but kanso functions are pure and effects are just descriptions, so a restart changes nothing you could observe. javascript engines are famous for the same move—guess the likely type, recover when surprised—but they do it at runtime, and kanso does it ahead of time, with none of the machinery: no on-stack replacement, no deopt metadata, no interpreter to fall back into. and it was built to generalize; the integer bet is only the first probable-but-unproven fact worth wagering on.
the data rulings, briefly
three smaller decisions each hand the compiler more structure while handing the reader less to remember. destructuring always wears braces ({ artist minutes } = song), so every field read is name-resolution against a known record and compiles to fixed offsets. a type may declare zero fields (type null), so a marker like json's null is ordinary library data, not a keyword. and a field may list several permitted types (title:null string), so reads dispatch on what the value is and the "is it null or a value?" conditional becomes unwritable—the compiler sees a closed set of tags it can enumerate and often erase.
rendering is dispatch
string interpolation looks like a language feature and compiles like a library. "{x}" is a call on an ordinary dispatch group named to_string, whose arms for the primitive types ship with the primitives themselves. so giving your own type a custom rendering is not an api to learn—it's the language's one mechanism, applied:
type money
cents:int
pub play =
price = money 350
print "that costs {price}"
fn to_string (money cents)
"${cents / 100}.{cents - (cents / 100 * 100)}"
that program prints that costs $3.50. no import, no interface declaration, no registration—an arm for a type you own joins the group, and dispatch does the rest. the ownership rule from the import model is what keeps this safe: nobody can re-arm to_string for int, because no module owns both the group and the primitive, so "{42}" renders identically in every program ever compiled. and that guarantee is also an optimization license: where the compiler can see a value is primitive-only, it skips the dispatch entirely and emits the direct renderer—custom rendering costs nothing on the paths that don't use it, which the cost golden pins as a constant.
the whole toolchain in a browser tab
a third engine emits wasm bytecode directly—hand-rolled, no llvm anywhere near it, because llvm is measured in hundreds of megabytes and a browser tab is not. the whole thing is about half a megabyte of wasm, and it's what lets the playground compile the program you type into machine code inside the page, tail calls and all. its value representation is deliberately odd today (every value is a small handle into a registry the toolchain owns), and that first cut is staged toward a self-contained module over time, every step of the migration pinned by the same corpus of byte-identical golden outputs the other two engines are held to.
6the deep end
the memory manager that isn't
the arena tier is the picture; here is the theorem it draws. the claim is that kanso's compiler never has to reject a memory bug—a leak, a use-after-free, a who-frees-it-last—because there is no way to write one down. the bug is unrepresentable, the same way "invalid states unrepresentable" works for a database schema, applied here to memory itself.
you cannot get here by building a smarter analyzer. whether a value is uniquely owned is a non-trivial semantic property, so by rice's theorem it's undecidable in general, and the usual consolation—that a human can see cases the analyzer misses—doesn't hold: a human "seeing" uniqueness is producing a proof, and where no proof exists, no one sees it. worse, a reject-what-i-can't-prove checker is fragile in the exact way that ended whole-program usage inference in ghc (wansbrough & peyton jones, 1999): a distant, semantically-irrelevant edit flips an inferred fact and rejects code three modules away. so kanso builds neither a smarter checker nor a runtime fallback. it keeps the undecidable case from ever arising, by never letting a value's ownership become something the compiler has to settle at runtime.
the first construct does most of the work for free: value semantics. nothing aliases—every value is its own thing, and "mutation" produces a new value—so a use-after-free cannot be written. the compiler never has to check for it and grant it; the danger is simply absent, the way an undeclared variable is absent, and so the compiler never rejects for memory at all. what's left is not correctness but performance: for each value, can the compiler prove it uniquely owned (and reuse its storage in place) or not (and copy it)? the fallback for "not provably unique" is a static copy, settled at compile time—no per-object header, no runtime count. a distant edit that flips a value from unique to shared makes a copy appear—one line got slower—it never breaks the build.
this is where being closed-world stops being a footnote. rust's borrow checker is local: it must survive separate compilation and unknown callers, so it makes you write the proof as lifetimes. koka's perceus gives up proving uniqueness statically and tracks it in a runtime count (reinking, xie, de moura, leijen, PLDI 2021); its fip keyword (lorenzen, leijen, swierstra, ICFP 2023) recovers a static guarantee but still rests on that count. kanso sees every call site—no separate compilation, no unknown caller—so it can try to prove statically what those systems annotate or count, and settle reuse-or-copy before the program runs. no count, so no header; freeing is a drop the compiler inserts at the statically-known last use.
value semantics reuses beautifully for tree-shaped data threaded in a straight line. the three cases where it would otherwise force a copy or a count each already have a construct in the language whose shape keeps them settled, so the problem pattern is never expressed—only its safe form:
state is a fold. persistent state—the mutable global of an imperative language—is in kanso one value, threaded single-file through pure update(state, action) → state by the one executor loop. one owner, hand to hand, always uniquely owned: the ideal in-place case, not the aliasing trap. caching, the classic memory hazard, is here the most reuse-friendly pattern there is.
local mutation is a build block. where an algorithm genuinely wants scratch mutation—a hash table filling, an array sorting, a graph wiring itself up—a build block permits it under one rule: writes happen only through set, set parses only inside build, and its target must be born in the same block. the last expression freezes to an ordinary immutable value on the way out; outside the block, mutation doesn't parse. no var, no let mut, no mutability annotation on any name—mutability is a property of the place, so auditing "where can state change?" in a codebase is grep build. it's haskell's runST without the rank-2 ceremony, because closed-world makes "nothing escapes" a syntactic check.
and the block quietly settles the one question every reference-counted language dreads: cycles. set is an identity-preserving field write—a stays the same node while its field changes—which is exactly what closing a cycle requires and rebinding can never do. that looks like it should break counting, and here is why it doesn't:
immutable values cannot point at younger values. a value that existed before the block ran was already complete; making it point into the block would be mutation of pre-existing data, which is exactly what the block forbids. so every pointer in the heap aims pastward—except inside a build block, and the block boundary contains the exception. it follows that a cycle can only exist among values born in the same block: cycles cannot cross birthdays. the strongly-connected cluster is always one block's birth cohort, so the runtime counts the cohort, not the nodes—the block allocates into its own arena, the frozen result carries one count for the whole cohort, interior pointers (cycles included) are invisible to counting, and the last outside reference frees the arena in one shot. no cycle collector, no weak annotations, no leaks. the tradeoff, stated plainly: keeping one node of a frozen graph keeps its cohort alive, the way a go slice pins its array. python bolted a tracing collector onto its counts for this; swift asks humans to annotate weak and they get it wrong constantly; kanso's answer is a scoping rule that was already there for taste reasons.
irreducible sharing is a region—was the original fourth construct, and the laziness deep-dive demoted it. regions demand that a batch share one static lifetime, and a deferred computation's true lifetime is a runtime fact; every system that parked laziness in regions leaked or gave up (jhc, ghc's compact regions, the ml kit's gc backstop—receipts in §07). what replaced it: the value graph is acyclic except inside frozen build-block cohorts, and a cohort counts as one unit—so reference counting stays complete, irreducible sharing is just counted, and a region survives as the back-end optimization for batches the compiler proves die together. tofte–talpin stays honored as the ancestor of the bench; it stopped being the answer for sharing.
so the four bind together: a pure value core that reuses in place, a fold for state, a transient block for local mutation, a region for irreducible sharing. every program sits inside the statically-settled fragment by construction: the constructs that would create a memory problem are not in the language to reach for, so a checker never has to reject anything.
the honesty tiers, because most of this is design, not code yet. built and tested: the pure value core; whole-program borrow-vs-consume inference that hands each function a printable ownership signature—ownership as a per-function contract you can pin and blame, never an ambient verdict a distant edit silently flips; in-place reuse for uniquely-owned list builders; and build blocks themselves—build/set, the block-born rule, and cycle construction run byte-identically on all three engines, pinned by differential goldens. designed, spec-reserved, unbuilt: cohort freeing (the arena-per-block story the birthday theorem licenses), regions, and the static drop/free/reuse codegen the whole scheme rests on—the one part that's memory-unsafe to rush, and the reason it waits for a careful hand rather than a fast one. the residual cost, plainly: the fine-grained persistent-sharing pattern (clojure-style shared subtrees) that value semantics copies instead of sharing—a real speed cost in collection-heavy code, rare in the parsers and compilers kanso is for, and never once a program that fails to compile.
the frontier
the standing directive for what lands next: a core change is on the table only if it doesn't tax the user and the payoff is game-changing. the queue, each entry named with the property that licenses it:
generalized speculation—the integer trick from §05, past ints: bet on any probable-but-unproven fact (a dispatch target, list-not-map, ascii-not-utf8) with bail-and-restart, jit-style, ahead of time, zero deopt metadata. cost-bound inference (the raml lineage)—type-system-driven polynomial cost bounds so kanso check can flag "this used to be o(n), your change made it o(n²)," and feed the same bounds to the parallelism scheduler. equality saturation—rewrite the ir by e-graph instead of ordered passes; pipeline fusion (map f . map g becoming one pass) is unconditionally sound under purity, so phase-ordering stops being a problem category. simd structural scanning—simdjson-class byte classification feeding the switchboards the backend already emits, behind bytes primitives with no language surface; this is the separate, harder frontier that beating serde on raw throughput actually requires.
where the decode cost actually sits, measured 2026-09-02. callgrind on jsonbench, every function attributed to where it came from: 1,849,923,051 instructions in emitted kanso against 639,250,897 in runtime.c and 43,795,809 in libc — 73.0%, 25.2% and 1.7% of 2,533,004,731. that is 150 decodes of a 188,698-byte document, so a decode is 16,886,698 instructions and the whole cost of the thing is 89.5 instructions per input byte. the largest single symbol is value_for at 25.58%, and it is a merged one: clang inlined parse_string, parse_array, parse_object, parse_number and skip_ws into it, and not one of the five appears in the profile under its own name — so that figure is the whole value-parsing path rather than one fat function. the largest runtime entries are k_b_find2 at 3.36%, k_utf8_bad at 3.28% and k_b_to_float at 2.51%. the runtime's share has been falling as the twins move its one-liners into the emitted code and its dead guards come out: it was 31.5% of a larger total a day ago, and the three in-place mutations that led that list are all inlined now.
that paragraph has been re-sat twice since, so the pair of figures below belongs to the sitting kanso#1221 landed on rather than the one above. the measurement the day before it read 1,729,183,050 emitted against 941,714,306 in the runtime, and k_b_append_mut was third on that list. it is not on this one at any position, because the in-place byte append now inlines into whoever calls it: the runtime half of a decode fell 99,775,050 instructions, the emitted half rose 54,739,500 as the twin’s body crossed the line, and the difference is what the decode saved. reading the emitted share as a bigger number and the runtime share as a smaller one is reading the same code from a different side of an inlining decision, which is why the total is the figure to quote and the split is a map rather than a score.
the same measurement on 2026-08-31 read 1,728,709,950 emitted against 1,078,786,257 in the runtime. the emitted half has moved by 473,100 instructions and the runtime half has fallen by 111,406,137, which is what the last two days of runtime work look like from the outside: the one-byte append straight-lined and outlined, k_die and its siblings marked noreturn so a tag check stops building a frame for the error path it never takes, a sixty-four-deep summary walk removed from every beat pop, and the utf-8 ascii test answering a seven-byte run with two overlapping loads rather than a walk. k_b_append_mut went from 7.05% of the decode to 3.80% and k_utf8_bad from 4.11% to 3.02%. the frontier entries above are the ones that can still move the larger half, because they change what the backend emits rather than what it calls.
the mined queue: kernels the profiler already names
beneath the moonshots sits a band of narrower work where a published algorithm maps one-to-one onto a line in kanso's own profiles. each entry here carries its paper, the profile line it exists to erase, and byte-identity as its acceptance test. they land in order; the status marks move as they do.
1. shortest-round-trip float rendering shipped — ryū (adams, PLDI 2018). the digit core forms the half-ulp interval, scales it by a generated 125-bit power of five, and picks the shortest digit string inside — pure integer arithmetic, no probing, no dtoa; the format layer follows the old %g byte format but for one deliberate difference: a whole mantissa keeps its point in the exponent form, so a double of 1e15 renders 1.0e+15 where %g writes 1e+15 (ruled 2026-09-06, pinned by a_whole_float_keeps_its_point). the interpreter renders through the same rules over rust's shortest digits, so all engines agree byte for byte (a latent large-exponent divergence found and closed on the way). fuzzed over fifty million doubles: zero failures, with the only divergences being 495 legal shortenings in the subnormal range — the true-shortest canon the probe could never reach. the dtoa family is gone from the encode profile entirely. dragonbox (jeon, 2020) is ranked rather than queued, and the reason moved on 2026-09-02. rendering used to be 3.8% of the encode profile, behind an append-and-copy pair at 19.5% and the encode walker at 13.5%. four twins inlined the appends that day and encodebench fell 13.6%; on the sitting that leaves — 2026-09-02, and every 7.75% on this page is that day's — render_ryu is 7.75% and k_b_append_wide is off the profile altogether, so what ryū sits behind is the encode walker at 16.17% and the escape path.
a noinline probe on 2026-09-05 split that 7.75% in two, which the estimate above had been guessing at: the digit core is 85% of it and the %g format layer 15% — 396,958,400 instructions against 85,485,600, over 849,200 renders. so dragonbox's margin applies to nearly the whole line rather than part of it, and would recover something near two per cent of encode rather than one. the same probe found the core's own tail worth fixing first: it wrote its digits one at a time, a 64-bit division per digit into a scratch buffer and then a reversal, where a 100-entry pair table writes two at a time straight into place after a length ladder. encodebench fell 1.2992% for 112 bytes of machine code, and the four benchmarks that render floats are the only four that moved at all. at the function level the pair table took render_ryu from 480,160,400 instructions to 412,359,600, a fall of 14.12% in the line itself. as a share of encode those two read 9.20% and 8.09%, both on the 2026-09-05 sitting; neither is comparable with the 7.75% above, because three encode changes landed between the two days and each moved the denominator. nothing here is declined for being small — the order is by ceiling.
2. float parsing at memory speed shipped — the eisel–lemire algorithm (lemire, "number parsing at a gigabyte per second," 2021), now inside gcc, chrome, and rust's core. the decode mirror of the rendering entry, verified against an independently written reference over thirty million doubles, and pinned by a presence counter: el_parses reads 318450 on the decode board and 2123 on the encode one, so a change that drops the fast path turns ci red.
3. utf-8 validation in vector registers shipped — the full keiser & lemire algorithm ("validating utf-8 in less than one instruction per byte," 2021): three nibble lookups classify every two-byte window, a saturating compare pins the 3- and 4-byte continuation runs, all-ascii blocks skip classification entirely, and a trailing zero block makes truncation need no special case. spec-strict — overlongs, surrogates, beyond-U+10FFFF rejected — which also closed a latent divergence against the interpreter's strict validator, pinned by a differential golden. verified against an independent reference over seventy million boundary-crossing sequences at zero mismatches.
4. branchless map lookup measured, declined — eytzinger layout (khuong & morin, "array layouts for comparison-based searching," 2017). built, benchmarked, and reverted: at json-realistic map sizes (tens of keys) the index costs slightly more than the binary search it replaces, and at ten thousand string keys the two tie — string comparison dominates and layout can't help it. the paper's regime is huge arrays of machine-word keys; kanso's maps aren't that. the measurement stays here so the idea doesn't get re-mined.
5. building results in their final place already won — tail recursion modulo context (leijen & lorenzen, ICFP 2023). the technique writes each cell into its final position instead of threading an accumulator, and its regime is cons-cell construction. kanso builds with flat arrays and a frontier push, which already sits at that endpoint, so there was nothing left for it to buy. read and closed rather than built; the queue slot went to the names the profile actually showed.
6. in-place as a guarantee, not a find queued — fully-in-place functional programming (lorenzen, leijen, swierstra, ICFP 2023), plus call-pattern specialization (peyton jones, ICFP 2007). the linear analysis currently discovers in-place opportunities; FBIP's discipline turns "discovered" into "stated and checked," and SpecConstr automates the accumulator-shape specialization the enumerable's typed fold arms do by hand.
7. constants that stop being recomputed shipped — constant applicative forms (peyton jones, implementing functional languages; ghc evaluates a caf at most once). a zero-argument definition is a constant, and kanso used to rebuild it per call: the json decoder's bytes_false = [102 97 108 115 101] compiled to a function that heap-allocated a five-element list every time a false was parsed. now the body emits under a build symbol and the definition becomes a load from a cell filled once, before main, into permanent storage — the only cache an arena rewind cannot invalidate. the decode gauntlet drops 1,874,992 allocations and 95 mb of allocation bytes, and one arena block; per-decode cpu falls 6–14% depending on how quiet the machine is.
the first version cached lazily, checking a ready flag on each call, and measured slower — a store on that path is an alias-analysis barrier and it sat inside the hottest dispatcher on the board. filling every constant before main leaves the read a bare load, which is where the win came from. the lesson generalizes past this one technique: a memo check in a hot path can cost more than the work it skips.
8. the fold-state shelf shipped — original to kanso, and until it landed the largest unclaimed number on this page. the beat analysis rewinds the arena between iterations when it can prove the iteration keeps nothing across the line, and an accumulator used to break that proof by construction: encode_items acc xs i hands (elem_onto acc xs[i]) onward, an expression, and an expression's result is assumed to be new. the analysis declined, and then nothing rewound — including everything the iteration allocated that was not the accumulator.
the license that fixed it is identity, never type. a bytes accumulator may cross a rewind when it is the very object that arrived at the loop's entry, threaded through appends the linearity analysis proved in-place — pointer identity, established by a greatest fixpoint over groups whose every arm returns a chain of its first parameter, through conditionals, guards, local bindings, folds whose folder chains its own accumulator, and calls to other chaining groups. the header is then below the mark; growth allocates outside the arena and a mut-grow frees its predecessor, so the payload is never above the mark either; and raw bytes hold no pointers, so nothing inside the accumulator can dangle. a fresh builder of the same type has none of those properties, and a golden pins the refusal: the fresh-builder loop reads zero rewinds, and under a deliberately type-only license it reads forty with output still correct by memory-layout luck — which is why the counter is the pin and the output cannot be.
the license also reads around a cycle. mutual recursion — a loop that steps through a helper group and returns — forms a tail cluster rather than a self-tail, and a bytes slot in one crosses when every inner edge feeds it a mut-append chain rooted at one of the caller's own chain-threaded slots: the same fixpoint, one license over. a chain-threaded slot is never carried, so a bytes-accumulator cycle becomes a plain rewinding ring, and the loop-plus-helper shape a width-conscious program naturally writes needs no rewriting to earn its rewinds.
sorted views made the same move for the same reason. the first licensed run took 362 ms against a 5 ms baseline, because every rewind freed every map's cached sorted view and each iteration rebuilt them. views are malloc-backed now: nothing arena-backed can dangle, the cache registry never fills, and the sweep each rewind used to pay walks an empty table. a transient map's view leaks with its map — the recorded trade.
what it bought, measured the day it landed: encodebench's arena blocks fell from 905 to 5 across five million rewinds — the encoder runs in constant arena space. kq pretty-printing the 1.9 mb document holds 47.5 mb against jq's 30.7, down from 211.9 before this page's uniqueness campaign began, and at 5.25 instructions per cycle it out-executes jq's 5.20 — the stall the old working set cost is gone, and the footprint bought speed on its way down: 29.5 ms against jq's 101.4, with the buffer-reuse shelf and one-build literals carrying the last leg. on the plain path query the footprints sit at parity or below jq's on both documents. the pretty-print gap closed for good when the printer learned to stream (2026-07-27): each top-level element is rendered, written, and dead before the next is built, so the output never exists as one string — 30.0 mb against jq's 30.8 on the same document, and kq now holds less memory than jq on every scoreboard row.
the carry path was measured and is not the mechanism. a probe that lifted the bytes exclusion from the carry gate turned all three encode loops into carry beats with byte-identical output and a footprint four hundred and ninety times worse — the copier hands a carried builder back at exact capacity, so every carried iteration is a copy followed by a doubling, the quadratic the exclusion exists to avoid. a threaded builder must be invisible to the carry, which is what pointer identity provides: never staged, never copied at a boundary.
8. appending without a new header shipped — the largest single source of garbage on the encode board, and it was not the bytes. a byte builder's fast path writes into its own spare capacity and then allocated a fresh twenty-four-byte header to carry the new length: 42,312,800 of them, 62% of every allocation encoding made. lists already avoided the equivalent — the linearity analysis proves a push is uniquely owned and codegen emits push_mut, which extends in place. append_mut is the same mechanism for appends, and getting it to select a site took four separate blockers off the path.
the wrapper was the first. all forty-two million appends happen inside the standard library's one-line text/append, so every call site showed the analysis a user call rather than an append, and inside the wrapper the accumulator is a parameter owned by a caller one frame away. inlining a wrapper whose whole body is one builtin call passing its own parameters in order undoes the rename before anything looks. then the byte builder had to count as a fresh value at the root of the chain, a conditional's arms had to count as one use rather than two, and a fold had to read as the accumulator loop it is — the encode path threads its builder through list/fold, and the folding lambda's parameter is unique only inside a lambda whose accumulator arrives first and appears nowhere else.
encode allocations fell from 68,640,508 to 26,327,708, bytes from 2.29 gb to 934 mb. kq's resident set went from 211.9 mb to 139.9 on the same document — the first third of a fall the escape fix, the fold-state shelf and the reuse shelf carried on to 47.5, and the streaming print finished at 30.0.
writing the adversarial spec for it turned up a divergence that had nothing to do with folds. importing a qualified name enrolls a bare-named clone so both spellings dispatch, and the clone carries the original's span; nobody calls the clone by its bare name, so every call-site test over it passes by default and it keeps a linear parameter the real declaration was refused. in-place sites are keyed by source position, which the twins share. an imported builder held by two owners was written through one of them — AB AC from the interpreter, ABC ABC from native. clones are now invisible to the analysis, which is what their declaration already claimed about provenance.
9. making the rewind's cache sweep cheap two attempts, both reverted — before every rewind the runtime walks a registry of maps holding cached sorted views, asking of each whether it survives the mark. with the byte shelf enabled and loops rewinding twenty thousand times, that walk was 1,596 samples out of about 1,600 — effectively the whole program. two fixes were tried and neither survived contact.
the first scoped the walk to entries registered since the mark, on the reasoning that older ones sit below it. the reasoning was circular: an old entry is only re-registered if it was reset, and it is only reset if it was walked, so skipping it leaves a stale pointer forever. the golden corpus caught it immediately. the second deduplicated the registry with a flag on the map, which worked and bounded the registry by distinct maps rather than by rebuilds — but the extra word grew every map, and the decode board went from four arena blocks to five for no demonstrated win. reverted on the ratchet's own terms: a regression without evidence of benefit is just a regression.
a counter settled it a fortnight later, and the answer was that neither attempt was aimed at the cost. a decode never asks a map for a sorted view at all: on the 188 kb board the runtime attempts 1,254,150 view insertions with a view behind none of them, and 632,550 beat pops carry nothing. both functions were being entered to answer a single load and a branch, which divides out to eighteen instructions per write and twenty-six per pop — the frame, not the work. moving each test to its caller drops decode 0.41% and encode 0.59%, for 1,024 bytes of machine code per binary.
10. retaining the region instead of copying the carry measured, declined — the cohort pop already refuses a copy that buys too little. it sizes the survivor before evacuating and keeps the region when that survivor is worth more than half of what grew. the loop's carry sizes the same walk to allocate its buffer and then copies regardless, so handing it the cohort's ratio is four lines.
a probe said first where the copies are. on the one-shot board, 63,967 evacuations split as one cohort pop doing about 31,986 and two loop carries doing 31,981 — and the cohort's own kept-counter reads zero, so that guard ran, priced the trade and chose to copy. with the ratio ported, decode saves three copies and 112 bytes, encode saves nothing and its arena peak goes from 4 mb to 5, and one-shot does not move at all, because both its carries hold a survivor worth less than half their region. the argument that retention is self-limiting — keep the garbage, the region grows, the rewind comes back — does not survive encode, where 352 retentions in a row grew the arena by a whole block instead of converging. the cohort tests a size floor before its ratio; adding that floor gives numbers byte-identical to baseline. the guard is a loss without it and nothing with it.
the two sites ask one question about different distributions. the cohort decides once, at a pop, where the region is large and the survivor is the whole result. the carry decides every iteration, where the region is one iteration's garbage and the survivor is the accumulator. so those 63,967 copies are not waste any policy here would remove, and avoiding them without keeping the garbage takes reference counting.
8. maximal sharing measured, declined — hash-consing, the ATerm lineage. the premise was that real-world json is deeply repetitive, so counting the repetition came before building anything. on the 188 kb board: 5,513 composite subtrees, 5,210 of them distinct. only 65 appear more than once, and sharing every one of them would avoid 2,831 bytes — 1.5% of the file. the three biggest repeats are [true], [null] and [false], six bytes each. paying a hash of every subtree during decode to save that is a straight loss.
the repetition is real, and it is all in object keys: 8,361 occurrences of 500 distinct names, 94% redundant, 38,537 duplicate bytes or 20.4% of the file. string values repeat 0% — every one of the 2,114 is distinct. so the shape that would pay here is key interning, not subtree sharing, and its downstream prize is pointer-compare map lookups, which the profile bounds at about 3%. recorded so the idea is not re-mined in the general form; a workload with genuinely repeated subtrees would reopen it.
9. the cheap experiments queued — post-link layout optimization (BOLT: panchenko et al., CGO 2019) on the emitted binary, and superoptimization (souper; STOKE, schkufza, ASPLOS 2013) of the ten hottest runtime kernels. bounded, verifiable, measured in an afternoon each — though the BOLT half waits on a linux box: it is ELF-only, and the development machine's toolchain ships no llvm-bolt and no perf to feed it.
profile-guided optimization was the first of these tried, and it is measured, declined. instrumenting the decode gauntlet, replaying it, and rebuilding against the profile moves per-decode cpu about one percent — floor 1.1%, first quartile 1.4%, median 0.6% over sixty interleaved runs, which is the size of the noise on a loaded box. the decoder's dispatch is a jump table and its inner loops are already vectorized, so there is little branch-prediction headroom for a profile to recover. the cost is a two-pass build and a checked-in profile that goes stale, against a compile story that currently fits in one pass.
examined and declined, so the queue stays honest: persistent tree collections (RRB vectors, HAMTs) — the flat-array-plus-arena model beats their constants at this language's working-set sizes, and purity plus static reuse already dissolves the sharing problem they exist to solve; and gpu/polyhedral work — the wrong workload class for a parser-and-tools language.
10. carrying the sorted view across a shared put built, measured, declined — the declination above has one measured edge. static reuse dissolves the sharing problem where the analysis can prove a map is uniquely owned, and there the read-write loop is now linear: ten thousand distinct keys cost 36 ms and 4.6 mb. where it cannot — a loop that reads the old map after writing the new one — every put copies the whole pairs array, and four thousand keys cost 368 ms and 351 mb. that is the quadratic a persistent tree exists to remove, and it is the one shape where this model loses outright.
the cheap half of the fix was built: the copying put already pays an O(n) copy, so it can carry the sorted view across rather than dropping it and leaving the next read to sort one. measured back to back on four thousand shared keys, 370 ms against 368 ms — the sort it saves costs about what the extra copy adds. the bottleneck is the pairs copy, which no view trick reaches. reverted rather than shipped, because code that buys nothing still has to be read. a real fix is a different map representation, which is a memory-model question rather than an optimization, and it is not queued as one.
11. the collection surface stopped copying shipped — a method rather than a technique from a paper. the welfare index used to read three benchmark programs and a compile, which is a narrow shelf: a string build that got twenty times faster and a sort that shed two hundred times its allocations both left the number exactly where it was. what a model leaves out it weights at zero. so the index gained a basket — string accumulation, map read-write over repeating and growing key sets, list build and index read, lazy map/select/fold, group_by and tally, arithmetic, records, join, slice and sort — and then the basket was asked, repeatedly, where its allocations were.
it answered ten times, each answer only visible once the one before it was fixed. the sort was an insertion sort rebuilding its run per element: four thousand numbers cost sixty-four million allocations and peaked at 4.4 gb. merging halves took that to 327,296, and passing index ranges rather than materialising the halves took it to 20,014. drop walked every element it discarded, so reaching the last ten of four thousand allocated 8,030 times; skipping over a cursor is arithmetic, and it became 51. a fold whose reducer pushed minted a list header per element because the mark that would have let it write in place was keyed by source position and a lifted lambda carried no file — 4,015 allocations became 15. a tally copied its whole map every element, because the check refused an accumulator that appeared in a sibling argument, when a builtin forces its arguments and the read is finished before the write: 12,024 allocations and 128 kb of map headers became 25 and nothing. reading a key then writing it built the key twice, and evaluating a repeated pure interpolation once halved it. a record update allocated a fresh record where the old one was finished. a string built by joining onto itself copied itself every append, at 162 ms for a hundred thousand of them, and now writes into one buffer at 10.
the basket went from 91,864 allocations to 20,107, and from 131 mb resident to 3.5. every item left in it is about one allocation per element, which is the data being produced rather than the machinery producing it.
two things are worth keeping from how it went. the wins came from asking a local question rather than widening the shared analysis: three attempts that changed the linearity fixpoint shipped corrupted programs, and the three that asked "does every caller hand this over, and is every mention inside this one expression" did not. and the string builder, which took four attempts across the day, needed the smallest change of all of them — the seed converted where it enters the loop, so its header predates every mark and survives without anybody widening a licence to let it.
the guards are fixtures rather than arguments. two of the corruptions lowered the allocation count, so a counting pin called them improvements; what catches them asserts bytes. reuse_guard and builder_guard read a value after the write that would have clobbered it and print the same bytes on a compiler that optimises and one that does not, because in each of those cases the analysis declines.
the faster-than-rust thesis, stated honestly: not "beat rust at a microbenchmark loop"—llvm emits the same instructions for both. the winnable claim is beating idiomatic rust on allocation-heavy real workloads, because the arenas eliminate the defensive clones the borrow checker pushes people into, fusion is unconditional under purity, speculation specializes what monomorphization can't see, and the cost model schedules parallelism the user never wrote. getting that performance out of rust asks a great deal of the programmer; the aim for kanso is to ask it of the compiler instead.
12. skipping the getters a program never reads built, measured, declined — every field of every type declares an accessor arm so the name resolves wherever it is read, and a later pass deletes the ones nothing mentions. compiling the json module builds 134 of them and deletes all 134, which reads like free work to skip: ask which accessors the program mentions, and build only those.
it is slower. jsonbench goes 8.79 to 9.24 ms and the basket 6.57 to 6.92, held across three interleaved before-and-after passes, best of nine runs of twenty compiles each. the reason is the shape of what was traded. an unbuilt arm saves a struct and a vector push; deciding not to build it costs a walk of every expression in the program, and across a module that walk runs once per file over the union of all of them. the deletion pass already walks the program for the same reason, so the second walk buys nothing the first had not already paid for.
there is a version of this that could pay: compute the mention set once and let both synthesis and deletion read it. that needs the set to survive inlining and fusion, which move calls around between the two points, and the saving on the other side is 134 vector pushes. recorded here so the arithmetic does not have to be redone.
the attempt found a real defect on the way, which is the usual reason to build one. suppressing an arm per file broke a getter of a sibling module's type used as a value — list/map ps _.x, where one file declares the type and another reads the field — because a file asked about itself alone cannot see its sibling's read. the module differential caught it and the union across files fixed it, but nothing in cargo test would have.
13. calling the escape body directly instead of through a fold built, measured, declined — the json encoder escapes a string by folding over its bytes with (a b -> esc_byte a b). that lambda is an eta-expansion of esc_byte and the compiler does not reduce it, because esc_byte dispatches on a byte and a group with a byte-discriminated parameter cannot be handed out as a first-class value. the profile makes the case look easy: the lambda's wrapper is 712,277,200 instructions, 13.59% of the encode board, one indirect call for every byte of every string that needs escaping.
writing the walk by index removes the closure, the boxing and the indirect call, and it is slower. the benchmark that watches the shipped library goes 5,231,282,203 to 5,388,806,908 instructions, +3.01%. every allocation counter is byte-identical — allocs, alloc bytes, arena blocks, peak, held peak — and one counter moves: beat_iters goes 5,032,401 to 16,691,201, exactly one new beat per byte. the index walk satisfies the beat analysis where fold's inner loop did not, so every step now marks the arena and rewinds it. removing the closure saved 279 million instructions; the beat cost 436 million, and its rewind reclaims nothing, because the iteration allocates nothing above the mark.
a beat around a loop whose body allocates nothing is the general shape, and the profile it exposed did pay. k_beat_iter was 28.65% of escapebench, and eight of its thirty instructions were a sanity check on the mark's own two words — a question fixed when the mark is written and re-asked at every rewind. moving it to the push took escapebench 7.36%.
recorded here so the fold does not get rewritten again on the strength of the profile alone.
14. refusing a beat to a loop that rewinds for nothing built, measured, declined — the entry above ends on a beat that reclaimed nothing, and the obvious next move is to stop granting it. escapebench is where to look: one group, filled, owns 1,206,000 of that benchmark's 1,206,002 rewinds, and a probe in the rewind says 99.75% of them find the arena pointer exactly where the mark left it. at about twenty-two instructions each that is 26 million of escapebench's 121, spent putting back something nothing moved.
the analysis already refuses a beat to a loop that allocates nothing — PureLoop, one line — so the question was only why filled is not in that case. refusing it by hand and measuring says the shipped benchmark loses 33,263,355 instructions, 27.58%, with allocations, arena blocks and peak byte-identical. at twenty-five times the inner length, still byte-identical. at a hundred and fifty times, where the growing accumulator's superseded buffers finally exceed one arena block, the picture inverts: one block and 1,048,576 bytes with the bracket, two blocks and 3,145,744 without it.
so the bracket is doing its job and the first two measurements were reading a workload that fits in one block either way. the cost is real and the benefit is real and they show up at different input sizes, which is the whole reason a rewind exists. three attempts on the runtime side went the same way: an early exit when the arena has not moved changes what gets flushed and moves sh_buf on two benchmarks; marking the rewind always_inline does inline it under lto and costs the live encode board more than it saves the guard; and an inline guard in emitted code would pay fourteen instructions on every rewind that has work to do, which on that board is four million of them.
what survives is a note about the corpus rather than the compiler. escapebench pins this bracket's cost on every run and its benefit on none of them, so a change that deleted the bracket would have read as a 27.6% win with every memory counter flat — the failure its own README was written to prevent, one level in. whether to raise its size is a live question and not a free one: the row is a welfare term and a bigger benchmark is a slower job.
15. skipping the clean run in front of the first escape built four ways, measured, declined — a third pass at the same board, and the plainest of the three. the encoder scans a string once for a byte that needs escaping and, finding one, folds over every byte of the string through the lambda item 13 could not remove — 712,277,200 instructions, and 13.83% of the encode board on the sitting after §01's pair table, which is the largest single line on it that is not the walker. the scan already knows where the first escape is. the bytes in front of it were proved clean by that scan and could leave in one copy.
the distribution says the run is worth having. of the 10,475 strings on the board 1,773 carry an escape — 16.93%, which is the profile's 709,201 fold entries over 4,190,000 encodes to the digit — and those strings average 16.44 bytes with the first escape at byte 4.34. so 26.38% of every byte the fold walks is a run already known to need nothing.
it works, and the shipped library's benchmark falls 3.08%: 5,208,664,610 instructions to 5,048,017,249, with the frozen twin byte-identical and both checksums unchanged. the objective still declines it, and the reason is the same in all four shapes it was written in:
livebench oneshot compile row welfare
two helpers, guarded run -3.0621% -1.4450% +390,025 -0.02
one helper, guarded run -3.0621% -1.4450% +268,455 -0.01
no helper, no guard -2.7457% -1.2955% +172,457 -0.00
no helper, guarded run -3.0842% -1.4555% +213,359 -0.01
read down the last two columns and the exchange rate is legible. the guarded run costs 40,902 more compile instructions than the unguarded one and buys 17,635,200 more runtime instructions, a ratio of 431 to 1, and the index prefers the compile side. the reason is the two satiations: compile cost satiates at 0.5 and this term sits near its baseline, where the curve is steep; the live encode row satiates at 2.0 and entered the corpus the same morning as a granted baseline, at its dimension's standing, where r is already near six and a three per cent fall moves the term score by 0.006. a library change that buys runtime by growing the library charges the least satiated term in the model to feed one of the most satiated.
that is a statement about the weights rather than about the change, and it was filed as one. it is closed: the gavel of 2026-09-06 put the objective's run side on one consolidated program, so nothing enters the corpus at its dimension's standing any more and the granted-baseline rule is retired. the welfare column above is what the objective answered on 2026-09-05, under a model that weighed the live encode row; today's weighs one program that includes the shipped library's encode path at its measured share. the code is reverted; the numbers are here so the next reading of the same profile finds them.
16. an escaped string read without a view of its bytes built, measured, declined — the encoder turns every string it writes into bytes, text/bytes s, only so find2_below can scan it for a byte that needs escaping, and on runbench that is 942,750 views of 32 bytes each. an experiment let find2_below read a string directly and skipped the view in the clean case, which is almost every string. allocations on the run program fell from 5,698,908 to 4,915,728 and the shared bytes from 41,290,272 to 22,493,952, but the arena's peak held at 5,050,064 bytes, so the views never set it. what is left is an instruction saving of about one per cent, and it needs find2_below to accept a string, which widens a public function in std/text. that is a change to the language's surface, and one per cent does not justify it.
17. compiling each shipped module once per process built, measured, declined — an entry that imports std/text through several libraries compiles std/text once per import, five times on the entry corpus. caching the compiled module by path removed the repeats and raised what the front end holds from 768,704 bytes to 1,030,180, because the cache keeps every module's tree alive for the whole compile. the development side weighs peak memory, and that rise outweighed the saved instructions. presizing the front end's hash maps to skip their rehashes failed the same way, at +1.09% on the peak.
7the lazy betv1 shipped
the newest ruling is the largest since the arenas: kanso is lazy. write foo = expensive_calculation and nothing runs; the computation is noted, and it happens when—if—the value is actually used. a value behind a conditional that turns out false is never computed at all. the motive is the one that runs through this whole page: don't pay for work that never reaches the answer.
the bookmark
the mechanism is a thunk: a small cell holding what to compute and the ingredients to compute it with—a bookmark into a computation the program may never open. forcing the thunk runs the computation once, writes the result into the cell, and drops the ingredients; every later reader finds the answer already there. that one cell buys three things. work that might not be needed costs nothing until it is. a shared computation runs once no matter how many readers demand it—okasaki's persistent data structures rest on exactly this, laziness plus memoization making amortized bounds survive sharing. and an infinite structure is ordinary data, because only the consumed part ever exists.
what the compiler proves, and what stays lazy
the compiler sorts every value by what it can prove about demand. provably never used: deleted, as dead code. provably used, immediately, and cheap: compiled strict—the exact code this page has been describing—because a thunk forced nanoseconds after it's built pays its cost and dodges nothing. everything else keeps its thunk, and that residue is where laziness earns its keep. one refinement came late and reshaped the runtime's plans: proven demand is a fact about whether, and says nothing about when. a value certainly needed but not for a while is guaranteed-useful work the scheduler may run early. that idea gets its own subsection below.
the two leaks, named
laziness has a reputation, and it's earned: haskell programmers debug space leaks. the failure has two shapes. chains: a loop that only ever promises additions builds a million bookmark cells before anything forces the first—linear memory where constant was wanted. retention: a bookmark's ingredients can include a large structure, which stays pinned until the bookmark is forced. the same rule bounds both—force what is provably demanded, eagerly—and neither is taken on faith: the prototype below measures the chain at its worst and watches the pinned structure release at the moment of forcing.
why the workbench can't hold a thunk
a rewinding arena is a region: everything in it dies together when the beat ends. a thunk breaks that bargain, because when a thunk runs is a runtime fact—a pending computation parked on the workbench could be forced long after its beat, and the sweep would tear the ingredients out from under it. this is a known dead end, with receipts. jhc, the one haskell compiler that made regions its whole memory strategy, documents its own programs as leaking. ghc's compact regions refuse to hold a thunk at all. even the ml kit, doing region inference in a strict language, had to bolt a garbage collector back on (hallenberg, elsman, tofte, PLDI 2002). regions and laziness don't compose, so kanso doesn't ask them to. the workbench keeps the strict data it already serves; thunks get a tier of their own.
counting without an asterisk
the thunk tier is reference counted: a cell whose frame can prove no reference survives it is freed at that frame's boundary and recycled through the free list shipped. the proof is a static classification — every use of the binding must target a callee position whose arms only force the value, ignore it, or return it bare — plus one runtime pointer compare for the returned-thunk case. cells the proof can't cover (returned upward, handed onward in a tail call, or reaching a position that could store them) stay live and are counted: the .mem goldens pin allocs, frees, and escapes per program, so reclamation is a diffed fact rather than a belief. the escape cases' full story belongs to the defunctionalized-thunk frontier, where ownership can ride the calling convention. counting normally comes with an asterisk—a cycle keeps its own count above zero forever, which is why python ships a cycle collector on top of its counts and why trial-deletion collectors exist (bacon & rajan, ECOOP 2001). kanso deletes the asterisk structurally. values are immutable, so ordinary data can't point at itself. the one construct that could—the knot-tied stream, ones = 1:ones, a cell whose tail is the cell itself—is banned in favor of the generator form, which allocates a fresh cell per step and never loops back. the prototype builds both: the knot leaks exactly one cell, on cue, and the generator version of the same infinite stream runs in two. the graph stays acyclic, so counting is complete. no tracing collector, no cycle detector, no pause.
the card comes back off the bench
readers of §02 will recognize this tier. perceus—counting with reuse—was demoted for kanso's strict data because the arena reclaims the same memory with one pointer reset. but strict data was never its best fit. thunk cells are small, uniform, and constantly churning, the exact population a free list serves: a cell whose count hits zero is handed to the next allocation, and the allocator stops hearing about it. the closest published work is first-order laziness (lorenzen, leijen, swierstra, lindley, ICFP 2025, distinguished paper), which grafts perceus-style counting and reuse onto a lazy fragment—with every lazy shape declared up front, and with open, library-extensible thunk shapes named in print as the unsolved problem. that problem is a fact about separate compilation. kanso compiles whole programs—no separate compilation, no runtime loading—so every thunk shape in a program is enumerable at compile time, and the compiler defunctionalizes all of them itself (reynolds' old trick, applied totally). the fragment the literature stops at is, in a closed world, the whole language.
the prototype's receipts
the cell design ran as an instrumented prototype before any engine work, every allocation and free counted. across every workload: 21.1 million cells allocated, exactly one live at exit—the deliberately leaked knot. on the shipped engine that workload—100k items, 5% used—runs 18x faster than a strict build of the same program, and within 38% of the rust programmer who restructured the code by hand. a thunk costs about 23 ns to build and force—the price the strictness analysis erases wherever demand is provable. and with the free list and defunctionalized cells, a 10-million-element infinite stream runs at 10.9 ns per element with two calls to the allocator, total: the first cell's memory becomes the third's, the third's becomes the fifth's, forever.
proven work on idle cores
a pool of pending thunks is a work queue. when a program blocks on io, the core it was using goes idle—and the scheduler can spend that idle time forcing thunks whose demand is already proven. the work is certainly needed; only its timing moves. an out-of-order cpu does the same with its instruction window during a memory stall, and this is that move, lifted to the language. the honest ancestors are optimistic evaluation (ennals & peyton jones, ICFP 2003) and eager haskell (maessen, 2002): both showed real speedups from running lazy code early, and both stayed research prototypes, because in an impure language a wrong guess needs rollback machinery and the rollback tax ate the winnings. kanso's functions are pure and effects are inert values, so forcing a thunk early cannot do anything that needs undoing—a mis-timed forcing produces a value nobody reads yet, and an err discovered early is stored in the cell, surfacing only if the value is ever demanded. purity removes the problem those systems died on, and the deterministic scheduler keeps early forcing invisible: every engine produces the same bytes whatever got computed during the stalls.
what's real today
and the lazy realm has a scoreboard of its own now. the workload: 100,000 items, each carrying a real computation (five thousand mixing steps), one in twenty actually used, every row producing the identical checksum. shorter is faster:
| program | seconds | why | |
|---|---|---|---|
| rust, hand-tuned | 0.08 | the human moved the work | |
| kanso, as written | 0.12 | the compiler moved it | |
| rust, as written | 1.73 | computes what you wrote | |
kanso --strict |
2.15 | the measured worst case |
read the rows top to bottom and the design argument reads itself. the tuned rust row is a human who restructured the computation into the branch; the kanso row is the same program in its clear form, within a whisker of that, because the demand analysis thunked the binding and the branch discarded nineteen of every twenty—the counters on the run say it plainly: thunk_allocs=100000, thunk_evals=5000, 95% of the work never born. rust-as-written pays for everything it wrote. and the bottom row is kanso's own --strict flag—the worst-case measurement mode—which forces every thunk and reports what this program would cost if laziness never saved a cycle: the bound you'd quote in a latency budget. the sources are in bench/; the checksums match on every row.
the honesty tier, in the house convention: v1 shipped in both engines; the pervasive form staged behind the same experimental gate. the fragment above—conditional-demand bindings, the cost gate, refcounted cells on a free list, forcing at scrutiny—runs in the native compiler and the interpreter today, held byte-identical by a new vein of golden tests. alongside every program's expected output, a .mem file pins its memory facts, and both engines must match it byte for byte, because evaluation counts are semantics, not implementation detail. the shipped receipts: the skipped computation evaluates zero times with unchanged output; the shared thunk evaluates once under two readers; the demanded thunk forces once and exits with zero cells live; and a skipped binding whose computation would have produced an err simply never births it—same output, because errs are values and effects are inert, so the classic lazy-versus-strict divergence has nothing to bite. the json gauntlet is untouched: its accumulators hit the cost gate and compile strict, and the §08 cost golden pins that at zero thunks and the same allocation count the ratchet holds. leak-freedom stops being a property we believe and becomes a constant the build diffs.
two dials came out of the same conversation, both awaiting the surface-syntax rulings: a strict mode (force everything, measure the worst case—a measurement tool, since forcing runs what laziness would skip) and a sync-style block (a scope guaranteed thunk-free, so peak memory inside equals strict memory). one gate in the compiler implements both.
the work-ahead engine design
the newest ruling on this frontier is scheduling, and it needs one picture. a fiber that blocks on io leaves a hole in the timeline. demand-driven evaluation cannot fill it — a blocked fiber demands nothing — so the hole is dead time, and the work it was saving up runs after the wait instead of during it. the strictness analyzer already knows which computations are certain to run. so the scheduler keeps a pool of them, and when a fiber parks, it works ahead:
what fills the hole is chosen by three rules, checked in order. that is the entire policy:
a fiber parks on io. the scheduler asks, in order:
1 proven work, inputs ready? -> run it. certain to be needed —
| early is free. program order.
no
v
2 a gate of proven work? -> run it. one branch decides its
| fate, and finishing it refills
no rung 1.
v
3 a free gamble, priced cheap? -> maybe. only here is anything
| speculative, and only when the
no rungs above are empty.
v
sleep until the io returns -> the driver actually rests.
determinism survives because the ladder keys off logical scheduler state, never the wall clock: same program, same seed, same work-ahead transcript, on every machine. purity is what makes the whole thing legal — evaluating early is invisible, because nothing can observe the order of effect-free work.
wall-credit: the substrate, shipped shipped
the engine's foundation is already live in both engines, and it has receipts. the deterministic scheduler used to track logical time only, so real time a fiber spent computing never counted against another fiber's pending sleep — a thread could grind for a second beside a sleeper and the sleeper still slept its full span. the scheduler now credits elapsed wall time against every deadline before it waits. the transcript and every counter stay purely logical, so replay and the goldens are untouched; only the physical wait shrinks.
# as-written sleep >> report min 2913 ms (stall, then work)
# overlapped report beside sleep min 2008 ms (work inside the stall)
# the sleep is 2000 ms. overlapped lands at 2008: the entire grind
# disappeared into the wait. min of 15 interleaved heats; the floor is
# physics — a run can never beat its own sleep, and none did.
cargo build --release
./target/release/kanso build bench/workahead/aswritten.kso --release
./target/release/kanso build bench/workahead/overlapped.kso --release
time ./aswritten; time ./overlapped
KANSO_SCHED_DEBUG=1 ./overlapped # watch elapsed + wait = deadline
today the overlap is written by hand — the two variants above are the same program arranged two ways. the engine's remaining job is exactly the gap between those rows: find the arrangement automatically, using the two-group heuristic, so as-written code gets the overlapped number. that closes the loop this section opened.
techniquesthe ledger
every named technique in the compiler, one line each: the idea, its source, and what kanso gets from it. queue items graduate here as they ship; the list only grows.
- rewinding arena allocation — the region lineage (tofte & talpin, 1997), reshaped around loop iterations. allocation is a pointer bump, freeing is a pointer reset, no per-object headers anywhere.
- compiler-proven beat boundaries — original to kanso. an analysis proves a loop's iterations keep nothing across the line, so the arena rewinds between iterations without a runtime check.
- carry evacuation — original to kanso. the values an iteration does keep are copied across the rewind, deep-copy bounded by a static slot count. deciding which values those are is bounded by nothing: the survivorship walk reads a node’s whole interior, so a loop carrying a list that grows asks a question that grows with it. section 33 has the measurement.
- static reuse analysis — the perceus reuse idea (reinking, xie, de moura, leijen, PLDI 2021) without its counting. provably unique values mutate in place; push and append claim their buffer's frontier.
- ownership signatures — whole-program borrow-vs-consume inference in the koka/fip lineage, printable per function, so ownership is a pinnable contract rather than an ambient verdict.
- strictness analysis over demand — the mycroft/ghc lineage. lazy semantics, strict compilation wherever demand is provable; the json gauntlet compiles to zero thunks, golden-pinned.
- whole-program force elision — closed-world escape analysis over thunk reachability; 133 spurious force sites deleted in one pass, the serde lead restored.
- defunctionalized thunk evaluation — reynolds (1972). a pending computation is a site id plus captured arguments dispatched through one generated evaluator, never a closure.
- refcounted thunk cells on a free list — counting scoped to the one dynamic lifetime in the language: frame-proven releases with a returned-thunk alias guard; escapes counted, never guessed. values never carry counts.
- two channels with opposite obligations — original to kanso:
noneis a value anderris the failure, and the rules run in opposite directions. anerrpropagates by itself and a function that accepts one must return one, so a failure cannot be absorbed quietly. anonepropagates nowhere, lives in a record field but never in a list or a map, and reaching an operation with no arm for it is an error. a lookup's “not found” therefore means one thing that nothing else can forge. - cohort counting for cyclic graphs — original to kanso (the birthday theorem): cycles cannot cross build-block birthdays, so one count frees a whole strongly-connected cohort. build blocks ship on all three engines. the cohort-freeing arena is measured and deferred: discarding two hundred thousand cycles in a loop holds at one arena block, because the beat rewind already reclaims each cohort — what is left needs reference counting to detect the drop, so it lands with the counted world rather than before it.
- fold fusion — deforestation (wadler, 1988) by way of fold/build (gill, launchbury, peyton jones, 1993). adapter chains collapse to a single scan with a composed reducer; lazy pipelines run faster than their eager ancestors.
- dispatch as jump tables — literal-byte arms compile to one indexed jump; the decode scanner and the escape table are switchboards, never if-ladders.
- whole-program monomorphization — one specialized copy of the code per concrete value shape that reaches it, type checks deleted rather than predicted.
- simd substring scanning — the serde/simdjson family trick: sixteen-byte wide scans behind
find2andfind2_below, with the movemask-free NEON reduction (the shrn-by-4 narrowing) on apple silicon. - proven-clean fast paths — one vector scan proves a string needs no unescaping (decode) or no escaping (encode); the common case is a single copy.
- the byte builder — a capacity-headed accumulator with frontier claiming, making a fold of appends amortized linear under plain value semantics.
- the zero-copy finish — a builder-owned buffer becomes the output string in place, terminator written into spare capacity and the frontier burned so no later append can write under it; the whole-output copy disappears (seventy-five megabytes per four hundred rounds on the encode golden, pinned).
- hot-accessor inline twins — the list-length header load inlines into every fold loop the way the tag tests already do; the runtime call survives only for the map and string cases that need it.
- inline twins — hand-emitted alwaysinline ir shims for the runtime's one-line helpers, recovering the cross-module inlining lto declines. the newest is the string literal's slot: a literal is built once into permanent storage and handed back thereafter, but the handing back was a call across the module line, 2,500,000 of them on the encode board alone. reading the slot's tag inline and calling the runtime only on the first miss drops twelve of the thirteen benchmark rows, widebench by 1.85% and three more by about one per cent, for 64 to 2,304 bytes of machine code apiece.
- always-inline bump allocation — the allocator's fast path folds into every caller; only the refill stays out of line.
- shortest-round-trip float rendering — ryū (adams, PLDI 2018): generated 125-bit tables, the half-ulp interval walk, fifty-million-double fuzz at zero failures; dtoa left the building. dragonbox (jeon, 2020) ranked rather than queued: on the encode profile of 2026-09-02
render_ryuis 7.75%, and 8.09% on the sitting of 2026-09-05; 85% of it is the digit core dragonbox would replace, so its usual margin is worth something near two percent — real, and behind the lines above it. (this line read 3.8% for a fortnight after the appends around it were inlined out from under it; §06 carries the sitting.) - small-integer fast paths — arbitrary precision that costs nothing until used: word-sized arithmetic with overflow checks, the bignum path behind them.
- register-return abi — two-word values pass in registers across the tailcc boundary; the canonical destructure costs no memory traffic.
- a guaranteed tail call between arms — llvm's
tailccwithmusttail, so handing off to a different function is a jump rather than a frame and a walk that dispatches across arms runs in constant stack. the convention is miscompiled on arm64 above -O0 once a call spills past the argument registers, so a release build keeps it exactly where the arguments fit x0–x7: the micro corpus is clean at a cap of eight and one sample breaks at nine, which is the register file itself. written for a program that ran out of stack rebuilding a record at depth, it also read 2.96% off decode. - self-tail calls compile to loops — a tail-recursive walk runs in constant stack, measured to twenty million deep. non-tail recursion still ends the stack, and the two engines end it at different depths, so a bound both engines share is queued rather than shipped; today the driver names the cause instead of exiting silently.
- the deterministic green-thread scheduler — cooperative description execution with a logical clock; transcripts are pure functions of the program.
- wall-credit sleep accounting — real compute time credits against pending sleeps, so wall-clock honesty never leaks into the logical transcript.
- work-ahead scheduling — original to kanso (designed): io stalls fill with proven-needed work first, speculative work only when the proven queue is dry.
- a worklist keyed to what reads what — the classical dataflow worklist (kildall, 1973) turned on the whole-program fixpoint's last blind spot. a function's answer changing has always woken only the functions that read it; a declared type's field growing woke everything, because nothing recorded who could care. the readers are static — a field set is read only where a constructor pattern destructures it — so an index built before the first round replaced the sweep, and checking the json library does 17,786 expression visits where it did 23,224.
- the chain floor — original to kanso. a bind chain is one bracketed loop, so a continuation that captures a large value would ride the carry buffers on every step; past a size threshold the bracket pops and re-marks instead, flooring the region under the survivor so later steps share it. it took kq's full print from 47.5 to 30.0 MB. the threshold is a real line rather than a formality: a document one row longer than the wide shelf's sixteen thousand elements is exempted and copies 3,648 bytes where the shelf copies a megabyte, which is why the shelf is sized to sit below it and keep the staging path measured. lowering the line was built, measured — the megabyte goes, seven shelves hold, peak does not move — and declined, because it would move the shelf out of the band it exists to watch.
- cost goldens — the performance ratchet: allocation counts, arena blocks, and beat iterations pinned as ci-diffed constants, so regressions are diffs rather than vibes.
- retired instructions as the third instrument — the work a process did, reproducible to a few tenths of a percent and immune to whatever else the box is doing. read beside cycles it separates how much work a program does from how well that work runs, which no clock can tell apart: it is what showed kq 3.65× ahead of jq in work but only 3.11× in cycles — a stall the fold-state shelf then bought back, to 3.45× in work and 3.63× in cycles: the same instrument, reporting the gap first and then its closing.
- memory-fact goldens — the .mem vein: evaluation and liveness counts pinned byte-identical across all three engines, because evaluation counts are semantics.
- differential engine testing — mckeeman-style differential oracles: one reference interpreter, two compiled backends, every golden byte-identical across all three.
- float parsing at memory speed — eisel–lemire (lemire, 2021): the decode mirror of the rendering entry, verified against an independently written reference over thirty million doubles and pinned by the
el_parsescounter, which reads 318450 on the decode board. - vectorized utf-8 validation — keiser & lemire (2021) in full: nibble-lookup classification with ascii-block skip, spec-strict, verified over seventy million boundary-crossing sequences.
- accumulating recursion as a loop — tail recursion modulo an associative operator, the ghc lineage of accumulator transformations. a group whose leftover work adds or multiplies something the pass can prove is a whole number carries an accumulator and descends in tail position, so
n * fact (n - 1)runs in flat memory on all three engines. floats stay out: their addition is not associative, so reassociating one changes the answer. - in-walk cycle guards — the path stack render carries, priced against a build with the scan removed. at depth 10 to 30 it sits inside the layout noise; at depth 1,000 it is 47% of render, and the linear scan crosses the cost of a whole pre-pass at about depth 1,400. json documents nest tens deep, so the guard stays as it is; a general cycle-safe walk wants a generation-stamped set with real removal instead.
- branchless map search — eytzinger layout (khuong & morin, 2017): built, measured, declined — string-keyed maps at json sizes don't reward it; the negative result is recorded so it stays declined.
- tail recursion modulo context — leijen & lorenzen (ICFP 2023). queued: results built in their final cells, no accumulator discipline at the source level.
- fully-in-place discipline — lorenzen, leijen, swierstra (ICFP 2023). queued: in-place-ness as a checked guarantee instead of a discovered optimization.
- call-pattern specialization — SpecConstr (peyton jones, ICFP 2007). queued, and reclassified: the enumerable's typed arms are already written by hand and already fast, so this buys source rather than speed — twenty-seven arms across three parallel sets (
fold,iter,next) that must each gain an entry when an adapter is added. a parity spec now fails when they disagree, which is the cheap half of the same guarantee. - maximal sharing — hash-consing, the ATerm lineage: built no further than the measurement, and declined. of 5,210 distinct subtrees on the decode board only 65 repeat, worth 1.5% of the file; the redundancy is all in object keys (94%), which is key interning rather than subtree sharing.
- post-link layout and superoptimization — BOLT (panchenko et al., CGO 2019); souper/STOKE (schkufza, ASPLOS 2013). queued as bounded experiments on the emitted binary and the ten hottest kernels.
- generalized ahead-of-time speculation — the moonshot tier: bail-and-restart specialization licensed by purity, jit tricks with zero deopt metadata.
- cost-bound inference — the RAML lineage (hoffmann et al.): polynomial cost bounds as types, feeding both the check verb and the scheduler.
- equality saturation — e-graphs (egg: willsey et al., POPL 2021): the ir rewritten by saturation instead of ordered passes, phase-ordering dissolved.
- simd structural scanning — simdjson (langdale & lemire, VLDB 2019): whole-document byte classification feeding the switchboards, the frontier past serde parity.
8the receipts
none of the numbers on this page ask to be believed. here is the scoreboard and the recipe. every engine reads the file at runtime—an embedded constant would let llvm fold the decode away—and the kanso harness accumulates a checksum so the loop provably runs.
# kanso, shipped default ~0.87 ms 4.2 mb rss, flat (arenas, §03)
# serde_json (hand-tuned) ~0.90 ms 6.8 mb rss
# reasonably-written rust ~1.04 ms 6.9 mb rss
# go encoding/json ~2.05 ms 11.8 mb rss
cargo build --release
./target/release/kanso run bench/make_jsonbench
./target/release/kanso build bench/jsonbench --release
time ./jsonbench # less ~3 ms startup, over 150
(cd bench/serde_bench && cargo build --release)
./bench/serde_bench/target/release/serde_bench bench/large.json
(cd bench/naive_json && cargo build --release)
./bench/naive_json/target/release/naive_json bench/large.json
the ratchet
the floors above are a dated sitting and the decoder has moved a long way since. a wall-clock floor is only worth publishing off an idle machine, and randomised-layout timing puts the spread WITHIN a single tree at about three per cent — larger than most single changes — so these rows are re-sat at a release and not on demand. what has happened since 2026-08-07 is on the record in the vein that can be read any day, and §38 and §39 are only the most recent of it: those two alone take 17.14% off what a decode retires. every row above is therefore a ceiling on its floor, stale in the direction that flatters nobody, and the ratios between the lanes are what that sitting was for.
a benchmark that drifts from one machine or one afternoon to the next isn't evidence of much. so the wins are pinned to counts that can't drift, not to wall-clock time. the 150-decode gauntlet is a deterministic program, so it performs the exact same number of allocations every run, on every machine, on both architectures. ci commits those counts to a golden file and diffs against it: 1,390,965 allocations, 2 arena blocks, 151 beat iterations—bit-identical on arm64 and x86_64. that arena count across 150 decodes is the flat-memory guarantee written as a constant; 151 beat iterations is one rewind per decode plus one for the program itself, on the record. a change that nudges any of those numbers fails the build before anyone has to squint at a flaky stopwatch, which is how the three ladder rungs stay put and how the next regression announces itself as a failing diff in the pull request that caused it.
a vein only watches the dimension it counts, and that is a thing to keep checking rather than assume. the compiler's own allocation traffic went unwatched until 2026-08-24, when a pass was rewritten to borrow the program's names instead of owning them — a quarter off its time — and every gate in the tree reported nothing. rounds and visits could not see it, because the compiler decided exactly the same work either way. peak could not see it, because the strings it stopped allocating were transient, gone long before the high-water mark. three gates blind to one shape is the argument for a fourth: the count comes off the counting allocator that has printed it all along, and it sits in a golden of its own now. what proves it is not a duplicate is the ratchet — put the pass back to owning its names and the other three stay green while this one reds. the allocator prints a byte total beside the count, and that one is deliberately left unpinned: it came out 26 bytes apart on two machines, which is twice the difference in the length of the directory the compiler ran in. two allocations hold the absolute path, so a golden row there would have been measuring the clone.
a pass that walks the same declarations twice is not therefore redundant built, measured, declined — the front end checks each dependency as it compiles it and then checks the merged program, and the merged program holds every declaration the dependencies brought. removing the per-dependency pass would have cut a third off the fixpoint rounds and an eighth off the allocations. it also silently stopped reporting a mistake inside a library reached through another library. the reason is that the rewriting passes run inside the dependency's own compile, before it returns: by the time the entry merges, a field read has been desugared into a shape the check that looks for it can no longer see. guard the rewrite and the diagnostic comes back. so a check that reads a syntactic shape can only fire in the compile where that shape still exists, and counting declarations says nothing about whether it can — which was the entire argument for the removal. the measurement is kept here, and the fixture that catches it is in the corpus, so the idea does not get re-mined.
that vein belongs to the rustc that built the binary, because a good part of the count is the standard library's: a hash map's growth schedule, a vector's doubling, a string's spare. so the file names its toolchain, and the gate refuses to read the rows anywhere else — the third vein to do that, through one script. a toolchain bump moves every row and none has regressed, which is a regeneration and a sentence in the log rather than a puzzle.
an interned symbol in the syntax tree measured, declined — a program uses around four hundred distinct names and the build 622, against roughly eleven thousand string clones, so replacing every name in the tree with a number into one table is the obvious move. what stopped it was the size of the conversion, counted rather than guessed: make the name type opaque — a newtype with a constructor and no way to read the text — change one field of the syntax tree to it, and rustc reports every site the table would have to reach. 365 of them, for one field of twenty-nine, and that is a lower bound because errors hide each other. the welfare index cannot see the win either: its compile terms count what the compiler decides — fixpoint rounds, expression visits, lines emitted — and the one term an interner touches is residency, which a table that holds every name for the whole compile would push the wrong way. the win is real and lives in wall time, which the index leaves out on purpose. what survives is smaller: two questions do the genuine text work, “is this qualified?” and “make a qualified name”, forty-seven sites between them, and a name carrying its module and its base answers both structurally for a fraction of the conversion.
the ratchet turned on itself, twice shipped — every gate here claims to catch something, and a nightly job applies each claimed defect to check the gate really goes red. two holes in that. the error corpus — 164 fixtures pinning byte for byte every diagnostic the language emits — had no row at all, so nothing in the tree had ever watched it fail. and one row's mutation had quietly stopped applying: it patches an exact line of the demand analyser that a later change rewrote, so the patch matched nothing, gave up, and the row reported green for a reason unrelated to its gate working. the nightly caught that one on schedule and named it precisely, to nobody, and it sat red for fifteen hours. the answer was not more watching. applying a mutation costs a text substitution where proving it reddens a gate costs a build, so the cheap half now runs on every change — twenty-nine mutations in seven seconds — and the change that breaks an anchor is the one that has to answer for it.
an arm cannot see an err its own hako raised retired 2026-09-15 — superseded by the ruling that a bare err is data: an (err …) arm matches it wherever it is written, and the foreign-only licence lives at the rescue word instead. the record of the build as it shipped stays as written. on all three engines, as dispatch rather than as a warning. an err records the package that raised it beside the trace line it already carried; native and wasm put both halves in one literal per raise site and split it in the runtime, because nineteen runtime signatures carry an origin and threading a second argument through all of them to move a package name was not worth it. the guard is emitted only where a pattern can hold an err at all, so nothing else pays for it.
the build turned up a question nobody had ruled: what a package IS. the old answer made all of std one, which was invisible until an err's raiser became part of dispatch and then wrong immediately — std/testing could not rescue a failure std/json raised, so the test harness could not report a test failure. the ruling is go's: a package is a directory, and its import path names it. that holds for a program's own modules too, and it is what makes the rule teachable instead of merely enforceable — a decoder module and the module that reports its failures are two packages, so the reporting arm sits exactly where a reader would put it.
the cost is the honest measure of the change, so: a package's own failure is now completely opaque to that package. a bare binder already refused failures and the two err-admitting patterns now refuse own-origin ones, which means you cannot hand your own failure to your own function at all. lib/json lost must, whose whole documented job was converting json's own parse failure — the caller writes it now, one package over, where that failure is foreign. nine corpus fixtures and four book samples moved with it. the front end got cheaper for the deletion: 16,818 expression visits to 16,806, 62,110 allocations to 61,974.
three rulings landed beside that work, recorded here because the log carries them and this page is what a reader is shown. welfare measures what compiling costs rather than what it counts. its compile-speed terms were fixpoint rounds, expression visits and lines emitted — counts of what the compiler decided to do — so a front end that got a quarter faster moved none of them. the measured veins, instructions retired and allocator traffic, are the terms now; the counts stay pinned as tripwires. built below. an arm never sees an err born in its own hako, and that was never advisory — the doctrine is dispatch semantics, so at match time such an err skips the arm and carries onward. built below. and the welfare floor is absolute against work that leaves the language alone, permeable to work that changes it: performance and implementation work may trade one dimension against another so long as the aggregate rises, but a language feature is never rejected for the score — it lands, the floor moves, and the fall is recorded against the change that spent it. packaging a compensating optimization into a feature's own change to hide its cost is gaming the index in either direction.
welfare reads the measured veins shipped — compile speed is instructions retired and allocator traffic, both counted, where it was three numbers the compiler kept about itself. the counts keep their goldens and still turn ci red when they move; they stopped being what the score is made of. the weight and the satiation did not change with the instrument, and that was deliberate: 0.28 and 0.5 were priced for the dimension — what teams shipping software pay for, 45% of the people who left rust naming compile times among their reasons — and the dimension had not moved. only the instrument improved, and an instrument is not a preference. the weight has since moved for a different reason and by a ruling: it is 0.32 now, funded from run memory. see below. the score falls 0.65 for the swap, because a term with no history enters where its dimension already stands, so compile speed gives up the credit its proxies had banked. nothing got slower, and scores either side of that entry do not compare.
the same change fixed what the goldens were pointed at. kanso check lib/json compiles lib/json's test file too, so every dependency the suite imports was being charged to the library — the row moved when a test changed its imports, which is a question nobody asked. the gates stage the library without its tests now, from the fixed path the instruction count already needed, and all three compile veins answer for the same program: 62,110 allocations where the golden said 65,543, 16,818 expression visits where it said 17,886, forty fixpoint rounds where it said forty-two.
a merged err came out of a hop nested on native and flat everywhere else shipped — three failures answer three reasons however they were grouped, which is the rule written in a comment directly above the struct that broke it. a failure passing through a function with no err arm only hops, and the hop rebuilt the err box setting four of its five fields; the arena bumps and does not zero, so the fifth read as whatever was there and the reason list stopped being a list of reasons. reading the rest of the family found the same shape three more times, including an evacuation copy that dropped the cause chain as well. one constructor takes all five now. the browser engine was never wrong — it calls the interpreter's own hop — so this was one engine against two, and it surfaced only because somebody went looking: no fixture in the corpus merged a failure and then hopped it. every binary grows forty-eight bytes for the repair.
three rulings landed beside that work, recorded here because the log carries them and this page is what a reader is shown. welfare will measure what compiling costs rather than what it counts. its compile-speed terms today are fixpoint rounds, expression visits and lines emitted — counts of what the compiler decided to do — so a front end that got a quarter faster moved none of them. the measured veins, instructions retired and allocator traffic, become the terms; the counts stay pinned as tripwires. an arm never sees an err born in its own hako, and that was never advisory — the doctrine is dispatch semantics, so at match time such an err skips the arm and carries onward, where today the program compiles with a warning and runs. and the welfare floor is absolute against work that leaves the language alone, permeable to work that changes it: performance and implementation work may trade one dimension against another so long as the aggregate rises, but a language feature is never rejected for the score — it lands, the floor moves, and the fall is recorded against the change that spent it. packaging a compensating optimization into a feature's own change to hide its cost is gaming the index in either direction.
how fast it compiles
the front end is quick enough to sit inside an editor's save loop. kanso check—parse, whole-program inference, and every diagnostic—finishes kq in 6.6 ms and the json decoder in 6.1 ms. each figure covers the standard-library modules the program imports as well as its own source, about a thousand lines of kanso in kq's case, so the front end clears a hundred and fifty thousand lines a second. an unoptimized binary takes 116 ms end to end; a fully optimized one takes 635 ms, and llvm at -O2 is nearly all of that. for scale, go on the same box builds a 28-line program against its already-cached standard library in 98 ms. (2026-07-25, loaded desktop, best of seven.) these are a dated sitting and they predate the work below: a quiet machine is a condition rather than an effort, so the clock figures are re-sat at a release and not on demand. what changed since is pinned in counters instead, which is the point of the paragraphs that follow.
the clock is the softer of the two measures. bench/compile_golden.txt pins the work itself—fixpoint rounds and expression visits per sample, beside the lines, calls and branches the emitter wrote—so a change that grinds a longer fixpoint to emit the same text stays visible even when machine noise hides it from the wall clock.
the pinned work is what caught the next one. inference is a fixpoint with a work list: a function's answer changing wakes the functions that read it, and nobody else. one line was defeating that. when a declared type's field grew—rect w h learning that w can be an int—inference had no record of who could care, so it woke the whole program. checking the json library ran four full sweeps of four hundred and seven functions, the last of them so that seven could move. the readers are static: a field set is read in exactly one place, the arm that destructures a constructor, so the functions that can care are the ones whose patterns name that type. an index built once before the first round replaced the sweep, and a change now reaches its readers in the round it happens rather than the round after. the library's front end dropped to 17,786 expression visits from 23,224—a fifth less work, for twelve more rounds of a loop that is now usually short. it reads 7,523 today, and that figure counts a different program: since 2026-09-08 the compile gates check a fixed corpus rather than the json library, because a term measured on a library moves whenever that library changes its imports. the two figures above are what they were when the change was measured, on the library.
that vein caught something else. bench/compile_allocs_golden.txt counts what the front end allocates rather than what it keeps, which no other gate in the tree can see, and one function answered for a fifth of the count. every walker over the syntax tree asked expr_children for a node's children and got a fresh vector back — 94,784 of them on the json library, each read once and dropped. handing the children to a callback instead takes the compile from 148,073 allocations to 91,185, a 38% cut, with rounds and visits identical on both sides because no pass moved and no node is visited twice. checking the library is about 13% faster, interleaved on one sitting.
the profile then said something unexpected. with the vectors gone, thirty per cent of every instruction the front end retired was the hasher. std picks siphash with a per-process random key because a server keying a map on a request header needs collisions to be unpredictable; a compiler keying a map on the identifiers in a file it was handed is paying for a defence against nobody. twenty lines of the multiply-rotate hash rustc has used for its own interner since 2015 took the check from 90.9 million instructions to 67.2 million, with the rounds, the visits and the allocation count identical on both sides.
the same change made the number pinnable. under the random key, the same binary checking the same sources retired 90,704,760 instructions on one run and 90,676,800 on the next — the seed moves the probe sequences and moves where the rehashes land. with the seed gone it reads 66,961,255 three times running. every other vein here pins an exact number and this dimension could not have had one.
the number then got a vein, which it could not have had a day earlier. bench/compile_instructions_golden.txt pins what the front end retires, because nothing else could see a quarter of its work leaving — the allocations, the rounds and the visits were identical across that change, and the peak moved inside its band. the gate compiles from a fixed path rather than from the checkout: the count moves about 160 instructions per character of working directory, and a row read where the repository happens to sit would pin the clone instead of the compiler.
one of those phrases is gone now, and the ruling that removed it is worth stating because it changes how every golden in the tree is read. the compile-memory vein was pinned inside a two per cent band, and clay ruled the band away: "that shouldn't be a tolerance. it should be a setting per platform." the band was the whole habitat of the bug it hid — two per cent of 871,649 bytes is 17,432 bytes of slack, granted for a host divergence measured in tens of bytes, and main drifted 376 bytes greener inside it with ci green the entire time. the gate asserts equality now, and where hosts genuinely differ the answer is a row per platform rather than a tolerance wide enough to swallow the difference. the host-pinning that landed the day before had already built the other half: every golden carries a measured-on line, and a gate refuses to read a row off a machine that did not measure it.
building it turned up how far the band had let the number drift. the stored figure was 871,649 and the runner held 864,300 — 7,349 bytes above reality, read out of the gate's own ci output, where it prints what it measured on every run including the green ones. the ruling had named 872,061 as the true number and twelve allocation merges landed after that was taken, so the recorded figure was itself stale; a container reads 864,274 against the runner's 864,300, which is the twenty-six-byte divergence the measured-on line is there to keep apart. the welfare score gains 0.02 for the correction, because the term had been scoring a peak the compiler left behind.
the gate also had no mutation, which is how it could rot in the first place. the ratchet proves every gate by breaking it, and compile memory had no row at all where the allocation and instruction veins had two each. it has two now, and the first is priced against the ruling: it moves the golden 1,026 bytes, which is 0.119% — red against an exact gate, and comfortably inside the two per cent the old one allowed.
three language rulings landed beside it, recorded here because the log carries them and this page is what a reader is shown. a demanded knot counts as a thunk allocation and the oracle moves to agree. >> stops at the first run-time failure — a question answered in july and re-asked in august, which is its own lesson about where a ruling has to live to stay found. and a hole in a build block is spelled _, filled by a field write exactly once: unfilled at the freeze is a refusal, and a second write to the same field is a refusal.
the hasher left malloc as the largest single thing in the profile, and the reason was one habit repeated in four places. a pass that reasons about names — which functions a declaration mentions, which types a public signature exposes, which bindings a statement makes, which bare aliases stand for a qualified original — was building a String for every name it wanted to remember, out of a program that already held all of them. the door analysis, the demand pass and the alias canonicaliser keep borrowed sets now, and check.rs turned out to have carried a borrowed twin of one of those walks the whole time. the json library's front end came down to 64,884 allocations from 91,185, and to 59,527,334 retired instructions where the seedless hash first read 66,961,255. the compile gates read 14,319 allocations and 24,998,478 retired instructions today — one row, one value, compared exactly, and §34 says why the per-chip key it used to carry is gone; it counts the compiler's own frame rather than the whole process, which §44 says why. it also counts a different program: since 2026-09-08 those gates check a fixed corpus rather than the json library, because a term measured on a library moves whenever that library changes what it imports. the four figures before it were taken on the library and are a comparison between themselves, not with this one. eight passes gave up their owned names over one sweep: the door analysis, the demand pass, the alias canonicaliser, provenance's fixpoint keys, the shadow checker's globals and the two arity maps beside them, the extern name set that had been copied whole once per file, and the bare-name walk that kept a string per identifier occurrence. the fixpoint keys were the largest single fall, at 1.8% of the front end's total work, because a fixpoint pays for its keys once per round rather than once.
the same habit is in the lexer and it cannot be fixed the same way, which is worth writing down so nobody spends another afternoon on it. Tok::Ident owns its identifier and the parser clones it into the ast, so every name in a program is allocated twice — lex and parse are 36.9% of the front end's allocations between them. the obvious move is to take the string out of the token rather than copy it, since the parser walks forward and a consumed token's payload is dead. inside one parser that audit passes: the position is monotonic, and the one backward read looks at a payload only for operators. it dies on the next question. a parser borrows the token stream, and the same tokens are handed to several parsers over overlapping slices — one place parses a whole line and then re-parses the halves either side of a marker. so "consumed" is not a property of a token at all, and emptying one robs a later parser of a name it still needs. removing the copy for real means the ast borrows too, which is a whole-compiler lifetime refactor rather than a local change; the ceiling is worth measuring before anyone starts.
a smaller ruling got built the same day, and it closes a hole rather than opening a door. this language refuses before anything runs, and reading a field off the wrong record did not: p = point 1 2 followed by p.name compiled fine and died at run time, because the check asked only whether ANY declared record has a field called name — and some record did. that is the right answer for a parameter, whose type is a question for whole-program inference. it is the wrong answer for a local bound to a bare construction, which is where the mistake actually gets written. three shapes now say their type without any inference at all: a construction written where it stands, a local bound to one, and a parameter that either carries an annotation or arrives through a constructor pattern. all three refuse at compile time with a span. everything else stays quiet rather than guessing, and the run-time refusal is still there behind it.
a bare parameter still says nothing on its own, and a census taken before the next increment showed what that leaves out. across the eight programs in the tree there are a hundred and forty-two field reads, and the three shapes above type none of them — real code reads a field off a parameter, which is the one case only whole-program inference can see. the rule that landed next asks a different question, and answers it with no inference at all. the interpreter reads . off a record and nothing else, so a value a body reads for two fields is some one record, and a record this program declares. when no declaration holds every field read off a value — m.to beside m.wanted, where to belongs to one type and wanted to another — the reads cannot all be of the same value, and the program is refused with a span before it runs. thirty values in the tree are read for two or more fields; every one of them is determined by that set alone, none is refused, and substituting a declared field name a value is not already read for is refused in ninety-four per cent of cases. a value read for a single field is out of reach, and so is the confusion that some third record happens to hold. the fixpoint is what answers those.
sizing that work turned up a number worth publishing on its own. the natural shape for the general case is the sibling of the value sets — a bitmask, one bit per declared type — and a census of every module in the tree kills it. what inference sees is a MERGED program, carrying its dependencies' types as well as its own, and those run to 87 types for the grammar checker, 79 for the fingerprinter, 67 for the site smoke test. eight programs here already pass sixty-three, so a sixty-four-bit mask would saturate precisely where programs are biggest. the per-file counts point the other way and that is the trap: lib/regexp declares fifteen types in its own source and carries forty-five merged.
the probe that produced those per-phase numbers is not shipped either, and the reason is a question the score cannot answer. it costs 92,873 instructions on every compile — about one per allocation — because letting the library read the binary's counter makes that static's address escape. three arrangements were measured and this was the cheapest: exporting the counter costs 173,229, a bare extra atomic 254,029. spending 0.14% of the compiler permanently to see where its allocations go is a trade welfare has no term for, which is one of the questions still on the ledger.
an interner would take the rest, and how big one would have to be is now measured rather than guessed. a census over the merged programs of the json build counts 394 distinct names in the largest and 622 across all four, against about eleven thousand string copies the front end still makes. a few hundred entries, allocated once, standing in for eleven thousand copies — and every name comparison in the compiler becomes a four-byte compare rather than a memcmp, which is what the 1.3 million instructions on the memcmp line are. the economics are not close. what is not measured is the churn: Ident holds a String today and every pass, both back ends and the interpreter read it, so the table has to be reachable wherever a name is printed. that is a question about how much code moves rather than about whether the move pays, and it is the reason this is written down rather than done.
finding the rest of them needed a better instrument. every one so far had been found by reading the instruction profile and guessing which pass behind a line was allocating, and that guessing was wrong at least once: a hot Set<String>::insert line looked like the checker's globals table and turned out to be 612 inserts. valgrind --tool=dhat records every allocation with its call stack, needs no code in the compiler at all, and agrees with the counter exactly — 82,788 blocks against a counter reading 82,776, the difference being the twelve made before the counting allocator is installed. a probe built into the compiler for the same purpose had been measured at 92,873 instructions per compile and declined; it was approximating something an external tool already does for free.
the first map it produced was still wrong, in a way worth knowing about. on an optimised build the innermost frame valgrind names is the ENCLOSING function, so everything inlined into a pass is charged to that pass. --read-inline-info=yes resolves them, and the same 77,261 blocks sort quite differently: a third of everything the front end allocates is String::clone, which had appeared nowhere in the coarse map because each clone was charged to whichever pass had inlined it. a profile that names functions is naming the frames the compiler left behind rather than the code that ran.
with a true map the largest cluster outside the lexer was provenance — the pass that decides which package's failures may reach a rescue — and it keyed its fixpoint on (String, usize), a declaration's name and arity with the name owned. a fixpoint pays for its keys once per round rather than once, and this one runs the whole program through up to two hundred. borrowing them is 1.8% of the front end's total work, the largest single fall of the sweep.
that pass had also never been pinned. collapsing every group's name, so provenance could not tell one declaration from another of the same arity, left all six of its advisories green and the json library's three unchanged — the central key of a whole pass, and nothing could see it move. the fixture that catches it is two functions differing only in name, one published and one private and uncalled; with the key intact only the published one is advised, and with the names collapsed the private one inherits what the other was fed and is advised for a rescue it never made.
the two counts do not move together. the alias change removes 1,413 allocations either way, but measured against the tree before the demand pass landed it was worth 531,217 instructions and measured against the tree after it was worth 344,565 — malloc's cost per call depends on the state of its free lists, so a leaner tree pays less for the same saving. allocation counts compose and instruction counts do not, and a ratio quoted from one measurement does not carry to the next.
four of the maps the front end builds know their size before they start filling, being keyed off program.fns. pre-sizing them cut the rehashing by 419,151 instructions and the total by 80,631, the rest going back out in the larger table. sizing each one to what it actually holds, so the table ends no bigger than growth would have made it, should have been strictly better by the bucket arithmetic and measured worse — a table at half load probes in fewer steps, and that is worth more than the memset it costs. the blanket hint stayed.
the allocator under all of that had never been the subject, and it was the largest single thing left. profiled on the corpus the compile gates check, glibc's malloc.c and arena.c come to 15.17 per cent of the instructions a compile retires — more than any function the compiler owns, against a hash map insert at 3.94 and inference's expression walk at 3.61. the compiler has kept a GlobalAlloc of its own since the compile counters were minted, but it is a tally that forwards to the system allocator, and what it forwarded to had never been asked about.
mimalloc goes under that tally rather than beside it: the wrapper still counts every call, every byte and the running peak, so the counters the objective reads stay the compiler's own demand. the three instruction veins fall between 8.7 and 9.1 per cent, and nothing else in the tree can see the change — allocations, peak bytes, passes, rounds and visits are byte-identical, the emitted llvm hashes the same either way, and the linker reads the same five shared objects because the c library is compiled into the binary. for scale, the ten compile-side merges before it moved these rows between 0.13 and 1.14 per cent each.
the run program does not want it, and the arithmetic is the reason to say so rather than to try. on the same tree the whole allocator bill inside runbench is 0.33 per cent, on a workload nine hundred times longer. a kanso program serves its values out of the arena and gives them back at the beat, so the system allocator sees the arena's block requests and almost nothing else; the compiler has no arena and asks for every string, every vector and every table it grows. this lever belongs to the compile side alone, and it is large there because that side never got the design the other one did.
one thing in the tree could see the swap that no counter could. mimalloc commits its first arena up front, so a process that has allocated almost nothing already holds about six megabytes; bind_chain_depth reads the operating system's number rather than the program's, and a constant floor under both of its readings squeezes the ratio it asserts from 3.16 to 1.97. the option comes off rather than the test, from the process's own constructor table because the top of main is already too late, and it costs under two per cent of the win.
a second option came off for a different reason, and the reason was a disagreement. ci measured the three compile rows twice on one commit and got two answers, 13 instructions apart on all three; the 2026-09-05 ruling is one row, one value, so the round went to hunting it. ten runs in a container gave the same digits every time and the two runners matched on chip, libc and compiler, which leaves what the process does once. the profile turned up something else on the way. mimalloc gives free pages back to the operating system on a timer — each arena carries a deadline, and a pass over one asks clock_gettime whether it has gone by. on the compile corpus it asks 163 times, and how many times it asks depends on how long the process has been running. in a container the machinery costs 8,288 instructions on the module row against 44,608 on the entry row and 48,487 on the library row: five times the cost for three and a half times the work, which is the shape of a term keyed to elapsed time rather than to anything the compiler did.
ci then settled it by disagreeing. with purging off the runner saved 8,301 instructions on the module row, 73,328 on the entry row and 66,857 on the library row. the short row agrees with the container to within 13; the two long ones saved sixty and thirty-eight per cent more. a fixed per-process cost would have carried across all three. a term counted out by elapsed time does exactly this, because the runner is slower under callgrind than the container and crosses more deadlines while it works. the projection written into the goldens for one round subtracted the container's figures and was 13, 28,720 and 18,370 out, in that pattern.
so purging is off, which leaves three unconditional reads at startup where there were 163 that answered to the clock. a compiler is a short-lived process that exits and hands everything back at once, so this removes work. the 13 is still unexplained, and the log says so rather than claiming the credit: a difference in how many deadlines expired would land on the long compiles and not the short one, where the gap was the same on all three. the wall-clock dependence was real and is gone; the disagreement stays open, with those three startup reads named as the next place to look.
9the surface a program needsshipped
a fast compiler with nothing to compile against is a demo. four pieces landed in one sitting, and each of them was a thing a real tool could not do without.
reads are applications. clay.name is a getter arm applied to a record, and _.name is that getter as a value — so list/map people _.name works, and one accessor serves every type that declares the field, because it is an ordinary dispatch arm. the expectation going in was that this would cost something. it did the opposite: field access was never a static offset. the runtime walked the record comparing field names, so binding the field by position instead is less work. twenty million reads went from 0.08s to 0.03s, interleaved, with the compile golden unmoved.
the getter carries a name no program can spell, which is what keeps a field from taking a name away from a type or a local. Haskell reached the same place from the other side — its selectors were monomorphic names, and twenty-five years of escaping that produced NoFieldSelectors and a field lookup that lives in the type system. kanso's dispatch was already there.
an os surface. before this the whole vocabulary was 39 builtins, lib/io exported five things, and none of them was stderr — so a tool's diagnostics went into whatever it was piped into. now there is stderr, an environment read where an unset variable is none rather than a failure, file existence, a directory listing, and a clock. lib/path came for free in pure kanso, since text/slice is enough for basename and dirname.
those names live in std/os now. the library takes go's split, which is the one the language agreed to mirror: os holds the filesystem, the environment, the arguments and the processes, and io keeps the surface a program reads and writes through. go answers most of the boundary cases and leaves one — its standard streams are files in os and the writing is done from fmt, and kanso has neither files nor a fmt. what it has is three verbs, and they stayed in io, because a module that keeps the read and write surface while holding nothing that reads or writes is a name with an empty room behind it.
the split paid for itself in code size, which was not the point of it. a program that imported std/io used to drag the filesystem, the environment, the arguments and the processes in with it; it now pulls a module with three names in it. the decoder emits 11,588 lines of ir where it emitted 11,603, the escape benchmark links 49,458 bytes of machine code where it linked 49,650, and every one of the eight benchmarks retires a few hundred fewer instructions before it reaches main. the welfare score does not move — a few hundred instructions in three billion is below anything a saturating term can see — which is the honest shape of this kind of win: real, pinned in three goldens, and too small to celebrate.
two of those decisions are about determinism, which is a running theme rather than a coincidence. a directory hands its names back in whatever order it likes, so a listing is sorted — otherwise a program's output depends on the disk. and a run that timestamps is unrepeatable, so KANSO_NOW pins the clock exactly as KANSO_SEED pins the dice. determinism also decides where a behaviour is tested: an unset variable reads the same on every machine and lives in the corpus, while a set one is asserted with the environment controlled.
starting a process. os/run cmd args answers a record of status, stdout and stderr. a non-zero status is what the process said and a caller reads it; only a process that never starts raises an err, which crosses a package boundary like any other std failure and so can be named in an arm. a browser tab cannot start anything, so the playground declines by name rather than pretending.
packages. imports are the manifest — there is no second file restating them. kanso install reads them, resolves each hako's highest release tag, fetches it into a content-addressed cache and writes hako.lock; kanso list and kanso update are the other two verbs, and there is no third. a dependency you need before its author has tagged it is kanso install --from owner/repo@branch, which writes the branch and its sha into the lock; list says the pin is interim and update walks releases past it, so the pin stays visible until a tag replaces it. the compiler never reaches the network: it reads the lock and the cache, so a build works on a train.
the resolver is where the design earns its keep. each import shape answers in exactly one way and never falls through to another, which is why a local directory called owner/repo cannot stand in for the hako of that name. the lock records a protocol beside the tag and sha, and the protocol is a protocol rather than a host: GitHub and GitLab both speak git, tag discovery and ref fetch are git standards, and the one thing a GitHub-specific fetcher adds — a tarball at a ref — trades a commit sha for a promise of byte-stability that broke on 30 January 2023 when a compression change altered archive bytes for identical files.
hako reads its own lock, in kanso. kanso list now exists twice: in Rust, which owns the CLI, and as a kanso module in the repo that reads the lock, reports what it pins and asks each remote through os/run whether the pin has fallen behind. a pin at a release is measured against the highest release its remote publishes; a remote that cannot be reached says so rather than guessing, and an interim pin sits off the tag series staleness is measured along, so the only thing worth saying about it is that it is interim. it is the first program kanso has that is neither a benchmark nor a book sample, and what it exercises is the whole of the surface above: a file read, a process started, text split and sliced, a list folded into one write.
a listing is one effect per pin, and that is where the port found the gap. a description binds a lambda, but a plain value re-enters the chain only by writing nothing to stdout — io/write "" . (_ -> x). every program that runs one effect per element of a list needs that function and io has no name for it, which is the next thing to settle.
arity is checked across a module, and carried by a function value. a module is a directory of files sharing one namespace, so the check that a call brings the number of arguments some arm can answer runs across the whole of it. a call to a sibling file used to reach the emitter unchecked and arrive as an undefined symbol from the assembler. bare names count, because no binding may shadow a declaration — adder = … beside fn adder is a name error — so a bare name matching a declared group is that group.
what the checker cannot know is a constant holding a function, whose arity is decided at runtime, so a function value now carries how many arguments it takes and the three engines agree on what to say when a call brings a different number. before that a one-argument closure called with two ran anyway: the extra argument sat in a register the callee never read.
a package can test its own failures. the two-universe rule forbids turning your own err into a value, and a library that cannot assert that its errors happen cannot be trusted. the way out needed no exemption: the harness is a foreign party, so a builtin doing the reading is a stranger rescuing. the rule's own mechanism settles it, because a rescue is attributed to the arm whose pattern names the err, and a builtin has no pattern to attribute. failed? is in scope in a _test.kso file and nowhere else, so the rule stands undiminished outside the harness.
the entry has no name a program can spell. a directory is run through the file main.kso, and a single file through its pub play. that is the whole of the mechanism, and a program may use the word main for whatever it likes. the compiler used to synthesise its entry under that name, in the reader's namespace, so a file writing main = … beside a pub play ran the binding and never ran play. the entry now carries a capitalised name on the same grounds as a getter's Get_: source identifiers are lowercase, so nothing a reader writes reaches it. an err arriving at the top says it reached the entry.
the gates are moving to kanso. the checks that guard this repo — the differential sweeps, the welfare index, the drift budget between the log and this page — were written in python, and the first of them now runs as a kanso program in CI. it is the cheapest real workload the project has. seventy lines written against the standard library turned up a text/slice that answered differently on two engines one position past the end of a string, which eight differential sweeps had missed: they probed far past the end, and an off-by-one only ever lives exactly one past.
a probe that never returns is the loudest finding a sweep has. two silences compare equal, so a differential sweep that only asks whether the engines agree will read a pair of hangs as agreement. every probe runs under a timeout, and a probe that outlives it is reported by name and fails the sweep. the sweeps that moved from python to kanso lost that property in the move and got it back — which is the more general lesson: a port deletes the thing it replaces in the same commit, so nothing compares the two, and a property can disappear without anything going red.
a segfault is not a diagnosis. a native program killed by the operating system reports that it ran out of stack, and deep recursion is only one of the things that ends a program that way. one reduced program spent an hour looking like runaway recursion and turned out to be a null pointer handed to text/join — a string carrying a length and no data, valid where it was built and empty where it was read. the sentence the compiler prints there names a cause it has not established, and the three engines were deliberately aligned on it, so correcting the wording is a change to all three or to none.
the gates that guard this repo are becoming kanso programs, and that is where the compiler's bugs are turning up. five of them have moved so far, and the move is worth more than the deletion. writing a real program against the standard library found a slice that disagreed between engines one position past the end of a string; a walk that accumulates while it performs effects that dies on the engine that ships and answers correctly on the one that reads; and a binding whose value becomes a thunk, in a function whose arguments can fail, that emits invalid machine code because every return releases every cell whether or not the cell exists on that path. none of those came from a corpus. they came from someone trying to write an ordinary program.
the trend gate is the first of them that is a program rather than a table of probes, and it wanted things a sweep does not: counters summed across lines with no map to fold into, so pairs are grouped by name and each group summed; names taken from the pairs rather than the map, because a map answers with its values and not its keys; strings joined by interpolation, because concatenation is for lists. it produces the same bytes as the python it replaces, on the path where a counter improved and on the path where a pure regression fails.
the native build refuses what it cannot represent, and there was exactly one place it did not. integers are arbitrary-precision in the spec and sixty-four bits in the machine, so the honest thing for a native program to do above that line is stop. addition, subtraction and multiplication all do. division did not: the least integer divided by minus one is one past the greatest, and the machine answered by wrapping — the same number back, with a successful exit. it raises the overflow now, like the others.
it was hiding inside a number. the sweep that compares arithmetic across the two engines reported four hundred and seventeen disagreements, which is a figure nobody reads twice. classified by what the compiler actually said, they are three hundred and ninety-six literals too wide to compile, one hundred and sixty-four overflows caught at runtime, and one wrong answer. a report large enough to skim is a place for a real finding to sit unread, and the fix for that is to say which kind each one is.
and the test that caught it made the suite lie. the numeric tests each get a work directory keyed by the program, except the key was the program's first twelve digits — and 9223372036854775807 * 2 opens with the same twelve as the division. two tests shared a directory, deleted each other's entry mid-run, and one reported the other's answer. the comment above that key already promised one directory per program; the isolation was described correctly and implemented on a prefix, and the suite was one test away from two of its members quietly trading results.
10a program written as a programshipped
kanso had one string form, and every gate in this repository is a program that writes other programs. a differential sweep holds a hundred small kanso programs and hands each one to three engines, so a hundred programs sat in the source as single lines with their newlines spelled \n and their splices spelled \{. that is unreadable in the file and unpasteable out of it, and it is where the last of the python was hiding: the scripts that had not moved were the ones whose payload was a program.
the form is swift's. """ as the last token of a statement's line opens a text block; content sits at the opening line's indent plus two; a """ alone at the opening indent closes it. every content line contributes its text and a newline, the last one included, and interpolation works exactly as it does in a quoted string. content lines are exempt from the eighty-column cap, because the width rule exists to keep a statement readable and a content line is not a statement.
two other shapes were considered and declined. ruby's <<~NAME carries a name that means nothing — the reader learns a label, uses it twice, and it says no more about the text than the quotes do. a per-line sigil, the | of yaml and swift's own multi-line comments in some styles, defends against a runaway fence by ending the block at the first unmarked line; kanso's indentation is already rigid enough that a runaway fence cannot get far, and the sigil costs paste fidelity, which is the whole point of the construct.
eight things are refused with a message each, and they are the ways a block goes wrong rather than a list of tastes: anything after the opening fence, a content line shallower than the content column, end of file before the fence, anything after the closing fence, \n inside a block where a real line break says it, \" inside a block, an interpolation that does not close on its own line, and a block of fewer than two content lines. the last one is the interesting one — a one-line block is a quoted string with more ceremony, and allowing it would make the choice between the two a matter of mood.
ninety-seven call sites moved across twenty files. the three heaviest are the differential gates, whose output is pinned byte for byte, and the migration's rule was that those goldens stay untouched: a regenerated golden during a rewrite of the thing that produces it is indistinguishable from a bug, so the only acceptable evidence that the rewrite preserved meaning is that nothing downstream moved.
11a header that outlives its storage
the strict tier rewinds an arena at every step of a bind chain, which is what makes an io walk cost new data rather than live data. three defects landed this week and they are one defect: a header that survives a rewind, holding a pointer to storage that does not. each was found by someone writing an ordinary program rather than by a corpus, and each needed a different part of the runtime to learn the same question.
the first is a list built before the chain's first bind and pushed to inside it. the header sits below the mark, so it outlives every rewind; the buffer the push allocates sits above it and does not. an in-place push now asks whether the list was born this beat, and grows a fresh buffer when it was not.
the second is the copy walk one level over. the walk prunes at any node that survives the mark — below the mark it outlives the rewind, so share it and stop — and never asked whether the storage that node points at survives too. a string taken out of a carry buffer and built into a fresh record leaves the record surviving with a field aimed at a buffer two carries from retirement. a node is shareable now only when it and its immediate interior both survive, asked in the sizer and the copier alike so the two agree about what they are doing. the exception is a string builder, whose storage is malloc'd exactly so that a rewind cannot reach it; copying one strips the capacity that says so, and the next append finds a plain string where it left a builder.
the third is laziness meeting the same edge. a call whose callee is a beat loop is bracketed with a mark and a rewind on the promise that its arguments are already evaluated and so live below the mark. a thunk is the argument that is not: it is forced inside the loop, its value is built above the mark, and the cell memoises a pointer to it that the first rewind invalidates while the cell still reads as forced. k_force declines a memo it cannot keep — when the value it just computed does not survive the innermost mark, the cell stays unforced, keeps its captures and recomputes on the next read. every cost golden is unchanged, which prices the shape: a thunk forced inside a beat whose result lands above the mark does not occur once in four benchmark veins.
the second fix is the only one that costs anything, and it costs a fortieth of a percent of encode allocations — the copies it now makes instead of sharing what it should not have shared. there is no version of it that does not pay them. welfare holds at 65.56.
what took longest in each case was the reduction. the copy-walk fault would not shrink at all: six self-contained attempts produced programs both engines agreed on, and a first generated-tree test passed with the bug because it named its tree tree — short paths mean short strings mean a different allocation sequence, and sweeping the root name's length showed the fault appears at eighteen characters and holds from there. the memo was found by building with kanso build, editing the emitted IR by hand and relinking it against the runtime object the compiler leaves behind. adding one force to a parameter answers in a single run what a week of reductions had not, and that is worth naming as a technique rather than as an anecdote.
12two things the compiler stopped letting through
an effect handed to a parameter nobody reads. an effect is a description and a description nobody forces never happens, so a call like ignored (io/write "…"), where every arm of ignored discards that position, wrote nothing and said nothing on either engine. the language's stated position was always that a body doing io must hand the io back or it abandons the effects above it — but that check reads a function's own body, so an effect abandoned inside somebody else's parameter walked straight past it. the new rule refuses only the case with no other reading: every arm throws the position away. a group where one arm reads the parameter and another does not is a real question and stays legal. it needed nothing new to decide, because inference already knows whether an argument describes an effect and the arms' own patterns already say whether all of them discard it.
a qualified name is its module's declaration. when a module's own pub shares a name with something a dependency exports, a consumer writing dep/join used to reach the dependency's arm through dep's name, and for a while the compiler refused the program as opaque rather than let it. the 2026-08-29 ruling settles it: dep/join is dep's own join. inside dep a bare join still dispatches over its own arms and the import's together, because the import's twin lives under a spelling no consumer can write. a list handed to dep/join when dep's arm takes an int is refused as a call with no arm for it, on all three engines.
the digest that finishes the story is std/sha256: FIPS 180-4 in ordinary kanso over std/bits, no builtin, because a hash is arithmetic on thirty-two bit words and the bitwise operations were already there. it is what the asset fingerprinter needed, and with the fingerprinter ported the only python left in this repository is the two gates that drive a headless browser.
13the language reserves no name
kanso has no main. a directory is run through the file called main.kso, and that is a filename rather than an identifier — a program may use the word main for whatever it likes. the same rule now covers play, which the compiler used to know: it looked for a pub binding by that exact spelling and synthesised the entry from it. one magic name is a rule; two is a habit.
play is an ordinary exported constant. what makes something runnable is a file of statements, and the way a sample stays one command away is a runner beside it:
# greeter.kso — an ordinary library
fn shout name
"hello, {name}!"
pub play = print (shout "kanso")
# main.kso — the runner
import "greeter"
greeter/play
the compiler is told nothing about either file. the runner names what it wants, and the convention lives where a convention belongs — in the harness that runs the samples, and in the playground, each of which generates the entry it needs. what this cost was one capability rather than a mechanism: a bare import used to name a sibling subdirectory, so a runner could not import the file sitting next to it. a module is a directory of files sharing one namespace, and one file is the smallest of those, so import "greeter" now reads greeter.kso the way it reads greeter/. both spellings present at once is refused: one name cannot answer two ways.
14a constant may name itself
read a data file describing relationships and you have to build the graph it describes, which means a node holding a node that holds the first one back. kanso used to refuse that outright. any constant whose body contained its own name got `graph` is defined in terms of itself, so it has no value, and the only way through was to build the shape at runtime out of ids and look each one up on every hop.
type node
name
peers
graph = { "a":(node "a" [graph["b"]]) "b":(node "b" [graph["a"]]) }
pub play = print "{graph["a"].peers[1].name}"
that prints b. the rule that admits it is about demand rather than mention. a constructor stores its arguments and never looks at them, and a list or map literal stores every element, so a name appearing inside one goes into a field and is read later, by which time the constant it names has an answer. a name appearing anywhere that scrutinises it, an operator or a guard or a call that forces what it is handed, is still a demand, and x = x + 1 gets exactly the error it always got.
admitting the program is the smaller half. the constant lives in a cell that is filled once and read thereafter, rather than recomputing on every read, which is what stops the definition chasing itself down the stack. a storing position inside it emits a value that waits instead of a value, and both the type inference and the code that decides where to force learn that a container read may hand back one of those. all of it is gated on the program holding such a constant at all, so a program without one emits and pays exactly what it did before.
the cell fills the first time something reads it. native used to fill every one of them before main, which meant a constant the program never asked for was built anyway — and the interpreter, which waits, disagreed with it on exactly that program. the rule settling it is that work defers until it is presented to IO, and building something early is a resource decision inside that contract rather than a difference an engine is allowed to show. so a knot nobody demands now allocates nothing on either engine, and one that is demanded costs a flag test on the read that finds it already built. the decoder links no such constant and does not move.
the browser carries it too, through a slot that answers when something reads it rather than a value sitting in memory waiting. all three engines run these programs and print the same bytes, pinned by differential goldens.
render and == ask different questions of such a constant. print "{x}" walks the shape and marks the place where it closes back on itself, so x = [x] prints [<cycle>]. == is asked which object this is, and a constant that names itself is a definition: asking whether two of those are equal is asking whether two formulas are the same formula. so it refuses, the way it already refuses a function or an effect — equality is not defined on a value that names itself. the refusal fires when the comparison arrives a second time at a pair of cells it is already inside, which is the cycle closing through one. an ordinary lazy binding is forced and compared like any other value, and a cycle a build block closed by writing a record in place keeps the structural comparison of §06.
the refusal holds wherever the comparison meets a function, at any depth. [k] == [k] with k a lambda refuses, as does a record or a map holding one, and a partial is a function like any other, so (&f n) == 3 refuses too. native refused both at any depth; the interpreter asked only of the operands and left partials off its list, and answered false.
15the connectives are words
a and b, a or b, not a. & | ^ stay with the bits, where they read as bit patterns rather than as questions, and there is no ! at all — a language that says and out loud and then punctuates its negation is mixing two metaphors in one expression.
none of the three costs an engine anything. the parser writes them as if at parse time, so no backend has ever seen an and node and none needed teaching. not sits on its own rung between the connectives and comparison: not a and b denies only a, and not a == b denies the whole comparison. that rung is what keeps the parentheses in not (a or b) load-bearing while calling the ones in (not a) or b superfluous.
flag == true is a compile error, and so are its three siblings. the comparison asks a question the value has already answered, and comparing two booleans to each other is exclusive-nor wearing a comparison's clothes — a spelling the language declines to offer. elsewhere this is what a linter warns about; kanso has no linter layer, and a language with one way to write a thing has nowhere to put a warning, so what a linter would flag is refused instead. the error names the replacement: the value itself, or not the value.
there is no nand and no nor. not (a and b) already says it, with the denial on the outside where it reads.
16a server is two ordinary statementsshipped
kanso speaks tcp now, and the syscalls were the easy half. a server that blocks on accept cannot share a program with the client that would connect to it, so the two would have to be two programs — which is the moment a language usually grows goroutines, channels and a select to arbitrate them. kanso already has the arbitration: adjacent statements are one parallel group, scheduled by a deterministic green-thread runtime. what was missing was for accept to answer “not yet” and go back in the queue instead of holding the runtime, and the same for starting a process. with those two as scheduling points, a program serves itself — the server on one line, the request on the next, byte-identical interpreted and compiled.
std/net/http sits on top in Go’s shape, with one difference that matters: a handler is a plain function from a request record to a response record. there is no writer to hand it and no recorder to fake, so a test calls it and reads what it answered, and injecting a different handler is passing a different function. the mux is arms on the path, and adding a route is adding an arm. many requests are a fold — the handler answers the response and what to carry to the next one, and carrying none closes the door.
processes gained the other half of Go’s pair. os/run starts one and waits; os/start answers the handle and os/kill ends what it names, because a browser told to render a page ignores its own exit budget and runs until something stops it.
17two failures, and which one you hear about
a failure in an operation propagates: hand an err to anything and you get the err back. when both sides fail, the answer carries both — the reason becomes the list of reasons, and neither failure is billed above the other, because neither caused the other. that merge is a fold rather than a pairing, so three failures answer three reasons however they were grouped, and the shape of the expression that produced them is not recoverable from the result. an err whose reason is itself a list stays one reason; the mark saying which is which is structural, because it cannot be read off the shape.
the bind is the other rule. .> is ordered, so the first failure is the answer and what follows it never speaks. adjacency accumulates because nothing there is first. short-circuit where there is order, accumulate where there is none — you never choose between the two, because each is the only behaviour its structure permits.
the two rules make the same pair of algebraic objects, over one carrier: adjacency associative and commutative, the wall associative with the first failure absorbing. that is what makes the grouping invisible, and it is also what a user-supplied version of either would have to obey.
18what a count can and cannot see
every performance claim on this page is pinned to a count rather than a stopwatch, because a count cannot drift between machines. that choice has an edge, and knowing where it sits is part of reading the numbers.
a count answers the question it was written to ask. bench/compile_golden.txt records what one inference pass costs — the rounds the fixpoint went round and the expressions it looked at — by resetting the counter and running exactly one pass. so it prices a pass, and says nothing about how many passes the front end asks for. those are different questions, and the second one is now pinned separately: the whole-program pass count is a watched number with the same standing as the golden, and moving it costs a sentence in the log naming the pass and the reason. the front end infers once and hands the result to the three checks that read it; kanso check over kq is 11.0 ms, of which about 2.4 ms is process startup.
the same edge shows up in memory, where one number is currently unpriced. a tail-recursive loop that allocates a temporary each time round mostly runs in constant memory, and one shape of it retains everything — 1.7 mb against 69.7 mb over 1.6 million iterations. the temporary need not be kept: a string that is bound, measured, and dropped costs what a string stored as a map key costs. a loop that brackets carries an arena mark and rewinds every iteration; the shape that retains carries none of it, so it has no rewind point at all.
which shape retains took three goes to state correctly, and the two wrong statements are worth keeping because each was killed by the next measurement. the page said the axis was what the loop carries — a scalar against a map or a list. it is not, and the fix it proposed on that basis — give a threaded container storage outside the rewound region, the way the byte builder already does — was built three times and reverted three times in one day, twenty-five days ago, every attempt moving the arena counter and none moving the process. the best of them got the counter down 69x while peak rss went from 71.8 mb to 75.2 mb.
the shipped compiler, at 1.6 million iterations, with the same string built and dropped every time round — sh_str reads 72,282,272 in all five:
carried recursion beat_iters arena peak peak rss
scalar self 1,600,000 1,048,576 1.7 mb
map self 1,600,000 1,048,576 1.7 mb
scalar through a helper 3,200,000 1,048,576 1.7 mb
list through a helper 3,200,000 1,048,576 1.7 mb
map through a helper 0 72,351,744 69.7 mb
one cell of five. a map carried by a loop that recurses through a second group is refused, and every other combination brackets — a map is fine when the loop calls itself, and a list is fine either way. the refused group draws no line at all from the beat report, so the analysis never reaches it. that is a much narrower thing than the sentence above it claimed for twenty-five days, and it is the reason the storage fix could not have paid: it aimed at what the loop carries, and the loop carrying a map is not by itself the problem.
the refusal is deliberate and it is priced. instrumenting the cluster selector on both loops shows they differ in one thing only:
map through a helper entries=1 edges_ok=true carried=[go[1], onward[1]]
list through a helper entries=1 edges_ok=true carried=[]
both clusters pass the edge check. list is in the threadable set, so the list slot is threaded and nothing is carried. map is not, so the map slot becomes a CARRIED slot — one the loop evacuates at every rewind — and a cluster that has both an entering tail call and a carried slot is refused outright. the comment on that rule gives the reason: a cluster reached only by a tail call is one whose cost nobody has measured, and the json string scanner pays eight gigabytes of copies for the licence.
so the forty-fold difference is a trade somebody already priced, on a different program. what is open is whether the price is right for this one. a loop carrying a single-key map would evacuate a handful of bytes per rewind where the string scanner evacuates a buffer, and the rule cannot currently tell them apart. the two ways out are to make a map threadable the way a list is — which returns to the sorted-view hazard the threadable set was drawn to avoid — or to make the carried-slot refusal read the size of what it would copy. the retention stays a pinned counter in the memory vein either way.
the sharpest version of that edge is a count that measures the corpus rather than the language. a survival check learned to memoise its answer, which took a sixteen-thousand element array from 11.6 billion retired instructions to 130 million — the streaming shape asks the same question of the same carried list once per element, so caching it is nearly free. the voting simulator asks a different question every iteration, and went from 14.2 seconds to 25.0. every allocator counter held still through it. so did the emitted-line count, the machine-code size and the welfare score, because the simulator is not one of the five programs the corpus runs. the memo now applies above sixty-four elements, which returns the simulator to 15.1 seconds and costs the array benchmark a tenth of a per cent.
three attempts to write a benchmark that fails without that guard moved not one instruction between them. the shape that reproduces it is not "a loop over small lists" — that was attempt two — and the difference is something about how deeply the carried value nests. until a program in the corpus has that shape, the guard is held in place by a measurement of a program outside it, which is a weaker thing than every other number here and is written down as such.
a count also answers only for what it is aimed at. the emitted-code golden exists because 7.6% of decode speed leaked away over eleven days with every allocation counter byte-identical — the decoder gained 20% more calls and 23% more branches for the same work, one or two per cent at a time. it watched the decoder. the same build step produces nine programs, eight of them carrying cost goldens of their own, and not one of those eight was counted on that dimension, so the leak that was caught once in the decoder could have happened in any of them unseen. they are counted now, in a golden of their own, because the decoder's file is also its history and summing eight programs into it would let a rise in one hide a fall in another.
widening a count turns out to be harder than adding one. the check that refuses a silent regression sums a golden's counters by field name across its samples, deliberately: one number for machine-code size across the binaries, with the per-file diff saying which one moved. a ninth binary joining that file therefore reads as the number going up, and the check refuses it as a regression on bytes no program grew by. wider and worse are the same shape to a sum. whether a summed vein should be compared per sample is an open question about what that check is for, and it is written down rather than answered.
the same edge caught this page itself. a check counts the log entries written since this file last changed and fails when they get too far ahead, and it read 0/3 on every pull request for days while twenty-two entries piled up. the count was honest about what it could see: the step above it in the workflow fetched the main branch one commit deep, which truncates the history for every step after it, and in a truncated history the boundary commit looks like the one that created every file — so the page always appeared to have moved with the tip. the fetch is a full one now, and the check refuses to answer at all on a shallow clone rather than reporting a zero it cannot stand behind. a gate that cannot see must not report success.
every counter this page names is recorded on every merged commit and drawn on the long view, one row each — what a run costs, what compiling costs, and the welfare score they roll up into.
19a module reached two ways is one module
a program with a dependency is a diamond. the entry imports shape and imports mid, and mid imports shape too. kanso used to build one shape per route, so the emitted code carried mid/shape/describe beside shape/describe and a value the entry built matched no arm the middle module had compiled. the call from the entry answered; the identical call one module in died with no overload of mid/shape/describe matches these arguments.
identity is the path a module lives at, and the route an importer took to reach it does not enter into the name. a name that already carries a qualification came from somebody's dependency and keeps the spelling its owner gave it. the diamond above now emits 429 lines of ir where it emitted 499, because the second copy was code nothing could call.
that has a consequence at the surface. geo re-exports list's sort under the name order, and no declaration anywhere is spelled geo/order — the declaration is list/order, and geo put it on its surface. so the qualified spelling is resolved where it is written: a module that surfaces a name it does not own has geo/order turned into the owner's spelling before anything looks it up. minting a declaration for it would put the two copies back.
import "../geo"
print "{order [3 1 2]}" # the bare spelling
print "{geo/order [3 1 2]}" # and the qualified one
the door opens only where the qualified spelling is free and one declaration answers it. a module that declares its own select beside an import's already owns geo/select, and two re-exports landing on one bare name leave a caller nothing to choose between.
20a recursion the compiler can reassociate runs as a loop
recursion is the only loop the language has, and a call in tail position costs no stack — a promise all three engines keep. a call with work left over when it returns is the other shape: n + weigh (n - 1) holds a frame per call, and a million frames is more than any stack holds. the compiler has always turned the simplest of those into a loop by threading an accumulator through a tail-calling helper, and the license asked for an integer literal as the leftover work. 1 + count (n - 1) qualified. n * fact (n - 1) did not.
widening it needs a proof that the operand is a whole number, because reassociating a sum of floats changes the answer. the proof comes off the shape rather than out of inference. the wrapper the rewrite already generates ascribes every counter position int, so a non-integer argument never enters the loop at all; requiring each recursive call to hand those positions arithmetic over counters carries that property down every level. an operand built from counters and literals is therefore an integer at every depth, and it is pure — a name, a literal and +/-/* reach no call and no effect — so computing it before the descent rather than after moves nothing a program can observe.
what it now reaches: n * fact (n - 1), n * n + r (n - 1), n * 2 + w (n - 1). what it still declines, and why: a float operand, because the reassociation is not exact; a call in the leftover work or in a counter position, because the pass reads one group and cannot see through a call; and double recursion like fib, because there is no single descent to thread. the rewrite remains an optimization rather than a promise, so a program that needs flat memory writes the tail call itself.
the evidence is an engine disagreement that closed. weigh 100000 — the sum of the first hundred thousand integers, written the way anyone would write it — answered on native, which has whatever stack the operating system handed the process, and died on the interpreter, which refuses unlicensed recursion at ten thousand frames. it now answers 5000050000 on all three engines in under a tenth of a second, and the fixture that says so runs on each of them. every allocator counter, both compile goldens and the welfare score held still: the rewrite fires on shapes the benchmarks do not contain.
a reassociation that is wrong does not fail — it answers a number — and no counter can see the pass at all, since it changes the shape a recursion runs in rather than what it allocates. so every shape is written twice: once plainly, and once with the leftover operand passed through a function, which the license refuses to read through, so that copy descends the way it always did. twenty shapes at four depths, both forms, both engines, six tenths of a second. three more shapes are there to be refused: their operand is a float, and the day somebody widens the license to reach them the two copies stop agreeing — 1.5000000000000002 against 1.5, which is the whole argument for the license in one line. the check's own liveness is a sum twenty thousand deep read off the interpreter, which refuses unlicensed recursion at ten thousand frames: delete the pass and that line dies rather than quietly comparing two unrewritten programs.
21bytes are a value, not a list that happens to hold numbers
the compiled engine has had a byte-string tag since it was written, and the interpreter had none. text/bytes "hi" answered a list of two integers there, and every function that wanted bytes took whatever list arrived and tried to read it as one. so text/append ["a"] "x" answered ["a" 120] under the interpreter and was refused by the compiled engine—a program that runs one way and dies the other, which is the one thing three engines are not allowed to do.
six programs disagreed, and in four of them the interpreter was the one answering. they are refusals now, in the same words on every engine: append, find2 and find2_below say they take bytes; utf8 and to_float say what they got instead. the interpreter has a real byte value, and a list of small integers is no longer quietly one.
which takes away the only spelling a program had. there is no bytes literal in the language, so [104 105] was how you wrote byte data down, and removing the coercion without replacing it would have left nothing. two things went in beside it. text/utf8 keeps its list acceptance—declared in the library, identical on every engine, rather than an engine quietly coercing—and text/to_bytes is the constructor, refusing loudly outside 0–255 instead of keeping the low byte. text/bytes covers strings; this covers numbers.
text/utf8 (text/to_bytes [104 105]) # "hi"
text/to_bytes [104 300] # err "to_bytes takes byte values (0-255)"
rendering is the one place the two still look alike, and deliberately: bytes print as their numbers between brackets, because that is what a reader wants to see. what changed is what a function will accept.
22the err rules are two tables
a failure that reaches an operation carries past it. three pattern forms can hold an err, and the rest decline one. two sentences, and between them they govern a few dozen places in the language, each written separately in three engines. asking the same question at each of those places in turn—what happens here—turned up five disagreements in one afternoon, every one a site where two engines agreed and the third did not.
none of them was hard to find. what was missing was the question being asked site by site, and somewhere to write the answer down. both sentences are tables now, and both tables are fixtures that run on every engine.
the first says what an operation does with a failure that arrives at it, and the answer is not the same everywhere. an operator, a comparison, a constructor and a & join MERGE—neither operand caused the other, so both reasons survive. a call with no matching arm, an interpolation and a .> bind take the FIRST, because those are ordered. a list literal and a map literal HOLD it, because a container is not its contents: the list has a length, it can be indexed, and printing it shows the failure sitting inside. a field read, an index and a builtin pass it THROUGH, because the operation does not happen at all.
operated: ["a" "b"] compared: ["a" "b"]
built: ["a" "b"] grouped: ["a" "b"]
called: a woven: a
sequenced: a listed: [err "a" 2]
mapped: { "k":err "a" } reached: a
indexed: a sized: a
the second says which patterns can hold an err at all. three can—the (err …) destructure, an :err annotation, and a typeset with err among its members—and the other seven decline one. that is ten pattern forms against two origins, and the two columns read the same: since the 2026-09-15 ruling made a bare err data, an arm takes a failure whoever raised it. until that day the own column was past in every row, which was what "your own failures only bubble" meant when written out; the licence that sentence named now lives at the rescue word, where a .? written in the package that raised the failure hands it on.
pattern own foreign
_ past past
v past past
v:err took took
(err r) took took
v:maybe took took
v:some past past
v:int past past
v:string past past
v:none past past
1 past past
took means the arm fired; past means the failure went by as though the arm were not written. two of the five disagreements were cells of this table. _:some took a foreign failure on two engines and was refused outright at build time on the third, which made an arm that names no err anywhere into an unmarked rescue. a typeset arm took its own package's failure on native alone, because the typeset branch of the emitter returned before it asked whether the annotation admitted err.
the other three were rendering, and rendering had a differential harness already—the project names it divergence-prone alongside float formatting and utf-8. the harness could not have caught any of them, because not one value in its corpus was a failure. it has eighteen now.
what the tables buy is narrow and worth saying plainly. a site that drifts on any engine goes red. a site nobody has thought about is a line missing from a file, which is a thing a reader can notice. and adding a pattern form to the language adds a row.
23who writes the plumbing
every language with tracked effects and carried failures meets the same fork: either the programmer writes the plumbing—bind, rescue, await, ?, a do-block—or the compiler does. kanso's answer took three decisions that arrived separately and only work together. each of them was nearly reversed at least once, and the argument for the whole has already been lost and reconstructed from the log once, so this entry is where it lives now.
the first: your own failures only bubble. as built on 2026-08-24, no arm could take an err its own hako raised—dispatch passed it as though the arm were not written—so the interior of a program had no use for a rescue keyword: there was nothing it would be allowed to catch. the 2026-09-15 ruling moved the line: a bare err is data and an arm names it anywhere, and what stays foreign-only is the rescue word, which hands on a failure its own package raised. this is also what cleared the road for the third decision. the standing objection to automatic lifting had been that it would coexist badly with hand-written local handling, and the first decision removes local handling from the language entirely.
the second: at a boundary, handling is dispatch. a foreign failure is a value that arrives at your arms, and an arm that names it—fn verdict (err e)—receives it the way any arm receives anything. a value arm converts; an err arm annotates. rescue did later become a word, and the arm is still what does the handling — the word routes a failure to a group and the group's arms decide, so what arrived was a way to say which channel you meant rather than a second way to handle. it is a library function and not a keyword.
the third decision was reversed on 2026-08-29, and section 32 has the ruling that replaced it. what it said: elaboration is signature-directed. the compiler tracks which values are plans, and the receiving position decides what happens—a position that demands a value gets the bind inserted, and a position that takes a plan receives the plan itself. retry (fetch url) hands retry the plan with no mark, because retry's signature says it takes one; parse (fetch url) becomes a plan whose parse step runs on the fetched body, because parse's signature says it takes the value. nothing runs at either call: plans only collapse when demand reaches them from io. the two classic objections dissolve under the rest of the language rather than under this feature. a plan you forgot to run is an unused value, and an unused value is a compile error. a plan that runs later than you intended cannot happen, because where order matters you wrote .>, and where you did not, there was no order to violate.
put together, that was a language in which the compiler wrote most of the plumbing: propagation threaded, effects elaborated by signature, and a surface of functions, arms, >>, adjacency and skip. three words were the exception, and they were the failure channel. bind, rescue and annotate became ordinary two-argument functions on 2026-08-26, for a reason this section did not anticipate: a program that can handle a failure two ways needs the explicit rescue either way, and once one of the three is spelled the other two may as well be. section 28 has the ruling and what building it turned up. the question of which machinery serves the failure half — the pattern matcher asking about provenance at every match, or the elaborator inserting the same calls it already inserts for effects — was overtaken: after section 32 there is no elaborator inserting anything, and the words are what a program writes. the book teaches the call-site half of this story in chapter 4, “nothing is asked of the signature”: the failing argument that never forces anything onto the function it was passed to. chapter 5 teaches the type, in “a box has a type”: <t>effect names a parameter that holds a box, and a box handed to something that reads the value inside is refused at check.
24a gate’s exceptions are claims
every diagnostic the compiler can raise is pinned by a golden. the error corpus is the only place a message’s exact text is checked, so an unpinned one can be reworded, weakened or lost with nothing going red, and a gate reads the compiler’s sources, collects the messages it finds, and fails on any that no golden covers. it was written after forty-one turned up unpinned at once.
the gate has two ways of letting a message past. both were wrong, and nothing in the tree could see either.
the first is a length floor. a message assembled from runtime values—format!("function `{name}` has no body")—offers only its leading literal run to match on, the text before the first interpolation, and that run has to be specific enough that finding it in the corpus means finding this message. the floor was twelve characters. three messages have ten: function , constant and the name , each counted with its opening backtick. each is raised by a program anybody could write—a fn line with nothing under it, a name = with nothing under it, a type called err—and none of them had a golden. the gate counted eighty-one diagnostics and reported the corpus complete. at ten it counts eighty-five, and the four sites it gains are the three gaps and one message a golden already covered.
the second is a list of messages excused as unreachable, each carrying a paragraph on which check speaks first. those paragraphs are readings of the source, and one had already been wrong once. writing the program that tests such a claim costs a few minutes, so each remaining claim got one. the claim that an indented line at the top level is always taken by the indentation rule or the blank-line rule holds wherever a declaration precedes the line. put the indented line first in the file and nothing precedes it: it reaches the top-level loop with a non-zero indent and is refused there by name. two other claims held. the fourth described a check that could not fire at all—both its loops tested for a leading underscore, and the lexer had stopped allowing those—so the check had been compiled into every run of the front end ever since, unable to speak.
the third way past the gate was not an exception at all. the gate finds a message by looking for Diagnostic::new( and reading the string that follows. the loader and the driver do not build a diagnostic — they write error: … as plain text, print it and exit — and there are thirty-one such sites in src/. every one was outside the scan. these are the messages a user meets before anything has been compiled: a module that cannot be found, an import cycle, a directory with no kanso in it, a build whose name collides with a directory beside it. the scan reads both openers now, and counts ninety-eight.
widening it walked into a trap. six of the fourteen new messages read as already pinned and four of those six were false. the corpus is the error corpus plus tests/*.rs, searched by substring, and a substring cannot tell an assertion about a message from a mention of the same words. no .kso files in matched the golden harness’s own assert. clang failed on matched a doc comment. cannot write matched the oracle’s unrelated refusal. cannot execute matched a panic the wasm spec writes for itself. all four run long, fourteen to sixteen characters, so no length floor could have caught them. what they share is the file they matched in.
so the two openers search different corpora. a diagnostic’s text is pinned by a .stderr file, or for the handful a corpus of single programs cannot express, by a rust test — tests/*.rs belongs in its corpus. for the driver’s messages that same file set supplied four false pins and no true ones, and its corpus is the error corpus plus the module differential, the surface that can put a tree on disk, which is what a loader refusal needs. a driver message pinned only by a rust test goes on the excused list naming that test.
after the split, six of the fourteen have a real pin. two were already covered by the module differential, the surface that puts a tree on disk, which is where the loader's refusals had been pinned by hand a few days earlier precisely because this gate could not see them. four are new: a fixture for a module that moved out from under an import, a differential case for a directory with no kanso in it, and two for the self-import guard.
the other eight are excused, and each excuse records what was tried rather than what the source appears to say. two are pinned by a rust test, which the driver's corpus excludes, so the excuse names the test. five fire on an io error from the host, which a container running as root with clang installed cannot produce. one fires when clang rejects what the compiler emitted, which is a compiler defect rather than a program anyone can write.
a ninth was going to join them. three attempts to reach the self-import guard had each been taken by a different check first, and the honest excuse looked like "unreachable, or i have not found the shape". asking a fourth time was cheaper than writing that: the guard tests a flag that is set around the whole of an entry compile, dependencies included, so no kanso run reaches it whatever the shape. checking the module root does. both arms answer there, and both are cases now.
then a third way of writing an error, and the one that hid longest. error[kind]: … is what a rendered diagnostic looks like on a terminal, so a message written that way as plain text reads to a user exactly like one the corpus pins. twenty-odd sites: the runtime’s endpoints, the stack-depth refusal, the exit-code refusals, the repl’s name lookups. keyed first on the constructor and then on error: , the scan saw none of them. with the third opener it counts a hundred and eight.
this family reads the wide corpus where the driver’s does not, and that difference was measured rather than assumed. the driver’s four false pins were short generic phrases a rust test can hold for a hundred unrelated reasons. an error[kind]: string is a rendered diagnostic, so a test holding one is asserting output. all three that matched here were checked by hand and all three were true.
four had no pin at all. two do now, in tests, because the corpus each belongs to cannot hold it: the repl’s nothing named has no corpus to live in, and the signal a killed program is named by cannot go in the runtime corpus, which asserts that both engines write identical stderr and both exit 1 — a signalled program does neither. the third was already pinned by a rust test, and its excuse names the test. the fourth is excused as unreachable on unix, and that claim is about control flow rather than about what looks unlikely: the driver calls the function only when the exit status carries no code, which on unix means a signal ended the process, which means the branch inside it that asks for a signal and finds none cannot be taken.
one message no opener can ever see, and a fix the counters refused. the scan matches on the leading literal run, so a message that opens with an interpolation has none, whatever openers get added. exactly one in the compiler is in that position — kanso test on a file declaring none answers with the file name first — and the same opening makes it the only driver refusal a reader cannot recognise as one, since every other starts error: . spelling it the same way fixes both, and was built and measured: it moves compile_instructions by 167 on the runner and by −37 on another host, which is layout rather than work, and the trend gate reads it as a pure regression — one counter worse and nothing better, the one trade it refuses outright. the counters cannot see message consistency, and a change whose whole gain is invisible to them does not get to spend them. so it is declined here and written down, rather than argued past the gate.
an exception in a gate is a claim about the program, and a claim here gets a program written to break it.
25what a green suite is evidence of
four bugs in one day shared a shape, and none of them was in code nobody ran. every one sat in a function the corpus exercises on almost every test: os/exit, io/write_err, the wall. what was wrong in each case was the endpoint around the call — what the exit code carries, which stream the bytes land on, which sentence names the fault, which side of a wall was judged — and in each case a golden pinned the FAILURE and nothing pinned the success.
os/exit is the clearest. one fixture in the tree touches it, and it hands the function something that is not a status, so the corpus pins the refusal. what a deliberate exit does — carry its code out and say nothing — was pinned on no engine, and the page’s endpoint had never learned to read one. a program calling os/exit 3 printed unhandled err reached the executor at its reader and answered 1. three of the four endpoints already knew better; the fourth was the only one nobody had written a program for.
the corpus walk would have caught it on the first run. it compares text and exit code against native for every program in three directories, and so does the browser harness. there was no program to run. that is the whole of the explanation, and it generalises: a corpus is evidence about the programs it contains and about nothing else.
io/write_err is the same shape one turn quieter. its fixture writes to the diagnostic stream between two writes to the ordinary one, and carries a golden for stdout only. the stdout golden does prove the diagnostic line never went to stdout, and the differential does prove the engines agree on the two streams concatenated. neither holds the bytes. drop the write on all four writers at once and both facts still hold. agreement between engines is a weaker thing than a pin, and the difference only shows up in the case where every engine is wrong together.
the reflex after three of these was to gate the other direction. one harness already walks every pub fn in the library and hands each one arguments of the wrong type, on all three engines; nothing walks the surface asking what a correct call does. finding the exports no program calls looked like the obvious next build. there are none — a hundred of a hundred are reached. the first answer was twenty-four.
qualified name only, across the corpus dirs 24 uncovered
plus the bare-enrolled form 1 uncovered
plus intra-library calls 0 uncovered
eight of the nine survivors of the second pass are written by their bare name after an import "std/list", which is how a program normally writes them. the ninth runs on every string a json encode touches. a name here has two written forms, and the twelve-character floor of the previous section was the same mistake in a different dress: both searches were measuring their own reach rather than the tree.
the fourth bug is where a green suite failed most completely. a bare name on either side of a wall reaches the run, because the check refuses literals and direct calls and a name is neither. teaching it about names is one line — the fixpoint answers a name at arity zero exactly as it answers a call — and that line refuses a working program, because a local may shadow a bare-enrolled import and the fixpoint’s row is the import’s. so the check takes the set of names the declaration binds, and asks the fixpoint only about names nothing nearer owns.
that walk shipped with three of the four binding forms. Lambda binds its parameters, Block and Build bind through their statements, and Guard carries a statement list of its own, so a binding written after a return line was invisible and its program was refused. seventy-three test binaries passed over it. no fixture in the tree binds a shadowing name inside a guard, so no amount of running the suite could have said anything about the case.
reading the walk against the Expr enum said it in a minute. the enum is a finite list of the things that bind, and checking a collector against that list is a question with an answer, where running a suite is a sample. the suite is worth more on almost every other question; on this one it was worth nothing, and it was green.
a golden earns its place by pinning text rather than outcome, and the same day showed why. removing the arm that refuses a bare name does not make the fixture’s program compile — it fails later, at run time, with a different sentence. a spec asserting that the program failed would have stayed green through the whole change.
26the sentences nobody was reading
the previous section ends on a corpus being evidence about the programs it contains. the next week said the same thing about the diagnostics themselves. a gate walks the tree collecting every error message the compiler can write and fails when one of them is pinned by nothing, and it knew four ways to write one. there are six. a RuntimeError in the interpreter is a struct literal with a bare message: field, matching none of the four, so all ninety-seven of the interpreter’s refusals could have been reworded, weakened or deleted with nothing going red — on the engine the differential law calls the oracle.
teaching it the two missing forms took the count from 175 literal diagnostics to 242. fourteen of the new ones were pinned by nothing at all. each was run rather than read: four earned fixtures, one was already pinned by another day’s work, one turned out to be a divergence, and seven are shadowed by a check that answers first.
those seven are the interesting ones, because each excuse is a claim about why a program cannot reach an arm, and a claim is a thing that can be wrong. one said a non-bool filter predicate is refused by list/select. it is not: select is a fold written in kanso and never calls the builtin, and what actually stops the arm is that filter cannot be named at all — it is not one of the fifty-five builtins the checker knows, and the hand-written internal spelling is refused by name. the same wrong mechanism had been written twice, against the two engines’ two spellings of the sentence, and the second copy was found only by re-running the probe that was supposed to have established the first.
the same sweep over the browser engine found the other half. twelve of the page’s twenty-five refusals were pinned by nothing, and one of them was a sentence worth reading twice: cannot read fields of this value, sitting behind a guard that fires exactly when the handle is not a value. the two engines that could be asked said <fn> and <io>. writing the fix produced the second half of the bug: routing everything through the value accessor named the closure correctly and broke the description, which is a deferred shape rather than a value, so the first draft traded one divergence for another.
then the sharper case. a description put into a list, a map, a builtin call or a record field is data on native and on the interpreter, and the page refused it:
d = io/write "a\n" .> (_ -> io/write "b\n")
xs = [(opaque d)]
pub play = xs[1]
native, interpreter a
b
the page error[runtime]: a bound description cannot be
used as data here
four sites read their elements through an accessor that answers a value for a value and a handle for a closure and refuses everything else, and all four are storing positions. the fix is four calls to a function already in the file. what makes this different from a wording mismatch is that the program succeeds on the oracle: the same source, run three ways, prints two lines twice and dies once.
the reason it went unseen for so long is worth stating plainly. the corpus test that compares engines on these programs runs native and the interpreter. a second walk covers the browser, and nothing in the corpus had ever put a description in a container. every instrument was working; none of them had been pointed at the case. that is the same finding as the section above, one level further in — there the missing thing was a program, here it is a program nobody thought to write because all three engines were assumed to agree about it.
one of the four fixtures pins something none of the others can. materializing a deferred description demands the right side of its sequence, so a fix that performed the effects at construction time instead of when the program asks for them would leave the other three green. the fixture prints a word before the description runs, and the word has to come first on all three engines.
27the answer, not the sentence
the section above ends on four storing sites and a fix that was four calls. the count was wrong, and how it was wrong is the more useful half. following the failing program found four. starting from the accessor instead — reading each of its call sites and asking what a program could hand it — found six, and then two more families the first pass had no reason to look for.
the first is the mirror of the storing one. a storing site carries a description onward and used to refuse it. a refusing site is meant to refuse, and the accessor refused first, in its own words: if d said a bound description cannot be used as data here where the other two engines say an if condition is true or false, got <io>. that sentence had been converged across all three engines the same morning it was found unreachable on one of them.
then the one that was not a sentence at all:
d = io/write "a\n" .> (_ -> io/write "b\n")
pub play = print "{1 + opaque d}"
native, interpreter error[runtime]: `+` is not defined
for these values
the page <io>
the operator read each side, found a slot that was not a plain value, and handed the operand’s own handle back. 1 + d evaluated to d. every gate this project has for diagnostics compares refusals, and this program does not refuse; the error corpus pins what a program writes to stderr, and this one writes to stdout. the walk that compares what each engine prints is the only instrument that could have caught it.
the fix hands the interpreter a placeholder description and lets it write the operator’s own words. that is sound because no operator succeeds on a description — equality was the one that could have, and it names them instead — and exact because every description renders <io>, so an error path that reports its operand cannot tell the placeholder from the real thing. three fixtures, one per arm a description can reach.
the placeholder has a boundary, and the two sides of it look identical at the call site. a site that refuses wants it, because building the real description would force the right side of a sequence and native does not evaluate one before refusing. a site that carries wants the real thing, because a placeholder handed back to the program is an effect the program did not write. err d is the carrying case: it runs to the endpoint on both other engines and died on the page. the two fixes now sit a few lines apart in the same file with the distinction written between them.
one more thing had to change before any of this could be watched. the gate that collects every error message the compiler can write had six openers by then, and none of them matched the browser engine — thirty-six sites spelling twenty-four sentences, on the engine a reader meets first, since the playground is what the website runs. a seventh and eighth opener took the count from 242 to 262. the comment claiming the gate already watched two hand-copied sentences for drift was written the day before, and was wrong when it was written.
the accessor’s own sentence has no program left that reaches it. it was reachable eight ways this week — an operator, a condition, four index forms, an err, two field reads, a destructure — and each is now answered by the site that owns the words, with the program that proved it in the corpus. it stays in the source, listed as unpinned with its reasons, because the arm is what makes the accessor total.
28and then bind became a word
section 23 used to be called “why there is no bind”. the argument it makes still holds: the compiler threads the failure channel, the programmer does not write the plumbing, and a chain reads as a stack of steps. what changed is the spelling, and the reason is worth keeping because the ruling went past what was recommended.
the recommendation kept the chain err-arm as the surface for annotating a failure — an arm whose shape told the compiler it was the failure branch. clay killed it in one sentence: “since you want to be able to handle an error in two different ways, you need the explicit rescue either way — and that kind of makes me think we should just be symmetrical and have all three forms explicit.” then, immediately after: “i think bind should stay parallel with rescue and annotate. be consistent.”
io/read_file path
bind (text -> json/parse text)
annotate (e -> "config: {e.reason}")
rescue when_failed
three words, and nothing about them is syntax. they are ordinary two-argument functions — bind effect callback — and a chain line that spells only the callback is the chain rule already in the language supplying the first argument, the same rule that makes (expect 1) . to (equal x) feed expect 1 into to. they can be called prefix-style outside a chain and mean the same thing. the dot retires from chain-step position; field access was never the monad and is untouched.
the part that pays for the extra verbiage is that the callback receives the err itself rather than an unwrapped reason. a dispatch group is therefore a legal callback, and its arms match reason types polymorphically, subtype rung and all — rescue when_failed is the whole generic-rescuer story, with nothing written specially for it. two more properties fall out of the spelling rather than being checked: annotate cannot resurrect, because its result is always re-wrapped as err with the original as cause whatever the callback returns; and rescue is the sole door, so the foreign-only license is enforced at the word instead of by enumerating sites.
the failure channel is now fully spelled and never inferred from an arm’s shape. that is the trade: a chain is longer to write and there is nothing left to deduce about which branch is which.
all three words work on all three engines as of 2026-08-29. getting there turned up a divergence: the page had been compiling rescue and then propagating the very failure the other two engines catch. nothing in the tree could see that: every fixture for these words needs the world to refuse something so an effect can fail, and a page has no world. the case that found it reaches the words through a subject that has already failed, which needs no filesystem and so runs everywhere. in chain position each word has had a fused spelling since 2026-09-09: .> is bind, .! is annotate, .? is rescue, the chain dot plus one character of channel, and the right-hand side is one function — a lambda, a name or a group. x .> f is bind x f on every input, and it desugars to the piped step the automatic bind has always been, so the beat and every other pass read it as they did; desugaring it to a call of the word instead lost the beat, and the decoder’s arena peak went from 2 MB to 251 MB. .! and .? desugar to the words. a . bind (f) step is refused with the fused form named, and the words stay prefix functions everywhere else. every step the compiler bound automatically is spelled .> now: infer’s own piped-over-description branch named 519 sites across the programs the sweeps compile, a script rewrote them, and a grep for a dot followed by a lambda or a name, each file it named compiled under the same instrument, found 44 more in the benchmark sources and samples no sweep reaches, 56 in the programs the Rust specs carry as strings, and one in the playground’s dice example. 599 dots in 132 files, and the grep census reads zero. a plain-dot step over a value keeps its dot, since under the effect type it stays an application and only a step over a box changes meaning.
since 2026-09-10 the plain dot opens nothing. x . f a is f x a, an ordinary application: a box handed through it arrives as a box, so held in math/random 6 . held receives the description itself and an interpolation of it renders <io>. holding is not opening, so a parameter that binds anything takes a box, and so do print, push, put and the words. what is refused, at check, is a provable box handed to something that reads the value inside — an operator, an index, a field read, if’s condition, a builtin that takes values, or a group none of whose arms binds anything at that position — with one sentence naming the reader and .> as the door. provable means a group whose joined return set is the description bit alone, a .> step over such a subject, a wall, or a constant holding one; a box that reaches a reader through a list or a lambda’s parameter is the runtime’s sentence, as it was. the book’s chapters are what the ruling still owes.
the type itself is spellable since the same day: <int>effect, the yield in front and effect as the head, written tight because < anywhere else is the comparison. a parameter declared e:<int>effect takes the box as data, and a box is a box whatever it will yield, so the annotation matches a description on every engine and two effect types are one shape to dispatch. bare effect is refused with the spelling offered, a yield in front of any other head is refused with the slice spelling offered, and the yield is checked as a type. nothing in the tree spelled it, so no golden moved.
building them turned up something the sizing had missed. rescue is new capability rather than a respelling, because an execution-time failure could not be handled inside a chain at all. a chain step over an effect wraps its expression in a one-parameter closure, and handing a failure to a closure returns the failure instead of entering the body — so the callback was never called, whether it was a lambda or a group with an arm ready for exactly this. chapter 05 of the book teaches that as the design, with a sample whose expected output is the endpoint message. so the err-arm half of the migration is empty for a stronger reason than a search of the fleet gives: the chain err-arm was never a working surface.
two smaller things the sample above still needs. an err had no .reason reader, and the argument against giving it one was that every operation on an err propagates it, so a callback holding one could not look at it. that was ruled the other way on 2026-08-29: an err gains .reason, .cause and .origin, and reading one is a second deliberate hole in infectiousness beside the one wrap_err’s second argument already has. a group callback still destructures. and bind, rescue and annotate are reserved names now: two fixtures in the corpus had defined functions called rescue and annotate, which is a fair sign of how ordinary the words are, and they were renamed.
29one home for a rule, two locks on the door
how many arguments does length take? the answer is one, and on 2026-08-29 the compiler gave four. write length x x inside a function nothing calls and the interpreter printed the program’s output and exited zero, the native backend refused the whole program with no span, the page compiled it and ran it, and kanso check said ok. a two-word program, print (wrap_err 1), did worse: the native backend indexed the second argument of a one-argument call and the process aborted with a rust backtrace.
none of that is exotic. a user function called with the wrong count has been refused at the site, with a span, identically on every engine, for as long as the arity walk has existed. builtins fell through it because nothing in the front end knew what one takes — and the counts existed the whole time, as a second field on the native backend’s emit list. that table answers a different question, which builtins get a direct c call rather than an inline expansion, and carrying the counts too is precisely what let the backend refuse a call the front door had waved through.
the counts live in the front end now and both backends read them there. that is the small half. the general half is that this was the third instance of one shape in eight days. the val accessor in the page’s runtime had a rule about what it answers and four sites spelled the rule out again instead of asking it. rt_maybe_bind kept its own inline list of the deferred shapes where everything else called the predicate, so the day the gavel added two shapes the page started handing a description to a callback as data. a module’s name printed as the resolved file on one engine and as the import path on another, from one line that formatted whatever it was handed.
a fact with more than one home has as many answers as homes, and the extra homes are the ones nobody maintains. how they got there is the less obvious half. not one of them was written as a copy. each was written as a local detail at a site that needed it, at a time when there was nothing to be a copy of, and became a duplicate later when somebody wrote the abstraction. the abstraction is what creates the duplicate, retroactively, and nothing points at the older code.
and then the opposite lesson, on the same day and the same rule. once the front end checks a builtin’s count, the backend’s own check looks like exactly the kind of second copy this entry is about. it is not. switch the front-end check off and print (wrap_err 1) still aborts, because three builtins are emitted inline and read their arguments by index before any guard runs. so the backend keeps a check, and it now covers every builtin with a count instead of only the ones it emits directly.
the distinction is between knowing a fact twice and refusing to proceed without it twice. the counts have one home; two places consult that home before doing something an absent count would make unsafe. a compiler that indexes an argument it never counted is one front-end regression away from aborting on somebody’s program, and the front-end check was three hours old.
the second lock was wrong on its first attempt. placed where the argument list first exists, it ran before the rule that says a declaration of the same name is that declaration — and the standard library declares its own bytes, with three parameters. a defence added a few lines too early in a long function turns a working program away. the only thing that caught it was a test that runs a real program and reads what it printed.
30the intersection nobody covered
a mutation sweep over the corpus — 8,724 single-token mutants across 490 programs, through both kanso check and kanso build — found nothing. that was worth knowing for a reason other than the zero: the two bugs it was built to catch live in shapes no program in the corpus has, and dropping or repeating a token cannot synthesise a shape. mutation explores the neighbourhood of what is written. so the corpus is the thing to grow, and the question becomes which shape is missing.
that is a claim about a number nobody had, so the next step was to count. every variant of the expression, pattern and statement types, over the 298 programs the playground runs, written to a golden the way the cost veins are. the answer: no construct is missing, and two are thin. the widening upcast (expr):type is carried by one program. blocks by two.
within the hour the thinness paid. the one program carrying the upcast runs definitions beside statements in a single file, so it is never reached as a module — and the pass that qualifies a module’s own names when somebody imports it had therefore never met an upcast. it renames a constructor pattern’s type, an annotation’s type, and every mention of a name the module owns. it descended into an upcast’s sub-expression and left the target name alone.
so inside a module, widening to a type that module declares held a spelling that no longer named anything, and the three engines said three things: the interpreter reported at runtime that a dog is not an animal while holding one, and both backends refused the whole module with unknown type. written module-qualified it worked everywhere, and so did widening to a builtin like int — which is what the one carrier does, and why it never showed.
seven fixtures in the corpus already exercise subtypes, and one of them carries a second golden for what it prints when it is reached as a module, showing that very rename in its output. none of them upcasts. the construct had coverage and the path had coverage; their intersection had none, and a per-construct count is what made that legible rather than a thing somebody might notice.
the census was wrong four times before it was right, and each wrong version produced a confident number. the walk over expressions reported that no program assigns a field inside a build block, because the child-walk hands such a statement to its caller as the value being assigned and its statement-hood is gone. seventeen programs were skipped by a let Ok(..) else { continue } — they parse through a different door — and the upcast’s one carrier was among them, which is how the first run reported the construct as carried by nothing at all, correctly, for the wrong reason. the walk recursed where the differential harness does not, counting two programs it never runs.
what the four have in common is that the instrument was built beside the thing it measures instead of out of it. every divergence between the copy and the original arrived looking like a fact about the language. the census lives in the differential harness now and calls that harness’s own corpus walk and gap list, and it fails on a program it cannot read rather than dropping it.
the golden moved by one line in the pull request that fixed the bug: the upcast went from one carrier to two, and the file names both. a construct that falls to no carriers reads NOBODY there, which is the failure the file exists for — the last program using a construct leaving the corpus, and the three engines quietly ceasing to be compared on it.
31the benchmark nobody had profiled
the welfare index averages each counter’s ratio against its baseline and then applies the saturating curve to that average. the curve therefore reaches no counter. a benchmark 138 times better than its baseline enters the mean linearly and unbounded, so it carries the term; the pretty-printing benchmark is 68% of run speed by itself, and the decode and encode rows the front page makes claims about are 1% between them. whether the curve belongs inside the average or outside it was ruled on 2026-08-29: each counter’s ratio is saturated first, and the equal-weighted average is taken over the saturated terms. a counter that has run away can then contribute at most its own share, and the two measured improvements that were held against the answer were re-scored under it, and both ship. the byte builder’s grow path was written twice, once taking arena storage and once malloc, differing in the allocator and the sign of a cap; one tail serves both, and encodebench falls 0.49%. the growth is then outlined, because one function holding both it and the fast path made every append push six callee-saved registers and build a frame before it could test anything — split, the common path pushes three, and encodebench falls another 3.06%. each cost welfare under the old order and pays under the new one, which is the difference the ruling was about.
what the arithmetic said, though, is that the row with the leverage had never been looked at. it had been in the tree for months as a number that went up or down; nobody had asked where its time went. eleven per cent of it was inside glibc, formatting integers.
a double that is a whole number renders as its digits and .0. the branch that handles it is guarded by a test that has already proved the number is an integer below 1e15, which is exactly the condition under which the cast to a 64-bit integer is exact — and it called snprintf with %.1f anyway. %f reaches a multiprecision routine that also drags in the arbitrary-precision multiply and divide. writing the digits directly is nineteen per cent of that benchmark.
the second was next to it. the utf-8 validator is vectorised, and the decoder calls it once per token: 265,950 calls carrying 41 bytes each, at 446 instructions a call. seven constant loads and two zeroed accumulators precede the first block, which a 41-byte token never amortises. reading eight bytes at a time answers the whole question for an ascii run of any length and never reaches the vector pass.
neither is a clever algorithm. both are a hot path paying for machinery sized for a case it does not have, and both were invisible to every counter the project keeps, because the allocation counts are byte-identical across them. the instruction vein saw them; nothing else could have.
the same week produced a bug of the opposite kind. reading a character by position remembers where the walk stopped so the next read resumes, and the remembered position names its string by address. a loop rewinds the arena, and the next string it builds can land where the last one sat — so the cursor pointed into a string that was no longer there, and native answered a character the interpreter did not. one golden pinned the shape it happened to take. the sweep that now runs asks the whole family: every pair of byte layouts, alternating between iterations, read by every hand that moves the cursor. sixteen of ninety-six cases disagree when the fix is removed, some answering a byte from the middle of a codepoint.
that is the difference between a golden and a sweep, and it is worth being explicit about. a golden pins what somebody thought to write down. a sweep enumerates a space and asks the two engines to agree across all of it, which is the only way the fifteen cases nobody would have written get asked at all.
32an effect is a type you can holdruled, not built
section 23 records a design with three decisions in it, and section 28 records the respelling that replaced part of the third. this is where the rest of it went. the sitting started at koka and ended with clay writing one sentence about consistency: “it might be ‘convenient’ to have bind be automatic and just not allow passing effects, but it’s inconsistent and threatens to make the language confusing, when our overarching goal is simplicity.”
<int>effect is a type. it is the unresolved outcome of an operation — will be an int, or a failure — and it behaves the way every other parameterized type in the language behaves. it can be bound to a name, passed to a function, stored, and received: a parameter declared e:<config>effect takes the box as data. holding the box is a different act from opening it, and a helper that annotates somebody else’s effect and hands it back is an ordinary function.
the three words are the only openers, and they are always written. bind, annotate and rescue take the effect first and a callback second, and there is no route from box to branchable value that avoids one of them. so signature-directed elaboration goes: passing a <text>effect into a position that wants a text is refused at compile time. propagation moves with it, from the call site to bind’s own contract — a failed effect handed to bind skips the callback and comes back out failed — and the err-in-err-out railway that used to run at every ordinary call retires with the sugar that implied it.
the derivation clay gave for allowing the box to be passed at all is the part worth keeping. with the combinators alone “we’d have Effect types that are basically indistinguishable. to get anything you can match/branch on, you’d need to call bind/annotate/rescue to get a type… but one might argue you should be able to match on something like <int>effect.” a language that has the type and refuses to let you name it is hiding a thing it already has. so the parameter spelling is admitted, and it opens no unmarked door, because it is an explicit statement that the author is holding the channel.
two boundaries stop the type acquiring a second meaning. the box answers “did it work”; a deferred description answers “has it run”. keeping those apart is what makes holding a result safe — a held box never delays work that was going to happen. and no arm may match a failed <int>effect against a succeeded one. an arm that could tell them apart would be a third eliminator with none of the ceremony the other three carry, arriving by the back door of the pattern matcher.
none of this reaches a signature. clay’s rider: “effect arguments should work just like any other type. no superfluous type annotations on functions. if you pass an effect to foo, and it does rescue or whatever, and it passes the value to a matching arm, then it Just Works.” a parameter that receives an effect is typed by inference like any parameter, and dispatch matches an effect-typed value like any value. the e:<config>effect spelling is there for an author who wants to state the shape, never as a toll.
what the ruling leaves owing. chapter 4 of the book taught the retired railway under the title “nothing is asked of the signature” — the failing argument that forces nothing onto the function it was passed to — and was re-premised on the explicit box when part 3 landed: the section asks its question at the call now, and says that a call the checker cannot see into is the case it cannot prove. the provenance hop a function used to accrue at a skipped call now accrues at a bind, which moves the ground under the eta-reduction argument and has not been re-run on the new ground. one thing that looked owed is not: what becomes of an effect nobody eliminates was raised as a hole and ruled backwards. an effect is a value in hand, the language already refuses an unused value, so a dropped effect is unspellable under discipline that predates the ruling — and unspellable visibly, where the railway’s version of the same guarantee was machinery a reader could not see.
33the directory a file sits in
a digest of eighty-eight kilobytes held eight hundred and forty-nine megabytes of resident memory. sha256 is a streaming algorithm — sixty-four schedule words, eight state words, everything a block builds dead the moment the next one starts — so the peak should be flat however long the message is. it was linear, and the allocated bytes and the peak agreed to within one per cent at every size, which is the whole diagnosis: nothing was reclaimed between blocks, so the peak was the total.
the cause was not in sha256. copy lib/sha256/sha256.kso into a program’s own directory and import it as ./sha256 rather than std/sha256, and the same bytes read 1,048,576 arena bytes where the import read 90,177,536. the emitted code says what differs: the local build emits four carry rewinds and the imported one emits none. same nineteen functions, same emitted names.
the beat analysis keeps imported groups out of the carry tier, and it decided which were imported by asking whether the declaration’s source path began std/ or lib/. that field is the one error origins are built from — a diagnostic, read as a semantic marker, by string prefix. so the same sources built from a directory called lib read 90,177,536 where any other directory read 1,048,576, and a program’s memory behaviour depended on where its author kept it.
five models of the shape were built before the bisection found it, and every one reclaimed: a two-function trampoline, a loop carrying a list rather than a scalar, an outer walk around an inner sixty-four-round loop, that shape with lists on both sides, and one long list grown by repeated push. each of them is the digest without the import, which is what makes them evidence rather than wasted work.
removing the exclusion turns the peak flat — 108,003,328 bytes to 1,048,576 at ten kilobytes — for four hundredths of a per cent more allocations, and costs eleven emitted lines in one benchmark. it is also refused, and by the right thing. a test in the analysis itself pins which loops may rewind, and its argument is not about cost: the two encoders it admits thread raw bytes, which hold no pointers, so nothing in the accumulator can dangle across a rewind. a scanner threading a list of records is on the other side of that line. the claim that no evidence for the rule existed was made here and was wrong — the search went through the log and the archive and not the test suite, and a pin lives in code.
with the pin read rather than assumed, the exclusion turned out not to be enforcing it. removing the prefix test admits four groups, all in std/list: three boolean predicates and a binary search, none of which accumulates. every group the json library itself declares is licensed identically either way. so the rule was buying those four predicates and the digest’s own cluster, and nothing else in the tree.
and then the removal was built and timed, and the trade is the wrong way round. the digest is linear in time today and quadratic with the carry: at a hundred and twenty-eight kilobytes the peak falls from 1,262,485,520 bytes to 4,194,320 and the wall clock rises from 1.3 seconds to 68. ci said it before the table did — the job that digests a 1.6 megabyte blob takes forty-nine seconds on main and sat past twenty-five minutes on the branch. so the exclusion stays, the directory defect is pinned with both its numbers rather than fixed, and the open question is where the quadratic lives: every deterministic counter doubles when the message doubles, so the cost is in work nothing watches.
three of the four timings that first said this change was free were measured wrong, and how is worth writing down. kanso build caches the native binary against the sources and the runtime, not against the compiler that emitted it. rebuilding the rust and re-running in the same directory re-runs the same binary. the two columns agreed because they were one column.
the quadratic was in the deciding, not the copying. the carry evacuation asks whether each list it might share lies wholly below the mark, and on an eight kilobyte digest it asked 33,024 times, walked 8,256 slots each time, and found no heap slot on any of them — 272,646,144 slot examinations to reach the same answer, against 33,827 allocations that actually evacuated. what the evacuation copied was linear in the message the whole time.
that made the change buildable, and it was built and swept: both arms in their own directories, every benchmark binary deleted first, the main arm reproducing all six allocation goldens byte for byte. the digest peak falls from 54,525,952 to 1,048,576 and its arena blocks from 52 to 1; digestbench retires 820,087,049 instructions where it retired 152,573,220. everything else in every vein moves by less than a tenth of a per cent.
and the two halves turn out to defeat each other. a lazy binding whose result the carry does not carry is exactly what the rewind reclaims, so the digest’s thunk evaluations go back to 8,256 — one per force, which is the state before the memo landed. the 52x memory win is bought by discarding the memos, and recomputing them is where the work goes.
welfare could not see any of that until it read the digest’s counters, so it does now, on both sides: retired instructions in run speed, peak bytes and arena blocks in run memory. pricing only the peak would rank any change that reclaims per block above one that does not however long it takes, which is the trade this section spent a day getting wrong in one direction and would then have got wrong in the other. with both sides priced the index reads 74.31 against 73.75 and the carry tier is declined by the model rather than by whoever measured it.
one thing about that answer is worth knowing before it is quoted. run both arms again with the digest counters’ baselines set to their own current values, and the verdict reverses: 70.14 against 72.99. a counter new to the index enters where its dimension already stands, so that it does not move the score on the day it lands — and on a saturating curve, where a counter enters also decides how much any later change to it can be worth. the decline stands on the model as written, and the reference is a separate argument that does not get made while a particular change hangs on it.
34the silicon a row was counted on
four of kq’s instruction rows moved between two ci runs. both job logs printed the same rustc down to the commit hash, the same llvm, the same runner image, the same glibc, the same valgrind. the one field that differed was the azure region.
glibc resolves memcpy, memcmp, strlen and their neighbours by ifunc at load time, reading cpu features, so one libc runs different code on different cpus. that is measurable from inside a single host. set rep_movsb_threshold to the 0x840 a runner reports rather than the 0x2000 an intel default gives, and one kq row retires 77,523,061 instructions where it retired 76,742,736 — 1.02%, byte-identical over two sittings, from one switch. the runner shift that started the investigation was 0.06% to 0.10%.
the first fix was a pin: record one host’s feature block, refuse anywhere it does not match, the way the toolchain gate refuses a moved glibc. ci killed it in two runs. the first refused and printed an amd epyc zen 3 milan, the second an intel ice lake-sp, the third an amd genoa; the container all of this was written in is a cascade lake. four cpus in four runs. a check that refuses every run but one is red for a reason no pull request causes, which is a gate nobody can act on.
on that third runner every kq row matched exactly, which qualifies the story rather than undoing it. most cpus land on the same memcpy and count the same, and now and then one lands elsewhere and does not. the pool’s heterogeneity is survivable, which is the difference between a vein that reports sometimes and one that is unusable.
so the script that shipped only reports. scripts/gates/dispatch.sh name prints this host’s cpu family and model on every run, which turns the next divergence into one line of a job log rather than an afternoon of version archaeology. differs answers 0 for the recorded silicon, 1 for other silicon with the differing lines named, and 2 when there is nothing to compare against. the instruction gates ask only about a row that already moved, because a row landing on its recorded value is right whatever counted it.
the first wiring exited green on answer 1, and that was wrong for a reason arithmetic settles. a block records one cpu out of at least four, so roughly three runs in four land elsewhere, and on those runs any moved row — a real regression included — drew a warning and a green tick. two of the ratchet’s mutations exist to redden exactly these gates; applied on a run that landed off the recorded cpu they would have proved nothing, and a row that proves nothing is the one failure the ratchet was built to catch. both gates now fail on a moved row whatever counted it, and the dispatch diff is a named cause printed beside the failure. whether the silicon accounts for a move is settled by a person in a pull request, with a re-run on the recorded cpu as the evidence.
reporting was enough for a vein whose rows agree across most of the pool. it was not enough for the front end’s own instruction count. a pull request that changed two documentation files and nothing the compiler reads moved that row 5,081, and four readings recorded earlier the same day as layout effects — +6,087, -416, +8,032, -3,289 — are two of them smaller than that and two the same size. nothing in those measurements separated the edit from the runner, so the row could not be read at all.
a tolerance was the cheap repair and it is the wrong one. held on one chip, with the binary’s sha256 printed beside both reads, the beat rewind split in §36 moves the same row -5,621. the noise and the signal are the same size here, so a band wide enough to swallow one swallows the other.
so the noise got a key, and the key is gone again. for a fortnight bench/compile_instructions_by_cpu.txt held one value per chip, named by family and model and nothing else, and a row was later allowed to pin two values, because one binary on one chip had been seen to count 41,831,767 and 41,832,275 with nothing to explain the 508. both retired on 2026-09-05. the 508 was two binaries read as one: the second value came off a build that had changed the compiler, and only two of the four rows had been re-sat. the drift the keying was built for turned out to be rust’s stack guard parsing /proc/self/maps at startup, which moves with where the linker happened to put things, and the gate counts kanso::main inclusive now and never sees it.
what stands in their place is one row, one value, compared exactly, in bench/compile_instructions_golden.txt — the file the welfare index, the trend gate and the prose gate already read. the evidence is eight within-binary sittings across two binaries, two vendors and four cpu generations, agreeing to the instruction, none disagreeing; the eight red ci rounds spent adding chip rows are what bought it. every move is attributed to the change under test and handled by the ordinary ratchet, like every other counter in the tree.
consistency is verified by reproduction rather than by keying: the same build on any runner counts the same number. a run that disagrees is a reproduction failure, and it halts the vein and is hunted to its source the way the /proc/self/maps term was — not written down beside the number it disagreed with, and not recorded as a mode. the gate prints the binary’s sha256 and a compile_sample line naming the cpu on every run, green or red, because that pair is where such a hunt starts: one sha counting two rows is the failure, and two shas is the change under test until the pair is built and both are read. tests/the_compile_row_holds_one_value.rs pins the rule the file has to keep, and a ratchet mutation writes a second value onto its line to watch it fail.
a gate that counts a command which WRITES has to earn that second reading. the codegen rows count kanso build, and a build leaves its output beside itself, so the second run of the identical command asked a different question of a box the first run had changed: five processes against six, and the dev row 9,273,832,919 and then 1,071,604,124 — an incremental build, reported as a reproduction failure about the compiler. re-staging the box between the readings fixed half of it. the other half was outside the box entirely: the runtime object is compiled once per profile and kept in the temp directory, the warm-up ran under the job’s environment and the measurement under an emptied one, and temp_dir() reads TMPDIR, so the two filled different caches and the first measured run paid for a compile the second found. the warm-up runs the identical command in the identical environment now. the general rule is the project’s own — external state is normalized before it is measured — and the specific lesson is that the profiles named the missing process all along, on a cmd: line nobody had read; the gate prints them by name now rather than counting them.
35the vein the objective leaves out
eleven deterministic veins feed the welfare index. a twelfth is measured on every run, pinned in a golden and diffed by ci, and weighed at nothing: bench/text_golden.txt, the size of each benchmark’s machine code. that looks exactly like the gap §33 closed for the digest, and giving code size a term is the obvious repair. a measurement says it would be wrong.
the list-index twin — the inline ir shim that answers xs[i] without a call, in the family §04 introduced — took encodebench’s .text down 144 bytes and its instruction count up 67,116,000, in one change, in a program that never indexes a list. inside d_list/fold_3 the shape is sharper: 4,083 bytes of machine code became 3,682, and the smaller function ran fourteen per cent more instructions. llvm had been specialising that dispatcher into four copies, and the twin’s inlined body pushed its caller past the cost where that paid, so one shared body with more branching replaced four specialised ones. forcing the twin back out of line with noinline restores 4,083 bytes exactly — the same number, not a similar one — and encodebench’s pre-twin count to the instruction.
a term rewarding smaller .text would have scored that regression as a gain, twice over. so code size is a diagnostic here. it says a kernel arrived or left, which is the job it did when the bit twins landed on the digest with every other row holding, and on this corpus it has been measured pointing the wrong way about what a program costs to run.
the exclusion is written as a spec rather than a comment, because four claims in this tree once rested on prose and none of them held. tests/the_objective_does_not_weigh_machine_code_size.rs goes red if the index starts reading the vein, and red the other way if the gate stops diffing it — leaving something out of the objective is not permission to stop counting it. both halves were watched failing.
that same index then declined the noinline repair. it buys encodebench 0.957% and costs digestbench 8.330%, deepbench 0.717%, basket 0.517% and scanbench 0.388%. the raw counts sum to 48 million fewer instructions and the score still falls, 75.09 to 75.03, because the benchmark paying most is the one the twin was built for. three of those five rows could not be seen at all until the same day: scanbench, escapebench and indexbench joined the objective four hours earlier, and without scanbench this trade would have scored five and a half million instructions cheaper than it is. the danger in a benchmark weighed at zero is what it does to the change after the one being measured, when nobody is reading the benchmark and everybody is reading the score.
36what an empty rewind costs
a beat loop is bracketed by three runtime calls, and the one that runs every iteration is the rewind: it puts the arena cursor back where the mark says, and flushes whatever the iteration registered — builder chunks, sorted map views, permanent slots, a shelf of donated buffers. on escapebench that call was 80,970,118 instructions, 43.66% of the program.
almost none of that was flushing. the common case writes three words and returns, and it was paying a call frame, four registry tests and a 96-byte memset to clear a shelf nothing had put anything on. the shelf is twelve pointers, and a program that never donates a buffer never writes one; a single dirty flag turns that memset into a test.
the rest is a split the runtime already uses in two other places: the empty case becomes an inline test and the frame is paid only by an iteration that actually has something to free. escapebench falls 22.68%, basket 8.48%, encodebench 2.82%, and the eleven benchmarks together 1.84% — 11,821,160,841 retired instructions to 11,603,420,615. the welfare index goes 75.73 to 75.95. it costs 7,008 bytes of .text across the corpus, which is the trade §35 says to state rather than score.
six conditions guard the fast path and only three of them can be caught by the corpus. dropping the buffer-shelf test segfaults; dropping the chunk-registry test moves the encode counters; dropping the block test moves encode and scan. the other three — the chunk spill, the view registry, the permanent slots — can each be deleted with every gate in the tree still green, because on every program in the corpus a rewind that finds one of those registries non-empty has also changed block, and the block test decides first. they are kept, and this is the record of what does and does not stand behind them.
37a gate that was red before anything was mutated
every job in ci carries at least one mutation whose whole purpose is to break it, and the nightly applies each one and refuses a gate that stayed green. the machinery rests on an assumption nobody had checked: that the gate was green to start with. a gate already red for some unrelated reason goes red under the mutation as well, and the run reads that as proof.
so the ratchet reads every gate once on an unmutated worktree before it mutates anything. on the first pull request that touched the runtime, it found two rows that had never proved anything. one named a golden that has never existed in this repository, so its gate exited non-zero from the day it was written. the other counts instructions under callgrind, and valgrind is installed in exactly one ci job — the one whose own gates need it. the ratchet runs every job's gates and had a base image. nothing in the tree related those two facts.
installing valgrind did not turn the gate green. the ratchet job and the cost-goldens job, on one commit, counted digestbench at 81,252,330 and 81,252,316: fourteen instructions apart, two runners, one image and one glibc. that is the first time those eleven rows have been sat twice within a single commit. every earlier reading came from one job on one runner, where a chip change and a code change arrive together and cannot be told apart.
two explanations are already dead. the benchmark binary is byte-identical whether it is built at the repository root or in a worktree on a shared target directory, and it retires the same count run from either place. what is left is the silicon, and one pair of runners is a data point rather than a design.
a gate whose golden is a set of exact counts is therefore declared host-bound. the baseline prints it, says the row went unproven on this runner, and carries on; the proving pass drops every row sharing that gate rather than applying a mutation to something already red. on a run that lands on the golden's chip the row is proved the usual way, so the cost is coverage only where coverage was never available. what a reader should take from a green ratchet is now exactly as much as the run actually established.
38a number was read twice
the json decoder walked every number's bytes twice. one walk asked where the number stopped, so the parser knew what to slice; the second walked the same bytes asking whether a ., an e or an E had gone past, because that answer decides to_int against to_float. the two questions have one answer each and the scan already sees every byte, so the mark rides along with it now as an argument. jsonbench retires 4.14% fewer instructions for it, measured 2026-09-03: 2,533,092,019 to 2,428,220,306 in the container, three runs, same digits.
the profile is the interesting part. value_for had been the largest symbol in a decode at 25.6%, and §36 already said that figure was a merged one — clang had inlined the whole value-parsing path into it. it reads 9.9% now, and the surviving number scanner stands beside it as its own symbol at 11.2%. what looked like one fat dispatcher was in large part two walks over 30,961 bytes of digits.
the trade goes the way this page keeps finding: the front end visits 2.2% more expressions and the compiler writes 89 more lines for the decoder, because two small walkers became one larger one carrying an extra argument. the machine code falls 48 bytes anyway, which is the second scanner leaving. no allocation counter moves at all — the same slice and the same conversion per number — so this is a change only the instructions vein could see.
the same idea declined, on the string path. a clean json string is found with one find2 and copied in one slice. a string with an escape in it is not: past the first backslash the decoder walks the rest a byte at a time. making that tail scan runs the way the head does costs 1.16%, and the distribution says why. in the benchmark's document 1,773 of 11,057 strings carry an escape, and they hold 4,562 runs totalling 16,895 bytes — a mean run of 3.7 bytes, a quarter of them empty because escapes come back to back. the fixed cost of a run is about 300 instructions against the 50 a byte the walk pays, so the run has to reach six bytes to break even and it does not.
the probe left something better than the change would have. two programs, one string clean and one carrying a single leading escape, at 2,000, 4,000 and 8,000 bytes: the clean one costs 25.6 instructions a byte and the escaped one 76.0, both linear. a byte inside an escaped string's tail costs about fifty instructions more than the same byte in a clean one, and none of that is a decision the library made. round the loop str_char spends six instructions indexing, ten boxing the byte and unboxing it again for the dispatch, eight dispatching, thirteen appending in place, four on bookkeeping. whoever wants that path faster should go after the ten.
39a byte switch that rebuilt the byte first
a function whose arms name byte values — 34 for a quote, 91 for a bracket, none for the end of the input — already crosses the call boundary as a raw machine integer, with 256 standing for none. §14 is where that came from. what nobody had checked is what happened on the other side of the edge: the entry block rebuilt a tagged value out of that integer, and the dispatch tree pulled the rebuilt value apart again to find the tag and then the byte. a comment above the reconstruction said the round trip folds back into a raw switch. it does not.
the loop in str_char, disassembled: cmp $0x4, cmove, cmp $0x100, sete, cmove, shl $0x2, test. seven instructions a byte spent taking apart a value that was never assembled, and every byte-dispatching function in the decoder paid it on every byte. the tree switches on the raw integer now, with 256 as the none case.
jsonbench retires 13.56% fewer instructions for it, on top of the single-pass number scan in §38 — that ratio is the container measuring itself twice, 2,428,220,306 to 2,098,859,754, because only ci may write the pinned rows — 17.14% for the two together, 2,533,092,019 to 2,098,860,167 on the runner — both figures are what that change measured, and the row has moved since. that is 74.2 instructions per input byte where §36 measured 89.5 the day before. oneshot falls 8.44%. the clean attribution is widebench at 3.85%: it carries a frozen copy of the json library, so the number scan cannot reach it and the whole of its fall is the dispatch. nothing rises. the seven programs that do not dispatch on bytes move by four hundred instructions or fewer, which is layout.
the machine code falls 8,080 bytes across the eleven programs and not a byte in the seven that were not dispatching on bytes. the counter that actually pins the change is branches in the emitted-code golden: it falls by 24 in the decoder and 76 across the eight beside it, because the tag test and the none test are gone from every one of those calls. there is no output that distinguishes a raw switch from a boxed one, so a branch count is the observable end.
one literal had to be guarded, and the fixture caught it before ci did. 256 is the sentinel, so a program that writes fn kind 256 — an arm no byte can ever reach — would have sent every read past the end of a byte string to that arm. the boxed tree is immune, because it tests the tag before it looks at the byte. native answered "a byte that cannot be" where the interpreter answered "some other byte", which is the differential law doing its job on a change that had already passed every counter. a group with any literal outside 0 to 255 stays on the boxed path.
the wider lesson is about the comment rather than the dispatch. an optimiser is entitled to change its mind between releases, so a claim about what it will do is not a claim a comment can hold. §31 made the same point about specs that read prose; this one was prose asserting an outcome, and the thing that could see it was a counter nobody had thought to look at.
40the refusal that could not be acted on
three compile veins carry a measured-on line naming the toolchain that counted them, and a host the line does not name is refused rather than compared. the refusal is right: a compiler's own allocation counts are partly the standard library's, so they belong to the rustc that built it. what the refusal then said was let ci measure it and copy the rows out of the job log. that instruction could not be followed. every compile gate asked the question under set -e, so a host that did not match stopped the gate before it measured anything, and the job log held no rows to copy.
the runner's rustc moved 1.98.0 to 1.98.1 and all three refused at once, each pointing a reader at numbers that did not exist. the only host allowed to produce them is ci, and ci had just stopped. a branch in that state could not be brought to green by anybody.
there are two questions where the gate asked one. may this host be compared against the golden, and may it be measured? a container answers no to both: a container's numbers going into a golden over the runner's is the accident the whole mechanism exists after, so it stops and prints nothing. ci answers no to the first and yes to the second, because ci's sitting is the only one that may ever be recorded — it measures, prints its rows, and still fails. §34's cpu block already drew that line the same way.
then the sitting itself. on the first chip ci landed on, under 1.98.1, the front end allocated 25,485 blocks and held 715,275 peak bytes — the numbers 1.98.0 recorded, to the byte — and retired 508 instructions fewer than the golden now carries. one run later ci landed on a second chip, on the same binary, and read 41,832,275, which is the figure the golden holds.
508 apart, and that gap is the one §34's row has been chasing all day: two runs on ONE chip and ONE binary produced both values earlier the same afternoon, so neither the silicon nor the toolchain picks a mode. the allocation and peak counters hold across all of it; only the instruction count is bimodal. what moves between the two modes is glibc's allocator — _int_malloc and memcmp shifting together, an alignment difference downstream of a heap layout difference — with every kanso symbol identical to the instruction. pinning the malloc and string tunables took the spread from 5,064 to 508 and did not close it.
so an exact per-chip pin cannot hold: a row set to either mode is red about half the time on a chip that produces both. that was filed as a design question about the vein, and the filing was wrong — it had a measurable answer, and the answer is that the row was never only measuring the compiler.
kanso check reads /proc/self/maps before it reads a line of kanso. glibc's pthread_getattr_np walks it on rust's behalf to find the main thread's stack, parsing each line with sscanf until it finds the one holding the stack pointer, and what that costs is a property of the host's memory map. adding one shared library to the process — four more lines to walk — moves the row +32,090 with the compiler doing identical work; at eight the reading has climbed monotonically over a spread of 226,714. the profile diff of a smaller version of the same perturbation names the parse and nothing else, with no kanso symbol moving at all. a startup term is global rather than per-chip, which is why two chips in three produced both values.
the tempting way out is to stop counting glibc, and it is the worst one, because a change to the compiler can make glibc do more work — a measure of the compiler's own instructions goes blind to exactly that. so glibc stays in the count and the row is made consistent instead. two candidate causes were tried and both failed: readdir order was already sorted where a compile reads it, and disabling address randomization changed nothing on the container or on the runners. what settles it is the binary's sha256, which the gate prints beside the count on every run — two runs on sha de5bfab22fbd and cpu family 0x6 model 0xcf counted 41,831,767 and 41,832,275. one binary, one chip, two values.
so the row pins a pair. both values are exact and a third is refused, which is the difference between this and a tolerance: a band wide enough to hold 508 also holds the 5,621 a runtime change moved this row by. the number welfare reads stays the first of the two, so a mode flip cannot reach the objective as a regression. what is still open is whether two runners' memory maps differ by the amount the 508 needs, and the row would have to print the map's line count beside the count to say.
the granularity question the refusal raised — whether rustc should be read as major.minor — stands on its own, because a patch bump that reddens three gates while moving no counter is the shape the rule was written against, and one observation is not enough to loosen a pin.
41a sitting, and the four things it settled
five rulings landed in one sitting, and four of them change what a reader of this page should expect. they are collected here rather than given a section each, because they were decided together and two of them only make sense beside each other.
fallibility is what the box is for, and io was never the criterion. foo["bar"]! raised the question: an err with no io underneath — a value, or an effect? treating a violated insistence as a defect that halts was proposed and declined. any operation whose answer can include an err yields <t>effect, io or not, and there is one failure system rather than two. foo["bar"] stays the data form — the value or none, absence as ordinary data, no box. the reason is enforcement: under explicit elimination the box is the only thing that makes propagation a contract instead of a convention.
the three combinators get fused chain spellings. .> is bind, .! is annotate, .? is rescue — the chain dot plus one character of channel:
config = io/read_file path
.> json/parse
.! (e -> "config: {e.reason}")
.? when_failed
each is sugar with a fixed desugaring, so the parser learns three operators and not three meanings; the semantics, the licenses and the auto-rewrap all stay at the words. the right-hand side is a bare function, so the common case needs no wrapper lambda. in chain position the fused spelling is the only one, which retires . bind (f) as a chain form.
the weights moved, and the argument is recorded because the doctrine requires it before the floor does. compile speed rises from 0.28 to 0.32, funded from run memory at 0.30 to 0.26. a developer feels compile latency on every edit and feels peak memory only when something falls over, so compile latency is an adoption gate where the memory model is an identity. compile memory stays at 0.12: the compiler's own footprint is a container guard rather than something anyone perceives. the run-speed term splits into an advertised half and a guards half. the objective's job is to punish regressions in the order a developer notices them, and until this sitting it did not. the split is retired — the gavel of 2026-09-06 made the run side one consolidated program with the shapes as phases at measured proportions, so there are no halves to weigh against each other. the four weights above are unchanged.
the bimodal compile row is made deterministic rather than excused. §40 left that row reading two values 508 apart on one binary and one chip, with the options being to pin both, band it, or stop counting glibc. the last was declined outright: if a change to the compiler can make glibc do more work, a measure of the compiler's own instructions goes blind to exactly that, so glibc stays in the count. the other two waited on two candidate causes, and both failed — readdir order was already sorted where a compile reads it, and disabling address randomization changed nothing on the container or on the runners. so the row pins the pair, which is what the ruling allowed once a residual survived both: two exact values, a third refused, and the number the objective reads still the first of them. the mechanism is still open. the count includes glibc parsing /proc/self/maps for the stack bounds before main, one more shared library in the process moves the row 32,090, and that cost belongs to the host rather than to the compiler; what is not established is that two runners' maps differ by the 508 this needs.
42what the split cost, measured
the arithmetic that fixes: nine guards against two advertised rows meant a guard carried nine elevenths of the run-speed term and the front page's own claims carried two elevenths, so a shape win scored as if a real workload had got faster. on the parity fixture a thousandfold win is now worth 52.99 on an advertised row against 49.11 on a guard, where before the split both read 48.48. the rule is "advertised versus everything else", so a benchmark added later joins the guards without anyone editing a list.
the score falls 76.17 to 73.06 and nothing about the compiler changed. compile cost is the weakest dimension and it just gained weight; run memory is strong and lost some. the fall is the objective being re-based, and it is the second such step in the recorded line — the first is the 2026-08-29 saturation ruling, 87.85 to 73.83. scores either side of a step were taken with different rulers and do not compare.
43the corpus was written around the defect, so the objective priced its repair at zero
inference read a chain's yield off the head's bare name. os/read_file hit the builtin table because the wrapper's name collides with the builtin's; os/read_file! — same body, one character more — missed it and fell to the top set, and so did seven other effect wrappers. a document read with the bang then had no type where the loop analysis reads one, and a loop past it ran on a grow-only arena.
the benchmark that reads a file at runtime had been hand-written around that, with a comment saying so. its main spelled the bang out into a pair of arms to get the type back. so when the yield was fixed, jsonbench read the same arena blocks, the same peak and the same beat count it had read the day before, every counter in every cost golden held to the byte, and the welfare index moved by nothing. the corpus had been measuring the workaround.
eleven benchmarks and not one of them read a file the way the repair fixes. the fix was for real programs and the corpus was not made of them.
the fix is a benchmark, not a weight. bench/readbench is the fixture at benchmark size: os/read_file! "bench/large.json" . go, threaded through a loop that allocates, two hundred rounds because the fixture already did two hundred. same program, same bytes, on the two compilers that bracket the repair:
arena_blocks arena_peak_bytes beat_iters
before the repair 41 42,991,616 1
after 1 1,048,576 201
its two memory rows enter the objective on a measurement rather than on the entering rule. a counter with no history is granted its dimension's standing, so landing day never moves the score — but this one has history, read off the commit before the repair, and granting it would have priced the repair at zero a second time. the work row is granted, because its history says nothing: 2,000,668,271 against 2,000,657,408, a fall of five ten-thousandths of a per cent. it is in the objective as a tripwire, so nobody can buy the arena back with instructions. the entering rule is retired, by the same 2026-09-06 gavel: nothing joins the objective a benchmark at a time now, so a counter with no baseline is refused by name rather than granted one. readbench's rows are diagnostics today and the repair they priced is in the run program's digest phase.
both trees under the same model read 70.83 and 73.52. the repair is worth 2.69 where it was worth nothing.
the round count is a free parameter and it decides the size of that number: the defective side is linear in rounds and the fixed side is flat. it was chosen by matching the fixture, before the score was computed, and not adjusted afterwards. a benchmark whose ratio is a dial is a benchmark somebody can tune to the answer they want, and saying so is the only guard that exists.
the general shape is the one this page keeps finding: a dimension a model leaves out it weights at zero. scanbench, escapebench and indexbench ran for weeks with their work weighed at nothing. this is the same failure one layer out. the corpus had been written to avoid the defect, so there was no counter to forget to read.
44the row counts the program, not the process
the compile row moved between binaries that do no different work. seven of them, built from sources differing only in code or data nothing reaches, and the whole-process count spans 3,963:
variant .text row maps program
baseline 2550854 42,344,081 112,580 41,878,959
+50 dead fns 2552534 42,348,024 114,845 41,879,987
+200 dead fns 2558486 42,347,128 112,586 41,879,361
+400 dead fns 2565174 42,348,044 110,341 41,879,922
+3 KiB .bss 2550854 42,346,221 114,720 41,878,959
+64 KiB .bss 2550854 42,346,221 114,720 41,878,959
+64 KiB .rodata 2550854 42,344,099 112,598 41,878,959
the movement is above main. before a rust program runs, the loader maps its shared objects and the runtime places a stack guard, which parses /proc/self/maps — and how long that file is depends on where the linker happened to put things. so the row reads kanso::main inclusive now, which is everything the compiler does. that frame spans 1,028 across the same seven.
the anchor is this crate's own symbol for a reason the first attempt supplied. it read std::rt::lang_start::, which is the standard library's, and ci refused: the toolchain the runners carry does not emit that frame. a name std owns can move under a version bump with nobody here touching a line. kanso::main carries an inline(never) that says why it exists, and sat exactly 10 instructions below the closure on every profile retained — the table above was read with the closure, and the spans are the same under either.
§38 ruled that glibc stays in the count, and it still is: every libc call the compiler makes is inside that frame. what comes out is the 465,122 instructions of process startup that happen before the compiler runs, which is a different thing from a measure of the compiler's own instructions going blind to work the compiler causes glibc to do.
the residue is real and worth saying. where the difference between two binaries is data, the frame is identical to the instruction. where it is code, it is not: 7,632 bytes of .text nothing calls moves the frame 402, and not monotonically — fifty dead functions move it more than four hundred do. an earlier reading of four binaries saw only the data cases and read invariance into them. against the smallest front-end change on record, a fall of 0.07%, the residual sits at a twenty-eighth where the whole-row term sat at a seventh.
a residue survived the anchor, and it is thirteen instructions. two sittings of one commit read compile rows exactly thirteen apart, and four pairs of them did it with the per-process floor — the frames whose cost is identical across all three compile workloads — reproducing to the digit between the two. so the thirteen is not in what those workloads share. eight runs of one binary in one container read the same number every time, so it is not the run. a branch whose whole diff was two test files read it too, so it is not the change. what is left is the build, and every job compiles its own compiler.
bucketing every frame's self cost by the length of its name found it. two sittings of one branch, one on the row and one thirteen under it, printed digests identical in thirty-one buckets of thirty-two and forty frames in both, so nothing appeared or disappeared: one frame among forty moved. three of those forty are process setup and cost thirteen, thirteen and twelve instructions — thread attributes and glibc's arena. a fourth counter on the same sitting moved six rather than thirteen, which is what one difference in once-per-process work looks like when two different programs reach it.
the thirteen is settled, and the answer was to stop counting it. five builds of byte-identical compiler source gave two faces thirteen apart on all three kanso check rows, while the start-up row — whose gate prints no result line — read one value on every one of them. that is the tell: kanso::main inclusive counted the line the run prints when it finishes, and under it linewriter runs memrchr over the formatted bytes to find the last newline, a frame whose cost moves with the binary's layout. not printing costs more than the line does: an environment variable asking for quiet costs 11,606 instructions against the 1,116 the quiet saves, because the compiler asks getenv about seven thousand times and each ask walks the block, and one more argv entry costs about 2,047. both move the initial process layout, which is the same class of thing the thirteen is. so the term is excluded rather than normalised, which is what §44's rule provides for: the three gates subtract std::io::stdio::_print inclusive from the anchored reading, each golden's header names the exclusion, and three builds then read one row where the old one could not manage two.
the gates say which it is from inside a job now. a row that parts from its golden is counted a second time on the same binary, and the job prints whether that read the same number, with the binary's sha beside it. those lines were being emitted into the job log, which is the one artifact a reader may not be able to fetch, and then past github's cap of fifty annotations when three rows parted at once — so a gate that had already settled the question reported nothing. they are annotations now, and the verdict comes before the explanation of it.
what the split goes blind to is the one thing a compiler change can do to move the dropped half: grow a dependency. one more shared object costs about 32,090 instructions of loading. so a golden holds the sonames ldd reports — five of them — and a new one turns that red by name. a number that moved 32,090 inside a thousand of linker luck is a puzzle; libfoo.so.1 appeared is an answer.
45a counter whose definition changed
§44 redefined the compile row and subtracted the same 465,864 from welfare's baseline, so the ratio stayed comparable. the trend gate read the pair and printed this into the record of that merge:
improved: compile_instructions 41,845,704 -> 41,379,840
every changed counter is priced (or improved).
nothing in the compiler got faster. the two readings are of different quantities, and a ratio between different quantities cannot say which way anything went.
the signal turns out to be exact. welfare --set moves the floor and never the baseline, so a baseline value moves only when somebody has decided the old reading and the new one are not the same measurement. a counter whose golden and whose baseline move in the same diff is therefore neither better nor worse: it prints as re-based, counts toward neither side of the pure-regression rule, and still owes a sentence naming it and the value it landed on. that is the third state §30 built for a counter the baseline does not carry at all, reached from the other direction.
matching a golden row to a baseline by name covered one vein and looked like it covered all of them. welfare renames on the way in — work_jsonbench in the gate is decode_instructions in the baseline — and for the memory terms it also sums the arena, held and perm peaks, so a *_peak_bytes counter has no row anywhere holding its value. joining on values was measured before anything was built: of 27 counters, six match no row at all and two match four goldens each.
so the link is written down. bench/objective_sources.txt carries one <objective counter> <gate key> a line and several lines for a sum, and a spec replays it — summing each counter's rows out of the goldens and asserting the total is what welfare --counters prints, for every counter welfare scores and no counter it does not. a rename on either side, a pool added to the sum, or a benchmark joining the objective turns that red instead of quietly unlinking a term. it held 41 rows for 27 counters when it was written and holds 7 for 5 today: the gavel of 2026-09-06 put the run side on one consolidated program. the spec replays the file rather than a count, so nothing here was ever unchecked — the prose was, and this sentence is the second time that has been true of a number in this file.
two things the work found on its way, both of the same shape. the first cut of the check classified nothing in the runtime vein, because it was handed the counter with its golden's prefix stripped — so it worked for exactly the four goldens whose prefix is empty, which is the vein it was written for. and bench/cost_golden_read.txt, holding two of welfare's terms since readbench joined the objective, was in the trend gate's list of goldens not at all: setting a row in it to 9,999,999,999 produced no output whatever, exit 0. both were found by probing rather than reading — move the number, watch what the gate says.
46where a program claims its stack
a call whose arguments go through the runtime needs them side by side in memory, so the emitter writes an alloca and stores into it. it wrote that alloca in whichever block was doing the storing — a record built inside an if arm put its array in that arm, a dispatch arm the same. sixty-eight of the decoder's seventy-three stack slots stood outside their function's first block.
llvm reads a slot placed anywhere but the entry block as a dynamic stack object, whatever its size. a function holding one keeps a frame pointer it would otherwise omit, restores rsp through that pointer on every return path, and claims the slot again on each pass through the block. on the decoder that came to 12.03% of all retired instructions in push, pop and leave — 252,410,292 of 2,098,864,471 — with 51,317,285 of it the frame pointer itself.
the slots are held back now and written at the head of the entry block. the decode falls to 1,914,624,416, 8.78%, and ten of the twelve benchmarks fall with it. the two that rise, deepbench by 0.314% and scanbench by 0.133%, are the two whose machine code grows most: a fixed frame gives llvm a different layout to schedule against, and it inlines differently in both directions.
no allocation counter moves, and neither does the emitted-line count — the same lines are written either way and only their order changes. the property is asserted where it lives, on the .ll the build hands clang: a spec compiles a record built in each arm of an if and reads back every slot's block.
what is left is the other nine and a half per cent. removing the frame pointer leaves 201 million instructions of callee-saved push and pop on the decode. the decoder is a chain of mutually tail-calling functions, and a tail call pops the whole frame before its jmp while the target pushes it straight back. that is a different question and this change does not touch it. the welfare number reads 73.73 after it, up from 73.53.
47what a seven-byte string paid to be checked
every string a json document holds is validated as utf-8 before it becomes a value. the check was one function: count the bytes, test a word at a time for the high bit, and where that fails, run a simd pass over the run. llvm will not inline a function of that shape anywhere. the wide pass loads seven constants and zeroes two accumulators before its first block and carries a table of three sixteen-byte lookups after it, so the ascii answer — two loads — sat behind a call. 1,571,250 answers on the decode, 83,092,800 instructions, 53 apiece for a mean run of seven bytes.
the front door is always_inline now and holds the counter and the word test. the pass is a function of its own, and a run reaches it only by carrying a byte with the high bit set.
turning the ascii runs away left that pass measurable on its own: 60,769,050 instructions over 203,700 calls, 3.29% of the decode. thirteen per cent of the document's strings are not ascii, and each cost 298 instructions to validate about seven bytes, nearly all of it the setup. the portable arm of the pass walked the grammar a byte at a time, and nothing but a host without simd had ever reached it. it is a function now and the door chooses on the length — 32 bytes or fewer walk the grammar.
the decode falls to 1,877,751,716, 1.93%, and four other benchmarks fall with it. seven are byte-identical, being the programs that never validate utf-8. the machine code grows 6,736 bytes across five programs, because the word test is written at each of the validator's four call sites instead of once inside the pass. the welfare number reads 73.77.
the differential harness extracts the validator's text out of runtime.c rather than carrying a copy, so each split broke it and each repair was to teach it the new arm. it reads four pieces and the length threshold now, and what it checks is the composition: 45,189,025 validator checks against an independently written reference, zero mismatches. dropping the surrogate bound on 0xED gives 2,048 of them, exactly ED A0-BF by 80-BF, so the fuzzer reaches the arm that moved.
one thing measured and declined: replacing the eight-comparison chain that decides a lead byte's width with a 256-entry table costs 1,428,150 instructions instead of saving any. the chain is predictable, and the ascii bytes that dominate a mostly-ascii run are tested before it and reach neither.
48a door that takes registers instead of boxes
a builtin the emitter calls takes and returns KValues, sixteen bytes each. the emitted module declares that type; clang compiles the same C function to the lowered one, {i64,i64} in and integers out, and lto sees two function types and keeps the call. so an always_inline written in runtime.c cannot reach a caller the compiler generated — marked on k_b_find2 it moved the decode by zero instructions and left the symbol in the binary. the machine abi does agree; this is a barrier to inlining and not a miscompile.
the lever is the prelude, a set of small llvm functions the emitter writes into every module and marks alwaysinline. a twin there tests the argument tags itself, unwraps what it needs, and calls a raw door taking scalars; anything the test turns away goes to the boxed function unchanged.
find2 scans for the first of two bytes. four KValues do not fit the six integer registers the abi has, so the fourth arrived as a pointer to the caller's stack and the callee loaded and unpacked it before it could splat the byte — fourteen of the function's fifty-four instructions, where the scan itself was ten. the raw door takes a pointer, a length and three integers. the decode falls 2.695%.
slice's bytes arm builds a view, a header and two words, behind three failure guards and two tag tests. the raw door takes four scalars and holds the four compares and the add. a further 1.126%, and widebench 0.659%.
the second one costs something, and the cost is legible. four benchmarks rise — the ones that slice a list or a string, where every call now pays three tag tests before falling through. scanbench is 0.175% of that, and the others are thirteen instructions, two, and two hundred and fifty-six. the welfare index takes the trade at 73.82 to 73.85.
and one function needed no twin at all. k_str_n copies n bytes into a fresh string; it is static and so is every caller, so the attribute reaches them. indexbench falls 8.585% and scanbench 1.508%, which more than repays what the slice twin cost there. readbench rises 1.886%: it reads a file and holds few short strings and many long ones, and an inlined copy loop loses to a called one when the copy dominates.
its price is machine code. all twelve .text rows rise, 34,320 bytes in total, where the two twins grew four rows and eight — the linker drops a twin nothing calls, and a copy loop written at every site that makes a string cannot be dropped. the welfare index has no term for machine-code size, so it scored the change without seeing that. whether it should carry one is a question about the weights, left open beside the measurement that raises it.
49the argument that would not fit
utf8 applied to slice is lowered as one call, and that call takes three KValues and a pointer to the wrapper's origin — seven register-sized arguments where the sysv abi has six. the last one spills to the caller's stack and the callee reloads it before it can put the origin in an err. the raw door takes a pointer, a length, two integers and the origin, and nothing spills.
the decode falls 0.445%, the one-shot 0.182%, and the encode 53,542 instructions. eight of the twelve benchmark binaries come out byte-identical to the ones before the change, so those rows could not have moved; of the four that did change, none rises.
it costs 32 bytes of machine code, 16 on each of the two programs whose binaries grew. the twin is internal and the linker drops it where nothing calls it, which is why the whole cost sits on the callers.
the fixture is tests/golden/micro/utf8_of_a_byte_slice_at_its_edges.kso. nothing in the corpus had reached this door over bytes: the one program that read an invalid byte through a slice passed a list, which falls through to the general slice and the general utf8. the fixture walks every range the clamp refuses and both sides of the validity test — a two-byte sequence taken whole, the same sequence cut in half, and its continuation byte alone.
50counting the doors instead of guessing them
three doors were twinned by reading a profile and picking what looked expensive. the question the mechanism actually asks is narrower and countable: which of the functions the emitter declares take more than the six integer argument registers the abi has? a KValue is two, so the answer comes out of the source with no measurement at all. it is five. k_call4 and k_b_find2_below take ten, k_call3 and k_b_find2 eight, k_b_utf8_slice seven.
two were already done and two are call doors no benchmark reaches. that leaves find2_below, which spills four arguments — more than anything else in the tree — and carries 481,347,200 instructions of self cost in the encode, 8.3% of it.
the twin takes 239,593,600 of them, about half. encodebench falls 4.129% and the one-shot 2.050%; widebench's binary changed and its count did not, and the other nine binaries came out byte-identical, so those rows could not move. nothing rises. four spilled arguments are four stores in the caller and four loads and unpacks in the callee on every call, in front of a scan that is a compare and a branch per byte.
the count is not a complete test for candidates, and slice is why: it fits in six registers and its twin won anyway, on the guards. what the count gives is a closed list for THIS mechanism. after find2_below the only doors that overflow are the two nothing calls, so the family is finished until a new door is declared.
51a lambda that captures nothing is one value
a string literal is the same value on every evaluation, so the emitter builds it once into a permanent slot and hands that back forever. a lambda that captures nothing has exactly the same property and never got the same treatment. escape_able_2 in the encode benchmark allocated a closure header and an empty environment 709,200 times to hand list/fold something that never differed — 1,418,400 of the encode's 16,249,027 allocations, and 48,225,600 instructions.
the emitter can tell: the capture list it already computes comes out empty. where it does, the alloca and the two allocations go away and the site becomes a load after the first visit. the encode falls 0.650%, the digest 0.539% and the one-shot 0.315%; the deep benchmark rises 5.672%, and the objective takes the trade at 73.99 to 74.00.
that rise is worth being precise about, because it is not the walk doing more work. the carry walk visits the same nodes to within 0.2% — k_copy_size is entered 1,534,544 times before and 1,535,314 after — and each visit costs twelve instructions more. removing the allocas moved the optimiser's inlining budget inside a function this change does not touch. two attempts to steer it made it worse: marking the literal closures as permanent survivors so the carry path shares them took the deep benchmark to 5.680%, and marking the door cold cut the encode's gain in half while moving the deep benchmark nowhere.
the fixture is tests/golden/micro/a_lambda_that_captures_nothing_is_one_value.kso. sharing closure identity is the new risk: two capture-free lambdas in one body each need their own slot, and a lambda that does capture must still get a fresh closure per call. routing a one-capture lambda through the literal path moves exactly the three capture lines and leaves the seven others alone.
52what the counters could not see about themselves
the compiler is watched by two sets of counters. the runtime ones ride on eleven benchmarks and say what a program costs to run; the compile ones say what the front end costs to decide. they are read by different gates, they refuse on different hosts, and until this week nothing ran either set as one command. a change was expected to remember which files it had moved, and a branch that remembered the memory vein and the code goldens missed nine of the eleven cost goldens; ci found them a round late.
scripts/gates/all_counters.sh now reads every runtime vein and names each one that moved, and scripts/gates/all_compile.sh does the same for the six compile gates. the second one has a job the first does not: three of its gates refuse outright on a machine whose glibc or rustc does not match the golden's measured-on line, and a refusal exits exactly like a regression. so it separates the two, and a session running it on a container can tell a vein that moved from a vein this host may not read.
the two sets are less separate than they look. lib/*.kso is compiled into the binary, so a line added to the json library is a line the compiler carries and compiles, and every compile counter moves. a twelve-line library change read as an improvement to the objective with those veins stale and as a cost once they were regenerated, which is the whole reason the second sweep exists.
the benchmarks had a matching gap. two of them vendor a frozen snapshot of the library on purpose, so that a move in their numbers is the compiler's and not the library's. what the freeze costs is that nothing then watched the library's own encoder under load. bench/livebench runs the same program against the library that ships, prints the same checksum, and the difference between the two is the frozen copy's drift: 10,149,724 instructions, 0.19%, at the sitting it was added.
and one counter turned out to have been reading a number from a compiler that no longer existed. the front end's instruction count is held per chip, because glibc resolves memcpy and its neighbours by cpu feature at load and one libc runs different code on different silicon. when a change moves the compiler's bytes every row goes stale at once, and ci re-sits them one chip per run — so for a while the table holds a mixture, and the single number the objective reads follows the first row. that row was one of the stale ones. two values 519 apart sat under one key and were recorded as the two readings a single chip had been seen to give, which is a real thing that happens and was not what happened here. five chips have since read the same number on one binary, across three amd generations and one intel.
53proving a check unnecessary costs what the check costs
the encoder validates utf-8 that it wrote itself. every byte of a finished document came from a structural character the encoder emitted, from the ascii digits of an integer or a float, or from a string that was already a str and therefore already valid; and then k_b_utf8 reads all 188,698 bytes of it and checks. that is 111,966,800 instructions on encodebench, 2.108% of the run, once per iteration at 1.4834 instructions a byte. the cheap ascii path cannot answer, because 5,865 of those bytes have the high bit set and it needs the whole run to be plain.
the obvious repair is a bit on the bytes header meaning "everything written here is still valid", held by an append of a string or of an ascii byte and cleared by anything else, with the conversion skipping the pass when it holds. the surface is small enough to be tempting: the header has eight spare bytes already, because the arena hands out 32 and the struct is 24; two places construct one; five places write into one; one place reads the answer.
it does not pay, and the counter that says so was already in the golden. the encoder makes 42,312,400 appends over its four hundred iterations — 105,782 an iteration for 188,698 bytes, or 1.78 bytes an append. maintaining the claim costs a load, an and and a store on that arm, so 126,938,400 instructions to save 111,966,800: a net loss of 0.28%. at two instructions an append it is 84.6 million against 112 million, a margin the fourth header field's own cache pressure would take back.
the reason generalises past this one idea. the validator reads every byte once, and the appends already touch every one of those same bytes. tracking cannot be cheaper than scanning the same bytes; it is the same work moved earlier and paid whether or not anyone asks for it. skipping a check means proving the thing the check would have proved, and the proof costs what the check costs. the variants die the same way: oring each byte into an accumulator and testing the high bit at the end is the ascii test again, and letting only strings hold the bit leaves it clear at every read, because the encoder appends bytes constantly. what would change the answer is a shape where bytes arrive in far fewer and far larger appends, and neither the encoder nor any benchmark here is that shape.
54the frame costs more than the work
after the bytes view inlined, the escape reducer became the second-largest line in the encode profile. w_klam17 — the c-abi wrapper k_closure_lit points at, called through a function pointer by list/fold — runs 11,658,800 times for 712,888,240 instructions. that is 61.15 instructions to decide whether one input byte needs escaping and append it.
callgrind with --dump-instr=yes puts the cost on named instructions rather than on a guess. seventeen run on every call and eleven more on every call but eight hundred; a clean byte takes twenty-nine more, and the escape cases another forty on 3.9% of calls. so the escape machinery is under a tenth of the line and the fixed cost is most of it.
the fixed twenty-eight are the frame. six callee-saved registers pushed at entry and popped at exit, the stack slot, and the return: twelve instructions of spill and reload is 2.84% of encodebench, fifteen with the frame and the ret is 3.55%. the tag work the function actually does — the failure test, the byte test, the range ladder — is about eleven instructions.
the frame is the body's doing, and the program's three lambda wrappers show the relationship directly: 692 instructions of body pushes six registers, 68 pushes four, 53 pushes three. the clean path is 84.3% of the calls and pays for the escape path's register pressure. the upstream answer is preserve_none, which wants llvm 19 where ci's clang is 18.1.3.
the split k_b_append_grow uses in c — a small hot leaf tail-calling a cold half — was the answer that looked available without the toolchain, and it was measured the same day. it does nothing. esc_byte's seven literal arms differ only in the byte after the backslash and each inlines two copies of the append twin, which is where the 694 instructions come from; giving the seven arms one shared body leaves livebench at 4,911,031,267, oneshot at 26,642,454, jsonbench at 1,732,114,716 and every .text byte-identical. llvm inlines the shared body back into all seven arms and emits the same code. the inliner reassembles what the source separates, so keeping the cold arms out of line needs a way to say so — and kanso has no noinline for user code, which makes it a language question rather than a patch.
three repairs died before any of them was built, which is the useful half of the section. the append twin normalises a builder's capacity with a neg, a compare and a conditional move, and that sequence appears eighteen times in the reducer's machine code. the first thought was to count eighteen; one byte takes one path, so it costs three. the second was to drop it; the sign of cap records whether the storage came from the arena or from malloc, so dropping it reads a negative capacity as no room. the third was to decide the regime once when the fold starts; k_b_append_grow rewrites that sign on every growth, from the beat depth at that moment, so there is no invariant to hoist. the ceiling on the whole idea was 0.7% and the floor was a miscount.
55ten shortcuts on the run program's hot paths
the objective measures one consolidated run program since 2026-09-06 (§53 says why), and the run program is where the runtime's own habits show up at scale: it decodes, encodes, escapes, digests and tallies in one binary, and a shortcut that saves twenty instructions on one runtime call is worth what the call count says it is. one day of profiling with --dump-instr=yes found ten of them, and every one is the same shape: a runtime function doing a fixed amount of work to learn it had nothing to do. an inner beat handed its 256 KiB tenure block up to the depth outside at every pop rather than freeing it, forty-nine blocks a phase for 78 KB of tenured bytes, and every membership ask walked all forty-nine; it opens its tenure in the block outside now. a beat pop with nothing to carry paid the six callee-saved pushes its calls need to skip a copy, migrate zeros over zeros and return at the hand-up's first line, 507,678 times of 507,685; the pop tests four flags and returns before the frame. a push off the frontier asked whether its header predated the beat by walking the block chain twice to find a pointer a few bytes below the bump pointer. the utf-8 scalar arm validated the ascii around a wide character a byte at a time. a token slice of four to seven bytes, a rendered number's digits, an indexed wide character and an empty map literal's no pairs each went through glibc's memcpy, which spends fifteen instructions choosing how to move five bytes. a buffer's size class was found by doubling from sixteen up to it, 1,183,000 times a run, where one trailing-zero count says it. and the sizing walk that decides whether a trip's result is worth copying asked two walks of the same chain of every node, one to learn the arena keeps it and one to rule it above the mark before reading the tenure blocks.
together, on the container with clang 19: runbench 2,736,140,571 -> 2,602,519,000 retired instructions, -4.8836%, the same bytes out, and no counter in the twelve cost veins or the lazy tier moved for any of the ten. the work vein is the only witness such a change leaves, which is why each one carries a ratchet row that puts the old shape back and asks it. three repairs measured worse and are declined in the log with their numbers: a byte loop bounded by the character's width instead of the call (+805,016), a word-at-a-time tail for the escape scan's two-byte search (+2.2879%, because a run between escapes is a few bytes and the mask's setup costs more than walking them), and deciding an immediate's zero before the sizing walk's frame (+1,777,234, because the guard at seventeen call sites cost more than the frames it spared).
three more landed on 2026-09-07 after the constant freeze and the walk reorder below, each with its row. s[i] on text hands back one character, and a character of two, three or four bytes was built in the arena every time -- 690,000 times a run on the index phase, whose subject is six characters doubled to 1.4 MB -- where an ascii one came from a cache of interned singles; wide characters come from a direct-mapped cache of 256 permanent strings now, keyed on their own bytes and never evicting, and the cursor answers the next character directly instead of through the general walk. the utf-8 wide pass tests four sixteen-byte blocks for ascii with one mask before it classifies any. text/slice of multibyte text counts the characters between its two positions a word at a time, by continuation bytes, instead of stepping through them. runbench 2,483,620,655 -> 2,466,456,226 (-0.6911%) on the container, indexbench -12.78%, the same bytes out, and the run program's arena peak 57,478,864 -> 45,944,528 bytes with 345,000 fewer allocations.
and one more the same day, from reading the token slice's sixty-two instructions: the test in front of the allocation counters was written as != 0 around a second if, a shape left from when the switch initialised itself lazily, and clang compiled it to a compare and two branches at every inlined allocation -- 7.3 million a run. it makes the same > 0 test every other counting site makes now: runbench 2,466,455,728 -> 2,453,160,233 (-0.5391%), jsonbench -0.82%, pendbench -1.06%, every counter identical, and with k_alloc a branch smaller clang inlined two of its callers into theirs.
and a fifth, from the chain step's sizing walk: the bind chain sizes its continuation on every step to choose between leaving it, staging it through the carry pair and flooring the region under it, and the walk's closure, description and subtype arms recursed on every slot they held, an int or a string as often as a pointer -- 176,113 of the 510,001 slots the run program's walk visited, each paying a frame to learn it was not heap. those arms ask first now, as the list arm has since deepbench, and the seen-map's probe is inlined at its six sites: runbench 2,453,160,233 -> 2,446,395,268 (-0.2758%), deepbench -7.45% and widebench -2.44% where the walk is a larger share of the work, the same sizes answered and every counter identical.
and a sixth, the step that walk was serving: a chain step whose region has not drifted a quarter megabyte past the last staged top now leaves without sizing at all. that alone segfaults, and the reason is a rule that held only because a stage happened every step -- the copy-out at a carried pop pruned at any survivor whose immediate interior survived, so an arena node two levels down holding a pointer into the depth's carry buffer was left for the caller's next stage to repair. let the chain leave for a few hundred steps and the buffer is grown, freed and reused underneath that pointer first; at that one call site the copy-out no longer prunes. deepbench 599,235,221 -> 410,388,149 (-31.5147%), widebench 49,141,298 -> 36,460,870 (-25.8039%), runbench 2,446,394,810 -> 2,418,520,678 (-1.1394%) on CI, the same bytes out. the deeper walk copies more where it now runs and the evacuation counters say so; the objective weighs that against the memory it saves and reads 65.86 -> 65.95.
that copy-out walks deep again only when it has to, which is almost never. a carry pointer reaches a node the walk would share by exactly one route: the program writes one in place, through set, a map insert or a list push, into a node that was already there. so the seven write sites say when they do it — one unsigned compare against a bounding box over the live carry buffers, and a walk of those buffers for the few pointers the box admits — and the pop prunes while nothing has. across the whole benchmark corpus nothing ever does; on scripts/trend_gate, the program whose segfault found the hazard, 1,634 writes do. deepbench 410,388,149 -> 389,214,232 (-5.1595%), runbench 2,418,520,678 -> 2,414,841,737 (-0.1521%) and pendbench 602,145,183 -> 598,215,444 (-0.6526%) on CI. asking cost 4,274,224 instructions of runbench before join’s ask moved behind its thunk test, because join forces every element and asked in front of a check that already tells it whether the slot changed. what asking still costs shows up on basket, whose row rises 2.02% while its evacuated bytes fall from 55,104 to 192: the walk that used to copy those bytes asks about them instead, and asking is what the row counts.
a seventh: an if over a comparison forced a value that is a boolean by construction. the emitter merges two arms into a phi for + - * and the six comparisons — the inlined fast path, and the call the slow path makes — and that phi carried no set, so every reader of it took the default, which is the top of the lattice, which admits a thunk. the force is a call and it went in front of the test. both arms are known: k_cmp answers a boolean or the failure it was handed, k_add and its two siblings answer an int, a float or that same failure, and none of the six can answer a thunk. runbench 2,414,841,737 -> 2,400,271,058 (-0.6034%), pendbench -0.2679% and digestbench -3.0230% on CI, the same bytes out. what moves with it is the emitted code rather than any allocation counter: 106 fewer calls in the run program, because forcing a value that is not a thunk allocates nothing and costs only the call. the front end never sees the phi, so all three compile counters are byte-identical.
an eighth, and it is the seventh's other half. once the phi carries a set, a value the emitter has PROVED is a boolean can skip the failure test in front of it as well as the force — testing a value that is not a failure answers true, so the test is a call and a branch that always go the same way. that is 27 of the run program's 1,223 tag tests, and all thirteen programs' emitted code falls with them: the run program's calls 6,026 -> 5,999, its lines 34,873 -> 34,819. runbench 2,400,271,058 -> 2,398,991,511 (-0.0533%) on CI, the same bytes out, and both compile veins fall with it because the compiler is one of the programs it compiles. every allocation counter is byte-identical: removing a tag test allocates nothing.
the same fold on any set with no failure in it is sound too, and removes 623 rather than 27. it is slower — runbench +0.1534%, and the objective declines it. the decode does what the profile promised, obj_key_start falling 4.12%; the encoder more than takes it back, in the pair that calls itself through another and in the slow force behind it, 8,640,030 instructions across four call chains. what a proof buys is not the check it removes but what the compiler can still see afterwards, and a check standing in front of a cycle is holding something up.
the narrow rule needed one more thing than soundness. a set with nothing in it satisfies every question asked of it, so “the inference reached no shapes here” read as “proved to be a boolean”, and a call handed a failing argument stopped hopping it out and refused the overload instead. the fold asks for a non-empty set now, and the one block that hops a failure asks with a twin that never folds: it runs BECAUSE a value is outside the set recorded for it, so no set can answer there, true or not.
a ninth, on 2026-09-09, and it is the encoder's: append acc "{n}" is how the json encoder writes every number, and the template rendered n into a string for the one purpose of copying its bytes into acc and dropping it — 379,530 strings a run, built to be copied once. the emitter fuses the pair now: one call writes the digits into a stack buffer and copies them into the accumulator from there, three words whatever the length where the accumulator has 24 bytes to spare, so no string is built and memcpy is not called for a handful of digits. the fusion did nothing on its own. the arm the tag switch lands in had never been told what the switch decided, so n:int's body still carried the whole group's set, forced n again and kept the to_string dispatch that a record would need; the arm reads the case's tags now. runbench 2,368,295,010 -> 2,347,625,055 (-0.8728%) on the container, the same bytes out, and every allocation counter identical — the string that is not built was arena garbage the rewind already freed for nothing.
56a constant is built once
a zero-argument definition is a constant, and the interpreter has always computed each one once, behind a cell: mention it twice and the second mention reads what the first built. native did that only for a constant whose body was a literal. every other constant — a table joined from literal lists, a map built by a call, a value computed from another constant — compiled to an ordinary function and ran its body at every mention. the archive saw this on 2026-08-30 and left it for a ruling, because a frozen constant was then built before main and a body that could fail would have failed too early; the 2026-08-23 ruling that a cell fills on first demand removed that ground, and the question was left standing with its answer already given by the oracle.
the profile found it before the record did. sha256's sixty-four round constants are eleven literal lists (six to a line, for the width limit) joined by text/concat, and compress reads the table once a round: ten concatenations and sixty-four pushes, 16,000 times a run on the run program, to build a value that never changes. since 2026-09-07 every constant freezes — built on first demand, copied to storage that outlives every rewind, read from a cell after that — and the runtime treats a frozen buffer as immortal, so a beat that carries a reference to a constant no longer copies the constant into its carry buffer, and a constant reached through another constant keeps its identity, which is what a self-referential value needs for its cycle to come back to the same list on both engines.
the freeze's own sitting left eleven small work rows up, and the cause was in the walk that places a node against a beat's mark rather than in anything frozen: k_where started at the head of the block chain and reached the mark's own block only after every block newer than it, and the node the sizing walk asks about is almost always in that block, above the mark if this lap built it and below if the last one did. since 2026-09-07 the mark's block is asked first, then the newer blocks, then the older: pendbench 608,937,982 -> 604,694,569 (-0.6968%), deepbench 686,441,869 -> 647,639,361 (-5.6527%), widebench -2.8793%, encodebench -3,009,187 and runbench 2,487,359,798 -> 2,483,621,153 (-0.1503%) on the container, the same bytes out, and the two rows the freeze had raised are under their pre-freeze values. deepbench was never in the freeze's list: its fold sizes a long list of survivors every pop from a block that sits under several newer ones. the obvious repair, sending a tenure-free ask back through the walk from the mark's block, measured runbench +2,936,479: that walk visits every older block for a node built this lap, and runbench's chain is fifty-four blocks at its peak.
runbench 2,602,519,000 -> 2,487,359,798 retired instructions, -4.4249%, the same bytes out; and the allocation veins, which the work vein cannot see, move further: runbench's arena peak 156,045,008 -> 57,478,864 bytes and its blocks 148 -> 54, digestbench's allocations 213,707 -> 23,841 and its peak 54,525,952 -> 3,145,728. the price is one permanent header and one copy per constant frozen, and a _build symbol with a cache in front of it in the emitted code, which rises by a define or two per program. welfare, which weighs both, rose 57.80 -> 64.05 on the container's sitting.
57what a profile looks like when the loops are gone
the sections above spent a fortnight taking loops out of the run program, and on 2026-09-07 the same instrument was pointed at both halves of the objective to ask what was left. for each function, the cost of its single hottest instruction over that function's own self cost: a tight loop puts a large share on one instruction, a straight line spreads it. across the twenty functions carrying seventy per cent of the run program, the highest such share in this project's own code is 0.063 — the escape scan's vector loop, fifty-five instructions long. everything else sits between 0.005 and 0.042. the one outlier is glibc's memcpy at 0.465, which is not ours.
the front end reads the same way. the top function in kanso check lib/json carries 3.79%, and it is a hash map insert. bucketing every function's self cost gives the compiler's own passes 53.4%, hash tables 16.7%, malloc and free 14.4%, rust's standard library 8.4%, libc's string and memory routines 3.9%, and process startup the remaining 3.2%. so a third of a compile is the data structures rather than the passes, and the largest removable-looking piece inside that is rehashing: growing maps costs 4.2% of the process across 1,470 rehashes. it is not one change. attributed to the pass that owns each map, that 4.2% spreads over twenty-odd call sites and the largest is 0.68%. the allocator itself cost about 124 instructions a call, which is what glibc cost; the compiler runs on mimalloc since the swap recorded at the end of §08, and that whole bill fell from 15.17 per cent of a compile to 7.39. the pinned rows beside this profile count a different program — 14,319 allocations and 24,998,478 retired instructions on the fixed corpus the compile gates have checked since 2026-09-08 — so read the shares above against the library and the rows against the corpus.
one apparent lead on that profile is worth publishing as a warning, because it is the kind that survives a careless reading. getenv shows up at 0.53% of the compile: the phase timer asks the environment whether reporting is on at every phase entry, and the inference fixpoint asks once per round. caching that answer is three lines and obviously right. it is also worth almost nothing, because glibc's getenv walks the environment block linearly, so its cost is a property of the shell that launched the profiler rather than of the compiler. the gate that counts this row runs under an emptied environment, and there the same work measures 0.06%. a profile taken at a normal terminal overstates it eightfold.
bucketing the same run a second way says who wrote the instructions rather than where they cluster, and that turns out to be the more useful question. of the run program's 2,399,081,635 retired instructions, 55.83% is code this compiler emitted from kanso source, 41.21% is the hand-written c in the runtime, 2.90% is glibc and the remaining 0.06% is the loader and a sort. every runtime shortcut landed over the fortnight above worked the 41% side. the larger half is the compiler's own output, so what a change there buys is bounded by how good that output already is — and two probes into it, one on the guards a match arm inherits and one on the failure checks the emitter writes at every call, both found work llvm had already removed. an instruction count read off the emitted ir is not a cost: encode_onto carries five failure checks in its ir and reaches its jump table after exactly one compare.
a second warning of the same kind, from the same profile a fortnight later. the emitter asks whether an ir line declares a stack slot, and it asked by searching the whole line for "= alloca " — a substring search over 144,261 lines, which read as 18.2 million instructions and looked like a clean lever. a stack slot is always declared at the line's first space, so the check became a read of one position. the measured fall was 1.8 million, a tenth of what the frame advertised: the search was finding its match early on almost every line it matched, and the frame's cost was mostly the lines where it did not match at all, which the new check still has to look at. an inclusive frame names an upper bound and nothing more.
what this leaves is per-call work spread thin over short paths, on both sides. the way down from here is emitting fewer instructions rather than finding a hot one, and anybody opening a profile of this project should read this section before spending a day in it.
58the carry tier, decided by a byte count instead of a path
the beat gives a loop back its own garbage at every iteration. the carry tier is the version for loops that thread a value through: the value is staged, the arena rewinds, the value is copied back. it costs a copy per iteration and it is the difference between a scan that holds one block and one that holds a hundred and eighty-nine.
which loops get it is decided by where their source file lives. beat_loops builds a set from d.file.starts_with("std/") || d.file.starts_with("lib/") and strips those groups' carries. the reason written beside it is real — a shared library driver threading its caller's source through a loop would copy an unbounded value every iteration — but a path is a poor way to ask that question, and this section is the attempt to ask it properly.
the case for trying: the run program's arena peak is 45,944,528 bytes, and one of its eight phases holds 38,797,312 of that on 4.87% of the instructions. the loop holding it is std/regexp's walk, refused by the prefix. let it carry and the same benchmark holds 1,048,576 bytes in one block and runs 2.4x faster. that is not an artefact of a benchmark whose pattern never matches: add a capture so the loop fills its slots at every position and the fall is 489,684,992 to 2,097,152, with allocations within two thousand of each other and 201 KB evacuated to buy 490 MB.
the property the prefix stands in for is the size of what the copy moves, and the runtime already knows it. k_beat_iter_carry sizes everything staged before it copies anything, so the bound is three lines at a point where the number is in hand: over the threshold, return without carrying, which leaves the arena un-rewound — grow-only, which is what these loops do today, so the fallback is correct by construction. at four kilobytes it admits the regexp walk and refuses sha256's, whose state and schedule are rebuilt every round: the digest benchmark goes from 4,569ms and 560 MB to 309ms and 2 MB.
the first version of it was eighteen times slower than the baseline on the run program, because the bound is checked after the sizing walk and a loop that declines pays that walk every iteration for nothing. latching the decision per beat depth fixes it — 7,334ms to 2,184ms with the peak and the evacuation byte-identical, which also says the latch changes nothing but the cost.
and then the trade does not close. sweeping the threshold on the run program against a 405ms, 45,944,528-byte baseline: 512 bytes costs 11% of the time and moves the peak 0.4%, because the phase that holds the memory declines at that size; 4,096 bytes takes the peak to 10,292,944 and costs 5.4x the time. nothing in between gives both, because the carries that buy the win are the same size as the ones that cost the time. 128 bytes makes the peak worse than the baseline, which is its own unfinished question.
so the prefix stays, and the objective is why: run speed carries more weight than run memory and is the term furthest from satisfied, so 5.4x the instructions for 4.5x the peak is declined by a wide margin. what would reopen this is a cheaper evacuation rather than a better threshold. two other shapes were tried and rejected on the way — removing the exclusion outright, which left the run program still running after ten minutes, and keying the decision on the inference sets of the carried positions, which separates these two programs only because one is well typed and the other defeats inference.
59why the whole-program check runs once per module
compiling a module compiles everything it imports first, and each of those runs the whole-program check over everything loaded so far. seven modules means seven checks over a growing prefix of the same program. doing it once at the end instead looks like free money, and an ablation priced it at 10,632,023 instructions, 20.38% of the compile row.
the smallest program that says why not is three lines. a module defines an accumulating tail call, an entry file imports it and runs it a million times, and with the check moved to the top the program dies:
error[runtime]: the program ran out of stack: recursion went deeper than
the stack holds
trmc::rewrite turns that recursion into a loop, and it declines any declaration whose name carries a slash. a dependency's declarations are qualified as they are merged, so count arrives at the top as trmc_count/count and the rewrite passes it over. trmc only ever reached that function because it ran inside the dependency's own compile, while the name was still bare.
it is a family, not an accident. six passes ask the same question the same way: trmc::rewrite, canonicalize_bare_aliases, typeset_constructions, foreign_constructions, and check_bare_ambiguity at two sites. each is about a module's own declarations — what this module constructs, which bare name its arms make ambiguous, which aliases it spells short — and a qualified name is how the merge marks a declaration as somebody else's. run once at the top they apply to the root and skip every dependency. four of the six are checks, so the failure they produce is a refusal that quietly stops being raised.
so most of the 10.6M is work that would stop happening. an ablation that drops those six passes measures a different compiler, and the saving it reports belongs to that one.
callgrind says what is actually reachable. against 51.6M on the fixed corpus, check_merged is 19,092,779 and infer::infer inside it is 11,803,600, none of which asks about slashes. the four guarded checks together are under 0.02%, so keeping them per module costs almost nothing. the two guarded rewrites are 1,651,829 and 504,558 and have to stay. that leaves about 8.5 million instructions, roughly 16.5% of the row, for a version that moves the whole-program check up and leaves those six where they are.
the slash is one mechanism of three, and the split was built to find the other two. with the three slash-guarded checks kept per module, scripts/module_differential reads 29 modules and 2 wrong: the arity message quotes m/one where the source says one, and the dispatch message quotes m/twice for twice. neither check is slash-guarded. what makes them per-module is that the message names a declaration, and at the root the merge has already qualified it — the spelling rule this compiler settled when a module gained the right to be named the way an import writes it. keeping those two per module returns the sweep to 29 and 0.
the third mechanism is where the saving actually goes. with five checks per module the suite is 122 binaries and four fail. one of them is the measurement showing up: a spec pins how many times the front end infers the whole program, and running the check once at the root is what moves it. the other three are one defect. a diagnostic raised on a dependency's declarations at the root loses the file, the span and the module suffix that the dependency's own compile supplied:
error[naming]: `silly` answers only true or false: name it `silly?`
--> deep_library_error/main.kso:4:8
error[naming]: `silly` answers only true or false: name it `silly?`
(module deep_library_error/deep)
the first points at main.kso, which does not contain the fault. detection is unaffected, because the dependency's declarations are all in the merged program; what is lost is attribution. so any check that can fire on a dependency's declarations has to stay per module unless a root-raised diagnostic can name where the declaration came from — and that is nearly all of them, with infer::infer, 11.8M of check_merged's 19.1M, running for any check that reads inference.
the figures above are a ceiling, not an estimate of a change. an env-gated skip of the whole check for every non-root module reads -18.33% on the module row and -23.30% on the library row, and reaching either needs provenance on merged declarations rather than a two-way split of check_merged. that is a larger piece of work and it is what this thread owes. kanso#1003 walked this route once and withdrew it as "the per-dependency check_merged is not redundant" without naming what made it so; there are three things, and they are written down now.
60a prefix belongs to a name, not to the head of a path
builtin_ names are how the standard library reaches the engine, and a program that writes one for itself is refused: builtin_length is the library's business. the resolver enforced that by stripping the prefix and asking whether what remained was a builtin it knows.
it asked the same question of a qualified name. builtin_shapes/circle lost its prefix, shapes/circle was not a builtin anyone has, and the reference was refused as internal to the standard library. every use of a module whose own name began with those seven bytes met the same answer, so such a module could be imported and never used. two files reproduce it:
builtin_probe/builtin_probe.kso pub hello = "hi"
main.kso import "./builtin_probe"
print builtin_probe/hello
error[name]: `builtin_probe/hello` is internal to the standard library
— import its module
the rule the check wanted is about a bare name. a qualified name is a declaration somewhere else, and what that module calls itself is its own business. so the prefix test applies to bare names, and a module may be named for what it holds.
61the compiler has two front doors and the objective watched one
kanso check <directory> is a module and takes compile_module_inner. kanso check <file> with a top-level expression is an entry and takes compile_parsed_entry, which merges the imports itself and runs its own whole-program check over the result. every compile gate in the tree checked a directory, so the second path was watched by nothing — and every kanso run takes it.
the gap cost three changes their measurement. all three touched the entry path and all three were read off a corpus staged in a temporary directory, where the count moves about 160 instructions per character of path; one of them projected a rise of 629 from a container where ci read a fall of 367. bench/entry_instructions_golden.txt pins it now, at 82,612,660 against the module row's 24,998,478, both counted on one binary in one job. its corpus names ten imports where the module corpus names four, because the entry path's own work is the merge and the check over everything the imports bring, and a corpus with one small import measures mostly the module path underneath it.
the ratchet asks the entry's whole-program check twice and the two veins separate cleanly: the entry row rises 15.25% and the module row is byte-identical. a vein whose defects another vein already catches would not be worth its callgrind run.
the objective reads both rows now. its compile term weighed one of the two compiles, which made a real improvement score as a loss. enroll_bare clones every exported declaration of an import under its short name, body and all — 145 of 882 declarations on this corpus — and the entry's whole-program check walks both copies. skipping the twins in three body walks costs the module row 3,306 instructions and saves the entry row 273,755, eighty-three times as much, and the term could see only the smaller number. so the skip was measured, held back a change, and shipped once the term could score it. summed, the two rows fell 212,644,590 to 212,374,141.
summing changes what is measured rather than what the compiler does, so the baseline moves with it and the score does not. the entry vein has no reading at the objective's epoch — its corpus did not exist then — so its baseline is imputed at the ratio the module row holds, which preserves the term to ten digits. that is the same rule this project has applied four times before, and it is the rule the corpus move of the same week did not get.
a change of measurement banks nothing. when the three compile gates moved from the json library to a fixed corpus, the counters roughly tripled because the workload is bigger, and the baselines they divide by stayed where they were. every ratio collapsed and the index fell six points for a change that touched no compiler code. the recorded reason was that the index has an arbitrary origin, so re-deriving a historical compiler's cost on a corpus that did not exist then would buy nothing — but the repair needs no history. it needs one measured factor per row, taken on the same head under both definitions, and that pair had already been measured. scaled by 2.7232, 2.7207 and 2.1047, the three baselines return every ratio to the digit it held before the corpus moved, and the chart goes flat across the change.
62a slash a person did not write
an import may be used qualified or bare. canonicalize_bare_aliases is the pass that settles that: it rewrites each bare use to the qualified name the import resolves to, so everything downstream sees one spelling. for most of the compiler that is the point. for two checks it is a problem, because both of them read a call's name and mean something by the shape of it.
the first is arity. its message quotes the head of the call, and a reader who wrote one 1 2 should be told about one — that was settled when a module gained the rule that a diagnostic names what the import writes. quoting m/one sends them looking for text that is not in their file.
the second is the opacity check, which refuses reaching across an import to construct another module's type. it needs no shadowing set, and its own comment says why: a qualified name can never be a local binding, so the slash is the foreignness. that holds exactly as long as every slash was written by a person. run the rewriting pass first and a slash also means the pass put one there, and a plain call of an imported function reads as a construction of the imported type that happens to share its name. the program compiles and the compiler refuses it.
the two want different repairs. arity wants the spelling the program used, so the pass hands it a record — line and column to the bare name it replaced — and the message quotes that. opacity wants no message at all: the site is not a construction, and no wording of a refusal would be right, so a head the pass rewrote is skipped.
the reason this is worth a section is where the fault was found. the entry path had cases in the module sweep, and both times the reordering was tried there the sweep caught it inside a round and it was reverted. the module path had the same reordering, done three days earlier, and no case at all — so the sweep read zero wrong while kanso check on a module refused a program that compiles. a corpus decides what a sweep can see, and a path with nothing pointed at it reads clean whatever it does. the module path has cases now, and so does the library path.
the library path took the same reordering last, and it went in with the record from the start rather than after two reverts. kanso check on a file of definitions is the third route through the front end, and until the day before it had no counter of its own — so the change waited on a vein rather than on a measurement. removing the record and running the sweep is what shows the record is load-bearing: 29 modules, 2 wrong, one of them the opacity check refusing a program that compiles and the other the arity message quoting m/one where the source says one. with it, 29 and 0. the vein that watches it fell 1,282,921 instructions, or 0.78%, landing at 162,840,377 in the same job as the other two compile rows, which did not move at all. two changes have moved it since: the fold-seed uniqueness condition in the linearity analysis took all three compile rows down together by about a thousandth of a per cent each with the allocation and peak counters byte-identical, and the allocator swap at the end of §08 took a further nine per cent off every one of them. the row reads 83,095,485 today. the ratchet asks that path's whole-program check twice, which turns the new gate red and leaves the module and entry rows alone.
the same habit turned up eight times across the emitter and the linearity analysis, and every one of them cost the whole program. a pass that reasons about names — which names a body mentions, which are used as values, which declarations a symbol table still needs — answered its question by walking every declaration, once per name. a thousand names against a thousand declarations is a million walks to answer a thousand questions. the repair is the same every time: walk once, build a set, ask the set. kanso build bench/runbench fell 60.00% when the linearity analysis was indexed both ways, 69.64% when prune_unnamed stopped re-walking, 31.38% on the beat’s tail-call question, 9.61% on the fixpoint’s group scan and 8.94% on the emitter’s, each figure against the head it landed on, so they compound rather than add. the emitted IR is byte-identical on every one: the compiler decides the same things and stops asking the same question a thousand times. the eighth was not a scan. three analyses were each building the same linearity result from the same program, and the emitter builds it once now.
63the chart of the project's own cost, and what it was drawing instead
the numbers page publishes a trend chart of what this compiler costs, commit by commit, over five hundred rows. the rows are written by ci on every push to main. each line is one of the counters the welfare objective weighs, so the chart is the long view of the score.
for most of a year it drew something else. the objective was rebuilt twice — the compile side in september, retiring fixpoint rounds, expression visits and emitted lines, and the runtime side days later when the benchmarks became one consolidated program — and each rebuild left the chart drawing counters nothing was optimising. five of its six lines read retired counters. the two the objective actually reads for a run, which every row had carried since the day of that gavel, were drawn nowhere at all, and one of the retired lines held two distinct values across every row that had it, so it drew as a flat line for a year.
the asymmetry is worth naming, because it says where a check was needed. the row side cannot drift: scripts/perf_record writes a row's objective counters straight out of the scoring program, unfiltered, so a counter that joins the model joins the row in the same commit. the page named its counters by hand. so bench/objective_sources.txt is the list both sides answer to, and a spec replays it against the chart in both directions — a term with no line turns it red, and so does a line reading a counter the objective does not weigh. the two lines that are deliberately not objective counters are named in that spec with the reason each one stays.
the colours had the same shape of fault. one palette was drawn on both the light and the dark page and had been chosen for the dark one; on the light surface every one of the seven lines sat under three to one, and three of them crowded the blue band closely enough that a reader with common colour-vision deficiency could not separate the nearest pair. colour on a categorical chart is computable, so it is now computed: each mode is stepped for the surface it is drawn on, and the order was searched rather than picked. of the five thousand and forty orderings seven hues admit, five hundred and thirty-six clear the adjacent-pair floors in both modes, and two hundred and sixteen of those also clear them among the three compile lines taken every pair against every other — which is the comparison a reader of this chart actually makes, and one the adjacent check cannot see.
three of the light steps still sit under three to one. that is allowed only when the numbers are legible some other way, and they were not: the panel beneath the chart carried a row for none of the seven series, and explained three counters the objective had retired. it now opens with the chart's own lines, in the chart's own order, built from the same list the chart draws from rather than from a second one written by hand. a faint line is readable when its number is written down under it.
64a file that is its own input
the perf history is written by ci on every push to main and read by ci on the next one: the job appends a row, rescores all five hundred under the current formula, and pushes the result back. that makes the file its own input, and a tool whose output returns to it as input has to be right about what it wrote.
on the ninth of september it was not, and main went red on a refusal that was working correctly. the rescore rebuilds a run count for the rows measured before the consolidated benchmark existed, and splices that rebuild onto the first row carrying a real count. it refuses when the rows either side of the splice did different work, because a rebuilt column spliced onto a row that moved slides the whole history smoothly and plausibly, and nothing downstream could tell. the guard fired. what it refused was the tool's own output: forty-eight rows had gained a rebuilt count on the previous run, nothing in the file said so, and on the next pass the oldest of them was chosen as the anchor and compared against a row forty-seven commits away.
two repairs, different in kind. the tool now chooses anchors among rows that carry a count and do NOT carry the full set of per-phase counters, which is what separates a rebuild from a measurement given only the data; and it recomputes rather than keeping a count it finds, so a corrupted file repairs itself on the next run with no hand edit. the property under all of it is that feeding the tool its own output gives the same answer, and a spec now holds that red in a three-row fixture.
the second repair is for a later reader. a row says where its run count came from — rebuilt, measured, or nothing at all when it has no count. the mark carries no factor, and that is the point of it. the compile side of the same file IS re-based by an exact divisor: the workload the compile term measures has been re-measured three times, each change of ruler has a per-counter factor recorded against the commit that made it, and a row is divided by the product of every factor from its own epoch forward. the two pure re-basings then flatten to nothing and what is left standing is the compiler having got faster. the run side has no such number. its rebuilt counts come from per-phase shares rather than from a ruler that moved, and a 1.0000 written beside them for symmetry would say the opposite of what happened.
the mark is sticky, and the first cut of it was not. it looked like the welfare column, which is recomputed from scratch on every pass, so it was built the same way and stripped before each rebuild. whether a number was computed is a different kind of fact: nothing re-measures a commit from three weeks ago, so a rebuilt count stays rebuilt, and a pass that cannot rebuild anything must not read the counts it finds as measurements. with the anchors removed from a fixture, the re-derived mark relabelled all forty-eight rebuilt rows as measured. the sticky mark leaves them alone.
65what the run program pays that the program did not ask for
by the middle of september the queue had spent a fortnight on the library: shapes in lib/json and lib/text, the scans, the folds, the appends. the profile that came back off the consolidated run program had almost none of it left at the top. what was there instead was work no line of kanso asks for. one instruction in ten is a register save or restore at a call boundary. every inlined fast path the emitter writes begins by loading a counter switch and branching on it. neither is the program. both are what the program costs to run.
the register saves are the compiler's calling convention meeting a hot recursive descent, and the lever on them is how eagerly llvm inlines. the obvious next move was a rule: find the functions whose bodies are small and whose call sites are hot, and mark those. that was built and it does not work, and the reason is worth keeping. an inline decision's value belongs to the whole optimisation state after the decision — what constants propagate through, which branches fold, whether the caller's frame still needs to spill. eighteen functions were marked one at a time and measured one at a time; the sum of the individual wins and the win from marking all eighteen disagreed by more than either. no rule written over functions can see a quantity that lives over programs.
so the lever stayed the global one. the release link's inline threshold went from llvm's default to 500 and then to 2000, and the second step took a further 1.2281% off the run program on ci. this container projected 2.08% for that step, which is 59% more than landed. the container orders the rungs of a ladder like this and it does not size them; the number that goes in the log is the one ci measured.
the counter gates are a different shape and a larger number. the cost goldens count allocations, appends, string scans and two dozen other events, and they can only count them if the emitted fast path checks whether counting is on before it takes the shortcut past the runtime call that would have counted. that check is three instructions at each of eight emitted sites and twenty-seven runtime ones, 303 sites in the linked binary, and it asks a question whose answer was fixed before main. folding it away by hand and relinking with the shipped recipe reads 1.8286% of the run program and 4,992 bytes of machine code, the emitted and runtime halves exactly additive in both.
kanso build --counters now keeps the gates and a plain build strips them, so the twelve benchmarks the cost goldens read are built twice and everything else is built once. two cheaper escapes were tried first and both are worse. marking the loads invariant is arguably sound — the switch is written once, before any emitted code runs — and it does common-subexpression them; it measures 0.0877% slower with 288 bytes more machine code, because keeping the value live costs more than the reload saved. and the gate cannot simply be deleted with the fast path counting for itself: the emitted path takes its header from the arena, where the slow path reaches the allocator and increments two counters the goldens pin, so dropping the gate blinds them and keeping them exact costs three unconditional read-modify-writes against the gate's two instructions.
66a fixpoint that asked everything, over and over
the advisory pass tells you which failure a function can hand back, so a reader of a signature knows what it is taking on. it works out the answer by fixpoint, because a function's failures include the failures of everything it calls, and a call can run in either direction down a file. the way it ran the fixpoint was the simplest one there is: ask every declaration, in source order, and keep asking until a whole pass changes nothing.
on the corpus the compile gates check, that is 315 declarations over six passes, 1,890 visits. 142 of those visits learned something. the other 1,748 walked a body, collected the same answer as last time, and returned it.
the fix is the standard one and the only question was whether this pass could take it. it can, because a body asks about a fixed set of names: the body does not change during the fixpoint, and neither do the dispatch groups or the type names it reads, so the declaration indices a visit consults are the same indices on every visit. the first pass records them as it takes them and builds the reverse map. after that a declaration is re-asked only when an answer it read has actually grown.
the compile term came down 1,032,618 instructions on that change, the entry path 2,517,319 and the library path 2,519,532. allocations fell 424, because each visit builds a set and there are fewer visits, less what the worklist keeps for itself: a reverse read map and a queue; that count reads 14,319 today, after the lexer stopped allocating per word and per number and nine front-end passes stopped building containers they fill once and drop. the three instruction rows have: they read 24,998,478, 82,612,660 and 83,095,485 today. the allocator change at the end of §08 took nine per cent off all three, and later changes have moved them since, including the address-blind memcmp and memcpy every counted run preloads from 2026-09-23, which added under one per cent.
this change was measured twice, against two bases, because another change landed in the same function in between. the three instruction rows fell less the second time — 6,069,469 summed where the first sitting read 6,434,951 — and the allocation row fell by exactly the same 424 both times. that is the difference between a counter that counts operations and one that counts a share of them. the earlier change made a visit cheaper, so the visits this one stops taking are worth less; it removed no set, so the sets this one stops building are worth what they were.
the answers are the same, and that was checked rather than argued: both binaries over 1,074 files and 45 module directories, nothing diverging. a fixpoint's result does not depend on the order its queue drains — the union is monotone and the loop runs until nothing grows — but the differential costs a minute and the argument is not the evidence.
67two encode leads that turned out to already be right
both of these are declines, and they are here so the ideas stay declined. the section before this one is a change that landed; these two are the same amount of work, spent finding out that the code was already doing the thing.
the first was a standing note that the encode path pays for a 32-byte arena conversion on every append, worth about 0.61%. re-attributing it on merged main refuted the note outright: every append the emitter writes already mutates in place, so there is no conversion to remove. the 80.5/19.5 split the note rested on was a reading of the wrong two rows.
the second was ryū's pair loop, which extracts two decimal digits at a time out of a table. the loop looked like it carried a division that a multiply-high could replace. it does not: the division is already compiled to a multiply-high, and that is visible in the disassembly rather than inferred. two reformulations were written and both measured faster on a first reading — and both had introduced a real divide that the original did not have, which the second reading caught. a reformulation only measures anything if it is still computing the same thing, and that has to be checked in the disassembly rather than assumed from the source.
both cost a day the queue could have spent elsewhere, and both are written down here so nobody re-opens them. the standing note claimed 0.61% and carried no measurement, which is how it survived long enough to be worth a day.
68seventeen digits in, four digits out, one at a time
a double carries about seventeen significant decimal digits. a float a program writes down — a price, a coordinate, a measurement someone typed — carries three or four. printing one means working out which of the seventeen the shortest representation actually needs, and that is done by taking digits off the end until taking one more would make the number ambiguous.
so the loop that takes them off runs about thirteen times per float, and it was taking one digit a trip. sixteen instructions each: three multiply-highs and three shifts, which is what dividing three 64-bit numbers by ten costs, plus a counter, a compare and a branch. measured on the encode corpus it ran 9.41 times a call and carried a third of the whole renderer.
a step that takes two digits at once was already sitting beside that loop, and it fired once. the loop that replaces both takes twenty-one instructions a trip and runs 5.35 times a float, against sixteen instructions run about thirteen times. seven of the twenty-one are register moves carrying the three bounds around the back edge, which is what the fused shape costs and what the next change on this function would go after.
the digits are the same digits, and that was checked rather than argued: all four binaries run, two programs, byte-identical output. the loop decides how fast the digits arrive and nothing else about them.
the renderer costs 468.5 instructions a float, and it cost that in both benchmarks that call it — 191,070 calls in the run program and 849,200 in the encode benchmark, agreeing to one decimal place. two corpora with nothing in common except that the numbers in them are numbers people wrote.
69a question with no call sites answers yes
the licence to write through a parameter has two halves. one is about the function: does this parameter thread through to what the function returns, so that writing into it writes into something the caller is handing over anyway. the other is about everybody who calls it: does every call site pass a value it uniquely owns. the second half is a walk over the whole program, gathering the call sites of a group and asking each one.
a folder handed to list/fold is mentioned and never called. the program writes list/fold xs seed step and step sits there as a value; the call happens inside fold, at a site no walk of this program will find. the walk gathered nothing, found nothing to object to, and returned yes. every named folder in the tree had its first parameter marked an accumulator the compiler may write through.
so the seed was written in place while the caller still held it, and two folds over one seed read each other's writes:
seed = { "k":0 }
first = (list/fold [1] seed put_step)["k"]
second = (list/fold [10] seed put_step)["k"]
the interpreter reads 1 and 10. native read 1 and 11, because second folded over the map first had already written into. lists and byte builders went the same way, and an inline lambda in the same position was correct throughout — the fold's own arm asks whether the seed is unique before it licenses a write inside a lambda, and only a folder with a name reached the licence by the other route.
the refusal for this exact case had been written, and its own comment said the thing: handing a group to a fold means it is called from somewhere no call site in this program describes, and a question about what every caller hands over then has no calls to look at. it sat in a function the granting path never called. the caller half existed twice — once with its refusals in front, once as a bare walk — and the licence asked the bare one. the two are one function now.
a walk that finds nothing has two possible meanings, and they are opposite: nobody does this, or nobody here can see who does. an analysis that cannot tell them apart will grant on the second and call it a proof.
70a type name inside a shell is several names
a file's imports are checked against the module qualifiers it uses, and the compiler decides which those are by reading every identifier in every body. it found them by splitting a name at its first slash. that works for os/process written on its own, and gives the wrong answer for the same name written inside a type shell: <os/process>effect holds one qualified name and one bare one, and splitting the whole spelling at its first slash answers <os, a qualifier no import can match.
the file was then refused for borrowing an import it had written, and in the same run that import read as unused. two diagnostics, both wrong, for a program that compiles. map[string os/process] and []os/process split the same way.
the check scans the runs of name characters now and marks the ones holding a slash, so a shell of any shape carries its names through.
this decision is made once per identifier per body, which puts the scan on the front end's hottest path at import-check time. the first shape of the fix built a split iterator for every name and cost the compile corpus 1.23 per cent. the overwhelming majority of names hold no slash and so can hold no qualifier anywhere inside them; answering those on one scan, before the iterator exists, takes about eighty per cent of that back. the residue is what a slash-bearing name costs to scan properly.
71the box is explicit, and a bare err is a value
a call that can fail used to hand its caller a value or a failure, and the caller found out which by looking. the box is written now. <t>effect is a type a signature can name, the three chain words are the only doors into it, and nothing lifts a value into one behind your back. an err that is not in a box is data: it can sit in a list, ride through a constructor, match an (err _) arm anywhere, and be handed from one function to the next without ceremony.
what changed at the check is the last clause. an operator, an index, or a call with no (err _) arm at that position does not compile when a raised err can reach it — exactly as it does not compile today when a none can. the rule reads inference's answer sets with one bit added for a raise, and it reads calls: price "soup" is not a raise when only price "flan" raises, and an as-pattern binds what its pattern caught, so fn taken e@(err _) = e hands the raise on. it does not read names. x = decode s followed by f x is accepted, which is the same blind spot the none rule has always kept, and chapter 4 says so rather than leaving a reader to find it.
the tree had to be written to the rule before it could compile under it. 191 sites were refused on the first run. std/regexp gained eight hand-back arms at its entries, std/json five, and fourteen scripts bind a raise before reading it; the three vendored decoders and the benchmark programs take the same shape. none of those arms is new work at runtime — they name a failure the caller was already going to see.
xs[i]! answers a box too, holding the element or the missing-index err, and the words are what open it. the plain read xs[i] answers the element or a none for an arm to handle, unless the ifs around it prove the index in range: inference reads a guard returning before an out-of-range read, a comparison against length xs on either side, a literal offset, and a literal list's own length. 745 sites in the tree lost their bang because the value is what was wanted; 208 guards were written where the bound had not been spelled; 41 stand, each opened by .> or annotated by .!.
the bound prover is what this costs. it runs per read and its answer is recorded per span, and the compile corpus pays for it; the run side takes some of it back where a proven bound emits no check at all. the floor came down by exactly that much under the rule that a ruled part of the language lowers it rather than waiting, and the history entry beside the number says which ruling.
72three instruments that could not fail
the ratchet is a list of pairs: a deliberate break, and the gate that must go red on it. a row whose gate stays green is a gate nobody has proved can speak, which is the thing the list exists to find. three instruments turned out to be looking the wrong way in one week, and two of them were the ratchet's own.
the first watches lazy binds. its mutation flips a per-arm fold from and to or, so a bind goes lazy when any arm of a group qualifies rather than every one. the verdict is keyed by group, arity and index, with no arm to disambiguate, so the emitter builds a thunk inside the arm that qualified and releases it at the group's merged tail — which the other arm reaches. the row's recorded claim was a counter moving, thunk allocations from zero to half a million. what happens today is that the module stops verifying: the release does not dominate its use, clang runs with the verifier disabled, and the compile dies in the register coalescer instead. the benchmark never builds, a setup that fails reads as unbuilt, and unbuilt scores as red. the row had been passing on a crash. it points at the micro corpus now, where the same defect is caught in a second and a half by a callable that answers differently as a library.
moving it left the scan counters with no row at all — it had been the only mutation naming that gate. so a second was written for it: clear the prefix filter that keeps library loops out of the carry tier, which the benchmark's own header names as the reason its peak grows with the subject. the header was stale on the numbers. clearing the filter moves the beat iterations from 15 to 16 and the surviving slots from 0 to 2, and leaves the arena blocks where they were; the sitting the header records had changed three things at once, and a third of a change is not a third of its effect. the gate diffs the whole golden, so two moved rows turn it red exactly as the predicted one would, and the row's claim is now written as what the mutation does.
the third instrument reads this page against the log. it counts the lines a commit adds that begin ## and fails when the log has run more than three entries ahead of what a reader has been shown. two entries landed with their titles hard-wrapped at the column limit, and the wrap put ## at the start of the second line as well. markdown renders that as two headings, the second a sentence fragment; the gate counted two entries where one had been written, so each of those changes spent two of its three and every later branch inherited a budget already half gone. what separates a real sub-heading from a wrap is what sits above it: a heading follows a blank line, a wrap follows the line it broke from. the rule is now that no ## line is immediately preceded by another, and the archive's 1,272 entries had zero violations before it was written.
73what a claim about a measurement has to survive
two claims went out on one afternoon, each resting on a number that was real, and each wrong in the same way: the number was never checked against the scope of what produced it.
the first said the compile rows could not have moved by layout, because this tree has measured that term twice — seven binaries spanning 1,028 instructions on the anchored frame, and a hundred-function pair reading 5,849 — and the change in question moved them 140,122. two orders of magnitude is a real objection, and it produced a bisection that attributed 145,472 of it to 74 lines rewriting two functions a check never calls. that attribution was then written down as a mechanism: rewriting unreachable code moves what sits around it. building the ladder for it settled the matter the other way. eight rewrites of one unreachable function, eight distinct binaries, and the row is identical to the instruction across all of them. the objection was worth making, the explanation that followed it was wrong, and the move is unexplained again.
what the three calibrated shapes say now is that the term is small in every one of them: about 402 for an addition nothing reaches, zero for a rewrite of the same, 2,733 for an addition that is reached. the move under discussion is fifty times the largest of those, which puts it outside the family rather than at the top of it. naming it wants a frame-level diff of the two profiles, not another argument from a table.
the second said a check rule was missing, on three programs that each bound a failing value to a name before the position read it. the rule reads calls and not names, deliberately, and the section describing the rule says so. a correction to that claim then said the book fails to document the exception, on a search for the phrase “blind spot” — the book documents it, in the words the checker reads calls, not the names they are bound to.
what the three have in common is the shape of the reasoning rather than the subject. a measurement bounds the thing it measured and not the thing it resembles; a fixture shows what happened and not what was supposed to happen; a grep finds a phrasing and not a fact. none of the three claims could be repaired by measuring harder, and all three were one paragraph away from being avoided.
74three rows for one build, and the process that waits
a build runs five processes. kanso reads the program and writes llvm ir; the clang driver reads the ir; a second clang compiles the convention probe; clang -cc1 does the real compile; ld links. for most of this project's life the objective weighed none of it. every compile counter ran kanso check, which stops before codegen, so the optimiser's cost sat outside the index while its output was paid for.
the first attempt at fixing that counted the whole tree and put the number in a golden. it would not reproduce. two readings of one binary in one staging, at -O3 -flto, came back 3,105 apart — and a row read by an exact compare cannot hold two values. the gate was taught to print a cost beside every process it names, and the next reading answered on the first try: the three clang processes and ld were byte-identical, and all of the difference was kanso's own process.
inside it, two frames of 1,346 moved. one was memcmp, and the cause was the process id: kanso puts its pid in the names of the temp files it writes, and an earlier round had pinned the width of that tag at seven digits without pinning the content, so a comparison over those paths stops at a different byte each run. a probe binary with a constant tag took that frame to zero and the difference to 112. what was left sat in one frame — build itself, self cost, every callee identical — which is the inlined loop that waits for clang. the development tier is the control: the same code waiting on a child that finishes seven times sooner reproduces exactly.
a loop whose iteration count belongs to the scheduler is external state, and the rule here is that external state is normalized before it is measured or excluded and named. it cannot be normalized, so it is excluded: the codegen rows count the child tree, and the exclusion is written into the golden's header with the measurements that put it there.
that leaves what kanso itself spends emitting, which is real work and had no row at all. it gets one, anchored at codegen::emit_ir — the frame the wait sits outside of. four profiles were already on disk from the reproduction work, two binaries and two readings each, and every one of them reads that frame at the same number while the process around it moves. so a build is three rows now: what checking costs, what emitting costs, and what clang and the linker cost, each counted where nothing else can reach it.
75one number could not tell a compile from a run
the objective weighed seven counters into one score, and the counters came from both sides of the compiler: what a program costs when it runs, and what it costs to compile. a single sum has to answer one question about them — is this trade worth it — with one exchange rate, forever. clay ruled on 2026-09-16 that it becomes two scores and a third over them.
the argument that settled it is start-up. kanso test pays it on every invocation and a shipped program pays it once; the same microsecond is enormous in one context and invisible in the other. a single objective has to pick a number for that microsecond and be wrong in one of the two places. so the terms are sorted by who pays them. production weighs run speed at 0.45, run memory at 0.40 and the release build at 0.15. development weighs compile speed at 0.30, start-up at 0.25, the dev build at 0.16, interpreter speed at 0.11, compile memory at 0.08, emitting at 0.06 and interpreter memory at 0.04. each side sums to one, because a weight is a share of its own score and nothing else.
that last sentence cost a round. the four pre-split weights were carried over without renormalising, production summed to 0.71, and the score read 40.97 where it should have read 57.11 — a plausible number, in range, wrong. what caught it was recomputing the three scores by hand before asking the program. the refusal that catches it now is a check that each side sums to one, and the program exits rather than scoring an unbalanced model.
the meta over the two is 0.70 production and 0.30 development, and it saturates rather than adding. a sum would fix the exchange rate at 2.33 for all time: a development gain would always have to be 2.33 times the production cost to be worth taking, whatever the state of either side. saturation makes the rate a function of position. from where the compiler stands today it is 2.82; at parity 2.33; with production far ahead 1.09; with development far ahead 4.98. a side that is already good is cheaper to spend, which is what the third number is for.
and the baselines are half of any such term. six counters joined the model and three of them arrived with an origin this container had measured against a golden ci had measured. the objective compared the two hosts and read the difference as work the compiler had done: +68.3% on the dev build, +77.4% on the release build, +3.3% on emitting, a rise of 1.30 points with nothing behind it. each of those goldens says in its own header that the two readings are not comparable. with the origins taken from ci's first sitting the score reads 76.13 against a floor of 76.12766770905162, which is where an unchanged tree belongs. a term whose origin came from one machine and whose reading comes from another prices the gap between the machines.
76one question, asked once per name, five times over
a compiler pass that wants to know something about a name — is this call in tail position, does this binding escape, which arms share this group — can find out by walking the whole program and looking. that is correct, and for one name it is cheap. the passes here were doing it inside a loop over every name, so the walk ran once per name over a program that never changed between walks, and the cost grew as the square of the program.
five passes had the shape, found one at a time by profiling a build rather than by reading for it. the emitter asked two whole-body questions per name, and indexing them took 69.90% off kanso build on the run benchmark. the beat's two questions took a further 31.38% of what was left. the linearity fixpoint's group scan took 9.61%, and the emitter's own group scan 8.94% after that. each one is the same edit: build a map once before the loop, read it inside.
what makes this a campaign rather than four coincidences is that none of the four was visible in the rows the objective watches. all four are reached from the emitter, and every compile counter the objective weighs runs kanso check, which stops before codegen. so the passes could grow without moving a number anybody was reading, and the profile is what found them. the rows did move while this work was going on, by amounts of the same order — those are the linker placing a bigger binary differently, and they are not the work.
the emitted ir is byte-identical across all four, which is the check the shape asks for: an index that answers what the walk answered changes what the compiler spends and nothing about what it writes.
77which machine counted it
every instruction row on this page is an exact compare: one number, no band, and a build that reads a different one turns the job red. that only means something if the two readings came from comparable places, and two facts about where they come from were not being checked.
the first is what a binary's checksum says. the gates print compile_binary sha256= beside every reading, and a calibration ladder on this page leans on it as the proof that eight variants were genuinely eight builds. that use holds. the converse does not, and it was being assumed. two ci jobs built source differing in a log file and a shell script — neither compiled in — and produced .text=2796866, .data=12672, .bss=29912 both times, with different checksums. a container builds the same tree three times and gets one binary each time. so the build is deterministic on a machine and the file varies between machines, and a checksum that differs is not evidence that the executed code moved.
the second is the silicon. glibc picks memcpy, memcmp, strlen and their neighbours at load time, from a feature block the cpu reports, and the gates pin the cache-derived thresholds through GLIBC_TUNABLES so that part cannot drift. what a tunable does not reach is which implementation the resolver picks. the gates named the cpu family and the model and stopped there; a reader that compares the whole block had been written months earlier and had no block to compare against, so it answered "cannot tell" every time and nothing called it.
ninety-odd job logs print that block, because the reader prints it whenever none is recorded. within one job every printing is identical. across jobs there are seven distinct blocks differing in 57 rows, and the basic family takes three values — 0x19, 0x1a and 0x6, the last of them intel. level-three cache spans 32 MB to 480 MB. Fast_Unaligned_Load, Prefer_No_AVX512 and Prefer_PMINUB_for_stringop flip between them, and those are switches the resolvers read. the block is recorded now and the five instruction gates consult it when a row has moved, before their verdict and without deciding it: a different resolver is a candidate explanation, not a ruling.
the thing that sent anyone looking was six instructions, and the new instrument ruled itself out of explaining them. the two jobs above printed the same block, all 123 rows, and their four short veins agree to the instruction while the interpreted run is six apart on 2,178,502,266. six in 2.18 billion is three parts per billion; on the other four rows, which run from 4.8 million to 128 million, the same proportion is a fraction of one instruction and could not be seen. that is the shape of a term proportional to work rather than a constant, and the interpreted run is also the allocation-heavy one — 5,313,434 allocations against a compile's 27,397. where the allocator's heap starts moves with the size of the file the loader mapped. that is an argument and not yet a measurement, and the page says so.
78the blank edge of the welfare chart
the welfare history is five hundred commits wide, and its left twelve per cent is blank. a ruling from the seventh of september explains a blank left edge: rows before the reconstruction's baseline carry no run counters and stay unscored. anybody who knows that ruling reads the blank as the ruling. the dates say the blank is something else.
the dates settle it. the file today opens on 2026-08-20 and the baseline it was ordered to score from is 2026-08-10, ten days earlier. the window is five hundred rows and it rolls, so the rows that ruling excludes rolled off the end of it some time ago. every row now in the file sits inside the scored range, and sixty-two of them carry no score.
what those rows carry is counter names run together. one key reads compile_alloc_bytescompile_allocscompile_peak_bytescompile_passes and holds the number 5. the next commit's row writes those four as four keys, and compile_passes is 5 there too, so the welded key kept the last name's value and the other three numbers are gone. nine run-side names are welded into one key the same way, and six basket names, and seven encode names. the rescore walks all five hundred rows, finds nothing it recognises in the first sixty-two, and stamps them scored_weight: 0.00 against 0.23 for the row after.
it is old and it is shrinking. the same block was 151 rows on the twenty-seventh of august, and it loses one row per commit as the window rolls forward, so about sixty-two more commits clear it with nobody touching it. what welded the names is outside what the file can answer, because the rows that would show it are already gone off the end. so the mechanism stays open here, and what gets written down is the boundary and the delta.
this is on the page because of the reading it invites: a defect that happens to look exactly like a ruling doing its job, sitting where a reader would go looking for the ruling. the sweep that found it read every ruling august and september recorded — fifty of them — and found nothing else unbuilt.
counting those rulings turned out to be its own small lesson. the log writes a ruling as ## <date> — gavel: … and a count of that shape returns fifty, twenty-nine in august and twenty-one in september. the convention is younger than the project. before it the log wrote GAVEL:, GAVELED:, GAVEL (syntax): and GAVEL, IMPLEMENTED:, and a count of those returns twenty-three more — fourteen in july, six in undated headings, and three in august that the newer pattern walks straight past. no heading matches both, so the two counts partition seventy-three rulings between them.
the twenty-three were swept the same afternoon, and they made the count worth having. they are language rulings almost to a one — none and err, partial application, accessors, text blocks, as-patterns, what a record field may carry — and this project ships every behaviour with a golden, so the corpus is the probe. the golden suite passes, and it holds a named fixture for most of them. one error fixture, no_any_type, pins three of the rulings at once and quotes one of them back: a record field carries no type — write name and let the compiler infer what it holds. so all seventy-three have now been read against a build, and nothing in either set is unbuilt.
79the list of what is owed called itself a floor for eight days
a ruling that nobody builds is a ruling nobody can see. the list of them was created on 2026-09-09 after five sat unbuilt across 296 pull requests, and its own preamble said the list was a floor rather than a total, because one sitting — twenty rulings made on 2026-08-29 — had never been audited. for eight days that sentence stood, and on 2026-09-17 the list emptied for the first time, which is the state in which a floor is easiest to read as a total.
the audit ran that hour. each of the twenty was probed against a release build rather than read off its own text, and nineteen came back built or declined. two of the nineteen needed running rather than grepping, and both would have been reported wrong otherwise. the ruling that a qualified name means its own module's declaration asked for a red spec against one measured hazard, a dependency that declares one arm of a name while importing another module's arm of the same name; built as a pair of hakos, the bare call inside the dependency says no 2-argument arm of `join` (arms take 1), and the clone does not enroll. the ruling that behaviour arms travel with the type needed the group's real name: an arm written render m:money does nothing at all, because interpolation dispatches to_string. spelled correctly, money's own arm prints in a module that imports money and declares no interface, and a to_string s:string arm beside it is refused at the declaration.
the twentieth is the one worth the page. block-born is the whole cohort widened a field write's target from a name bound directly to a construction to anything the checker can prove was born in the block — through an alias, through a field of a born node, through an element of a born list, through a node an if chose. it was built on 2026-09-09 with all four. seven days later the build-hole ruling landed, which says a construction's unfilled field is spelled _ and is filled by exactly one write through the record's name, and two of the four went: a name whose birth is either of two records cannot be shown to fill one hole once. the reasoning is right and nothing here asks for it back.
what went with the two shapes is the earlier ruling's reason for existing. its words were that cyclic structures sized by data — a graph parsed from input, N linked nodes from a map — gain a spelling. against a build of the tree today every route is closed: an indexed element cannot fill a hole, a field built with a value cannot be written at all since the hole ruling, birth does not flow through a call, and N nodes cannot carry N names. the narrowing is recorded twice, once as one of seven refusals with its fixtures and once as a note about which lines of a golden a merge deleted. neither says that an earlier decision's purpose had been given up, and the two were never set beside each other.
the same afternoon produced a second row, and it arrived the way these usually do. a pull request changing two documentation files went red on interp_instructions, a counter the meta objective weighs: six instructions out of 2,178,502,266. the section after this one is what that turned into, and it is worth reading for the instrument rather than the finding: the six sent someone to build a reader for the whole cpu feature block, and the reader then ruled the silicon out.
the row acquired three explanations before it had one, and all three were published. host divergence, from two job conclusions. then a release build that does not repeat, from two checksums. then, when the other ten counters in the same two jobs turned out to agree to the instruction, the reading that those rows do not carry the exposure. the last is the one to watch: three parts per billion on a row of four million is a fraction of an instruction, so none of them could have shown it either way, and their agreement was never evidence of anything. each version reached the log, this page and a pull request comment before the next replaced it.
so the rest of august got the same treatment the same evening. counting the rulings turned out to need two greps, because the log has written a ruling two ways and this section said 56 and 35 before either was run. a count of — gavel: returns 50; a count of the older GAVEL:, GAVELED: and their variants returns 23 more, and no heading matches both, so the two partition 73. august is 32 of those and 29 are swept here, the three left over being the ones still spelled the old way. the five oldest sittings came back clean — three of their nine entries are build records rather than rulings, and the six rulings are built or superseded, with the bytes gavel's three claims run on both engines rather than read off its text. the last nine, on 08-24 through 08-31, gave up the second finding.
a knotted constant that something reads reports thunk_allocs=1 on native and 0 on the interpreter. both engines force it, both evaluate it, both print the same answer; one counter disagrees. that divergence was found on 2026-08-23, sent to the ledger in those words, and ruled the next day with the side that moves named explicitly — the oracle, whose builder makes the cell without touching the counter. twenty-four days later it still does.
it survived because nothing in the tree compares the two. the memory vein runs each fixture on one engine, and no differential script mentions the counters at all. the comment four lines above that loop names the extension that would have closed the gap, and is still written in the future tense. the gavel's own last line names the fixture that would have caught it, and that fixture was never written either. a ruling, a missing spec and a missing fixture, all three recorded and none of them done.
two unbuilt rulings in the twenty-nine, then, and twenty-one september entries left, two of them probed in passing and built. the list still calls itself a floor, and the word now stands on a count anybody can re-run instead of one carried in prose.
the habit underneath all of this is cheaper than any of the findings. a list that describes its own coverage — a floor, an audit, a sitting not yet swept — is telling you where to look. so is a gate that prints a binary's hash on every run: somebody wrote that line down because they expected to need it, and reading it first would have saved an attribution that was wrong.
80a quarter of start-up was hashing a constant
the interpreted engine's start-up became a weighted term on the 2026-09-16 gavel, so kanso play on a program holding one print was profiled. it retired 4,837,246 instructions, and callgrind put 25.35% of them in one function: the hasher.
two cache keys ask for it. one decides whether a staged runtime object may be reused, the other whether a linked binary may be; both must change when src/runtime.c changes, and both got that by handing the whole file to a hasher. runtime.c is 450,100 bytes, it is hashed twice, and 900,200 bytes at about 1.36 instructions a byte is the entire frame. the one-line program's own work is rounding beside it.
a constant's digest is a constant. the compiler that builds this one computes it, and the running compiler folds in eight bytes:
main 5,081,099
the digest 3,955,899 -1,125,200 -22.14%
the profile above was taken before seven whole-program scans landed underneath it, which raised the row to 5,081,099; the table is ci's sitting against that base. the saving is the same 1,125,200 either way, because it is the hashing that stops happening and it does not depend on what else the start-up does.
the digest is two FNV-1a accumulators with different primes and offsets, folded together so the key holds 128 bits rather than 64. a collision here would not be a slow build — it would be a runtime object reused against ir compiled for a different one, so the second accumulator is cheap insurance paid once at build time. the loop steps eight bytes at a time because const evaluation is interpreted and rustc refuses a long-running one by default.
the shape is worth naming because it is not about hashing. a cache key has to change when its input changes, and the cheapest way to write that is to hash the input — which is correct, and which quietly moves the cost from build time to every single run. what the input is fixed at compile time is the question worth asking before reaching for a hasher.
81what the encoder's frame is for
the run program's top frame is the json encoder's dispatch: 392,547,176 instructions over 2,380,860 calls, 21.5% of everything the program does and 164.9 instructions a call. eighteen of those run on every call before the dispatch has decided anything — six callee-saved pushes, an 88-byte stack frame, the failure test, and the jump table.
the dispatch has eight arms and only two of them loop. counted off the per-address costs, every call is accounted for:
string 942,750 39.6%
int/float 379,530 15.9%
map 248,490 10.4%
list 247,590 10.4%
true 189,990 8.0%
false 187,920 7.9%
json_null 184,590 7.8%
so 79.2% of calls take an arm that appends something and returns, and pays for a frame the looping arms need. the runtime already has this shape solved in one place: k_beat_pop splits from k_beat_pop_slow so the pop with nothing to do returns without the six pushes the migrates need. the reading was that the encoder wants the same split, worth about 1.45%.
it was built and it is worse. marking the two looping arms noinline in the emitted ir, linking both copies against the same runtime object with the same flags, output byte-identical:
control 1,823,406,531 text=318,482
noinline 1,835,914,985 text=319,042
+12,508,454 +0.686% +560 bytes
the release build sets -inline-threshold=2000, eight times clang's default, and that number was found by measuring rather than reasoning — it is worth 2.09%. it is right here too: inlining the loop arms saves more than the six pushes cost, because outlining adds a call and a return per list and per map element on top of the argument shuffling.
the more useful half is that the premise was wrong. outlining shrank the frame from 88 bytes to 56 and left all six pushes. they belong to the arms that make calls, which is nearly all of them — the string arm calls the escaper, the number arms call the renderer, and each needs its values preserved across that call. only true, false and null are literal appends, and those are 23.6% of calls rather than 79.2%. the ceiling was about 0.43%, for a change more invasive than the one that just measured 0.686% against it.
the shape is worth keeping even though the change is not. a frame is sized for the worst arm, and a dispatch whose arms differ a lot in what they need is a place to look — but which arms need what is a question for the disassembler, not for the source.
82a width bound cannot see past the first step
rendering a double is 4.58% of the run program — 84,209,220 instructions over 191,070 calls, 440.7 each — and a quarter of that is one loop walking digits off the significand two at a time. it takes 5.35 trips a float on the encode corpus, because the significand starts with seventeen digits and the shortest form needs six or seven.
so compute the count instead of searching for it. the two ends of the rounding interval agree above the first place where their difference has a digit, which makes declen(vp - vm) - 1 a lower bound on how many digits can come off, and one division by a power of ten takes them all at once. written guarded by the same test the loop uses, the step can never take a digit the loop would have left.
it costs 38.3 instructions a float. the function goes to 91,522,440 and the whole run program rises 7,313,220, 0.397%, with every instruction of the rise inside that one function.
a counter says why. over 190,890 calls the step fired every time and the count it computed was 1 on 58,680 of them and 2 on the other 132,210. never more. the bound reads the interval at full width, and the interval renormalises after every removal: dividing both ends by ten narrows the absolute gap but leaves the loop's test true for many more steps than the starting width predicts. the step pays sixteen comparisons and three divisions to take 1.69 digits where one trip of the loop takes two for twenty-one instructions.
declined. the reason belongs to the quantity rather than to the code, which is why no amount of tightening the step would rescue it.
the same probe priced the idea that would have come next. of the 191,070 doubles this corpus renders, 152,460 need seven significant digits and 32,940 need six — and none of them is a whole number. a fast path for small integral doubles, the standard trick in every float renderer, would fire zero times here. measuring the corpus cost one afternoon and saved writing it.
83a fixpoint round that rewrote nothing
the alias canonicaliser runs to a fixpoint: rewrite the bare names that stand for a qualified original, then go round again in case a rewrite exposed another. it went round a second time over a program the first round had left alone, and a round that rewrites nothing still walks everything.
compile_allocs 27,397 -> 27,313 -84 -0.307%
compile_instructions 35,869,355 -> 35,441,049 -428,306 -1.194%
entry_instructions 127,872,255 -> 126,348,616 -1,523,639 -1.192%
library_instructions 128,010,052 -> 126,804,150 -1,205,902 -0.942%
all three kanso check routes fall by about the same proportion, and that is the tell for a pass whose cost is proportional to the program rather than to anything the program contains. the entry route gives back four times the instructions of the module route at the same 1.19%, because it is four times the program. the run-side rows, the machine-code row and both codegen rows are byte-identical; two rows rose and both are layout, the larger being a hundredth of a per cent on 2.18 billion.
the useful part is the shape rather than the number. eight of these have landed now — a pass that answers a question about names by walking every declaration once per name, or by walking the whole program once per round when once per program would do — and every one of them was found by profiling rather than by reading. the compiler does not get slower in a way anybody notices; it gets slower one habit at a time.
84two leads in the emitter, priced and closed
both of these were written down as open, with a number beside them, and both numbers were wrong in the same direction — larger than the thing they described. that is the failure mode worth naming: a lead is worth what it measures, and a lead that reads twice its size sends somebody after twice the prize.
the first was substring search. the profile said 7,929,096 instructions inclusive, 9.32%, all of it reached from Backend::emit. that sums two frames that do different things. next_match is a CharSearcher, so most of its 5,806,878 is .contains(char) and .find(char) — single-character scans, already the cheap idiom, with nothing to give. the substring cost proper is <&str as Pattern>::is_contained_in: 3,260,397 inclusive, 3.82%, over 25,374 calls.
the sites did not converge either. two candidates instrumented on a real build account for 2,520 of those 25,374 calls; the rest is inlined into Backend::emit from somewhere a source grep does not reach, and the four whole-body questions beside it are hash-set lookups rather than searches. a release build with RUSTFLAGS=-g still annotates as ???:, because the hot code in these benchmarks is clang's, from runtime.c and the emitted IR, so a rustc flag was never going to give line information there. declined at 3.82% with the sites unfound.
the second was the unreachable blocks. a build of the run benchmark emits 649 of them across 599 defines, written down as “roughly 3.6% of emitted lines, paid by clang and ld on every build”. 649 lines of 36,085 is 1.798%. the 3.6% is the pair: each unreachable together with the call void @k_die(...) above it, which is 1,298 lines and 3.597%.
and that pair is the reason there is nothing to take. an LLVM basic block must end in a terminator. all four die entry points are declared noreturn in the emitted preamble and carry __attribute__((noreturn, noinline)) in the runtime, so the block after one of those calls has no fall-through and unreachable is the terminator it is required to have. 621 of the 649 are mandatory and the other 28 are ordinary block tails after a label or a ret.
emitting fewer means emitting fewer die sites, and each one is a check a running program can reach: an arity mismatch, an overload with no match, a destructure of the wrong shape, integer overflow. what clang and ld pay for those lines is real, and it is the price of the checks.
85a hop needs a name to print
in july an optimisation was declined on the differential law. rewriting (a b -> f a b) to f made the native engine print a provenance line the interpreter did not — passed through greet — and the two engines disagreeing is the one thing this project does not permit. the entry recorded it and moved on.
the 2026-09-15 explicit-bind ruling moved the ground under that argument in its own words: the provenance hop now accrues at binds rather than at skipped calls. so the argument was re-run rather than re-asserted. three spellings of one call, each on both engines:
label "flan" seasonal passed through label
lab "flan" seasonal (fn lab d c = label d c) passed through lab
(d c -> label d c) "flan" seasonal no hop line at all
both engines agree in every one. the divergence that killed the optimisation is gone, and what the three readings describe is a rule rather than a bug: a hop names the function the err was about to enter, so a named function records one and an anonymous one records nothing. a lambda has no name to print.
the optimisation stays declined anyway, on different ground. eta-reducing the third spelling to the first no longer makes the engines disagree — it makes the trace GAIN a line where the lambda printed none. that is still a change to what a program prints.
the direction is worth noticing, though, because it now points at source rather than at a rewrite. wrapping a call in a lambda silently drops its provenance line, on both engines, and a reader would not expect (d c -> label d c) and label to trace differently. whether that is the right rule for an anonymous function is a question about the language, and it goes to the ledger rather than being settled by an optimisation that wanted it one way.
86a refcount histogram is not a price
the interpreted run became a weighted term on the 2026-09-16 gavel, so it was profiled for the first time with a counter on it. the largest single thing in it is not interpretation: memcpy carries 18.49%, and 7.76% of the whole run is Vec<u8>::clone called 43,572 times from one function — about four thousand instructions a call. the byte builder copies its entire accumulator before extending it, 180,081,360 bytes over the corpus, accumulators up to 9,906 bytes rebuilt one append at a time.
the obvious repair is to extend in place when the buffer is unique. it was built at both sites and both were declined, and the refcounts are why. at the smaller site every one of the 6,606 calls found three holders — the fast branch never fires, and the 298,420 instructions the change appeared to save were the call to clone rather than the copy inside it, with memcpy coming back unmoved. at the larger site:
strong=1 1,326 3.6%
strong=2 11,880 32.1%
strong=3 13,206 35.7%
strong=5 9,228 25.0%
strong=6 1,326 3.6%
3.6% of the appends could extend in place and 96.4% could not: six and a half megabytes of the hundred and eighty. that share was measured and the fix was not, and the two turn out to be different numbers. run, on push, put and append together, it is 14,194,625 instructions — 0.64% of the interpreted row, with memcpy falling 12,529,663. the histogram answers for the calls it sampled at the instant it sampled them; what a fix is worth is a question only a run answers.
the histogram was right about the shape it saw, and that shape is most of them. an accumulator threaded through a name the caller still holds — stack (push xs n) (n - 1) — is pointed at by that frame too, and two fixtures of that shape read identical allocation counts with the fix and without it. an accumulator that arrives as another call's answer — push (push xs n) n — is pointed at by nothing else, and reads 22,362 allocations copying against 21,162 extending in place. lib/list's own put acc k (push (bucket acc[k]) x) is the second shape, which is why the corpus moved.
so the copy is a symptom and the holders are the thing. a buffer partway through a fold is held by the argument vector, by the wrapper's environment, by the caller's, and by the fold's own state, and the interpreter cannot know that all but one are about to go away. the compiled engine can: the linearity analysis proves the accumulator unique and the append extends the builder in place. that proof is the difference between the two engines here, and no amount of reference counting at run time substitutes for it.
87the interpreter worked the same answers out again every time
the interpreted run became a weighted term in september, and the first profile of it found memcpy at the top with interpretation nowhere near it. the second profile, taken after that copy was fixed, found something plainer: the interpreter was deriving facts about the program over and over that the program had settled at parse time.
a name is the clearest case. eval_ident answered “what does this name stand for” by walking a ladder — the environment, then the function table, then three descriptor names, then the type table, then the function and type tables a second time, a sixty-name array of builtins, and an allocation to build the reference it returns. counters on each rung, over the corpus: 1,056,329 calls, 724,304 of them locals that stop at the environment walk, 332,024 reaching the type probe, 325,412 returning a reference. so a third of all resolutions walked the whole ladder to reach an answer that cannot change during a run, because the tables are fixed once the program is parsed. remembering it took 171,322,972 instructions off the row, 7.90%, and 6.12% of every allocation the run makes.
the call sites asked the same question differently: call_named had a ladder of its own, and the tail-call loop spent two table probes deciding one bit about every callee. a second memory took another 33,279,216, 1.67%.
then the frame. a raise records two things — the trace line a report prints, and the package to hold responsible — and the interpreter was building both on every entry into every body. the corpus enters about half a million bodies. remembering that one per declaration, with three vectors opened at sizes the code already held, took 407,935,771 instructions off the row in one change: 20.77%, and 930,784 allocations, 19.35%.
four changes, and the row went 2,186,500,133 to 1,555,890,579. 28.84%, none of it from making the interpreter cleverer.
three things about the work are worth more than the numbers.
the frame is not only a diagnostic, and the first draft of this section said it was. half of it is: the trace line is read when a report prints and not otherwise. the other half decides what the program does. a failure carries the package it was raised in, the frame carries the package of the body currently running, and rescue compares them — its licence is foreign-only, so a rescue declines a failure raised in its own package and passes it through untouched. the frame is an input to control flow, and remembering it per declaration is safe only because the answer depends on the declaration and nothing else.
the way that got established is worth more than the claim. a fixture was named as the one that would catch a wrong frame — it raises from two functions and prints the name each failure carries — and then the memory was deliberately broken, every declaration made to share one entry, to watch the fixture go red. it did not. it is byte-identical under the broken build, because it exercises the readers rather than the memory. what does catch it is a chain-word sample, and it fails by printing nothing at all: with every declaration wearing the first one's package, a rescue declines a failure it should have handled and the failure leaves the program instead of its output. the fixture that looked right was wrong, and only running it said so.
the size of a table row decided part of one win. the remembered answers live in an enum, and its first shape had an arm holding a descriptor and an arm holding a value — 80 bytes and 32 bytes, so every row of a table with a few hundred entries was sized at 88. spelling the three descriptors and four literals as arms of their own took a row to 24 bytes and the run down a further 15,567,153. nothing about the work changed; the table stopped spilling out of cache.
a vector opened empty beside a length the code already holds is worth reading for. four of them, in the dispatcher's parameter match and the two argument lists, paid 47,438,686 between them without a line of logic changing. and the last one is the one to remember: Vec::clone allocates capacity exactly equal to length, so the interpreter's copy-when-shared handed back a full vector and the push immediately after it reallocated and copied the whole buffer a second time. every shared push and every shared append was paying for its contents twice. sizing that clone for the growth that follows takes a further 6.82% off the row, and it takes 7.92% off the peak as well, because the reallocation it removes was a doubling.
and one measurement about measurement, which got sharper the day after it was written. each change here was read twice: once on a container with a different rust and a different libc, and once on the runner whose numbers the goldens keep. for the first four, the instruction deltas came within a fraction of a per cent of each other — 0.18%, 0.12%, 0.63% — and it was tempting to write that down as a rule.
the fifth broke it. the container projected the clone-sizing change at 94,998,647 instructions and the runner read 295,627,669: three times as much, on the same source. and on that same pair of runs the two machines agreed to the byte on what the change does — 36,066 fewer allocations and 76,113 fewer peak bytes, the same integers on both. so the disagreement is not noise and it is not the change being misunderstood. what a copy of a thousand elements costs in instructions is a property of the compiler that built the interpreter; how many copies happen is a property of the program, and only the second of those crosses a machine boundary intact.
the practical rule is the one the goldens already enforce and the first four readings nearly talked us out of: a counter of operations can be compared across machines and a counter of instructions cannot, not even as a difference. the four close agreements were four changes whose copying happened to compile similarly on both hosts. a fifth showed what the rule is actually worth.
88the gate that accused a stable binary
the instruction rows are counted rather than timed, and a counted row can still disagree with the number the golden holds. when that happens the gate takes a second callgrind pass, because two things look identical from outside and are settled in opposite ways. if the same binary counts two different numbers, that is a reproduction failure: the vein halts, the cause is hunted, and no value is pinned. if it counts one number the golden does not hold, that is an ordinary ratchet — regenerate, and say in the log which way it went.
on 2026-09-18 two pull requests sat blocked on the same red check and the same sentence: this binary counted two numbers in one job: 1138001437 and then 1138004452. one of them carries two divisions in a float renderer, the other caches a pointer in the beat rewind, neither goes anywhere near the interpreted corpus, and they ran on different runners. both drew the same pair.
the difference is 3,015. two screens above the error, both jobs had printed interp_printed=3015 as a notice.
what the row excludes is the cost of the line the corpus prints. _print’s subtree ends in a memrchr over the formatted bytes, and what that frame costs moves with the binary’s layout rather than with anything the program does, so it is taken off — the rule is that a term which cannot be normalised is excluded and the exclusion is named. the first reading had it taken off. the second was the raw frame. two readings of the same quantity, less one subtraction, and the gap between them was the subtraction.
the three sibling gates do it correctly, and the reason is small enough to be worth naming. the change that introduced the exclusion wrote the reader as a function and called it on both profiles; the change that brought the same exclusion to the interpreted row a day later open-coded the pipeline instead, and reached one call site of two. nothing about the arithmetic was subtle. what was missing was that the second reading is a call rather than a copy.
the defect lives entirely on the failure path, which is the only path the second count exists on. so it was invisible while every row agreed, and it fired on exactly the occasions when a reader most needed the second count to be honest — and said, in words, the opposite of what had happened.
the cost was nearly larger than two wasted rounds. the branch-side diagnosis had already been written down: a +7 that had supposedly landed three times on three different absolute values, attributed to layout jitter, committed to the log before the job log was read. a spread of 3,015 on a row whose allocation counter never moved is not layout, and the gate’s own notice line said so. the entry was withdrawn, which is cheap; a page section quoting it would not have been.
the spec that now holds the property reads for the call rather than for the text of the pipeline, because one call site out of two is the shape the defect took and a spec looking for a missing subtraction would have found nothing to object to. it reads by the same marker as the spec that checks the exclusion itself, so the two agree on what subtracting means by construction — a looser first draft, matching the frame’s name rather than the read of it, failed the one gate whose program prints nothing and whose script explains why in prose.
89the compiler already knew, and the interpreter never asked
the interpreter grows a container by taking it by value and calling Rc::try_unwrap, which hands back the contents when nothing else holds them and copies when something does. counters over the corpus: 44,006 calls into push, put and append, and 42,680 of them — 96.99% — copy.
the second holder is not a surprise, and it is not removable. a bound name lives in an environment node, the environment is a persistent singly-linked chain, and unlinking the node to drop its reference costs more than copying the container does. so the refcount is honestly two, the copy is honestly required, and no amount of care at the call site changes either.
what changes it is a fact the compiler settled long before the interpreter ran. the linearity analysis exists to tell the code generator which reads of a container are its last — the ones after which the binding is dead — and it marks 54 such sites across the library. the compiled engine has used that answer for months. the interpreted engine had never been shown it.
it asks now, keyed by file, line and column so the answer survives a re-parse, and where the analysis says the binding is dead it takes the value out of the Rc in place. the runner's sitting:
interp_instructions 1,138,001,430 -> 1,075,174,600 -62,826,830 -5.52%
interp_allocs 2,539,998 -> 2,486,376 -53,622 -2.11%
interp_peak_bytes 884,985 -> 833,130 -51,855 -5.86%
this container had projected 53,327,307 and said in advance to expect more, for a reason the previous page section spells out: the change removes memcpys, so it moves bytes, and what a copy costs in instructions is a property of the compiler that built the interpreter rather than of the program. the runner read 17.8% above the projection, in the predicted direction. the allocation count, which is a decision the program makes, travelled exactly: −53,622 on both machines, to the unit.
four compile-side rows also rise — 715,346 instructions between compile, entry, library and emit. none of them can be paying for this change: kanso check stops before the interpreter runs, and the analysis sits behind a cell only the interpreter opens. they move because src/eval.rs is the compiler — it grew, and the layout moved under it. three of the four move by the same 0.19% on three different routes through the same front end, which is what that looks like.
the gate is the whole of it, and an ungated version proves that rather than asserting it. the take is a mem::take through a raw pointer derived from a live Rc — sound only because the analysis proved the other holder dead, and catastrophic otherwise. so the take was built first without the gate, and the suite was run: five specs went red, three of them written for exactly this defect. with the gate consulted, all five are green and both engines print the same answer. the difference between those two runs is what the analysis is worth, stated as a measurement instead of an argument.
one more thing came out of it, and it is about the corpus rather than the change. of the 16 distinct sites the interpreter actually reaches at runtime, every one is inside the 54 the analysis marks. so over this program the marking is complete as well as sound. that holds for this corpus; a program reaching a seventeenth site would say whether it holds more widely.
and a pin was missing. five specs catch the optimisation firing where it should not, and every one of them stays green if the gate silently stops matching — if a key changes, if a span moves, if the analysis is handed a different program. a per-round allocation count now fails when the optimisation quietly stops happening, which nothing in the tree could previously do.
90a byte beside the name, and why it cost eight
the interpreted run spends 4.14% of itself inside libc’s memcmp, and the largest caller is the walk up the environment chain that resolves a local name — 724,304 calls, the same figure section 87 publishes for locals that stop at that walk.
the walk compares one string against another. rust’s string equality checks the length first and then calls memcmp, so length is already free, and what reaches libc is every binding in the chain that happens to be the same length as the one being looked for. a single byte of the name, compared before the string, rejects most of those without a call. that is a standard trick and it works: the memcmp fell by 2,465,720.
the run rose by 10,409,061.
the node is why, and it was measured rather than reasoned about. size_of on the environment node reads 64 bytes without the byte and 72 with it. there is no padding to hide it in: the name is 24 bytes, the value is 32, the parent pointer is 8, and that is exactly 64. the compiler already orders the fields to pack them. so one byte of payload costs eight bytes of node, and the interpreted run allocates about two and a half million of them — twenty megabytes of extra traffic to save two and a half million instructions of comparison.
where to look next follows from that. a discriminator that fits in space the node already owns would keep the saving and drop the cost, and there is nowhere to put one: the name’s own 24 bytes are full, holding twenty-two bytes of text plus a length and a tag. what remains is changing how a name or a value is represented, which is a larger question than this lead was, and 2,465,720 is the ceiling on what answering it could be worth here.
declined on arithmetic. the idea is sound and the node is the wrong size for it.
so the other version was built too, and it loses by more. the name’s inline buffer is zero-filled before the text is copied, so two inline names hold identical twenty-two-byte buffers exactly when they hold the same text — the padding makes the length implicit, and an identifier cannot contain a NUL to blur it. the walk can therefore build a padded key once per lookup and compare a fixed-width block per frame, which the compiler inlines. nothing grows; the key lives on the stack for one lookup.
base row 1,115,996,775 memcmp 48,165,663
key row 1,150,537,278 memcmp 35,577,082
it takes 12,588,581 instructions off memcmp, twenty-six per cent of the whole figure, and the run rises 34,540,503. three times worse than the byte, with no node cost at all. the interpreted corpus prints byte-identical output before and after, which is what makes the two counts comparable — that is two builds on one engine, not the differential check across both, which a declined change did not earn.
two sound schemes losing in different ways is worth more than either result. the byte lost to the node’s size; this one has no node cost and loses by more, so what is left to blame is the shape of the trade: the setup is paid per lookup and the saving arrives per frame. that only pays when the chain is long, and here it is not.
which reframes the lead. the 48 million is real, and it belongs to the walk rather than to how the walk compares. anything that keeps the walk and makes the compare cheaper is buying a per-frame saving with a per-lookup cost.
and that ratio is measurable, so it was measured rather than left as a reason. a counter in the walk, over the same corpus:
lookups 1,056,329
frames visited 2,662,536 2.52 per lookup
misses 332,025 31.4%, and a miss walks the whole chain
the instrument lands on two numbers this page already carries. 1,056,329 is the resolution count in section 87; take the misses away and 724,304 remains, which is section 87’s count of locals that stop at the walk and also the caller count read off the profile. three routes to the same pair.
2.52 is the whole explanation. a scheme paying setup per lookup and saving per frame has two and a half frames to amortise it over. the padded key cost about forty-seven instructions a lookup against a saving of the eighteen-odd a small memcmp costs, times 2.52 — so it lost, and would have lost at any setup above about forty-five. the byte had no setup and lost to the node instead.
it also prices the change that would work. eighteen instructions a visit across 2,662,536 visits is the 48 million; resolving a local to an index at parse time removes the visit, and an index costs two or three instead of eighteen. the ceiling is around forty million, near 3.6% of the interpreted row — a number rather than a hope, and larger than anything tried here, because slots have to survive closures, which capture an environment rather than a frame. and section 91 retires it. the forty million divides a counter by a count; measured by removal, the walk’s comparisons are worth about two and a quarter million, and its real expense is the value it copies out of the frame at the end.
then the distribution was measured too, and it names a cheaper change first. an average hides a shape, so here is the shape — hits by depth:
1 267,899 37.0% hit frames 1,590,733 2.20 deep
2 199,016 27.5% miss frames 1,071,803 3.23 deep
3 130,524 18.0% total 2,662,536
4 101,095 14.0%
5 25,770 3.6%
6+ 0
the chain never exceeds five, so nothing here waits on a pathological case. and the split by outcome looked like the point: misses are 31.4% of lookups and 40.3% of every frame visited. a miss walks the whole chain and finds nothing, because the name was never a local — it is a function, a descriptor, a type or a builtin, and the resolution ladder carries on afterwards. multiplying those visits by the eighteen above put the waste at about 19.3 million.
that multiplication was wrong, and section 91 measures by how much. the visits were counted; the eighteen is an average over every visit, and the two halves differ by thirty-three times. a miss visit costs 0.89 instructions of memcmp and a hit visit costs 29.68. so the 19.3 million is really 956,485, and the forty million below is sized from the same average and moves the other way. read both figures in section 91.
91an average applied to the half it does not describe
section 90 ends by sizing a lead. misses never find anything, they are 40.3% of every frame the interpreter visits, and at the eighteen instructions a visit the section derives that is 19.3 million instructions of walking that cannot succeed. it reads like the cheapest thing on the page. so it was built.
the filter is small and exactly safe. one function makes every environment frame, so it records the shape of each name it binds — the first byte and the length — into a 256-entry table, and the lookup answers not here without walking when the shape is absent. the table only ever gains bits, and a set bit admits the walk, so a collision costs a walk that would have happened anyway and can never change an answer. over this corpus, 35 distinct names are ever bound as locals and the 332,025 lookups that reach no frame are spread over 73 others; one of the 73 collides, so the filter skips 326,733 of the 332,025.
the run rose by 24,320,374, 2.18%. four binaries, built in one worktree and staged into the same box in turn, say where every instruction of that went:
base 1,115,996,775
the reference threaded, never read 1,132,586,649 +16,589,874
the shape recorded, never tested 1,140,109,198 +7,522,549
recorded and tested, walks skipped 1,140,317,149 +207,951
the last row is the whole of what the idea is worth, and it is a rise. memcmp makes it plainer still: 48,165,663 in the first three binaries and 47,209,178 in the fourth. removing every one of the 1,071,803 miss visits takes 956,485 off the figure — 1.99% of it, from 40.3% of the visits.
so a miss visit costs 0.89 instructions of memcmp and a hit visit costs 29.68. thirty-three times apart, and the eighteen is the average of the two. the mechanism is visible in the names themselves: string equality checks the length before it calls libc, every local this corpus binds is seven bytes or shorter, and the names that miss are builtin_append, json/parsed, text/utf8. a miss walk is rejected on length at every frame and never reaches libc at all. the misses were cheap for the same reason they were misses.
and threading the reference cost more than the idea was worth. 16.6 million of the rise is a fourth argument to the binding function that nothing reads — 534,994 bindings, all from one dispatcher, 31 instructions each. recording into it adds 7.5 million more. moving the table to a thread-local would remove the first half, and the second half alone still exceeds the 956,485 ceiling, so no arrangement of this wins. declined on arithmetic, with both halves isolated rather than guessed at.
the lesson is about the sizing rather than the change. a count multiplied by an average is a measurement only where the average describes the thing counted, and here it described a population thirty-three times more expensive than the one it was multiplied by. the check costs one build: remove the visits and read the counter that was supposed to fall.
and the lead it was competing with dies of the same error, one level up. the first draft of this section said the 47.2 million left over must be in the hit visits, which is the whole counter attributed to one caller again. so that was built too: a per-site inline cache keyed by the address of the name in each node, holding the depth the site’s local was found at last time. it jumps to the remembered depth, checks the name it lands on, and falls back on a mismatch, so a stale entry costs one comparison and can never answer wrongly.
it is correct, and it costs 31,209,169. which counters moved is the answer:
eval_ident 115,147,991 -> 152,742,778 +37,594,787
dispatch 130,726,726 -> 130,726,726 byte-identical
match_one 68,846,453 -> 68,846,453 byte-identical
memcmp 48,165,663 -> 46,830,275 -1,335,388
so the walk’s comparisons are worth about two and a quarter million — 956,485 for the misses and 1,335,388 for the redundant compares on a hit, against a counter of 48,165,663. that figure was doubted for half an hour on a misread caller tree and then confirmed by a third route, which is the part worth keeping.
the annotated source settles it and needs no interpreting. a profile carrying line numbers gives a cost per line and per call site, and inside the walk it reads:
909,375 if frame.name.as_str() == name {
2,897,216 return Some(frame.value.clone());
31,514,002 => <Value as Clone>::clone (724,304x)
every comparison the walk makes, over all 2,662,536 frame visits, costs 909,375. the clone at the end of it costs 31,514,002 — forty-three instructions on each of 724,304 hits. thirty-five to one.
which explains every scheme on this page at once. the byte, the padded key, the shape filter, the per-site cache: all four aimed at the comparing, and the comparing was never the expense. a slot index would remove the visiting and the comparing together and leave the clone exactly where it is.
what would touch it is the walk answering a reference into the frame instead of a copy of the value, so a caller that only reads pays nothing. that is a change to what eval_ident returns rather than a trick inside it, and no number is offered for it here: 31,514,002 is what the clone costs, not what removing it would save, and the distance between those two is what this page keeps learning.
five figures for this counter were withdrawn getting here — 19.3 million, forty million, 47.2 million, 0.205%, and a withdrawal of the last one that was itself wrong. four came from dividing a measured total by a measured count. the fifth came from reading a caller list as belonging to the entry above it rather than below. the readings that survived were a differential and an annotated line, and those are the two instruments worth using.
so the day ends on a number nobody had looked for. the walk’s expense is Value::clone, and what that clone does depends entirely on which variant it copies. of the interpreter’s value kinds, maps, errors, lists, bytes, records, subtypes, function references, closures, partials, descriptors and thunks all sit behind a reference count, so copying one is a bump; floats and the five nullary values copy nothing. two variants allocate: a string, and an integer, which keeps its magnitude on the heap.
counted by hand over the corpus, with a Clone written in place of the derived one:
value clones 1,555,866
integers 610,763 39.3% allocates
strings 137,732 8.9% allocates, 4.1 bytes each
everything else 807,371 51.9% a bump, or nothing
half of every copy this interpreter makes goes to the allocator, and the integers are four and a half times the strings. every one of the 610,763 fits in a machine word — not one of them needs the wide representation it pays for. the strings average four bytes, so what their copy costs is the allocator rather than the copying.
the walk’s own hits have the same shape: 41.0% integers, 9.5% strings, 45.8% a bump, 3.7% nothing. which settles an ordering. an inline machine integer with the wide form on overflow reaches 41% of the walk’s copies and the same 39% of every other copy in the run, because it changes the value rather than a path. answering a reference into the frame instead of a copy reaches the 46% that are already only a bump, where the saving per copy is smallest, and costs a signature change through every caller that needs an owned value. integers first.
and writing the fixture before the code found the thing that would have bitten it. the native build is int64 and refuses past the boundary with a clear diagnostic; the interpreter answers exactly. that divergence is allowed, and the book pins native’s half of it. the interpreter’s own answers there are pinned nowhere and cannot be: the behavioural corpus runs every fixture on both engines and requires them to agree, so it cannot hold a program one of them refuses. the exact place a machine-word fast path goes wrong — promoting one step late, printing a wrapped number rather than raising — is a place nothing in the tree watches. that is a gap in the corpus rather than a fact about integers, and where such a golden should live is a question the project has to settle before the change is worth making.
so the cheaper change is not a faster walk but no walk. whether a name can be a local at a given site is decidable where that site is compiled, and skipping the walk for the ones that cannot needs no slots, no upvalue analysis and no change to how a closure captures — the three things that make the larger change large. it is worth about half as much for a small fraction of the work.
that does not retire the index: the 1,590,733 hit-frames stay, and only an index removes those. it orders them. skip first, measure what is left, and let the remainder make its own case.
92a fixture that emptied the library it was reading
one of the interpreter's allocation fixtures went red on two pull requests in the same hour, on two different machines, on branches whose diffs touched neither the interpreter nor the fixture. both rounds were written off as flakiness and left.
the panic said what it was, and reading it was the whole diagnosis:
the interpreted run failed: error[name]: unknown name `builders/run` --> /tmp/kanso-unique-container/run_300.kso:3:1
the library is there and has nothing in it. writing a file is a truncate followed by a write. the fixture's two tests staged into one hardcoded directory, and the test runner runs the tests in a binary on parallel threads — so one test's staging emptied the library while the other test's compiler was reading it. a zero-byte library is a valid library with no definitions in it, which is why the run got as far as resolving a name before it died.
the failing test is also not the one the file is named for. both rounds were recorded against the file's name, because that is the target the runner reports, and the test that actually died was the other one in the same file — which passed in the very same run.
each test stages its own tree now, named for a tag it passes, and every tag is the same length so a reading taken under one is comparable with a reading taken under another. the directory's name grew seven characters, the absolute allocation counts moved with it, and the figure the fixture pins — the cost of three hundred extra rounds, taken as a difference — did not move at all. that is the subtraction doing the job it was written for.
two things about the spec are worth keeping. twenty-five runs of the pair on an idle machine did not reproduce the race, so the spec forces the interleaving rather than waiting for it: it truncates one tree's library, leaves it truncated, and runs the other tree. and the first version of that spec passed against the shared directory, because it called the helper that re-wrote the library on the way in and healed the window it was testing. moving the staging out of the run helper is what keeps the window open across the read.
this page has a standing rule that external state is normalized before it is measured. this is the same rule read from the other side: a directory two threads write is not normalized, and it is nobody's.
93two copies the dispatcher was making for nothing
section 91 left the environment walk sized and the schemes aimed at it declined. the next profile of the interpreted run put a different frame on top: the dispatcher, at 130,726,726 instructions of its own, 11.70% of everything the run does. two of its callee counts had answers, and both were things being built and immediately thrown away.
the first showed itself by arithmetic. <str as Display>::fmt was called 119,535 times from inside the dispatcher, for 23,936,233 instructions. the dispatcher's tail-hop line ran 119,542 times. two numbers derived from different places, agreeing to within seven.
the dispatcher is a trampoline: a tail call to a named group re-enters arm selection rather than growing the rust stack. it held the name of the group as an owned String, and each hop wrote name = next.to_string(). but next is already an Rc<str> — the tail carries it that way — and to_string reaches Display, which reaches the formatting machinery. two hundred instructions and an allocation, a hop, to copy a name one pointer away. holding the name as an Rc<str> makes the hop a move. every use of the name inside the loop already read it as a &str.
the second was in arm selection. match_params built two vectors sized to the parameter list on every candidate it tried, and gave up the moment a pattern refused — so a candidate that failed on its first parameter had already paid for both. selection tries every overload in the group and keeps one, so most candidates fail. one pair of buffers now serves the whole list, and the candidate that wins hands its vectors to the next one rather than leaving it to grow from nothing.
measured together, against a tree that already had the value-position if:
interpreted run 1,070,343,718 -> 1,038,405,822 -31,937,896 -2.98% allocator work 34,485,525 -> 25,548,225 -25.9% allocation calls, dispatcher 1,471,992 -> 502,377
the corpus prints the same answer on both binaries.
one figure on the way there was wrong and is worth saying so. the allocations were first sized at 65,283,877, which is what __rust_alloc costs inclusive when called from the dispatcher — everything underneath it. the function's own cost across the whole program is 36,857,970. reading the first as a budget for the second is the arithmetic this page has withdrawn figures for before. what the change was worth is the differential, and the differential is smaller than either.
94a scope that was as deep as the argument list
section 93 took two copies out of the dispatcher and left the environment walk where section 91 had found it: 2,662,536 frame visits to serve 724,304 hits, 3.7 nodes a lookup. that ratio counts nodes, so what moves it is how many exist.
binding a name built one heap node holding that name, its value, and a pointer to the rest. so a call with four parameters pushed four nodes, and every name the body mentioned walked past all four before it reached the caller's scope. the depth of a scope was the length of an argument list.
a scope is two shapes now. one node for a name bound on its own — a deferred value, a pattern variable, a step inside a build. one node for a whole call's parameters together. the second scans its slots backwards, which is what keeps shadowing exactly as it was: binding three names as three nodes leaves the last one outermost, so a single node holding all three has to answer the last one first.
isolation, two binaries on one container interpreted run 1,038,405,822 -> 1,022,961,605 -15,444,217 -1.49% allocator work 25,548,225 -> 20,621,100 -19.3% the gates, measured on ci interpreted run 995,837,536 -> 975,944,763 -19,892,773 -2.00% allocations 1,740,991 -> 1,412,516 -328,475 -18.87% peak bytes 833,333 -> 834,117 +784 +0.09%
the two hosts agree on the direction and disagree on the size, and the disagreement is the usual one: how many allocations a change removes is a count the program decides, and what the removed work cost in instructions is the rustc that built the binary. the peak moving the other way is recorded and not explained. a peak is a high-water mark where the count is a total, so the shape of what is live at the run's widest moment can grow while the number of allocations falls; which shape that is would take a build per hypothesis, and none has been run.
most of why it is cheap is that it composes with the change underneath it. arm selection already fills one vector of bindings and hands the winner's on; the frame takes that same vector. the allocation selection had already paid for becomes the scope, instead of being freed and replaced by one node per parameter. a call that binds nothing pushes no node at all.
the reason it is a small change at all is that the scope type was sealed: one function built a node, one function read a chain, and nothing else in the interpreter named the type. three functions and one call site.
section 91 declined four schemes aimed at this walk — a head byte beside the name, a padded key, a shape filter, a per-site cache — and every one of them cost more than it saved. all four aimed at the comparing and left the structure alone. this one aims at how many nodes exist, which is the allocation count and the walk depth at the same time. that is a reason to try it, not a reason to expect it to win; the four that failed looked sound too. what settles it is the two binaries above.
95a frame built for the path that is not taken
section 94 made scopes shallower and the profile re-ranked. the name lookup came second at 115,147,991 instructions of its own, 10.61% of the interpreted run. the two obvious guesses were both wrong and both cost nothing to check, because the answers were in the same file: the name table has used a seedless fx hash since §34, so hashing is not where the time is, and the thing cloned out of that table is an enum of reference counts, so cloning it is an increment.
the annotated build says where it goes, and it is not in the body:
16,901,264 (1.58%) fn eval_ident(&self, name: &str, span: Span, ...
7,122,532 (0.67%) if let Some(value) = lookup(env, name) {
1,448,608 (0.14%) return Ok(value);
996,075 (0.09%) match named {
8,450,632 (0.79%) }
25,351,896 instructions, 2.37% of the whole run, on the opening line and the closing brace. a release build annotates every line as ??? unless it is built with debug info, which is why reading this took a second build.
a name that a scope binds is answered by the walk and returns. a name that no scope binds takes a borrow, a map probe, a match, and on its first visit the whole resolve ladder. both lived in one function, so the frame the second path needs was built and torn down for the first one as well. moving the second into a function of its own, marked never to be inlined back, leaves the first as a match over the walk's answer.
isolation, two binaries on one container base 1,022,880,858 split 1,007,027,010 -15,853,848 -1.5499% the gate, measured on ci interpreted run 975,944,763 -> 957,583,234 -18,361,529 -1.8814%
every other vein in that job is byte-identical — the three compile rows, what emitting costs, start-up, and both interpreter memory counters. this is the second change in a day confined to the runtime path that left every layout row alone, and the prior that editing the compiler's own rust usually moves the compile row has now missed twice running.
the base was read at two different paths of the same length and printed 1,022,880,858 both times. this row moves with the length of the path it is given, and the first pass ran the two arms in directories that differed by one character. the effect is four orders of magnitude larger than that could be; re-running was cheaper than arguing about it.
the same shape failed earlier the same day. an inline hint on the walk itself left the row identical to the instruction across two binaries that genuinely differed. the entry cost a profile puts on a function is not automatically call overhead waiting to be removed; there it was the function's own work and no hint could reach it. what is different here is a frame sized for a path that is usually not taken, and splitting removes that. neither outcome followed from the attribution, which is the argument for building both.
the same argument one function up was built and declined. eval is the caller, 92,306,061 instructions of its own with 16,623,552 on its signature line, and a match of fifteen arms where several build maps, lists and records while Ident and Int return in a few instructions. four of those arms — the widening cast, the map literal, the field read and the index read — were moved out of line the same way. same corpus, same harness, and the same base binary at 1,022,880,858, so the two readings compare directly.
base 1,022,880,858 split 1,026,292,184 +3,411,326 +0.3335%
declined. what separates the two is which arms are rare in the corpus rather than which look expensive in the source. a name that no scope binds is genuinely uncommon when a program is decoding a document, so the lookup's second path really is the one not taken. reading a field and reading an index are how a decoded document is used, so moving them out adds a call to a hot path and saves a frame the application arm goes on requiring anyway.
three builds on one hypothesis now: an inline hint on the walk changed the row by zero, the lookup split bought 15,853,848, and this cost 3,411,326. that a profile puts instructions on a function's entry line is a reason to build the experiment, and one time in three it was also a reason to expect it to work.
96five builds at one line, and the one that worked
§95 shipped one change out of a family of five. the other four are here with their numbers, because an idea nobody writes down as declined gets built again.
they all aim at one push. the annotated profile puts match_one's variable arm at 723,255 value copies for 29,241,307 instructions — 46.5% of every copy the run makes, off a single line. the call counts come out of the profile file rather than being inferred from instruction ratios: 1,555,866 copies, 1,100,726 pattern matches, 1,056,329 scope lookups, 332,024 of those missing, 55,711 dispatches, and 34,332 matches that recurse into a record. the lookups that hit come to 724,305, which lands on the 724,304 §91 counted from the other end.
score without binding. selection walks every candidate to compute a score, and a score needs the patterns matched rather than the values bound. so the walk bound nothing, and the winner was walked a second time to bind.
base 1,007,027,010 split 1,059,328,533 +52,301,523 +5.19%
declined, by nearly twice what the copying costs in total. the second walk is not only the copies: matching carries 44,851,891 of its own beside the 29,241,307, so walking the winning arm twice swamps what every losing arm's copying saved.
move the winner's arguments out. no second walk. matching recorded a position for a parameter at the top of the list and a value only for a name bound inside a record; the winner's positions became bindings by taking the values out of the argument vector, which the dispatcher owns and clears a line later.
base 1,007,027,010 moved 1,017,171,110 +10,144,100 +1.01%
declined. what a position costs is a tag beside the value it replaces, so the record pushed on the hot path got bigger rather than smaller, and the winner's positions are walked into a second vector that did not exist before. which of those two dominates was not measured.
so the line resists both routes. copying a parameter is thirty-two bytes and, for every kind of value but a big integer and a string, a reference count. deferring that costs more than doing it; indexing it costs more than doing it.
inline hint on the scope walk 0 lookup frame split -15,853,848 shipped eval arm split +3,411,326 declined score without binding +52,301,523 declined winner's arguments moved out +10,144,100 declined
five builds, one win, and the win was the smallest change of the five. each is two binaries on one container, one corpus, directory paths of the same length, identical output.
the draft of this section said the split flattened the profile, and it did not. the figure reached for was 11.70%, which was the top frame of a profile taken three changes earlier. measured on both sides of the split instead: the top frame goes 6.44% to 7.70%, the top five 23.34% to 24.86%, the top twenty 49.76% to 48.95%. twenty functions to reach half the run either way. the profile was already flat, and the split raised the top frame's share by taking work out of the function below it. the flattening belongs to §93 and §94.
the comparing has not moved all day either: 47,672,229, then 47,631,349, then 47,602,341, across three changes to how scopes are built and walked. §91 declined four schemes aimed at it. with the four here, the two largest named things in this profile have nine declined builds between them, and picking the top frame off a list has a record of one in five.
97the same two lines, one function apart
§96 ended by saying the next attempt should come off a different axis than the profile's self-cost list, which had gone one for five. so this one asked which callers reach the allocator, rather than which functions carry instructions.
456,133 __rust_alloc <- RawVecInner::finish_grow 471,517 finish_grow <- RawVec<T,A>::grow_one 237,980 grow_one <- kanso::eval::match_one 225,431 grow_one <- kanso::eval::Interp::dispatch_loop
463,411 of the run's 471,517 reallocations come from two call sites, and both are the same pair of buffers. dispatch declares a score buffer and a bindings buffer at the top of each call, and the candidate loop already recycles them — a winner hands its vectors back as the working pair instead of leaving the next candidate to allocate from nothing. what they do not survive is the dispatch: the winner's bindings become the environment frame and its score is kept beside the winner, so the next call starts at capacity zero and climbs 1, 2, 4 from nothing. 225,431 growths against 55,711 dispatches is four reallocations a call.
so ask for the arity up front. two lines, placed after the two clears at the top of the per-candidate matcher:
base 1,007,027,010 reserve 1,019,044,912 +12,017,902 +1.19%
that is where the day's fifth decline would have been written down. the matcher runs once per candidate, and arm selection tries every arity match, so the reserve was paid several times per dispatch to fix an allocation that happens once. the clears it sat behind are housekeeping on buffers the loop already owns; they are not where the buffers come from. (this said “around twenty times per dispatch” until 2026-09-19; the figure behind it was matcher CALLS a dispatch, which bounds candidates from above rather than counting them. §106 has the measurement.)
the same request, moved one function up to where the pair is created:
base 1,007,027,010 hoisted 985,444,659 -21,582,351 -2.14%
33,600,253 instructions between two placements of one idea, and the winning placement is the largest gain of the day — larger than the frame split in §95, which took 15,853,848.
and the saving is not fewer allocations. tallying the allocator edges on both arms of the same pair:
base hoisted delta __rust_alloc 1,361,559 1,362,891 +1,332 __rust_dealloc 1,349,431 1,350,763 +1,332 __rust_realloc 35,956 32,638 -3,318 grow_one 464,507 112,013 -352,494 finish_grow 489,542 137,048 -352,494
the allocation count goes up. what the reserve removes is the growth path: 352,494 fewer grow-and-copy pairs, and with them the capacity arithmetic, the doubling branch and the element copy each regrow makes. 21,582,351 over 352,494 is 61.2 instructions a growth, which is a plausible price for that work and not a plausible price for an allocation. the first draft of this section said the reserve allocates less, which would send the next reader after allocation counts that are flat to a tenth of a per cent.
this retracts none of the five declines in §96. each was a different change and each was measured. what it retracts is the inference that was forming around them: that the dispatch path had been read out, and that a correctly-measured quantity failing to become a saving was the shape of this code rather than the shape of five particular attempts. one of the five sat a function away from the one that works.
the base row reads 1,007,027,010 for the fourth time, from two binaries with different content hashes built from trees carrying the same interpreter at different paths. on a row known to move with path length, that is the clearest statement so far that the differential reads the code.
the tally has three entries nobody has been after yet: dropping from the dispatcher 229,625 times, allocating for a string copy 149,093, and the 8,106 reallocations that are neither of the two call sites above.
98what the instruments are made of
the sections above measure the compiler. this one is about the things doing the measuring, because two of them turned out to have inputs nobody had looked at and one of them was carrying a claim that was wrong.
the counted row takes two values, and which one you get depends on where you stand. the a/b discipline here says both arms must sit in directories of the same path length, and the reason written down was that the row moves with path length. one binary, one corpus, ten directories differing only in name and length:
/tmp/p len 6 1,007,027,010 /tmp/pathaaaa len 13 1,007,027,010 /tmp/pathbbbb len 13 1,007,027,010 /tmp/pathaaaaaaaa len 17 1,007,027,010 /tmp/pathaaaaaaaaa len 18 1,007,027,010 /tmp/pathaaaaaaaaaa len 19 1,007,027,010 /tmp/pathaaaaaaaaaaa len 20 1,007,027,010 /tmp/pathaaaaaaaaaaaa len 21 1,007,004,925 /tmp/pathaaaaaaaaaaaaaaaa len 25 1,007,004,925 /tmp/pathaaaaaaaaaaaaaaaaaaaaaaaa len 33 1,007,004,925
two values, 22,085 apart, one step between 20 and 21, flat across six lengths below it and three above. the two equal-length directories with different names agree, which rules the name out. and the longer path reads fewer instructions, so whatever this is, it is not a cost that grows with the string.
the environment is a second and smaller term: one extra variable costs 87 instructions whatever its size, so what the term counts is the number of entries rather than the bytes in them.
neither is loose in the gates, which already empty the environment and run from a fixed directory. what the measurement adds is the margin. there are two of those fixed directories, one 21 characters and one 18, and the step is at 21 — so they sit on opposite sides of it, and renaming the first one character shorter would re-base five rows by 22,085 with no compiler change behind it. a spec pins each name's length now. its first draft asserted there was one directory and went red naming the second, which is how the pair came to be measured rather than assumed.
and a golden's own header was wrong about the thing it was written to explain. the release codegen row is counted across a child tree, and last month's note describes that tree as “three clang processes and ld, every one of which came back byte for byte across two readings”. that was true of its two readings. three jobs on one day showed it false of the linker:
5,160,407,609 then 5,160,407,598 -11 probes 1,816,463 -> 1,816,452 5,139,582,528 then 5,139,582,517 -11 probes 1,822,415 -> 1,822,404 6,824,133,291 then 6,824,133,280 -11 probes 1,819,373 -> 1,819,362
every clang process byte-identical each time, the whole difference inside the linker, and all of it in one frame: the bucket probe of llvm's string map during the link-time-optimised link. three different probe counts and the same eleven. one of those three was the branch that carries this correction, which cannot touch codegen by construction.
and the cause was already known when that correction was written, which is the part worth publishing. the branch that fixes this had measured it: twenty-two object names, one binary, one corpus, every other input held and the output path absent each time. twenty read 5,163,341,031 and two of them — 4b8c1a and fedcba — read 5,163,341,042. eleven apart, deterministic per name, about one name in eleven. clang writes its link-time object to a path with fresh hex every run, the linker's plugin puts that path in a string map, and how far the bucket probe walks depends on the string.
so the frame is the probe, the residue is always eleven because there are two buckets, and a job draws when its two readings fall on opposite sides — twice one in eleven times ten in eleven, about one job in six. all three of those were treated as open puzzles for an evening and every one of them is a consequence. the fix pins the name, and what is left to settle is only which of the two tiers the pin covers.
the log's own rule covers how this went wrong twice. a report that something is absent is worth what the search for it being present was worth: the first round ruled names out on three samples, then five, of a one-in-eleven effect, and the second read a number that agreed with the experiment and called it a rival explanation.
the correction is appended to the header as a dated entry rather than written over the old one. what that sentence recorded was true of the readings it had; what is wrong is the general claim a later reader takes from it, and the difference between those two is the whole reason the log works the way it does.
a fourth job reproduced, and so did a second run of the first branch on the same commit. one branch on both sides of that list is what settles the draw as a property of the job rather than of any diff.
99the buffer nobody kept
§97 gave the two dispatch buffers their arity up front. the question left over is why either is allocated per dispatch at all.
the bindings buffer has an answer: it is moved out and becomes the environment frame, so its allocation is still doing work after the dispatch ends. the score buffer has none. it exists to compare candidates, it is held beside the winner while one is winning, and then it is dropped. so it was declared above the tail-call loop instead, and the winner's buffer is handed back to the working variable rather than falling out of scope.
base 985,444,659 kept 980,371,488 -5,073,171 -0.51%
the allocator edges say it is the change and nothing else:
base kept delta __rust_alloc 1,362,891 1,243,349 -119,542 __rust_dealloc 1,350,763 1,231,221 -119,542 grow_one 112,013 112,013 0 finish_grow 137,048 155,543 +18,495
119,542 allocate-and-free pairs at 42.4 instructions each. the growth path is untouched, which is the point: §97 took the growths and this takes the allocations, and the two costs came apart cleanly. the 18,495 added to the growth count is the retained buffer growing when a later dispatch brings more parameters than the one that sized it, and that price is inside the measured total.
119,542 is not the dispatch count, and it is not the iteration count either. there are 55,711 dispatches, so this is 2.15 pairs each, and the reason it exceeds one is the loop the buffer now lives above: a tail call goes round again without leaving the dispatcher, and every one of those was allocating and freeing a score buffer too. the frame builder runs once an iteration and is called 175,246 times, so iterations are about three per dispatch.
which leaves 119,542 removed against roughly 175,246 that could have been. a draft of this paragraph said one per tail hop and had not looked at the second number. about two thirds of iterations were allocating, and what accounts for the other third is not established here. the obvious candidate is arity zero, since asking for a capacity of nothing allocates nothing, and it stays a candidate: nothing in this profile counts dispatches by arity.
what the runner read. this container projected 5,073,171 and ci reads 6,858,880 — the same direction, larger, on different silicon. the allocation counter beside it falls 1,410,530 to 1,309,483, and the peak 833,466 to 833,463. every compile vein is byte-identical, which makes this the fourth runtime-only change in a row that moved no layout.
one claim about those counters was published and then half withdrawn, which is worth the space because the half that survived is useful. the allocation counter reproduces across hosts: the container printed 1,309,483 for this tree and the runner read 1,309,483, and on a second tree both read 1,410,530. the peak does not. on that second tree the container reads 833,458 against the runner's 833,466, eight bytes apart. so it is the traffic count that is host-independent rather than the memory rows as a family, which is why one of those gates carries no allowance for the machine and the instruction gates refuse to compare at all on this box.
a fourth change leaves the comparing alone. memcmp reads 46,087,562 then 46,118,166 across the §97 pair, +0.066%. a draft of this paragraph had it falling 1,484,175, which was a debug-info profile read against a release one — two build configurations rather than two trees. compared inside its own sitting it has not moved, and the three changes in §96 become four.
100one row, read twice
a golden carried the same row twice. two branches did it the same way: a merge brought main's value forward, wrote a header block saying which reading it was, and appended it under the row already there instead of replacing it.
the reader adds rows into a map keyed by name, with +=. that is right and has to stay — a cost golden holds one row per sample and the gate sums them, and one objective counter can be made of several gate keys. what the += cannot tell apart is a second sample from a second copy. the gate read 13,665,824,705, exactly the two values added.
the counter sweep, the compile sweep and all three page gates ran green on both branches. none of them reads a golden for a repeated key: they compare a measured row against the file, and when the file answers with a sum they compare against the sum.
what did speak was the spec that replays the objective's source list, and what it said was that the LINK was wrong or a pool had joined the sum. the link was fine and no pool had joined anything. that spec compares totals, so a doubled row arrives as a wrong total and it names the last thing that could produce one. two branches sat red on a message pointing three directories away from the duplicated line.
the new spec reads the trend gate's own list of goldens and asserts no counter name appears twice in one file, reporting the file, the name and the count. it was watched red by planting the fault, and the invariant was checked against the tree first: all twenty-nine goldens the gate names already hold at most one row per counter. it does not check values. whether a row holds the right number is what the gates are for; this checks that there is one of it to read, which is the property a merge breaks and no measurement can restore.
101the frame handed back
§99 kept the score buffer because nothing consumes it. the bindings buffer is the other half and cannot be kept the same way: it is moved into the environment frame, and the body needs it. what can be taken back is the frame itself, once the body has finished with it.
whether that is ever possible was worth measuring before building. an instrumented binary took a weak reference to the frame before the body ran and tried to upgrade it afterwards:
frames made 173,922 frames dead 173,921
one frame in the whole run outlives the body it was given to, and it is a lazy binding holding the environment it was defined in. a weak reference rather than a second strong one, on purpose: a strong clone would make every uniqueness check inside the body fail and change the thing being measured.
so the dispatcher keeps a handle beside the one the body gets, and afterwards asks for the frame back. the vector inside goes into a pool above the tail-call loop and the next dispatch takes it instead of allocating.
base 980,371,488
frame 977,583,095 -2,788,393 -0.28%
base frame delta
__rust_alloc 1,243,349 1,123,808 -119,541
__rust_dealloc 1,231,221 1,111,680 -119,541
grow_one 112,013 94,847 -17,166
drop_slow 262,323 160,154 -102,169
the projection off the frame count was seven to ten and a half million and the build paid 2.79. that gap is the useful part. the projection priced 173,922 pairs at the 42.4 and 61.2 instructions §99 and §97 measured; what arrived is 119,541 pairs at 23.3 each. fewer pairs than frames, because a pooled vector that already has capacity does not allocate when it is reserved again. and a cheaper pair than either earlier measurement, because this one buys its saving: a reference count up, a reference count down and an unwrap check on every dispatch, against an allocate-and-free it no longer makes.
the 102,169 fall in the reference-count teardown is the frame being unwrapped instead of dropped through it, and is the clearest single sign the change does what it says.
the node around the vector is still allocated. every dispatch that binds anything still calls for one, and reclaiming what is inside does nothing about it. that is the other half of the projection, untouched.
a spec caught this one too, the same spec, the third change running: an exact per-round allocation count that fell by two a round again. this is the change most likely to disturb what it measures, since it holds a second handle to a frame whose values are the containers the builders check for uniqueness. it does not, because the frame was already alive for the whole body — the extra handle moves when the frame dies, from inside the body to just after it, and uniqueness is decided while the body runs. the builders still answer what they answered before, and the number moved down.
102the node around the vector
§100 took the vector out of a dead frame and let the node around it go. that was the tool's doing rather than a choice: Rc::try_unwrap reaches the value by consuming the handle, so the only route to the vector ran through freeing the box, and the next dispatch called Rc::new for a fresh one. Rc::get_mut reaches the same vector through a handle that stays alive.
so the pool stops being a vector and becomes the frame itself. it lends its bindings to the candidate list at the top of the loop and takes them back once the body has finished.
base 977,646,574
node 972,776,892 -4,869,682 -0.50%
base node delta
interp_allocs 1,183,336 1,063,795 -119,541
interp_alloc_bytes 84,810,613 75,247,333 -9,563,280
interp_peak_bytes 834,079 834,079 0
the allocation count is the same 119,541 §100 reclaimed vectors for. not a number near it — the same one, which says the two changes take the same set of frames from two sides, and that set is now fully accounted. the bytes divide exactly: 9,563,280 over 119,541 is 80.0, the size of the box around an environment.
the price per pair is the undiscounted one. 4,869,682 over 119,541 is 40.7 instructions, against the 42.4 §99 measured for a small allocate-and-free pair. §100 got 23.3 for the same count because it bought its saving with a reference count up, a reference count down and an unwrap check on every dispatch. this change swaps one call for another at the same point in the same code and adds no traffic, so what arrives is close to the full price. together the two take 119,541 frames from two allocations each to none.
the peak does not move, where §100's rose 616 for the vector it kept. a retained 80-byte node is not resident at the high-water mark, because the node it replaces was resident there in the base arm too.
one thing in the per-symbol table needs saying rather than publishing: the allocator's own frames fall by about 5.5 million and the reference-count teardown rises by about 3.06 million. that is attribution moving. the bindings a frame holds used to be dropped through a bare vector, where the compiler inlined the work into the dispatch loop, and they are now dropped out of a vector living inside the frame, where it lands in drop_slow's symbol. the same value drops happen either way, and nothing here isolates the shift, so it is recorded as a place the profiler filed work differently and not as a second effect. the three counters above are what the claim rests on, each read twice.
the uniqueness test got stricter and that is safe here. try_unwrap reads the strong count alone; get_mut reads the weak count too, so a frame with a live weak handle would be reclaimed by the first and refused by the second. nothing in the compiler takes a weak handle to an environment — the instrumented build in §98 that did was reverted — so the two agree today, and the difference is written into the source, because the day something takes one this becomes a quiet loss of reuse.
the spec caught it again, the fourth change running: the per-round allocation count fell by two a round once more, 4,203 to 3,603, with the builders still answering what they answered before. four changes to one loop, the same two a round each time. the file declines to say which iterations those two belong to, and that restraint is the point — a decomposition guessed there is what the next reader would check their change against.
what is left is the loser's buffer. the candidate loop swaps the working vector with the outgoing best's, so when the winner's goes into the frame the other falls out of scope at the end of the iteration. the same trick a third time, bounded by one pair per dispatch that had more than one arity-matching candidate, and that count is not in hand.
what the runner read. 929,300,332 to 923,151,727, a fall of 6,148,605 (-0.66%), read twice in the same job with the same answer. the allocation counter falls 1,183,336 to 1,063,795 and the peak does not move. every other vein reported success — the twelve cost veins, all eight compile-side rows, both codegen tiers, emitting and start-up — which makes this the fifth runtime-only change in a row to move no layout.
this container projected 4,869,682 and the runner read 6,148,605. same direction, larger, different silicon, and the third time this family has landed that way: the ratios are 1.26, 1.35 and 1.35. three is enough to say the container under-reads this row's improvements rather than that one reading was unlucky, and not enough to say by how much.
the allocation count reproduced on both arms. the container read 1,183,336 for the base and 1,063,795 for this tree, and the runner read the same two numbers — a two-arm confirmation of the property this vein has and the instruction vein does not.
the interpreted row has now fallen from 1,075,174,600 to 923,151,727 over ten builds and five wins, 14.1%. the objective moved by a hundredth of a point for it, for the third entry running: that term is more than 136% improved against its baseline and sits under a satiating curve, so the next six million buys almost nothing. a change wanting to move the number has to find production work or an unsatiated term.
103the third pooling, which does not pay
§102 left the third of these. the candidate loop swaps the working vector with the outgoing best's, so when the winner's goes into the frame the other falls out of scope at the end of the iteration. pooling it is the same trick a third time. it costs.
base 972,776,892
spare 976,125,124 +3,348,232 +0.34%
base spare delta
interp_allocs 1,063,795 1,041,355 -22,440
interp_alloc_bytes 75,247,333 69,333,733 -5,913,600
interp_peak_bytes 834,079 834,303 +224
three readings of each arm, interleaved, each arm repeating its own figure exactly. and it does save the allocations it set out to save, which is what makes the result worth keeping: 22,440 pairs gone, and the row still up by 3.35 million. at the 40.7 instructions a pair §102 measured those pairs are worth about 914,000, so the bookkeeping cost something over four million.
the asymmetry is the whole answer. the spare has to be taken out and put back on every dispatch — 175,254 of them — to serve a reuse that fires on 22,440. about 24 instructions a dispatch, against a saving on one dispatch in eight. the two poolings that worked pay their bookkeeping on the same frames they save, one for one.
the ceiling was measured before the build and was still too kind. an instrumented binary counted 38,294 dispatches leaving the working buffer with capacity, which projected about 1.56 million; the real saving was 22,440 pairs, because some of those buffers had capacity they recycled within the dispatch rather than allocated. so the count was 70% high before the overhead was counted at all.
a ceiling computed from a count bounds the saving and says nothing about what collecting it costs. the two changes before this one happened to have negligible collection cost, and that is a fact about them rather than about the technique.
three poolings attempted, two kept. the score buffer, the frame's vector and the frame's node took the interpreted row from 1,075,174,600 to 923,151,727. the loser's buffer is the one that does not pay, and this section is here so it is not tried a fourth time.
104what a name lookup walks
the dispatch-loop pooling is finished — three attempted, two kept — and the next frames in the interpreted profile are match_one at 68,147,655, lookup at 67,319,929 and Value::clone at 54,332,739. this measures the second before anything is built on it.
look_calls 1,056,329 look_probes 2,662,536 2.52 name comparisons a call look_depth 1,155,900 1.09 frames walked a call look_bytes 5,291,628 1.99 bytes a comparison look_miss 332,025 31.4% of calls answer nothing look_miss_probes 1,071,803 40.3% of all probes look_sameptr 0
the chain is shallow and the scan is not. 1.09 frames a call means a lookup almost always answers in the frame it starts in, or fails in it, so the cost is the linear scan of one frame's slots rather than a walk up a long parent chain. anything aimed at chain depth is aimed at nine per cent of the calls.
the names are tiny: 1.99 bytes a comparison, two-character names compared with a length check and a memcmp call whose overhead dwarfs the two bytes it reads. __memcmp_avx2_movbe is 34,234,622 in the same profile, and 34,234,622 over 2,662,536 is 12.86 — the right size for these comparisons, and not evidence that they are these comparisons. nothing here isolates memcmp's callers, so that share stays open. the probe count is measured; the attribution is not.
two fifths of the comparisons are made by lookups that find nothing: 1,071,803 of 2,662,536. a miss averages 3.23 probes against a hit's 2.20, because a hit can stop early and a miss cannot stop at all — it scans every slot of every frame before falling through to the global table. a name the compiler could tell was global would skip the walk entirely.
the last counter measures the representation rather than an opportunity, and it was nearly published as a finding. a name is not a pointer: it is held inline, one length byte and twenty-two of payload, in the ast node itself. so each name is its own buffer, two names reading the same text can never share an address, and look_sameptr could only ever have been zero whatever the program did.
interning is also already ruled on, and that module's own header says so in its third sentence: it was measured and declined at 365 conversion sites for one field of twenty-nine, and the 2026-08-29 ruling took the other road — the one that needs no table, no lifetime and no id. the inline representation is that road. a first draft of this section proposed interning as the larger of two leads, against a ruling recorded a fortnight earlier, because the file that says so was one file away and was not read.
the 1.99 bytes a comparison is that same header's measurement from the other side: 89.8% of identifier occurrences are seven bytes or fewer. a comparison already that short is not where the instructions are.
nothing built. the lead is the misses, and it is the only one these numbers support: their saving is bounded below by work that is provably wasted.
105a filter on the frame, and a premise already refuted
the misses were the lead, and the cheapest way to skip one is to let the frame say no: a one-word membership filter over the binding names, tested before the scan.
base 972,776,892 filter 983,593,036 +10,816,144 +1.11%
two readings of each arm, identical both times, allocations byte-identical. worse than the third pooling, and for a related reason.
a membership filter is an optimisation for a long scan: it earns its keep when rejecting costs much less than walking. the measurement that suggested the lead had already said the scan is 2.52 comparisons over 1.09 frames. that is not a long scan, and there was never much for a filter to skip.
the two fifths is a share, and it was read as though it were a magnitude. 1,071,803 wasted comparisons is most of the name-comparison work and still a small number of instructions, against a filter computed once per lookup on all 1,056,329 calls and rebuilt on every frame that binds — both paid whether or not anything is skipped.
that is the second decline in a night, after the loser's buffer, and both come from one habit: pricing a saving from a count while leaving the unit cost unmeasured and the overhead out of the model. so the rule earns a sharper form: before building on a count, divide it. 1,071,803 probes over a run is a share; 2.52 probes over a call is a scan length, and the second is what decides whether skipping the scan can pay. both numbers were in hand both times.
what remains true is that the misses are 31.4% of lookups and walk frames they will never match in. what is now also known is that each walk is short, so anything that helps has to remove the call rather than shorten the scan. that is the global-name skip, and it needs a per-ident answer with no hashing in it: the whole-program side table the demand analysis uses is asked once per statement where this would be asked a million times, and the ident node has nowhere to keep a bit — giving it one is 278 match sites across sixteen files. that is the size of the thing, measured rather than guessed.
106twenty candidates a dispatch was two point six
the dispatch loop's own comment, the entry about the misplaced reserve and this page all said arm selection tries around twenty candidates a dispatch. measured on the interpreted corpus:
dispatches 175,254 candidates 456,967 2.61 a dispatch match_one 1,135,058 6.48 a dispatch, 2.48 a candidate of which bind 817,017 72% of matcher calls push a binding
the useful part is where the twenty came from. the entry shows its working: 1,100,726 calls to the matcher against 55,711 dispatches. that is 19.75, and it is matcher CALLS a dispatch. the matcher runs once per parameter of each candidate tried, so calls a dispatch bounds the candidates from above and equals them only if every arm took exactly one argument. the denominator was right and the numerator counted something else.
the old figure's own arithmetic agreed with the smaller number. the misplaced reserve cost 12,017,902 instructions; over 175,254 dispatches that is 68.6 a dispatch, which is 26.3 per reserve at 2.61 candidates and 3.4 at twenty. no reserve costs 3.4. the contradiction sat in the same paragraph for a fortnight.
that is the third count today used with the wrong denominator, after the loser's buffer and the frame filter. those two divided a run total by nothing and read a share as a magnitude; this one divided by dispatches and labelled the answer candidates. a count is worth what its denominator says, and the denominator has to be the thing being counted.
what stands is the conclusion: the reserve belongs once per dispatch, and moving it was worth 21,582,351. what changes is the size of the explanation — it was misplaced by a factor of 2.61, and its twelve million came from a reserve call that is not cheap rather than from being paid twenty times.
and it retires a lead. a loop trying twenty arms to find one looks like an obvious place for a pre-filter on the first argument's type. at 2.61 tried and 202,987 of 456,967 matching, the loop already tries barely more arms than it accepts, so there is little to filter. closed by the measurement rather than by an attempt.
107where the production run spends
three changes in one night took 15.8 million off the interpreted row and moved the objective from 77.25 to 77.27. that is the model working: the interpreted row sits on the development side and is 136% improved against its baseline, so it is deep into its satiating curve. the objective says where the money is in one line — run speed carries a satiation of 2.0, which is late, and 0.45 of the production side.
three entries that night ended with the observation that a change wanting to move the number has to find production work, and the next change went back to the interpreter each time. this is the first look at the other side.
1,840,368,355 PROGRAM TOTALS
392,547,176 21.33% json/encode_onto
154,484,748 8.39% json/obj_key_start
97,131,375 5.28% json/parse_value
91,704,604 4.98% runbench/tally
85,720,338 4.66% json/array_delim
84,209,220 4.58% render_ryu
77,645,700 4.22% json/scan
65,074,779 3.54% json/str_escape
61,672,983 3.35% k_beat_iter
49,334,986 2.68% k_b_at
and divided, which is the rule the two declines that night bought. a share is not a magnitude, and a total says nothing about what one call costs:
encode_onto 2,380,860 calls 164.9 instructions a call obj_key_start 784,179 calls 197.0 parse_value 1,100,187 calls 88.3 k_b_at 690,000 calls 71.5 k_beat_iter 2,685,021 calls 23.0
the encoder is the largest frame in the production workload by a factor of two and a half, and 164.9 instructions a call is a lot for a function whose job is mostly appending a few bytes.
the obvious suspicion is wrong, and reading the output said so before anything was built. encode_onto is an overloaded group of eight arms, all of arity two: true, false and null, then int, float, string, list and map. a document of mostly strings and numbers looks as though it must reach its arm past the three nullary tests every call, which would make arm selection a share of the 164.9. the emitted ir settles it:
%t9 = extractvalue %KValue %x1, 0
switch i64 %t9, label %L7 [
i64 2, label %arm0 i64 3, label %arm1
i64 0, label %arm3 i64 1, label %arm4
i64 6, label %arm5 i64 9, label %arm6
i64 10, label %arm7 i64 7, label %L8
]
a jump table on the value's tag. eight arms cost one switch, arm order decides nothing, and the only linear step is the record case, one check for tag seven. the dispatch is not where the instructions are.
so the 164.9 is the appending. per call the common path is one failure check on the accumulator, the switch, and an arm body: a string is two byte-appends around a whole escape pass, where an int or a float is one rendered append. that is real work rather than overhead, and it is why the frame is large.
two changes died that night for want of this step. the loser's buffer and the frame filter were both built on a suspicion the code would have refuted, and this one was refuted for the price of reading eighty lines of ir. the lead stays open and its shape has changed: what to attack in the encoder is the appending path, and nothing here says that path is wasteful.
108the object gets a name the run chooses
the release-tier codegen row measures a whole build, clang and ld included, and it had been drifting. pinning the linker plugin’s thread count took the drift from millions down to eleven instructions, and the eleven turned out to be a string: clang writes its link-time object to /tmp/<stem>-XXXXXX.o with fresh hex every run, ld’s plugin puts that path into a hash map, and the probe walks a different number of buckets depending on the name. twenty-two names on one binary and one corpus — twenty read 5,163,341,031 and two read 5,163,341,042, reproducibly, the same name giving the same answer every time.
-save-temps=obj derives the object’s name from its input, so the string is the same on every run and the draw stops. that is the change, and it is a measurement change rather than a compiler saving: nothing a user runs gets faster.
and the container was wrong about what it costs. here the pinned name read 6,811,830,244 twice against 6,813,182,505 three times without it — a saving of about 1.35 million. the runner reads the same change as a cost of 9,653,055, seven times the size and in the other direction. so the pin’s effect on the row’s level is a property of the host; only its effect on the row’s steadiness reproduced on both.
that is worth stating on its own, because the container has now been right and wrong about the same kind of projection in one fortnight. it travelled for a change that moved a count — how many allocations a program makes — and it failed for a change that moves where bytes land. a projection from this box is worth what the quantity it projects is: a count travels, a layout does not.
the pin covers both tiers. the flag reaches the release path and not the development one, and the development row moved anyway by 3,868, which read for a day as a move nobody could explain — and re-basing a row on an unexplained move is the thing the goldens exist to catch. the explanation was in the gate. it sets the variable on all three of its env -i lines, so the development run’s environment block grew by nineteen bytes, and a child process inherits that block on its initial stack. this row is the child tree; the compiler’s own process is excluded from it. so the block is the only thing the change alters for the processes being counted.
three arms at that tier, one tree, each pass reproducing byte for byte:
KANSO_FIXED_TEMPS=1 595,943,218 the gate as written
nothing 595,947,057 without it
KANSO_FIXED_TEMPX=1 595,943,490 nineteen bytes, read by nothing
a variable the development path never reads moves the tree 3,839, and a variable nothing reads moves it 3,567. the 272 between those two arms is the block’s content, one letter apart at the same length. the container’s figures are its own and an instruction delta does not travel between hosts; what travels is that the unread arm moves nearly as far as the pinned one.
so an environment block is a term to hold still like any other. adding a variable to one tier’s command re-bases that tier’s row whether or not anything reads it, and the 2026-09-15 rule covers it: put the state into a known one rather than explain it afterwards.
109a variable nothing reads, and 117 instructions
the interpreted row has been disagreeing with itself. two ci jobs on one commit read 2,178,502,266 and 2,178,502,272 — six instructions apart, three parts per billion, each stable across the gate’s own second reading. the standing explanation was an argument rather than a measurement: that where the allocator’s heap starts moves with the size of the file the loader mapped, so the residue is a term proportional to work rather than a constant.
here is the half of that which can be measured on one box. two arms, one tree, the gate’s own anchor and exclusions, each read twice and byte-identical both times:
env -i PATH=... GLIBC_TUNABLES=... 972,776,892
env -i PATH=... GLIBC_TUNABLES=... KANSO_FIXED_TEMPX=1 972,777,009 +117
nothing in the tree reads that variable. it is one more entry in the environment block, and the row moves 117 instructions for it. the allocation counters do not move at all: traffic, bytes and peak are byte-identical across the two arms. the interpreter asked the allocator for the same things in the same order, and the row still moved.
and the carrier is named, because the term is exactly linear in the number of variables. four arms, each read twice and byte-identical:
0 extra 972,776,892
1 extra 972,777,009 +117
2 extra 972,777,126 +117
3 extra 972,777,243 +117
117 a variable, to the instruction, for variables nothing in the tree reads. layout does not do that: adding bytes to a block moves an address once and by whatever the alignment says, and it does not charge the same toll four times running. a scan does. diffing the two profiles frame by frame names it: the frames that move are getenv, __strncmp_avx2, and the allocator’s own _mi_prim_getenv, _mi_strnicmp and _mi_toupper. mimalloc looks its options up by name, each lookup walks the environment block comparing as it goes, and one more variable is one more comparison in every walk.
most of that is start-up and the row excludes it. the 117 is the part inside the anchor, which means the interpreter’s own thread resolves an allocator option while it runs. so this is not a term to explain: it is the 2026-09-15 rule’s own case, where the state can be put into a known one and the reading stops depending on it.
it is still not the cause of the six two jobs disagreed by — between two jobs on one commit the environment is identical, and six is not a multiple of 117. what it settles is that this row has a term of that shape, twenty times the size of the disagreement, with a mechanism rather than a suspicion behind it.
and the arm that looked like the candidate’s own claim was contaminated. growing the corpus file with comment lines the interpreter never runs moved the row too, 65 instructions for 62 KB and 4,484 for 124 KB. that reads as the mapped file doing exactly what the candidate said. but the bigger file costs the loader more allocations — traffic up five, bytes up 187,488, peak up 78,649 — and the anchor excludes the loader’s frames while sharing its allocator. that arm mixes whatever the file’s size does with work the loader really did, and cannot carry the claim. the environment arm can, because nothing reads the variable, the allocation counters hold, and the toll is the same 117 every time it is charged.
110six and a half per cent, getting in and out of functions
the production run spends 120,988,262 instructions on frames — the callee-saved registers a function pushes on entry and pops on the way out, and the stack it reserves. that is 6.57% of the whole program: 3.52% going in, 3.05% coming out, and it is a floor rather than a total, because the stack pointer’s own restore is not counted. more than rendering every float (4.58%), more than the entire arena beat machinery (5.09%), and until now it had no name.
30,952,350 json/encode_onto 2,380,950 calls, 7 prologue instructions
14,303,721 json/parse_value 1,100,286
10,760,607 json/obj_key_start 827,739
encode_onto is 21.33% of the run on its own, 164.9 instructions a call, twelve in and eight out. reading its jump table and counting the entries at each target says who pays that. the number arms take 379,530 calls and are four register moves and one call. true, false and null take 562,500 between them and are five instructions each into a shared append. strings take 942,750, maps and lists the rest. so two calls in five into that dispatcher pay a frame sized by the arms they do not take.
the obvious fix is the compiler’s own: sink the pushes into the paths that need them. llvm calls that shrink-wrapping and it is on by default, and it never fires here — compiling with it switched off gives a byte-identical prologue. with a jump table to eight arms and register uses in several of them, the entry block is the only place that dominates them all.
the other obvious fix is to split the dispatcher: emit the cheap arms so they need no frame, and tail-call the heavy ones. that rests on the frame being sized by the arm bodies inlined into it, and linking the ir by hand says it is not. marking the string arm’s callee noinline — the heaviest thing in there — takes the stack reserve from 0x58 to 0x38 and leaves all six pushes. marking every call site inside the dispatcher noinline, so no arm body is inlined at all, leaves all six pushes and 0x38 again.
the registers are the dispatcher’s own, then. it holds both of its arguments live across the calls it makes, and a value here is two words, so the arguments alone are four registers that have to survive a call. there is nothing for an arm split to take away.
what does move the number is inlining, because an inlined call has no frame at all. the threshold handed to the linker is already the lever, and pushing it further was measured on a program whose emitted ir is byte-identical across every arm:
225 1,938,999,983 343,128 bytes 22,317,482 calls
2000 1,840,276,313 424,088 18,812,341
4000 1,823,291,354 510,152
8000 1,821,134,592 567,496
225 to 2000 — the step this compiler already takes — removes 3,505,141 calls and 98,723,670 instructions. that is 28.2 a call, which is a frame plus the call and the return, and it prices the mechanism.
and the next rung is declined. 4000 buys another 0.923% of the run and costs the release build 3.64% more work: 6,827,333,184 instructions against 7,075,918,942, two passes, each arm byte-identical. the development tier does not move at all, because it compiles at -O0 and never sees the flag. scored against the goldens, the run gain alone is 77.27 to 77.33 and the pair together is 77.26, one hundredth under the floor. the break-even was written down before the second arm finished — a release build 3.2% dearer — and the measurement came in at 3.64%.
the reason recorded beside that flag had been binary size, and binary size is not scored at all. now the reason is the one the objective can see.
111what a point of the objective costs
the compiler has an objective function, and for a fortnight the way to decide what to work on next has been to read its weights and satiations and reason about them. the function can be asked instead. take the goldens, scale one row by a tenth, score the model, and the difference is what that tenth is worth. fourteen rows, one at a time, on 2026-09-19:
per 10% counter golden row
+0.6644 run_instructions instructions:runbench
+0.5091 run_peak_bytes cost_golden_run:arena_peak_bytes
+0.1952 codegen_instructions_release codegen_instructions_release
+0.0923 startup_instructions startup_instructions
+0.0692 codegen_instructions_dev codegen_instructions_dev
+0.0446 interp_instructions interp_instructions
+0.0233 compile_peak_bytes compile_memory:compile_peak_bytes
+0.0214 compile_instructions entry_instructions
+0.0197 interp_peak_bytes interp_memory:interp_peak_bytes
+0.0197 compile_allocs compile_allocs
+0.0094 run_peak_bytes cost_golden_run:held_peak_bytes
+0.0068 emit_instructions emit_instructions
+0.0060 compile_instructions compile_instructions
+0.0003 run_peak_bytes cost_golden_run:perm_peak_bytes
the same sweep at one per cent gives the same order with every figure about a tenth of these, so the curve is near enough straight over that range and the table ranks reliably even though each number is an average over its own step.
two rows are two thirds of the board. the run program’s instruction count and its arena peak come to 1.1735 of a 1.6808 total. everything else together is worth less than the first row alone.
and the arena peak has had no work at all. it is worth 77% of what the run’s instruction count is worth, and nothing in the last fortnight’s log would suggest it: every entry in that window is instructions — the dispatch pooling, the beat rewind, the digit loop, the frames. the second most valuable row in the model has not been named once.
the table also explains a fortnight of small moves. the interpreted row is worth a fifteenth of the run row, and three changes that took 1.7% off it moved the objective by two hundredths. the compile row is worth a hundred and eleventh of the run row, and it is the row this project has spent the most rounds arguing about.
the figures are marginal at today’s ratios and move as the terms improve, which is why they carry a date. what would keep them current is a flag on the scoring script itself, printing this table out of the model it already holds.
112three quarters of the memory, in 4.9% of the work
the table above puts the run program’s arena peak second on the board. this is where it lives. one phase count at a time, taken to its floor, everything else left alone:
baseline 38,604,496 36 blocks
decode = 1 38,604,496 36 unchanged
encode = 1 38,604,496 36 unchanged
decode = 1, encode = 1 38,604,496 36 unchanged
deep = 1 38,604,496 36 unchanged
escape = 1 38,604,496 36 unchanged
pend = 1 38,604,496 36 unchanged
index = 1,000 35,651,584 34
digest = 1 35,458,768 33
split = 1 9,244,368 8
the split phase holds 29,360,128 of it, three quarters, and split is 4.87% of the program’s instructions. decode and encode are 69% of the work between them and hold none of the peak: taking both to a single round leaves the number unmoved to the byte.
scored, that is +4.75 welfare — 77.27 to 82.02. every compiler change merged in the two days before this section moved the objective by two hundredths together.
a cause was written down eleven days ago in the benchmark’s own header: codegen reads a list of beat loops, that list drops every group whose file begins std/ or lib/ from the carry tier, and clearing the filter gives a peak of one block with the allocation count unchanged. the measurement is the header’s and stands. the mechanism it names does not: instrumenting the filter shows exactly eleven imported groups losing a carry there, and the regexp walker has none to lose: it is reported grow-only because another group tail-calls it, and the filter only ever reaches a group that got a carry in the first place. what clearing the filter actually changes is somebody else’s loop, and which one is open.
clearing it wholesale is not the fix and was measured not to be — every library loop then begins evacuating, and the program that ran in 0.4 seconds was still running after ten minutes. the header names the real shape: the carry tier is decided by a path prefix rather than by the property that prefix stands in for, which is whether the loop threads its caller’s invariant source through itself. the machinery for deciding that already exists one function away, as the threaded-slot fixpoint the cluster analysis runs.
so the lead is not new. its size is: the largest single move on the board by two orders of magnitude, against a fix whose shape is already written down and whose crude form is already priced.
113the first cut at the prefix filter, declined at seven times
the section above says the carry tier is decided by a path prefix rather than by the property that prefix stands in for, and that the machinery for the property already exists: the cluster analysis refuses to carry a slot it finds threaded. so a carry that came out of that analysis has already been checked for the thing the filter guards against, and the filter could let it through. four lines.
arena_peak_bytes 38,604,496 -> 35,458,768 -3,145,728, worth +0.41
allocs 5,730,653 -> 6,550,655 +820,002
alloc_bytes 459,964,461 -> 494,316,813 +34,352,352
run instructions 1,840,276,313 -> 13,618,672,806 7.4 times
declined. the peak gain is real and the instruction cost is not survivable, and it is the 2026-09-01 catastrophe in miniature: that removal left the program running after ten minutes against a 0.4-second baseline, and this one, over the cluster subset alone, costs seven and a half times.
what it rules out is the useful part. the cluster analysis excludes threaded slots from a carry before it returns, so those carries had already passed the test the filter’s own comment describes — and exempting exactly them still blows up. so that fixpoint is not what the path prefix is standing in for. the next attempt has to find the difference rather than assume the two agree.
the split phase’s 29 megabytes are untouched by this cut, which is its own evidence. the regexp walker is reported grow-only for an outside tail call, so its carry comes through the demoted-entry path and the crossing-position test, neither of which has a threaded fixpoint at all: a bare parameter is accepted there only when its inferred set is narrow enough, so one carrying ordinary heap becomes a crossing position and would be copied every iteration. giving that path the fixpoint is where this goes next.
114the width of a carry, and a test that says not yet
the eleven groups the prefix strips were priced one at a time. alone, every one of them reads the baseline on both columns — peak 38,604,496 and 0.26 seconds — so the cost is a pairing rather than a group, and the pairs separate:
sha256/compress + sha256/turned 35,458,768 1.01s
sha256/blocked + sha256/digested 35,458,768 0.26s
the other seven 38,604,496 0.27s
all eleven 35,458,768 0.95s
the whole saving sits with the cheap pair. blocked and digested give the entire 3,145,728 bytes at baseline wall time; compress and turned give the same bytes and all of the cost. counted rather than timed, the cheap pair costs 778,170 instructions, 0.0423%, against 3.1 megabytes of peak. scored: 77.27 to 77.68.
the first pair carries two positions each and the second one. so the width of a carry reads the property the path prefix was standing in for: evacuating one slot a lap is what the tier is for, and evacuating several is where a library loop starts copying its caller’s work. it is a proxy for bytes copied per iteration, which nothing in the pass can measure, and it is a proxy the loop’s own shape supplies rather than its file name.
it does not ship yet, because a test goes red: the rule as written also admits two of the json decoder’s scanners, and a unit test keeps scanners threading records or lists on the grow-only arena while only the byte-builder encoders rewind.
what that test pins is worth reading carefully, and the first reading of it here was wrong. it looks like a judgement about freeing memory under a live reference. the prefix it protects is a cost guard and says so in its own words: carrying a library driver’s threaded source copies an unbounded value every iteration. safety is established elsewhere and still is — the licence that lets a value cross a rewind, the map exclusion, the byte-chain rule — and a neighbouring test pins a carried list directly, where a fixed-shape rebuild carries and the evacuation handles it.
the measurement says those two scanners contribute nothing either way, so the whole 0.41 is available without touching them. what is owed is a rule that admits the sha256 pair on a property, leaves the scanners where that test wants them, and says why the difference is real.
115a folder called lib was deciding a program's memory
the compiler decides which library loops may hand their arena back between laps. it decides it by asking whether the declaration’s file begins std/ or lib/. that field is the one error origins are built from, a path meant for a diagnostic, and it is deciding how much memory a program holds: the same package under a directory called lib compiles to one that never reclaims a block, and one directory over to one that does. a test has pinned that since it was written, as a defect, with instructions for the day it goes.
the prefix stands in for something real. a shared library driver threads its caller’s invariant source through the loop, and evacuating that copies an unbounded value every lap. removing it outright was built and measured on 2026-08-31: the digest went quadratic, 1.3 seconds to 68 at 128 KB.
eleven imported loops lose a carry at that point on the run program. priced one at a time, every one of them reads the baseline on both columns, peak 38,604,496 and 0.26 seconds. so the cost is a pairing:
sha256/compress + sha256/turned peak 35,458,768 1.01s
sha256/blocked + sha256/digested peak 35,458,768 0.26s
the other seven peak 38,604,496 0.27s
the expensive pair carries two slots each and the cheap pair one, and the whole 3,145,728 bytes sit with the cheap pair. that made the width of a carry look like the property to read, and a rule admitting a carry of at most one slot was built and measured. it is declined here, and what declined it is worth writing down.
under that rule the ratchet program dies. it runs clean on today’s compiler and runs out of memory on the new one. the runtime says out of memory, the driver translates that into the program ran out of stack, and neither sentence is what happened. printing the request at the failure gives it:
CARRYOOM need=18446744072171062384 depth=3 carry_n=1
that is 264 less 1,538,489,232. the walk that sizes a staged carry read a length of about minus one and a half billion out of a node it was asked to copy, the sum wrapped, and malloc refused it. a carried value whose interior the walk cannot read is the thing the prefix was keeping out, and carry width does not see it.
the failing group is list/holds_any?: it appears in every coalition that dies and in none that survives. it is the driver behind any?, and the backtrace at the failure has one any? running inside another’s predicate. excluding the three list drivers that take a predicate runs the ratchet clean and still reads 35,458,768, so a rule shaped that way would ship the whole saving.
that rule is not shipped, because there is no small program that fails without it. a nested any? over two 800-element lists reads the same peak either way. the coalitions are not monotone either: two loops together die where either alone survives, and a third added back survives again, which is what a failure sitting near a memory limit looks like rather than a rule being broken. a guard swept out of one program is a guard nobody can check. the next step is the small program that makes the sizing walk read a garbage length; the rule follows from that, and not before.
116the loop carried a value it only needed at the end
rendering a double is 4.58% of the run program, 440.7 instructions a float over 191,070 calls, and a quarter of that is one loop taking digits off two at a time. it divides three values a trip: the two ends of the rounding interval, to decide whether another pair comes off, and the significand, because the significand is the answer.
only the first two decide anything. the third is carried through the loop and read once at the end, so it comes out: count the pairs with two divisions a trip, then take it down in one step.
that was worth 924,584 instructions, 4.8 a float — and the number did not fit. disassembled, the loop went from 21 instructions with three multiply-highs and seven register moves to 15 with two and five. six a trip, and at 5.35 trips that is 32.1 a float. the tail was giving twenty-seven of it back.
the tail divided twice by a table entry, and a divisor the compiler cannot see is a real division rather than a multiply-high. but the two entries are a hundred apart, so dividing by the smaller one first leaves both remaining steps with a constant divisor:
uint64_t q = vr / RYU_POW100[pairs - 1];
round_up = (uint32_t)(q % 100) >= 50;
vr = q / 100;
one variable division instead of two, and the fall triples:
runbench 1,840,368,292 -> 1,837,534,164 -2,834,128 -0.1540%
render_ryu 84,209,220 -> 81,375,750 -2,833,470 -3.365%
a float 440.7 -> 425.9 -14.8
.text 319,346 -> 317,954 -1,392
the lesson is about where a saving goes rather than about ryu. taking work out of a loop that runs five times only pays if what replaces it costs less than five times what was removed, and the first shape of this change failed that test while looking like a win. the disassembly is what said so; the instruction count alone would have been read as a small success and left there.
117the same name, allocated twice
§87 left the interpreted row at 1,555,890,579 and a lead open: the profile still had the allocator near the top. two more changes took it to 1,138,001,430. the first sized a clone for the growth that followed it. the second is smaller than either and is the one worth reading, because the cost it removes was never anybody's decision — it was two pieces of code each doing the right thing on its own.
calling a function binds its parameters. the matcher walks the pattern against the argument and pushes what it matched into a vector of name-and-value pairs, allocating a fresh string for each name. the caller then drains that vector into the environment, and the environment's constructor took a borrowed name and allocated its own copy of it. so each binding allocated the same bytes twice, and freed the first copy a line later. half a million body entries, six bindings each.
the fix is two steps, and taking them one at a time is what says where the cost was. giving the constructor the name by value removes the second copy for the bindings that came through a match: the corpus loses four allocations a round. storing the name as the front end's own identifier type removes the first copy for every binding whatever route it came by, and that is six a round. the type is twenty-four bytes, which is exactly what a heap string costs, and it keeps a name of twenty-two bytes or fewer inside the value. across the shipped library 99.77% of identifiers are that short and 89.8% are seven bytes or fewer, so for almost every binding the remaining copy is twenty-four bytes of stack and no allocator at all. no structure grew.
the row fell 122,261,480, 9.70%. allocations fell 1,303,589, 33.92%, which is the largest single move that vein has recorded. the campaign that §87 opens now reads 2,168,428,538 to 1,138,001,430 — 47.52% of the interpreted run, and 52.17% of its allocations, none of it from making the interpreter cleverer.
storing a shorter name has one consequence worth a spec: the boxed path, for names past twenty-two bytes, is now load-bearing and nothing in the corpus exercised it. the fixture binds a forty-five-byte parameter, a thirty-nine-byte local, and two parameters whose first twenty-two bytes are identical. the last pair is the one that matters. a store keeping only the inline prefix would give both the same key, and a single long name would still match a lookup truncated the same way, so one long name alone proves nothing. it was watched red before it was green, with the constructor truncating deliberately: the interpreter stops with an unknown name and the compiled engine prints the right answer, because compiled code binds nothing through an environment. what catches it is the corpus running both engines rather than the corpus running at all.
and a sixth reading of the measurement question §87 ends on. that section says a counter of instructions cannot be compared across machines, not even as a difference, and explains the four earlier agreements as four changes whose copying happened to compile similarly on both hosts. this change was projected on the container at 127,890,665 and read on the runner at 122,261,480: 4.40% apart, on a change that moves no copying at all.
the rule survives and the explanation does not. an instruction delta is still not something to rely on across a machine boundary, which is why the golden is the runner's reading and this box's number never goes in it. but the earlier agreements were not luck. this change removes allocations, and an allocation is a count the program decides, so the two machines differ only in what one call to the allocator costs. the clone-sizing change moved bytes per copy, and what a copy of N bytes costs is the compiler that built the interpreter. a projection travels when what it projects is a count of operations and does not when it is a price per byte. two readings support that and it is a prediction rather than a law: the next change that moves bytes should be expected to break the box's projection again.
three days later that failure would not happen again. the rule was rebuilt from the same commit — one file, src/beat.rs, is all that separates it from today’s compiler — and the arena peak confirms it is doing its work: 38,604,496 bytes over 36 blocks becomes 35,458,768 over 33, which is the figure this section quotes above. with that control passing, the ratchet exits cleanly three times from kanso run, three times from a dev build, three times from a release build, and once more under a sanitizer that would have caught the bad read. memory pressure does not explain it either: the rule’s ratchet peaks at 28,220 KB against the old compiler’s 28,176 KB, and the whole program fits in 28 MB.
so the reason written here is withdrawn. the rule is still not shipped, because the other objection stands untouched: it admits the two json openers, and the test that keeps scanners out of the rewind tier had to be edited to let them in. what has gone is the crash, and with it the case that the matter was settled — a measured 0.41 of the objective is a live question again rather than a closed one.
118a test that could not fail
one of the memory fixtures ended its header with a promise. it builds a shape the runtime's comments call dangerous — a node below a beat's mark holding a pointer into a tenure block, which a rewind can free — and it said that a change which stopped handing those blocks to the depth outside would turn the test red rather than turning some later program into a crash. that promise was written on 2026-09-08 and was false for a fortnight.
the runtime has two counters for tenure storage: blocks claimed and blocks freed. handing a depth's blocks up to the depth outside does not free them; it moves them, and the depth outside frees them when it pops. so the block is claimed once and given back once whether the hand-up happens or not. replacing the call with an outright free left all sixty-seven memory goldens byte-identical, and the fixture that promised to catch it passed.
the missing counter is the hand-up itself, and minting it is a one-line change with a long tail: a new counter is additive, so twelve cost goldens and every memory golden move in the same commit. with it there, three fixtures go red on the swap instead of none. one of them also moves the block count from five to three, which is the whole of what any counter could see before.
the fixture was carrying a second claim, and the experiment that checks it is cheap. the header said the hand-up is what makes the dangerous shape harmless here. the beat's pop does three things in order: it deep-copies the result out of the beat's storage, it migrates three registries, and it hands the tenure blocks up. writing garbage over the depth's 39,200 live tenured bytes at each of those three points gives:
before the deep copy dies, reading a length out of 0xAB
after the copy, before migrates correct output
at the hand-up correct output
the copy is what reads the tenured bytes out. after it, nothing below the mark still points into tenure, so freeing the block early costs this program nothing — which is also why it runs clean under a sanitizer with the hand-up removed. the hand-up is load-bearing, but on a different fixture: an inner beat that opens its tenure inside the outer depth's block, where the blocks really do have to travel.
three poisons one step apart, with the answer flipping across one of the two gaps, isolates the step. a saving that merely arrived alongside a change leaves the mechanism open, which this page has had to withdraw once already.
what closes it is the mutation script. the ratchet now applies exactly the swap the header describes and requires the memory vein to go red, so the promise is checked on every push rather than by whoever happens to read the comment.
119a profile you cannot diff
two builds of the same source read three instructions apart, and the job logs could not say where. each gate that counts instructions printed its profile's top frames as --threshold=90 | head -40: forty lines of the hundred and twenty-five that threshold has, fifteen functions once the headers come off. all fifteen were equal. the number is real and the instrument stopped one screen short of it.
the whole table is about a thousand rows and reaches functions that retire a single instruction. printing it costs a few kilobytes of job log. six gates now do, and the one property that is easy to get wrong is that they do it on every run rather than on the failing one — a comparison needs the side that agreed, and the side that agreed never takes a failure path. a spec derives the list off disk so a seventh gate cannot skip it.
the first thing it was pointed at was a row that has drifted by single digits between machines for a week. five builds on one machine, each a distinct binary:
main, unpadded .text 2,862,578 row 973,143,830
main, read again 2,862,578 973,143,830
+64 KiB of rodata no code reads 2,862,578 973,143,830
+40 never-called functions 2,862,690 973,510,133
+38 of the same functions 2,862,690 973,510,133
sixty-four kilobytes of data nothing reads leaves the row alone to the instruction. a hundred and twelve bytes of code nothing calls moves it by 366,303. the whole-table diff puts 366,005 of that in one frame, glibc's __memcmp_avx2_movbe, and the rest in eleven frames moving by four instructions or fewer. the program is deterministic and its input is fixed, so the same comparisons happen in every arm; what changed is where the compared bytes sit.
that does not explain the drift it was built for — two binaries with identical section sizes read an identical row here, and the drifting builds have identical sections too. what it does is strike a candidate off and hand back a frame to look at.
and the frame turned out to be worth looking at for a different reason. callgrind_annotate will not say who calls it: at full threshold the callers it lists come to three parts in a thousand of the frame. the raw profile will, because callgrind records every call site beside its caller with the call's cost on the next line, and summing those accounted for the frame exactly — 47,999,433 instructions over 2,617,709 calls, eighteen each, on identifiers of twenty-two bytes or fewer. a twentieth of the interpreted run was spent entering and leaving a function that compares bytes which fit in two registers. the next section is what came of that.
120nine instructions to find a pointer that never moved
a beat loop is one the compiler proved gives its arena back between laps, and k_beat_iter is what it calls to do that. the run program calls it 2,692,766 times for 61,672,983 instructions — 3.35% of everything it does, and 22.9 instructions a call against a fast path that is six stores and three tests.
disassembled, the path was 23 instructions. nine of them turned k_beat_depth into &k_beat_stack[depth - 1]: a load, a decrement, a range test, a zero-extend, a shift and two leas. that address cannot change for the life of the loop. the c compiler cannot know it, because the loop body calls other functions and any of them might push a beat.
two changes. the registry summary was a parallel array indexed by depth, so the rewind — holding the mark pointer already — had to turn it back into a depth to read the flag; it is a field of the mark now, read at a displacement, and the shelf flag beside it folds into the same branch. then the innermost mark is cached beside the depth, and the eight remaining address instructions become a load and a test.
k_beat_iter 61,672,983 -> 40,204,728 -21,468,255 -34.81%
k_beat_pop 14,517,216 -> 18,521,968 +4,004,752
k_beat_push 15,017,844 -> 15,518,439 +500,595
runbench 1,840,367,648 -> 1,823,406,517 -16,961,131 -0.9216%
fourteen benchmarks carry a row in that vein, and the spread across them is the finding. a row falls in proportion to how much its program beat-loops:
escapebench 84,780,592 -> 75,228,606 -9,551,986 -11.2667%
basket 33,678,746 -> 32,776,834 -901,912 -2.6780%
runbench 1,821,933,936 -> 1,804,998,570 -16,935,366 -0.9295%
livebench 2,825,430,323 -> 2,805,024,580 -20,405,743 -0.7222%
encodebench 3,497,149,260 -> 3,476,743,520 -20,405,740 -0.5835%
deepbench 347,289,236 -> 347,635,275 +346,039 +0.0996%
widebench 33,516,094 -> 33,644,020 +127,926 +0.3817%
escapebench is the extreme because escaping a string is a tight beat loop with almost nothing else in it, so the fifteen instructions are most of what a lap costs. the two that rise are paying the layout term without beat loops to spend it on — every binary here grew between 368 and 560 bytes, since the mark carries a field more and there is a new global beside it. the objective weighs the run program alone, so the score moves on one of these rows; the other thirteen are watched rather than scored, which is why the vein is diffed whole.
the three frames are a container's, because that is where a profile can name them. the row the project pins is ci's, on its own machine and its own baseline: 1,804,998,570 against 1,821,933,936, a fall of 16,935,366 — 0.9295%. two different boxes, two different baselines, and the two deltas are 25,765 apart, which is 0.0014% of the number. that is as close as this vein gets between machines, and it is why the attribution above can be trusted even though the absolute figures beside it cannot be compared with ci's.
the change costs something too, and one row carries all of it. src/runtime.c is include_str!'d into the compiler, so seven compile-side rows move by four thousandths of a per cent or less — they carry the file's bytes without compiling them. the eighth compiles them: the release-tier codegen row rises 11,227,515, 0.1645%, which is clang at -O3 -flto over a mark that grew a field and a global beside it. paid once per release build against 16.9 million given back on every run of the program, and the objective — which weighs run speed at 0.45 and the release build at 0.15 — takes the trade: 76.65 to 76.71.
the cache is maintained at seven sites and two of them are hot, which is the part worth reading. a push writes the mark it just built, so its range test is dead code and the site costs one store. a pop may land at depth zero or past the top of the stack, so its cmov stays and the site costs eight. the loop iterates five times for every pop it does, and that ratio is the whole trade: five instructions off the iteration buy eight onto the pop.
a cached pointer is an invariant, and this one fails quietly. a stale k_beat_top rewinds to an outer loop's mark, which frees memory the inner loop is still reading, and what a reader sees is a wrong answer somewhere else. so the counting build asks at every iteration whether the cached pointer is the one the depth names, and dies by name when it is not. the shipped binary compiles the question out.
the spec for that took two tries and both failures say something about the language. the first program accumulated strings, which compiles to a carry beat — that one computes its own mark, never reads the cache, and the spec passed with the maintenance torn out. the second carried a scalar and allocated nothing, and a loop with nothing to reclaim emits no beat at all. what it takes is both at once: laps that allocate, and a carried value small enough to cross one. each lap builds a padded string and keeps its length.
121a row that moves seven, and what two readings are worth
the interpreted row sits at about 1.138 billion instructions and it has moved by seven three times this month, on trees whose compiler source the interpreted corpus never touches. seven is not a number anybody would chase. what it is good for is a question this page has had to answer repeatedly: when a counted row disagrees with its golden, how much evidence does it take to say what happened?
the honest answer this time was two readings and no mechanism.
two branches read the row on the same day. one caches a pointer in the beat rewind, the other changes two divisions in a float renderer, and they ran on different runners — one of them an intel part whose feature block the repository does not have on file. both read 1,138,001,437 to the instruction, against a golden holding 1,138,001,430 from the tree before the merge. and the allocation counter beside it read 2,539,998 on both and on main.
that last row is the one carrying the weight. allocations are decisions the program makes; instructions carry the layout of the binary that made them. a change in the first would mean the interpreter was doing something different. there was no change in the first.
so the row is re-based on two agreeing readings. and pulling the third job log — the one whose tree set the golden — bounds what the cause can be:
tree silicon .text .rodata interp row
kanso#1522 PR AMD 0x19 / 0x1 2,840,050 803,856 1,138,001,430
kanso#1502 Intel 0x6 / 0x6a 2,840,050 805,776 1,138,001,437
kanso#1504 AMD 0x1a / 0x2 2,840,050 806,288 1,138,001,437
the silicon is not it. the two trees that agree to the instruction ran on an intel part and an amd zen 5 part, whose glibc feature blocks differ in fifty-odd rows — cache sizes, the rep-movsb stop threshold, the xsave sizes. both resolved to the same avx2 memcmp. the instrument that answers this was built two days ago for a standing question about a residual on this very row, and the resolver has been its leading suspect since; on this row the suspect has an alibi.
and .text is not it, which is the surprising one. all three trees emit byte-identical .text, 2,840,050 bytes, and the rows still differ. so a row that moves while the code section holds cannot be code layout in the ordinary sense.
.rodata is the only section that moves: 803,856, 805,776, 806,288. the smallest reads 430 and the two larger read 437, which looked like a correspondence over three points.
so it was tested, and it is not one. main, built twice on one box, the second build carrying 4,096 bytes of non-zero immutable data that nothing reaches:
.rodata=823,832 .text=2,860,898 row=1,173,233,661
.rodata=827,928 .text=2,860,898 row=1,173,233,661
the section grew by exactly 4,096, the code section held byte for byte, and the row did not move by one instruction. three points lining up was two coin flips.
that leaves all three candidates dead — not the silicon, not the size of the code, not the size of the constants. what is left is the one thing the table cannot separate: equal .text size is not equal .text content. three branches put different functions at different addresses with different alignment, and the totals happen to match. that is layout in the narrow sense of addresses rather than sizes, and nothing here isolates it.
the probe took four minutes and killed a claim that had already reached a log entry and this page. the general lesson is not about this row: when a section names the experiment that would settle something, that is the experiment to run before the careful sentence, not after.
the draft of this section said something stronger, and it was wrong. the first write-up had three readings of seven, across three trees at three different absolute values, and concluded a pattern. the third reading was not a reading. it came from the gate’s second count, which was subtracting the printed line from its first pass and not its second — so a binary that counted one number was reported as counting two, exactly 3,015 apart, with interp_printed=3015 sitting in the job log two screens above the error. that is the subject of section 88.
the shape is worth naming because it is not a slip about arithmetic. a claim was assembled out of three data points, one of which was an instrument reading itself, and the instrument had printed what it was doing. what makes three readings better than two is that they are independent; a number the tool generated is not a third opinion.
122the same trick, one layer out, declined
two changes had just taken a fifth of the interpreted run's memcmp traffic away by comparing short identifiers as machine words instead of calling out to the C library. the next two callers on that list are hash-map probes: an environment of globals and a table of resolved callees, both keyed by String and looked up with a &str, which hashes the bytes and then compares the key by calling memcmp. key them by the identifier type instead and the comparison should be the same two word loads, for free, since the callers already hold one.
it removes exactly the calls it was supposed to — 497,907 of them, which is the two lookup counts to within sixty-eight — and the row goes up by 4,315,544.
+10,456,358 <Q as hashbrown::Equivalent<K>>::equivalent 2,984 -> 10,459,342
-9,776,759 __memcmp_avx2_movbe
+3,485,161 Interp::call
+2,149,368 __memcpy_avx_unaligned_erms
a map keyed by String and probed by &str compares along a path that ends in a memcmp call. a map keyed by the identifier type and probed by the same compares through hashbrown's Equivalent, and that shim did not inline: 21 instructions a probe against memcmp's 19.6 plus its call. the word compare is in there and never got the chance to pay. the rest is two cold sites that now build an identifier where they used to pass a string slice.
reverted. the cost sat in the machinery the library wraps around a comparison once you change what is being compared, rather than in the comparison itself — which is what the previous two sections were about. the same code, one layer further out, stops being free.
123two terms under every row
the tables from section 119 sat in the job log where nothing could read them. a log api hands back the tail of a job and caps it at five thousand lines; the job that prints them runs twenty thousand, and the tables are most of that, so two of the twenty came back. each table is packed now — gzip and base64, two hundred columns wide — which turns seventeen hundred rows into a hundred and ninety lines and puts all twenty inside the tail together. the block carries the digest of what it unpacks to, so a reader can tell a whole one from a cut one.
the first change to arrive under that was a counter added to the c runtime: a declaration, an increment, and a line of output. the runtime is include_str!'d into the compiler, so the compiler's own bytes move with it. six rows count instructions on this tree, and all six moved. diffing each table against the same table from a build without the counter:
table total delta = constant memcmp memchr rest
compile -246 = -2,235 +2,023 -34 +0
entry +3,212 = -2,235 +5,481 -34 +0
library +4,699 = -2,235 +6,968 -34 +0
startup -1,488 = -2,235 +781 -34 +0
emit -628 = -2,235 +1,637 -30 +0
interp +361 = -2,235 +2,626 -34 +4
three terms account for every row exactly, and none of them is in the compiler. not one named function in the kanso binary changed its self cost by a single instruction, across all six tables.
the first term is twenty-nine frames that move by the same amount everywhere — the same −2,235 whether the workload is thirty-seven million instructions or nine hundred and forty-seven million. work that scales with the input cannot do that; a fixed cost paid once per process can. the frames name themselves: getdelim, sscanf, strtoul, and the _IO_* family, which is what reading a text file line by line is made of, with pthread_getattr_np moving alongside them. that function's job is parsing /proc/self/maps. so the constant is the c library reading this process's own memory map once at thread set-up, and what it reads differs because the binary's layout differs.
that is as far as a flat profile goes, and the limit is worth stating: a flat table has no caller edges, so this is the signature of that parse rather than proof of it. the raw profile's call records would settle it.
the second term is __memcmp_avx2_movbe, between seven hundred and eighty-one and six thousand nine hundred and sixty-eight depending on the row. it scales with the workload, which is what a code-layout effect does, and it is the same frame the previous section handed back.
what this is for: a row on this project has drifted by single digits between builds for weeks, on trees that cannot reach any decision the compiler makes. the answer here is two carriers, both in the c library, and one of them is already under a standing rule that external state gets normalised before it is measured. neither is a cause yet. the map parse is a property of one pair of binaries rather than a constant of the project, and −2,235 is three orders of magnitude larger than the three instructions that started the search. what changed is that the next occurrence can be taken apart the same way, in one command, instead of counted and left alone.
124a compare that does not read the address
section 123 left its second term open: __memcmp_avx2_movbe cost a different number of instructions on two trees whose compiler did the same work. the line-level profile says why. for a comparison shorter than thirty-two bytes, glibc first asks whether either operand lies within thirty-two bytes of a page end, by or-ing the two addresses and masking the low twelve bits, and takes a longer branch when one does. forty lines of runtime text, which the compiler embeds, moved strings in .rodata across page offsets, and 478 comparisons on the compile corpus changed branch with every call count the same.
the counted runs now preload their own memcmp, bcmp, memcpy and memmove, whose branches depend on the length and the bytes and never on where the bytes sit. on one container, with the compare replaced, the trees before and after that runtime change read the same compile row to the instruction. memcpy joined after a prediction failed on ci: with only the compare replaced, three rows still moved when the runtime change was merged underneath, and all of the move was in glibc's copy, which chooses part of its large-copy path by the distance between destination and source.
the replacements are checked for correctness first, because a wrong memmove would corrupt the compiler under measurement and the row would count whatever the corrupted compiler did. the gate that builds them runs all four against libc on 175,083 cases, including every length to 300 at every overlap distance from −140 to 140, and refuses on the first disagreement. it was watched failing with memmove forced to copy forward. two probes then compare the same bytes at two page offsets and copy the same bytes at two distances, and the gate refuses unless each pair costs the same.
they also have to cost about what libc costs, or the rows weigh the wrong work. the first compare stepped eight bytes at a time and cost three times libc's per call on the interpreter's name comparisons. the version that shipped takes libc's shape: an overlapping head and tail for short inputs and 32-byte vector steps for long ones. five of the six rows it touches read 0.24 to 0.61 per cent above libc and the sixth 0.60 below, and the welfare baselines for those rows were scaled by the same ratios, so the objective does not score the change of instrument as the compiler getting slower.
the runtime benchmarks were checked for the same exposure and do not have it. runbench built with its data section grown by one, three and four kilobytes counts 1,820,479,435 instructions each time. the runtime compares values in its own arena, and those addresses do not follow the binary's layout.
125the scan gives its memory back
sections 112 to 115 put three quarters of the run program’s arena peak in one phase, the split phase, and priced reclaiming it at +4.75 welfare. they went after it through the carry tier, the machinery that copies a loop’s live values out before each rewind, and found that tier closed to library loops by a path prefix. none of that turned out to be needed. the loop that holds the memory needs to carry nothing.
find_all tries a match at every start position of its subject. the loop that does it is three functions calling one another in tail position, and the cluster analysis refused it a rewind for two reasons, both in the library’s spelling. the walk’s answer, a match or none, went to the next function as an argument, so a record crossed from one position to the next. and the position itself came out of a hit’s field at every entry. a field read is not typed, so the position’s set was everything, byte builders included, and the analysis will not let a slot that might hold a byte builder cross a rewind. the walk’s answer is a local now, and the loop is entered once with at | 0, which is typed as a whole number. what crosses a position is the compiled pattern, the subject and an integer, and the cluster rewinds at each one.
before after
scan benchmark peak 161,480,704 1,048,576 154 blocks to 1
run program arena 38,604,496 9,244,368 36 blocks to 8
welfare 77.37 82.09
the welfare figure is ci’s, with the instruction rows it read. the rewind’s test for the seek cursor, below, costs the run program 0.41% of its instructions, and the peak pays for that many times over.
the first build of it made prose_check run past five minutes where it had taken twenty-eight seconds. reading a character by position in text that is not all ascii resumes a walk from one remembered place, and the runtime forgot that place on every rewind, because the arena hands back the addresses above the mark and the next string there is a different one. the subject a scan walks arrived from outside the loop and sits below the mark, where nothing is handed back, so forgetting it was never needed. with a rewind at every position it was ruinous: every position walked the page from the front. the rewind now forgets the place only when the string lies in the range it hands back, and prose_check takes fifteen seconds. a new counter, seek_resumes, counts the walks that resumed, since forgetting the place moves no output and no allocation and nothing else could show it.
126a phase gives its garbage back
a construction cohort is a bracket the compiler puts around a call whose arguments cannot be grown by the callee: mark the arena, make the call, and rewind to the mark afterwards, copying the answer out first when it lives in the heap. the license asked whether the callee's module name extended the caller's by one segment. that was how nesting was spelled before a module's identity became its canonical path, and after that change the test matched only calls from the root module. an archive entry from 2026-08-18 recorded that it had stopped matching and said what it should ask instead. nothing rebuilt it.
the run program pays for it in its phases. runbench/tally calls index/total, escape/total and split/total, and each one builds its own strings and answers a number. none of those calls was bracketed, so each phase's strings stayed live under everything that ran after it. the license now asks whether the callee is declared in a different file from the caller. a direct call by name can only reach another file through an import, so that is the relation the archive entry asked for.
admitting every such call cost 13.1 per cent of the run program's instructions, all of it in json. its recursive descent calls its own modules millions of times, and every call paid a push and a pop. so a caller that a cycle can reach gets no bracket. the compiler finds every cycle in what the function bodies mention, then everything those cycles reach, and a caller outside that set runs a number of times fixed by straight-line code, which is where a phase starts and ends. excluding only the members of cycles was not enough: json/number_done is in none, scan calls it once per number, and jsonbench rose 5.3 per cent until reachability was added.
with the regexp scan and the digest changes on main, the run program's peak is the index phase, and it falls from 6,098,640 bytes to 5,050,064, with four cohort frees where there was one, for 0.11 per cent more instructions. oneshot pays 11.3 per cent for one pop around json/decode that sizes a survivor nearly as large as what the call grew and then keeps the region. oneshot had that pop when the license was first generalized, and its peak does not move.
127the release build stops recompiling the runtime
a release build hands the linker the program as bitcode, and the linker optimizes it and generates its machine code in one pass, which is how a helper in the runtime gets inlined into the loop that calls it. the runtime used to go through that pass too, on every build, and it is three quarters of the code the pass generates: on the codegen corpus the linked runtime is 59,082 bytes and the program 19,135. the runtime is now compiled once, as machine code, and cached.
three runtime helpers are the exception. the two byte scans behind the json scanner and the arena rewind every beat loop runs were being inlined into the program by that pass, and without them the run program rose 3.38 per cent. they are compiled into a second, small bitcode object, built from the runtime source's own text so the specs that read those functions keep checking what ships, and the pass inlines them as before. the release build's codegen row fell 56 per cent and the run program moved 0.31 per cent on ci. compiled on its own the runtime had also lost the program's inlining threshold, which the link used to apply to both halves; giving it back cost the release row 0.08 per cent and gave the run program 0.29.
the program was also optimized twice: once as clang turned its ir into bitcode, and again at the link. on linux the first pass now runs at -O1 and the link stays at -O3. that took the release row down another 39.6 per cent and the run program up 3.4 per cent on ci; skipping the first pass entirely cost the run program 13.7 per cent on this container, because the link's pipeline expects its input already simplified. the objective prices a release build at 0.15 of production against run speed's 0.45, and at these ratios the build's saving is the larger. apple's linker has no spelling for the link's level, so other hosts keep -O3 for both steps.
128a short float is found by dividing back
the encoder prints every float through ryu, which finds the shortest decimal that reads back as the same double using 125-bit multiplies and a loop that strips digits it does not need. that took 430 instructions a float on the run program. every float in the encode corpus has seven significant digits or fewer, and for a decimal that short there is a cheaper way to find it.
the renderer now tries the decimal places in order. at place p it rounds the float times 10p to an integer m, and accepts m × 10-p when m divided by 10p gives back the float exactly. that division is one correctly rounded operation on two exact doubles, so a match proves the decimal reads back. a match is also never missed: while the gap to the next double, times 10p, is at most a quarter, only one decimal with p places can read back, and rounding the product finds it. so the first place that passes gives ryu's own answer. a float outside 2-20 to 250, or one with more digits than the bound allows, goes to ryu as before.
with m and p in hand the text needs no digit buffer: it is m divided by 10p, a point, and m's last p digits. a differential fuzzer compared the new renderer with the old one on 1,495,188,243 doubles weighted toward short decimals and their neighbours, with no difference, and the float spec gained a check it had lacked. it tested that the text reads back and that no shorter decimal does, and a renderer choosing the wrong neighbour of the right length passes both. it now also checks that the nearest decimal of that length is chosen whenever that one reads back. together the two changes take a float from 430 instructions to 222. the first moved the run program 1.54 per cent on ci.
one lead from the same profile was declined. encode_onto saves six registers on each of its 2.4 million calls a run, because the link inlines the list, map and string encoders into it and every arm then pays their frame. keeping those three out of line saves 1.4 per cent of the run, but no rule the emitter could apply without a profile finds them: every rule measured either cost more elsewhere than it saved here or saved nothing.
129a dev build calls its helpers
every module the compiler writes starts with thirty-three small functions: tag tests, the fast arms of append and index, the closure-call twins' guards. they are marked alwaysinline, which is right for a release build, where copying each one into its call site is how the loop that calls it gets fast. at -O0 the always-inliner still honours the mark, and the instruction selector then walks every copy. a dev build now defines the helpers without it and calls them. on ci the dev tier's codegen row fell from 473,933,874 instructions to 431,017,582, 9.06 per cent. the dev text is made when the compiler is built, because checking each line for the mark as a module was written cost the start-up row 10,508 instructions.
what still costs the dev tier most is the two-word value every function passes and returns. the fast instruction selector at -O0 handles neither an aggregate argument nor an aggregate return, and a return it cannot lower takes its whole block to the slow selector. on the codegen corpus 164,838,950 of clang -cc1's 338,572,876 instructions were spent there.
130a module carries the helpers it calls
nothing at -O0 removes a function nobody calls, so each dev module was still compiling all thirty-three. on the codegen corpus twenty-four of them were called by nothing. the compiler now reads, when it is built, which helpers each helper calls, and a module defines only the ones its program reaches through that graph. on ci that took the dev row from 431,017,582 to 396,949,978, and the start-up row, which emits a module for a one-line program, from 866,508 to 775,903. a release module is held to the same rule. its optimiser would have dropped the unused helpers after parsing them, and leaving them out of the text took the release compile from 555,145,625 instructions to 545,749,531 with the link's count unchanged.
moving each return into a block of its own, so that only the return took the slow path, made the compile slower: 287,802,692 instructions became 297,839,204, because each block the slow selector takes pays a fixed cost to set up.
131a record slot with a wildcard arm stays boxed
a group whose arms all take a record at some position passes that position as the record's two words rather than a pointer to it, and the caller turns what it holds into those words. that is sound when nothing but a record of that type or a failure can arrive. a group with fn total (point x y) and fn total _ was given the convention anyway, because the wildcard arm names no record to disagree with. total 5 then ran the native build out of stack, reading fields off an int, and total (pair 1 2) answered 10 where the oracle answered 7, reading the pair's fields as a point's. a slot now keeps the convention only when every shape the inference sees reaching it is a record, a failure or a thunk, and no arm there takes any value at all. the json decoder's slots meet both conditions, and no benchmark moved.
132what the fast selector can read
at -O0 clang picks machine instructions with a fast selector that handles a narrow set of instructions and hands everything else to the slow one. an instruction it cannot handle sends the rest of its block to the slow selector, and on the codegen corpus that was most of the dev build's time. four changes each removed one kind of instruction it could not handle.
the three predicates the dev build calls most, the failure test, the truth test and the record test, took the two-word value as an argument, and the fast selector cannot pass one. a dev module now pulls out the value's tag and payload and calls forms that take them as two plain integers. a literal's tag and payload used to be read off the literal with an instruction the fast selector refuses, and the emitter now writes the numbers themselves. the four functions the runtime uses to print a record and read a field by name were switches over the type id. they are loads from constant arrays now, written outside the text the emitter's own passes read, which also took the start-up row down 6 per cent and the release build's codegen down 1.1. a dispatcher on int literals or on a tag compiles to a switch, and a dev module now compares the cases in turn instead. release keeps the switch for its jump table.
on this container the four together took the dev codegen row from 396,836,531 instructions to 359,341,981. the returns remain: a function that returns the two-word value still sends its return block to the slow selector, and there are 164 of those on the corpus.
133what holds the run's memory peak
the run program's arena peaked at 5,050,064 bytes: two 1 MiB blocks, the index shape's 1,572,880-byte subject string, and the 1,380,032-byte slice cut from it. an earlier reading of this section put it at four 1 MiB blocks and one of 855,760 bytes, which adds up but cannot happen, because a block is 1 MiB or exactly one allocation larger than that. section 136 says what happened to the slice and to the block after it. zeroing the program one phase at a time names what holds each. the top-level decode keeps two blocks for the whole run, and decoding "[1]" there instead leaves one. the decode loop adds two, because each decode starts partway into a block. the pending-cell shape adds one, the index shape adds the oversize block, and the other phases add nothing. smaller blocks barely change it: 512 KiB reads 5,050,064, 256 KiB 4,787,920 and 128 KiB 4,845,568. giving the top-level decode a cohort, as an experiment, kept the region at the pop with the peak unchanged, because the run goes on to use most of that tree.
three ideas from the same profile were measured and declined. a shared cache for four- to seven-byte tokens halved the bytes strings allocate and made the run 0.49 per cent dearer with the peak unmoved, because a token's arena path is one bump and the lookup costs more than the bump. keeping escape_rest out of line to drop escape_onto's frame cost 1.9 million instructions. moving to_float's strtod fallback out of line saved 0.006 per cent.
four more were measured later the same day and declined. reading a number's digits eight at a time made the run 0.23 per cent dearer, because the run's numbers have three or four digits either side of the point and a byte loop reads those for less than the eight-digit conversion costs. giving a list four slots on its first push instead of eight cost 0.78 per cent and moved no block of the peak. arena blocks of 256 KiB instead of 1 MiB took a quarter of a megabyte off this program's peak for 0.70 per cent of its instructions, which is the objective's weights nearly cancelling on a saving that depends on where one program's live set falls against a block boundary.
the fourth marked bytes known to be valid UTF-8, a string's view and the slices of it that cut at character boundaries, so that utf8 of them would skip its check. the check is worth 34 million instructions on the run program at most, and nearly all of that is spent on builders. carried through every append, the mark cost more per byte than the check it saved and made the run 1.8 per cent dearer. carried only on views and slices, it made the run 0.78 per cent dearer, because the slices it could vouch for are the decoder's short ascii tokens, and checking one of those is cheaper than proving where it was cut.
134a view read where it was made
text/bytes of a string borrows the string's bytes and writes a three-word header for the view: its length, a pointer to the data and a capacity. json's escaper makes one for every string it writes, 942,750 times on the run program, and does nothing with it but read: a length, a scan for the next byte that needs escaping, a byte by index, slices to copy out. the header went into the arena, and every one of those reads loaded it back.
the emitter now writes such a header into the frame of the function that made the view, where LLVM takes it apart into registers. it does this only when every later use of the view reads it while that frame is still there. the function must sit on no cycle of the call graph, and every mention of the view must be the first argument of length, find2, find2_below or slice, the base of an index, or an argument to a parameter that is itself only read that way. returning the view, putting it in a list, capturing it in a closure or binding it to another name keeps it in the arena. a function holding a framed view makes its tail calls as ordinary calls, because a tail call gives up the frame before the callee reads the header. the argument must also be proven a string: when the header might come from either the frame or the arena, LLVM keeps it in memory and most of the saving goes. on the run program this saves 28,956,330 instructions, 1.64 per cent, and 942,750 allocations.
135a map whose keys arrived in order
a map keeps its pairs in the order they were put and builds a sorted view the first time something reads it: a copy of the pairs, sorted, with repeated keys collapsed to the last one put. the view lives outside the arena and stays for as long as the map does, and on the run program the views of the document's 2,761 objects were the whole of held_peak_bytes, 728,040 bytes. every one of those objects had its keys put in ascending order with none repeated, so every view was a copy of pairs that were already sorted.
building a view now checks the order first, and when the keys are already ascending it points the view at the pairs. the one write that changes a view's order is a put into a map that is being read as it is written, and that put keeps sharing only while the new key sorts last; otherwise it copies the view out, at the size the view would have grown to by then. on the run program no view is allocated at all, held_peak_bytes falls to 416,312 and the run takes 0.27 per cent fewer instructions.
136a long slice shares its text, and a big string leaves its neighbour open
text/slice of a string used to copy the characters it kept. a slice of sixty-four bytes or more is now a sixteen-byte header pointing into its parent's bytes. a builder's slices are still copies, since a builder's storage moves as it grows; every other string's bytes live as long as the string does, and evacuation copies a view's own bytes and nothing around them.
the byte after a view is its parent's, so a view has no terminator. the runtime needs one only where it hands a string to the C library: opening a file, reading the environment, starting a process, and printing its own error sentences. those places take a terminated copy when the string is a view.
on the run program the index shape cuts 1,380,000 bytes out of a 1,572,864-byte string that nothing reads again, and both were live at the peak. sharing the slice takes the peak from 5,050,064 bytes to 4,718,608, and the run allocates 1,396,400 fewer bytes.
what was left at the top was a fresh 1 MiB block pushed straight after the subject, while the block before the subject still had room. an allocation larger than a block gets a block of exactly its size, and that block used to become the place small allocations were bumped from, with nothing left in it. now the older block's unused tail is split off and kept as the bump region above the oversize block, and a rewind hands the tail back to the block it came from. the peak falls to 4,194,304, which is what the run reads with the index shape cut to one character, so the index shape no longer sets it.
137an arm no value reaches
std/list's next has an arm for every lazy adapter the library declares, eleven of them. A program that maps once reaches next, so it emitted all eleven arms and everything each one calls. The codegen corpus is sixty-two lines, and 1,560 of its 4,628 lines of IR were adapters it never builds and what they call.
A value of a declared record type only exists if some expression names the type: a construction, a partial of it, the constructor passed on as a function, or an upcast to it. The runtime builds one record type of its own, the entry that entries returns. So an arm whose pattern names a type nothing builds can never be chosen, and the compiler now drops it before emitting, along with whatever only that arm calls. Deciding what counts as built needs reachability too, since std/list declares a builder for every adapter. A builder only counts if the program can call it, and calling it is itself decided by which arms stay. The two are worked out together until neither changes.
On the codegen corpus the IR falls from 4,628 lines to 3,068. The dev build's compiler work falls by 30 per cent and the release build's by 47 per cent. The program prints the same thing, the interpreter never sees the change, and the differential corpora are the check that the two engines still agree.
138what an empty literal, a map walk and a loop entry cost
The decoder writes [] for every array it opens and {} for every object, 272,349 and 273,339 times in one run of the run program. Each went through the general literal builder with a count of zero, which asked for a size class at run time, took the header and the buffer in two separate allocations, and ran a copy loop over nothing, about fifty instructions a literal. An empty literal now compiles to a call that takes both in one allocation from a size class fixed at compile time. The run program falls by just under one per cent and the decode benchmark by two.
entries already built its records in one block. The list it returns needed an item buffer and a header, two more allocations, and they now come out of the same block after the records. A beat loop's entry mark is taken inline, the way each iteration's rewind already was, rather than through a call.
Two costs came out of the compiler itself. A function with several checked integer operations wrote one overflow trap for each, three identical lines apiece, and now writes one that they all branch to. The parser counted the blank lines between each pair of neighbouring lines by scanning every blank line in the file, which is quadratic in the file's length. The list is in order, so two binary searches give the count, and checking a program with ten library imports got close to four per cent cheaper.
139a large object in a release build
A release build keeps the tailcc convention only on functions whose arguments fit in the registers, because arm64 miscompiles it past its eight argument registers. That rule was applied on every architecture. The JSON decoder's obj_key_end, which skips the whitespace before a key's colon, takes nine words of arguments, so every key of an object entered it through an ordinary call, and the calls did not return until the object closed. The stack held a frame for each key, and a release build decoding an object of 300,000 keys died on a segmentation fault.
x86-64 lowers a tailcc call whose arguments spill onto the stack correctly, so the limit there is a matter of cost rather than correctness. At nine words the decoder keeps its tail calls and runs in one frame however many keys the object has. At ten and above std/regexp's wider functions keep theirs as well, and copying more than three stack arguments on each tail call costs more than the frame it saves, so nine is where x86-64 stops. arm64 keeps eight, and with it the frame per key, until the decoder's loop takes fewer arguments.
140one signature for a tail cycle
The JSON decoder reads a document through a cycle of functions that end in tail calls: parse_value jumps to string_scan, which jumps to obj_key_start, and so on round the loop. Under tailcc each of these saved the callee-saved registers it used when it was entered and restored them before every jump out. A function with several exits cannot move those saves off its entry, and on runbench the pushes and pops in four of the decoder's functions came to about 55 million instructions.
LLVM has a calling convention with no callee-saved registers, preserve_none, and the closures already use it where clang 19 is installed. It could not be used here because LLVM lets a guaranteed tail call change arity or parameter types only under tailcc, and the decoder's functions take from two to nine words of arguments. A release build on x86-64 now gives each such cycle one signature. Every parameter is split into 64-bit words, every function in the cycle takes as many words as the widest one, and the narrower ones are padded. The widest decoder function takes nine words, and x86-64 passes twelve in registers under this convention. With the rewrite in place, a function keeps its tail calls up to twelve words rather than nine, because none of its arguments go on the stack.
The program's other functions take the same convention, so a function saves nothing and its caller saves only what it still needs after the call. The runtime calls a few program functions by name, such as the one that evaluates a deferred binding, and those keep the C convention; the build reads the list from the runtime's own extern declarations.
Four functions in the runtime that the program calls often take it too: the one that appends a rendered number, the one that lists a map's entries, the one that turns bytes into a string, and the one that reads a float out of a slice. Each was measured on its own, and two others that were tried, indexing and taking a string slice, cost more than they saved, because their callers keep more live across the call than the callee used to save. They stay as they were.
Together these take runbench down 5.42%, jsonbench 7.50% and livebench 6.29%. pendbench and digestbench rise by less than half a per cent: their calls into the cycle outnumber the jumps inside it, and a caller now saves what it keeps live across such a call where the callee used to. arm64 is left as it was until the convention has been measured there.
141the interpreter keeps what it was handed
When the interpreter calls a function with several arms, it tries each arm against the arguments and keeps the best match. Every arm that named a parameter bound it to a copy of the argument, and a copy of an integer or a string allocates, so the arms that lost paid for copies nobody used. A named parameter now holds nothing while the arms are compared, and the winning arm's arguments are moved into place once it is known. The argument list was already discarded before the body ran. The interpreted corpus runs 3.2% fewer instructions and makes 5.5% fewer allocations, and prints the same thing.
A smaller change on the native side: appending to a byte buffer that has no storage yet always grows it, and the check for that case now comes before the function that fits bytes into existing room. That function used to save five registers on every call, including the calls that went straight on to grow.
142a list of five
One decode of the benchmark document allocates about nine bytes for every byte of JSON it reads. Sorting the allocations by size found the largest share in list buffers of sixteen slots. A list starts with four, and growing one took the smallest power of two that held the new element and then doubled it, so the fifth push went to sixteen slots. Most lists in a JSON document are short, and this one's average three and a half elements.
Growth now doubles again only past sixteen, so the steps are 4, 8, 16, 64 and 256. The run program's peak memory drops from four arena blocks to three and a half, for a fifth of a per cent more instructions: a list of nine to sixteen takes one more step to reach its size. Stopping the extra doubling at eight gave the same peak, but it moved every longer list onto the sizes 32, 128 and 512, which leave twice as much unused past a thousand elements.
143a buffer that has already left
A list that a loop builds from outside the loop keeps its storage out of the arena, because the rewind at each iteration would otherwise reclaim what the loop is building. That storage comes from malloc. When such a list outgrew its buffer, the runtime allocated a bigger one, copied the elements across, freed the old one and registered the list's buffer again with the beat that will free it. The escape shape of the run program does this five times for each of its 3,872 lists, and each grow cost about 658 instructions.
A buffer that has already left the arena now grows with realloc. The list keeps the field the buffer hangs from, so the registration made on its first grow still names it and is not repeated. The run program runs 0.43% fewer instructions on this machine. The old and new buffers are never both held, so the permanent peak holds one buffer where it held two: 20,512 bytes becomes 16,400 on the run program, and the basket benchmark falls from 5,308,464 to 4,259,872. A map's pairs have the same grow, but no program the beat analysis brackets threads a map through a beat, so that path was left as it was.
144three letters at once
The JSON decoder recognises true by checking the three bytes after the t: cs[p + 1] == 114 and cs[p + 2] == 117 and cs[p + 3] == 101. Compiled read by read, each byte paid for an overflow check on p + k, a test against each end of the bytes, a merge of the byte with the none an out-of-range read gives, and the compare, about fourteen instructions a byte. The run program matches 612,500 of these literals.
When every conjunct of an and chain reads the same bytes at the same position plus a constant, and compares with a byte literal, the emitter now tests once that the whole window of reads lies inside the bytes and then compares plain loads. Inside the window no read can miss and no sum can overflow. Outside it the chain is false, since some read is none, but the compiler keeps the general path for that case, because a position near the top of the integer range would make one of the sums overflow, and that traps. The decode benchmark runs 3.46% fewer instructions and the run program 1.55% fewer.
145a loop that allocated nothing
Every iteration of a beat loop ends by rewinding the arena to the mark it took at the top. The fast path of that rewind asked three questions: whether anything had been shelved or registered, whether the arena's current block was still the mark's block, and whether the arena pointer still stood at the mark. The run program takes 2,708,992 beat iterations, and most of them allocate nothing, so the three answers came out the same every time.
The rewind now asks about the pointer second. Arena blocks never overlap, so a pointer equal to the mark's is already inside the mark's block, and the block no longer has to be loaded to say so. A loop that did allocate takes the same path as before. The run program runs 0.495% fewer instructions and the escape benchmark 4.93% fewer.
146a backslash and a letter
The JSON encoder writes every escape as two appends to the same builder: a backslash, then the letter that names the character. Each append of a small integer to a builder the function owns is a call that checks the builder is still owned, that its bytes end at the front of the buffer, and that one more byte fits. The second call then asks the same three questions about the builder the first call just returned.
After a function's body is emitted, the compiler looks for two such calls where the second appends to the first one's result and nothing else reads that result. It replaces them with one call that checks for room for two bytes and stores both in order. If that check fails, it makes the two single appends as written, so a builder that has to grow still grows the same way. The live-encode benchmark runs 0.712% fewer instructions and the run program 0.235% fewer.
147a list known to be a list
The compiler infers, for many parameters, the set of kinds of value that can reach them. The JSON encoder's encode_pairs acc es i is only ever handed the list entries builds, and then itself, so es is known to be a list. Until now the emitter used that kind of fact for bytes and nowhere else. length es still asked whether es was a list or bytes before loading the length, and es[i] asked the index's kind, the container's kind, and then list-or-bytes a second time before loading the element.
length of a value known to be a list or bytes now loads the first word of its header, which is the length for both. A plain index of a known list checks that the position is in range and loads the element, and answers none outside the range as before. The strict form, es[i]!, answers a box around the element, so it keeps the general path. The run program runs 1.1% fewer instructions, the live encoder 2.8% fewer, and every program whose count moved also got smaller.
148the letters a match must hold
Searching text built from the letters a to m for [a-z]+zzq finds nothing, and finding that out used to be quadratic. At every start position the run of letters was taken to the end of the text and given back one letter at a time, and each step asked whether zzq came next. The run program spent about 58 million instructions on it.
A compiled pattern now carries the longest run of plain characters that sits side by side in its top-level sequence, which every match must contain. For [a-z]+zzq that is zzq. A class, a repetition, an alternation or an anchor ends a run. The first scan of a subject looks for that literal with a byte search, and when the subject does not hold it, the answer is no match and no position is tried. Later scans start after a match, and a match holds the literal, so they skip the question. The run program runs 3.92% fewer instructions.
Two small memory tests used that same pattern to watch the scan itself: whether a scan that keeps nothing holds a flat peak, and whether it keeps its place in the text. With the literal asked for first they stopped scanning, so they now search for [a-z]+[x-z], which ends in a class the text never supplies.
149what a call used to ask the long way
The interpreter asks three things on every call: which function a function value names, which frame a declaration runs in, and whether a push at this spot may write into the list it was given. Each answer lived in a hash table. The first two were keyed by an address, and the third by the file's path together with a line and column, so a container call hashed the whole path. Together they cost about 39 million of the 711 million instructions the interpreted corpus takes.
Each frame now keeps its own file's in-place sites, gathered once. The other two tables each have a small array in front of them, indexed by a few bits of the address and holding the last answer for that slot; a hit is one compare. The interpreter also asked, on nearly every value, whether it was a lazy cell waiting to be forced, and asking was a function call. That test is now inline, and only a cell makes the call. The interpreted corpus runs 4.3% fewer instructions, and its peak memory rises by 17,824 bytes for the arrays and the per-frame sets.
150reading a string one character at a time
A string of mixed-width characters cannot be indexed by arithmetic: the tenth character might start at byte ten or at byte forty. The runtime keeps a cursor, the last character an index found and the byte it starts at, so a loop that reads a string in order resumes from where it was. The run program's index phase reads a 690,000-character string that way. Each index still went through the general search first, which asked whether the string was all ascii and whether the cursor's own character was the one wanted. An index of the next character now steps from the cursor directly, and it costs 59 instructions where it cost 71.
The same phase builds its string by joining it to itself until it is long enough, asking its length after each join. A string remembers its character count once something has counted it, but a join produced a string with no count, so each new string was scanned from the front. When every piece of a join and its separator already carry a count, the joined string now carries the sum. The pieces are asked in a pass of their own that stops at the first piece nobody has counted, because a join of four thousand rendered numbers asking each one cost more than the count saved. Together the two changes took the run program down by 10.7 million instructions, 0.76%.
151four letters in one store
The json encoder writes every true, false and null by appending a literal to the text it is building. The append already skipped the runtime, but it still loaded the literal from memory, read its length and copied it. A literal of up to eight bytes is now written into the call as one number the compiler worked out, and when the text has eight bytes of room at its end the number is stored in a single instruction and the length moves by the literal's own length. The bytes past the literal land in room the text already owns, beyond its length, where nothing reads them.
When the room is short, a separate function builds the literal and appends it in the place it was going. The first version handed that case to the append that makes a new copy of the text's header, and the encoder, which keeps its text by identity through its loops, went on writing into a header the loop had already given back. It crashed in the counting build, and a memory test now runs that shape. On the run program the change saves 6.6 million instructions, 0.47%, and the encoding benchmark 4.4%.
152a second play of the same file
kanso play builds a native binary and keeps it, so running an unchanged file again costs no clang. It found that binary by hashing the program's IR, which meant compiling the file and writing the IR on every run. For a file holding print "x", the compiler spent 615,803 instructions before handing over to the binary: 172,583 lexing, parsing and checking, and 404,031 emitting text whose only use was to name a file that already existed.
A play file may import only the standard library, and the compiler carries its own copy of every module but one. When a file loads nothing else, its text and the compiler decide its IR, so the binary gets a second name under a hash of what decides it: the file's name and text, the compiler's path, length and modification time, the runtime it links, the closure convention the installed clang takes, whether counters are compiled in, and any KANSO_ setting. A later play reads the file, forms that name, and runs the binary without compiling anything. The name exists only if a play compiled the same inputs cleanly, so there is nothing a second check could find. A file that reads std/expect, which comes from disk, keeps using the IR's key, since that module can change while the file stays the same. When a warm play's program dies by a signal, the file is compiled then, so the message can name what went wrong.
The start-up row counts exactly this case, a warm second play, and it fell from 615,803 instructions to 51,616 on the container that built it. The front end and the emitter keep rows of their own, which compile the benchmark corpora directly; neither moved.
153small values in the interpreter
The interpreter held every map as a B-tree and every int as an arbitrary-precision integer. Both are right for the large case, and both charge the small case for it. A B-tree leaf has room for eleven entries whatever it holds, and most maps a program builds are JSON records with a few keys. The interpreted corpus decodes 220 records of four keys, and their leaves made up 158,400 bytes of the interpreter's peak memory. A big integer keeps its digits in a heap vector even when the number is 1, and copying the value copies the vector.
A map of up to eight entries is now a vector kept in key order, and a ninth key turns it into a B-tree, so a large map built one key at a time still inserts in logarithmic time. An int is a machine word while it fits one. Addition, subtraction, multiplication, division and remainder try the word first and move to the big integer when the word overflows. A result that fits a word always goes back into one, so each number has exactly one form. Equality, ordering and map keys rely on that, and a spec checks that max + 1 - 1 equals max.
Together the two changes cut the interpreted run's instructions by 9.8% and its peak memory by 10.1% on the container that built them. Putting the big integer behind a shared pointer and changing nothing else was tried first. It saved 4.5% of the instructions and allocated 27% more often, because every new int then took two allocations. It was not kept.
154linking a dev build
A dev build is compiled at -O0 and linked, and the link was a large part of what it cost. lld is fast at its own work, but it is a shared-library build of LLVM, and the dynamic loader spent more resolving its symbols than lld spent linking the program. It also computed a SHA-1 build ID over the output, which nothing in a dev loop reads. A dev link now goes to gold when gold is installed, with no build ID. A probe asks once per pair of tools whether clang can drive gold, and a machine without it keeps lld. Release builds still link with lld, which runs their LTO. The dev codegen row fell 11.7% in CI, from 142,163,302 instructions to 125,549,128.
The same change turned up a start-up row that read 52 instructions apart on two CI runs of one binary. The compiler names its cache files after a hash of the tools it found, and clang's modification time is part of that hash, so the name changes from one runner image to the next. The names were written with {:016x}, which writes a number's own digits and then one padding character at a time. A hash whose top hex digit was zero cost more to name. Every such name is now written sixteen digits at a time from a fixed loop, and the process id in temporary paths the same way. With a copy of clang whose modification time was set by hand, five times gave three different readings before and one after.
155room from the start
Three changes to the runtime take work out of the run program without changing what it computes. Rendering the shortest decimal for a float searches for the fewest digits that read back as the same number. The search began at zero digits every time, and the floats a program prints in a row tend to need the same number, so it now starts where the previous one ended and walks up or down from there. That took runbench's instructions down 0.97%.
Three containers started too small for what they were about to hold. An empty map, {}, had room for four pairs, and the decoder opens every JSON object with one. The benchmark's objects hold one to five keys, so each five-key object grew its map on the fifth key: 56,430 grows a run, each a new buffer and a copy. The seed is five pairs now. An empty list, [], had room for four elements, and the benchmark's arrays hold one to six, so every array of five or six grew on its fifth push; it opens with six. A list that outlives the loop building it moves out of the arena and grows by realloc; it used to leave at eight slots and pass through sixteen, sixty-four and 256 on its way to a thousand. It leaves at 256. On the container that built them the three cut runbench by 0.84%, 1.14% and 0.52%, with the arena, held and permanent peaks unchanged. Every empty map and every empty list is 32 bytes larger, and a short list that leaves the arena now takes 4,112 bytes where it took 272.
The decoder unescapes a string by appending its pieces to text/bytes "", a view of the empty string that owns no storage, so each of runbench's 175,527 escaped strings began by growing its builder from nothing. The emitter now writes that literal as an empty builder with 64 bytes of room, which is the capacity the first grow used to choose, so every later grow is the one it was. Runbench fell another 1.69%. A first version tested for an empty string wherever bytes ran, and it cost encodebench two instructions on every string it escaped, so the test moved to the one place the answer is known in advance, the literal.
One fix came out of testing these. The browser engine runs compiled programs inside a WebAssembly interpreter, and a program that recursed too deeply ran out of that interpreter's stack in the middle of a runtime call. A trap in WebAssembly unwinds nothing, so a Rust RefCell borrowed at that moment stayed borrowed, and the next program's setup panicked on it. The runtime's cells are now rebuilt at each entry point when a dead program still holds them, and a spec runs the deep recursion followed by an ordinary program on one engine.
156one spelling for a step that comes after
The wall is gone. a >> b and a .> (_ -> b) were the same program, and Clay ruled on 2026-09-26 that the language keeps the second. The lexer now refuses >> where it is written and names the spelling to use. Every sequence in the tree was rewritten, in the library, the scripts, the tests, the examples and the book, and the sections of this page that describe the wall describe the language before that date.
The operator took a good deal of machinery with it. The parser's handling of wall lines, the checker's refusal of a wall operand that could never be an effect, the stack-exhaustion hint for a function that recursed on a wall's right side, and a description node in each of the three engines are all removed. The failure rule the wall carried belongs to the bind, which already short-circuits on the first err. A group of bare lines still accumulates its members' failures, and a step after a group is written by naming the group and binding after the name.
One feature had to be rebuilt rather than removed. --plan listed every step of a wall chain because a wall held its right side as a description. A bind holds a function, and the plan has no value to hand it. A callback written _ -> step ignores its argument, though, so the plan now calls it and renders the step, and examples/effects.kso plans exactly as it did. A callback that reads its argument is still shown as a continuation.
157a mark for each map
A loop whose iterations allocate takes a beat: a mark in the arena when the loop is entered, a test at the end of each iteration that rewinds to the mark if anything was allocated, and a pop when the loop returns. The JSON encoder has two loops, one over a list's items and one over a map's pairs, and both were beats. The only thing either allocated was the list of entries a nested map builds as the encoder descends into it. So every list and every map in the document paid for a mark, a pop and a test per element, to reclaim storage that only the nested maps allocate.
The mark now goes on the descent. When a recursive function calls back into its own cluster with an argument that allocates, and the function answers only numbers or bytes, the compiler marks the arena before evaluating the arguments and pops after the call. The pop gives the storage back when the result cannot point into it, which for the encoder is always: the result is the byte builder, which existed before the mark. With the entries accounted for there, neither loop allocates anything, and neither takes a beat. Runbench fell 2.40% and livebench, which encodes the same document every round, 6.75%. Peak memory did not move on either. A region is kept only where it lets its loops drop their beats: a copy of the encoder whose loops allocate for other reasons got one at first, paid a mark per map on top of its beats, and ran 6% slower.
The first version put a mark on a loop's own back edge in a test program, because that edge also passed the three tests. A mark after a call means the call can no longer be a tail call, so a loop of a hundred thousand trips kept a hundred thousand frames and ran out of stack. An edge that closes a cycle of tail calls never takes a mark now.
158guards, pushes and the beat's null test
Three small tests came out of the run program's hottest paths. A guard, return x if c, compiled its condition as a value: a comparison or an or became a tagged boolean, and a call into the runtime asked whether it was true. A tail if already branched on its comparisons directly, and a guard is a tail if whose else is the rest of the body, so it now takes the same path. The JSON encoder opens every list and map with such a guard. Runbench fell from 1,192,666,206 instructions to 1,190,403,668, and encodebench 3.4%.
A list's buffer kept its capacity with the storage regime in the sign, negative for a buffer that came from malloc. Every inlined push asked for the magnitude with a negate and a conditional move before it could compare. The capacity is now stored doubled, plus one for malloc, and room for n slots is 2n against that word in either regime: one shifted add. The beat's cached mark was a null pointer past the stack's sixty-four marks, so every iteration tested it for null before rewinding; it now points at a mark that stands for none, whose regime bit sends the rewind to a slow path that returns at once. With both, runbench reached 1,176,309,124 and escapebench fell 11%.
Two more were built, measured and declined. Reading four digits at once in the number parsers lost 0.36% on runbench: the corpus's floats have three integer digits, so the four-digit test failed on every float. Keeping a word on each mark that equals its pointer only while nothing needs flushing made the rewind's common case one compare, and escapebench fell 6% in every version of it, but the word has to be maintained at every push, and the encoder's beats are short enough that encodebench rose about a percent while runbench gained a quarter of one. It was declined on that trade.
159failure, length and room, each asked once
A group that returns a small record hands it back in two words, and a failure travels in the same two words. Its low byte is the error's tag, and that byte is where the record's value field keeps its own tag. So when a clause takes the record apart with a name or a wildcard in the value position, the test on the value field already refuses a failure. The compiler used to test the whole word first and the field second, which asked one question twice for every token the JSON decoder hands on. It now asks it once. Runbench fell from 1,176,309,124 instructions to 1,168,439,677.
Three more came from the same kind of reading. An index xs[i] is tested 1 <= i <= len against a length LLVM knew nothing about. The compiler now tells it, with an assumption after each load, that a length is never negative, and LLVM folds the pair where nothing upstream had already settled it. Writing the test as one unsigned compare directly was tried first. Runbench liked it, but encodebench rose 7%, all of it in the encoder's escape loop, and that version was declined. Where no guard had settled them, the index's two compares were then split into two branches, with the length loaded between them so LLVM could not fold them back together. A slice's bounds, 1 <= from <= to <= len, did become two unsigned compares, in the runtime and in the emitted append. And the in-place push, put and append each asked two questions about the buffer, whether it ends at the frontier and whether the new length fits, joined into one branch. At -O3 LLVM computes a joined condition as flags and combines them; as two branches it is two compares. An empty slice, a fifth of the decoder's runs between escapes, now returns the builder from between those two branches instead of travelling through the copy as a zero length. Together these took runbench down another 2.0% on this container, the decoder 4.6% and escapebench 2%.
160the nightly in shards
The nightly ratchet applies every mutation in the table and checks that its gate goes red. It had not finished since 2026-09-04. For nine nights its baseline was red, and from 2026-09-14 the job ran out of its ninety minutes every night: the baseline took eleven and a half minutes, the rows ran for the rest, and the job was cancelled. The report is written once, after the last row, so each of those runs printed nothing about any row. The table had grown from 62 rows to 250. The nightly now runs eight jobs, each proving every eighth row from a different start, and a spec holds the workflow's list of jobs to the count its command divides by. The ratchet also names each row on stderr as it starts, so a run that stops early says which row was running.
The first sharded run finished in 23 to 38 minutes a shard, and it found one row blind. The welfare gate's row raised jsonbench's instruction count and expected the index to fall. The index has read only the consolidated run program since 2026-09-06, so the jsonbench figure no longer moved it, and the row had been unable to fail for three weeks. It now raises the run program's count, and the gate goes red.
161two counts without a popcount
The utf-8 validator counts the characters in the text it checks, because that count becomes the string's length and length never scans again. Its vector pass counted continuation bytes with a mask and a population count. The runtime is built for SSSE3, which has no population-count instruction, so the compiler wrote one out in shifts and masks: 22 of the 51 instructions in every sixteen-byte block that held a non-ascii byte. The count now stays in a vector register, summed across lanes with psadbw, at four instructions a block. Runbench fell 0.42% on this container, livebench 1.29% and encodebench 0.88%.
The differential harness for the validator had never checked that count. It tested the validator's yes or no and a separate counting function, and a pass that stopped counting one continuation byte got through 53 million cases without a mismatch. It now checks the count from both passes over 400,000 valid strings, and the same defect fails it 45,136 times.
The slice walk had the same popcount on a word whose only set bits are the byte-high ones. Shifting each to its byte's low bit and multiplying by 0x0101010101010101 sums the eight into the top byte in three instructions. Runbench fell another 0.21% and indexbench 2.8%. The same multiply in the general character counter cost pendbench 0.9%, for a reason not isolated, and that counter keeps the builtin.
162numbers read a word at a time
The json decoder reads each number out of the document with to_int or to_float on a slice, and both parsed a digit at a time, about ten instructions a digit. A slice ends inside the document, so the eight bytes that end at its last digit are in the buffer. An integer of up to eight digits is now read as that one word: the bytes in front of it are filled with zeros, one test checks that every remaining byte is a digit, and three multiply-shifts combine them. A decimal of the form 466.1234 is read as two such words. The length of the digit run at the top of the last word says where the dot is, and the integer part must fill the bytes between the sign and the dot exactly. The two words give the same significand and exponent the digit loop would have built, and they go to the same Eisel-Lemire parse, so the answer cannot change. On this container runbench fell 1.48% and jsonbench 3.38%.
A four-digit version was built earlier the same day and declined. It did not know where a number's digits ended, so it tried four and failed on every float whose integer part is three digits long. Reading from the end avoids that, because the slice's length is known before anything is read.
Two things were checked and recorded without a change to the compiler. The wall's rule that two failures merge into one answer was removed on purpose on 2026-08-06, in kanso#783, the same day it was measured, with a sentence in the log saying why. An early return in k_b_utf8_slice_raw for a span outside its string was built and declined: runbench rose 0.07% and jsonbench 0.17%.
A pull request that changed src/runtime.c waited hours for its ratchet job, which proves every mutation whose file the branch touched, 78 of them for that file alone. The job now runs in four parallel shards, each taking every fourth row, and the required check fails unless all four succeed. kanso#1676, which touched the runtime, finished its four shards in 22 to 33 minutes.
163a map written from its columns
The encoder used to write a map by asking for entries m, a list of entry records, and taking each record apart. A ruling on 2026-09-27 added keys m and values m, two lists in the order entries gives, so that position i of each is one pair, and a micro golden pins that alignment on every engine. The encoder now indexes the two lists and builds no record. Measured by CI, runbench fell 0.455%. livebench rose 0.314%, because it encodes many small maps, and each map now makes two calls where it made one. The interpreter rose 1.91%: it copies each column into a fresh list, and the walk's two index reads each go through its dispatch where one pattern did.
A beat loop decides, for each value a lap hands to the next, whether it can cross the rewind as it is. A value handed back unchanged on every lap crosses by identity, and the runtime never looks inside it. Lists, records, strings and closures were treated that way. Maps were not, because a map's first read used to cache a sorted view in the arena, and a map older than the mark would have held a pointer the rewind freed. The view has been malloc'd since the view registry was built, so maps now cross by identity too. A program that walks a 64-key map for twenty thousand laps fell from 209,505,643 instructions to 95,905,706; half the old figure was the runtime asking the same 128 slots whether they survived. livebench fell 0.256% and encodebench 0.200%.
Three more were built and declined. Reading a JSON number's span and value in one runtime call lost 2.01% on runbench, because the answer came back as an arena record that the decoder then had to take apart, where main's path inlines the parse and returns in two registers. Walking the first bytes of a number one at a time before the sixteen-byte classifier lost 1.94%: the run program's numbers average seven characters, and the classifier is already cheaper than seven byte tests. Reading a map's pairs in place, through two builtins only the standard library can name, lost 1.59%. The two reads cost about what the two lists had, and the encoder's output builder lost the region section 157 gives it, because that region is opened by a descent whose argument allocates, and keys m and values m were those arguments. Without them the builder grew in the arena by copying, 24 million instructions of it.
What remains for a program that walks a map with entries is the record itself. On a 64-key map walked twenty thousand times, entries read 95,904,950 instructions and the same walk over the two columns 66,424,441, about 23 instructions a pair. A compiler pass that keeps the record out of memory when the code receiving it only takes it apart would recover that. Its payoff on a real program was measured by doing the rewrite by hand on encodebench's frozen encoder, which walks maps the way a user would: 1.00% fewer instructions. The same rewrite in lib/json cost livebench 0.314%, because a small map pays for two calls where it paid for one, and a pass cannot see how large the maps will be. The pass stays on the list with that trade beside it.
164a clean token is not scanned
The JSON writer asks of every string it writes where the first quote, backslash or byte below a space is. The decoder hands out each token of four to seven bytes as one shared permanent string, so on the run program the writer asked that question 942,750 times a run of the same few hundred keys. Each shared token now carries a byte, written once when it is first made, saying whether it holds any of the three, and the scan answers "none" for a clean token without reading it. 764,370 of the run program's scans are skipped. Measured by CI, runbench fell 1.73%, livebench 5.17% and encodebench 3.07%. The release build's own cost rose 1.28%. The flag costs one permanent byte per shared token, 632 bytes on each benchmark. A micro golden writes shared tokens that hold a quote, a backslash and a tab, so a flag computed wrong shows in the output.
When a call asks for both keys m and values m of one map, as the encoder's map path does, the two lists are now built by one runtime call into one allocation. runbench fell 0.978% and its allocations 1,955,773 to 1,707,283.
Two changes were to the measuring rather than the compiler. mimalloc counts the host's NUMA nodes by probing /sys the first time a thread allocates, and the interpreter's thread was that thread, so its row would read about 800 instructions more for each node a machine has beyond the first. The allocator is now told there is one node, and the row fell 804 instructions. The cost-goldens job also uploads the interpreter's and start-up's callgrind profiles beside the compile rows' now, so the next time the interpreter row reads differently on two runners of one commit there are two profiles to compare.
165programs written at random
The differential law says every engine answers every program the same way, and until this sitting every program that checked it was written by hand. A scratch generator now writes small modules of bindings, list and map calls, typed dispatch groups and countdown loops, with the loop counts read from os/args so that the compiler cannot fold them away. Each module runs on the interpreter and as a dev and a release build, and the three outputs, error streams and exit codes are compared. About 1,600 generated programs ran on all three engines. Two of them disagreed.
The first failed everywhere. A binding whose value was a list adapter, read once after a return ... if, lost its name: enumerable fusion inlines such a binding into its one reader, counted the readers with a walk that entered the guard, and rewrote them with one that did not. The interpreter answered "unknown name" and the native backend refused the program. The rewrite now uses the walk that covers every form.
The second agreed on every byte and took 8.4 seconds interpreted where the dev build took 5 milliseconds. A loop whose seed list the caller reads again after it copied the whole list on every lap in the interpreter, because the proof that lets a push write in place has to hold for every caller. The native engine checks at run time instead, with a length stamped on each buffer. The interpreter now checks its own way: where the write's list is a name its function reads once, and the value has no holder besides that name, the push extends it in place. The first lap copies and the rest do not, and the same loop takes 35 milliseconds.
166what the generated programs found next
The generator ran six programs at a time, and six of the first sixteen builds failed with clang crashing on the runtime's C source. The runtime is compiled once per set of options and cached in the temp directory. Every build that found no cached copy wrote the source to the one shared path, and the write truncated the file under whichever clang another build had reading it. The object was already written under a name of its own and renamed into place; the source now is too. A test starts eight builds into a fresh temp directory, 150 milliseconds apart, and failed in five sittings of five before the change.
Once the generator learned to build lists of maps, four programs in six hundred finished interpreted and timed out or were killed as native builds. Printing a list joined each element's text onto everything before it, and each join copied the whole string into the arena, so the cost grew with the square of the output: 4,000 small ints allocated 32 megabytes to print eight thousand characters. The native renderer now writes a whole container into one buffer and makes it a string once, as the interpreter always has. A memory fixture renders 4,000 ints and 500 maps and allocates 207,280 bytes, where it allocated 38,215,552 before.
167five more disagreements
The generator kept running after section 166, and five more programs found places where the engines disagreed.
math/round of 1e30 printed int64's largest value on the interpreter and its smallest on both native builds. A kanso int has no width, so the interpreter now answers the integer the double holds exactly, and a native build refuses a float past int64 with the diagnostic arithmetic already gives on overflow. What NaN and the infinities should answer is a question for the ledger, and until it is ruled they answer what the interpreter always did.
push [1 2] bad, where bad is an err, stopped with the err on the interpreter and printed a list holding it natively. Every builtin's arguments are infectious in this language, and the C push asked only whether its list had failed. Probing the other builtins with two failing arguments found a wider form of the same gap: seventeen of them answered the first failure where the interpreter merges all of them. Each now tests all of its arguments at once and merges what failed in argument order.
A function nobody called failed its dev build with a clang type error. A lazy binding's cell tail-called a group that returns its record in two registers, and the path for that tail call returned the pair as though it were one tagged value. An ordinary caller never reaches the case, because the analysis that sizes a tail caller's return covers it; a lazy cell's site is built outside that analysis.
9007199254740993.0 * 0.1 printed ...099.3 interpreted and ...099.2 natively. The double lies exactly halfway between those two shortest decimals. The native renderer takes the even digit, as Python and JavaScript do, and the interpreter now does the same.
The last came from a batch that began calling std/render, std/json and std/regexp, and it accounted for all ten divergences in that batch. Interpolation calls render/to_string on a value of any kind, but inference typed that group's parameter from the program's explicit calls. A program whose only explicit call passed an int had its interpolated records arrive at the renderer as their type's tag, and one whose calls all passed strings put a false assumption about the tag into release builds, which trapped. The group's parameters now start out holding any value.
Each fix has a golden that fails on the code before it and a ratchet row that restores the old code and watches the golden fail again.
168values the compiler took for failures, and the reverse
Five more programs disagreed, and two of the fixes are in what inference believed about failures.
A list literal keeps a failing item in place, where push, put and a record constructor hand the failure on. Inference typed every element read as never failing, so a compiled build called the next function with no test of the err and told the optimizer the argument held bytes. The err went in anyway, and a line went missing from the trace. Inference now notes, once for the whole program, whether any literal can hold a failure, and every element read admits one if so. No benchmark has such a literal, and none of their code changed.
The opposite mistake was older. Since 2026-08-10, dividing by zero answers a value, the text “division by zero”, and inference had gone on typing the answer as an err. A dispatcher skips an err before any arm runs, so a function handed the answer took its parameter to be a number, and the compiled body multiplied the string’s address. The interpreter refused the *. Division now answers what the value is.
The other three were smaller. text/to_int of a number past int64 answered an err on native builds, which std/json turned into “invalid number”; a compiled build now refuses it with the diagnostic arithmetic gives on overflow. list/cycle [] never ended on any engine. And an unhandled err whose reason held a NUL was printed only up to it.
A sixth program did not build in release at all. The link runs LLVM’s dead-argument elimination over the whole program, and the pass narrowed some of std/regexp’s predicates to return one word where each still ended in a musttail call returning two, which the verifier refuses. A program only had to declare a subtype for the inlining to fall that way. Rewriting the emitted calls cost the run program 1.82% or more, because what the pass will narrow depends on what inlining did first. The release link now runs LLVM 19’s own pipeline with that one pass removed, for 0.052%.
A larger generated program found a seventh. Reading a missing key answers none, and a list holding that read was handed to json/encode. The interpreter refused it because no arm of the encoder takes a none. A release build died on a segfault. A compiled dispatcher that matches no arm passes a failure on and dies on anything else, and that test still counted a none as a failure, two months after the ruling that made it a value. It now counts only an err, and every dispatcher that can reach the test is a compare shorter.
An eighth came from a curried function. A parameter read only as the head of an interpolation can be grown in place as a string builder when every caller hands its value over, and each caller converts the string it passes into a builder first. The analysis finds the callers by walking the program, and it passed &name by, so a function reached only through a partial application seemed to have no callers, and qualified. The partial application then converted an int, and a native build died. &name now counts as the function escaping as a value, which rules the parameter out.
A ninth changed only a trace. A program that gives render/to_string an arm of its own has every record it interpolates or prints rendered through that group, so the arm is found. A native build sent an err there as well, and the group handed it on with render/to_string recorded as a place the err had passed through. The interpreter returns an err from an interpolation or a print before rendering anything, so its trace had no such line. A native build now tests the value before the call when it can be an err, and hands an err on untouched.
A tenth came from text/join. A list literal keeps a failing item in place, and a join answers the first item that fails, but the compiler typed its answer as text. A guard then read length of the err as a number and returned early, and a function whose only arm takes a string was called with the err. A join now admits what a literal can hold, the same way an element read does.
An eleventh was a map indexed by a none. The interpreter indexes a map by an int or a string and refuses anything else, and a native build answered a miss. Finding that turned up a disagreement inside the interpreter: every builtin, at among them, reads through a subtype to the value it wraps, but the interpreter’s index syntax did not, so it refused a subtype of int as a position where at accepted one. An index now reads through a subtype on every engine, and a map index refuses a key that is neither an int nor a string.
A twelfth came from teaching the generator subtypes, which the book says flow wherever their parents flow. A native build writes a few builtins out itself instead of calling them, and those paths read a subtype as the type it was proven to be: json/encode of a subtype of string crashed, and a slice refused one as a position. A native build refused to build a subtype of bool, some or done at all. Once it could, a subtype of bool turned out to be refused as a condition by every engine, which each explained by saying a condition is true or false and that it had got true. text/join and text/to_bytes refused lists of subtypes the same way. Each of those places now reads through the wrapper, and the checker refuses a subtype whose parent names no type.
One disagreement was left alone. The interpreter says a NaN equals itself and a native build says it does not, and the interpreter’s answer depends on the NaN’s sign bit, which differs between x86 and ARM. How floats compare is a question about the language, so it went to the ledger, and until it is ruled the generator does not make NaNs.
169arms that could never run
When the generator learned to write subtypes into dispatch arms, it began writing arms that never ran. The engines agreed about every one of them, so the differential could not call them a disagreement. What it could do was point at a family of programs the checker let through, and probing that family by hand found the rest.
A constructor pattern that takes the wrong number of fields matches nothing. (id n) for type id int takes one field of a value that has none, and (pt n) takes one of two. The checker now follows a wrapper to the record at the bottom of its chain and compares the count there.
An arm can also be dead because another arm takes all of its values first. Arms rank literal, then type, then generic, and within a rank the one written first wins. _:pt and (pt a b) both take every pt. So do none and _:none, and _:[]int and _:[]string, since a list arm asks only whether its argument is a list. The overlap check compared these by spelling and now compares what an arm can test.
For the same reason, an earlier arm that takes every value a later one does leaves it nothing. (pt n _) above (pt 1 _) sends pt 1 2 to the first on every engine. Written the other way round, both run. The check now compares two such arms field by field. Two typeset arms rank alike too, so an arm for type wide id tag word written above one for type narrow id word takes every narrow first. And several arms together can leave nothing for a later one: arms for true and false leave nothing for _:bool, and arms for each member of a typeset leave nothing for the typeset, because a literal and a member type rank above it. The check asks both questions now, and says which arm to move when moving one would help.
The first version of that last question looked at every declaration in the module for each arm. The corpora that check a module merged with the standard library paid about eleven million instructions for it, which CI reported before anything was banked. The check now sorts the declarations once and compares an arm only with its own group, and the corpus that held eleven million extra instructions checks in two million fewer than it did before any of this. None of these rules changes the verdict on a program in the standard library, kq, vse or kanso-json.
170checks that read one file of a module
A module is a directory of files that share their declarations. The checker reads it twice: once a file at a time, and once as a single merged program. When the generator began splitting records from the code that uses them, it found checks on the per-file side that asked only their own file's types.
A constant can name itself inside a constructor, as in ring = node 1 ring. That stores a place for ring rather than asking for its value, and the cycle check lets it through because node is a type. With node declared in the next file, the check did not know it was a type, and every engine refused the constant as defined in terms of itself. The check that a constructor pattern can match had the same gap the other way round: (tall (rect _ _)) hands a subtype of a two-field record one field, and in a module whose records lived elsewhere it passed kanso check and the arm never ran. Both now read every file's type declarations.
The merged side had the opposite problem. It sees every type, but most of what it raised carried no file, so through an import a refusal printed the module's name and no line. Each diagnostic is now placed in the file of the declaration it is about.
Probing the same shapes by hand turned up one more. A subtype is built from the one value it wraps, and the arity checks left subtypes out without checking that count. tall 1 2 passed the checker; the interpreter then refused it at run time and the native backend refused the build, each in its own words. The checker now refuses it before anything runs.
codathe mushroom test
every proposal on this page passed the same two filters before it earned a section. the first is the house rule you've met everywhere else on this site: polymorphism over conditionals—if the design answer is "the user writes an if," the design is wrong, because dispatch makes the arms checkable and the conditional unwritable. the second is the mushroom test. a mushroom looks like a new organism; dig, and it's the fruiting body of a mycelium that was already there. that's the question every ruling faces: does this add a concept, or reveal that an existing one already covers it? zero-field types looked like a feature and turned out to be record types with nothing left to remove. enums looked like a feature and turned out to be typesets of markers. the integer tiers looked like machinery and turned out to be one case of speculation under purity. proposals that add a concept wait; proposals that reveal one land.