commit f6aa2f2ab094620c0b1ac8bef34ef6b1381c7463
parent 13b62e96a6d28334558d8ed64d924d9e641b9960
Author: Ryan Sepassi <rsepassi@gmail.com>
Date: Sat, 13 Jun 2026 14:39:31 -0700
doc(plan): refresh PERF-O0-CODESIZE.md to post-campaign state
Drop the landed levers' design sections (rationale now in the commit messages);
record what landed (§2), refresh the comparison + per-mnemonic diagnosis to the
current 1.16x tcc / +55,698 excess (§1, §3, with fp-vs-sp shown resolved — both
are fp-relative, kit now beats tcc on far addressing), and re-rank the remaining
levers (§4): args-into-arg-regs, spill/value-residency reduction, byte/half far
slots, signed load-with-extend, rv64 L4-P1.
Diffstat:
1 file changed, 163 insertions(+), 367 deletions(-)
diff --git a/doc/plan/PERF-O0-CODESIZE.md b/doc/plan/PERF-O0-CODESIZE.md
@@ -7,67 +7,30 @@ instructions emitted** — making kit's `-O0` machine code as dense as tcc's.
The two tracks compound. Every emitted instruction removed is also one fewer to
emit, relocate, and write to the object — so codesize wins shrink the
-emit/objwrite/assemble phases *multiplicatively* (`PERF-TCC-SLIM.md` §2.1). This
-track is most of the remaining gap to parity after the per-operation slimming.
-The goal is **smaller and faster output while holding -O0 compile speed** — and
+emit/objwrite/assemble phases *multiplicatively* (`PERF-TCC-SLIM.md` §2.1). The
+goal is **smaller and faster output while holding -O0 compile speed** — and
because un-emitted instructions cost nothing to emit, the well-chosen wins here
*improve* compile speed rather than trading it away.
The arch focus is **aarch64** (the reference backend); x64/rv64 carry analogous
-taxes and follow once the aa64 design is proven.
+taxes — the shared-NDT levers (below) already help them, the aa64-specific ones
+follow once the design is proven.
---
-## 0. Landed (branch `o0-codesize`, measured 2026-06-13)
-
-Five levers landed and gated (determinism + sqlite e2e O0/O1 golden+vs-clang +
-toy/parse/smoke-x64/smoke-rv64/dwarf/debug + alloca/far-slot/narrow/p1-x64
-clang-differential probes). **aarch64 sqlite `-O0` `.text`: 455,546 → 398,433
-insns, 1.3291× → 1.1625× tcc (−57,113, half the excess gone).**
-
-| lever | change | insns | where |
-|---|---|--:|---|
-| **L2** | 2-insn top-record epilogue (`mov sp,x29; ldp [sp],#16`) | −2,633 | aa64 |
-| **L1** | positive-scaled far-slot addressing (Fix B); alloca→x28 anchor | −25,579 | aa64 |
-| **L4-P1** | scalar call result stays in the ABI result reg (gated `ndt_result_reg_stable`) | −8,969 `mov` | shared, aa64-on |
-| **L4-P2** | copy source materialized straight into the destination register | −16,151 | shared |
-| **L3** | elide the zero-extend after a narrow (zero-extending) load | −3,587 | shared |
-
-The three shared-NDT levers also help other arches for free: **x86-64 sqlite
-`.text` −3.18%** (L4-P2 + L3); rv64 likewise (smoke-green). L1/L2 are
-aa64-specific (the concept ports, the mechanism does not). L4-P1 is gated to
-`ndt_result_reg_stable` arches — aa64 (x0 is call-only-clobbered); **x86-64 is
-explicitly excluded** because RAX is an implicit div/mul operand (a result left
-there is clobbered before its consumer — caught by the p1-x64 probe).
-
-**Deferred frontier** (high value, but frontend/CG-API-invasive — left as
-follow-ups, not rushed): **L4-P3** args-into-arg-regs (~30 K `mov xK,reg`; needs
-a call→arg-producer destination hint threaded through the frontend);
-**L3-signed** `ldrsb`/`ldrsh` (~1.3 K; kit's CG integers are sign-agnostic, so
-the load can't know it feeds a sign-extend — needs a frontend widening load);
-**rv64 L4-P1** (a0 is call-only-clobbered, so sound — pending an on-target
-div-after-result check before flipping `ndt_result_reg_stable`).
-
----
-
-## 1. Original state (measured 2026-06-13, pre-campaign)
+## 1. Current state (measured 2026-06-13, branch `o0-codesize`)
`build/release/kit cc -c sqlite3.c` vs `tcc -c sqlite3.c`, arm64-macOS, on the
3.50.2 amalgamation (`tmp/projects/sqlite-amalg/sqlite3.c`, 9.28 MB).
| metric | kit | tcc 0.9.28 | ratio |
|---|--:|--:|--:|
-| **machine code (`__text`/`.text`)** | **1,822,184 B / 455,546 insns** | **1,370,940 B / 342,735 insns** | **1.33×** |
-| object file | 2,216,680 B (Mach-O) | 2,108,134 B (ELF) | *not comparable¹* |
-| compile instructions | 1,968 M | 667 M | 2.95× |
-| compile wall | 0.18 s | 0.06 s | 3.0× |
-
-¹ tcc emits ELF, kit Mach-O — different format/symbol/string-table overhead.
-**The honest size metric is the `.text` machine code**, and there kit is **1.33×
-tcc, +112,811 excess instructions.** (This is well down from the 1.57× recorded
-when §6's history was written — the value-stack-residency + lazy-homes +
-coalescing commits landed since and reshaped the picture; `mov`, not `sub`, is
-now the #1 excess.)
+| **machine code (`.text`)** | **1,593,732 B / 398,433 insns** | **1,370,940 B / 342,735 insns** | **1.16×** |
+
+**The honest size metric is the `.text` machine code**, and there kit is now
+**1.16× tcc, +55,698 excess instructions** — down from 1.33× / +112,811 at the
+start of this campaign (half the excess removed; §2). The object file (Mach-O vs
+tcc's ELF) is not comparable, and compile-speed lives in `PERF.md`.
**Reproduce:**
@@ -83,350 +46,183 @@ SRC=tmp/projects/sqlite-amalg/sqlite3.c; SDK=$(xcrun --sdk macosx --show-sdk-pat
---
-## 2. Diagnosis — where the 112,811 excess instructions live
+## 2. Landed
+
+Five levers landed and gated (determinism + sqlite e2e O0/O1 golden+vs-clang +
+toy/parse/smoke-x64/smoke-rv64/dwarf/debug + alloca/far-slot/narrow/p1-x64
+clang-differential probes). Design rationale lives in each commit message.
+
+| lever | change | aa64 insns | where |
+|---|---|--:|---|
+| **L2** | 2-insn top-record epilogue (`mov sp,x29; ldp [sp],#16`, was a 3-insn x16 dance) | −2,633 | aa64 |
+| **L1** | far fixed slots addressed positive-scaled off sp (`ldr/str [sp,#+ofs]`, was `sub xN,x29,#off; ldur`); alloca anchors a callee-saved x28 frame base | −25,579 | aa64 |
+| **L4-P1** | scalar call result stays in the ABI result register (was `bl; mov cachereg,x0`) | −8,969 `mov` | shared¹ |
+| **L4-P2** | copy source materialized straight into the destination register (was scratch-load + `mov`) | −16,151 | shared |
+| **L3** | elide the zero-extend after a narrow zero-extending load (`uxtb`/`uxth`/`ubfx`) | −3,587 | shared |
+
+¹ L4-P1 is gated to a new `NativeRegInfo.ndt_result_reg_stable` capability —
+**on** for aa64 (x0 is clobbered only at calls) and sound for rv64 (a0), **off**
+for x86-64 (RAX is an implicit div/mul operand: a result left there is clobbered
+before its consumer; the p1-x64 probe catches it).
+
+**Cross-arch:** the three shared-NDT levers (L4-P1/P2, L3) apply to every backend
+for free — **x86-64 sqlite `.text` −3.18%**, rv64 likewise (smoke-green). L1/L2
+are aa64-specific (the *concept* ports — positive-scaled / fold the epilogue —
+the *mechanism* is per-arch).
+
+Prior partial work this built on: value-stack residency, lazy transient homes,
+lazy-dup + copy/convert coalescing (`[[o0-value-stack-residency]]`,
+`[[o0-lazy-transient-homes]]`, `[[o0-copy-convert-coalescing]]`), which took the
+ratio 3.45×→1.33× before this campaign.
+
+---
+
+## 3. Diagnosis — where the remaining 55,698 excess lives
-Per-mnemonic signed diff (kit − tcc), from the opcode histograms:
+Per-mnemonic signed diff (kit − tcc) on the current `.text`:
| mnemonic | kit | tcc | **excess** | what it is |
|---|--:|--:|--:|---|
-| `mov` | 73,824 | 3,701 | **+70,123** | reg-reg copies — **#1 lever** |
-| `stur` | 62,819 | 15,899 | **+46,920** | frame spills (negative-offset, unscaled) |
-| `ldur` | 80,853 | 37,994 | **+42,859** | frame reloads |
-| `sub` | 33,758 | 6,713 | **+27,045** | **29,379 are `sub xN,x29,#imm`** — far-local address building |
-| `cbz` | 8,942 | 2,648 | +6,294 | compare-branch lowering |
-| `movk` | 4,360 | 357 | +4,003 | building frame offsets > 4095 |
-| `uxtb`/`uxth`/`sxtb`/`sxth` | 5,434 | 333 | +5,101 | sub-word extends after narrow loads |
+| `mov` | 48,561 | 3,701 | **+44,860** | reg-reg copies — **#1 remaining**: arg setup into x0–x7 (~30 K) + residual materialize/value-stack copies |
+| `ldur` | 71,934 | 37,994 | **+33,940** | frame reloads — register pressure |
+| `stur` | 48,913 | 15,899 | **+33,014** | frame spills (≈3× tcc) — register pressure |
+| `str` | 22,668 | 8,211 | **+14,457** | "" (more spills) |
+| `sub` | 11,484 | 6,713 | +4,771 | residual `sub xN,x29` (byte/half far slots + &local) |
+| `movk` | 2,371 | 357 | +2,014 | far-offset builds in the residual `sub` paths |
+| `sxtb`+`sxth` | 1,341 | 99 | +1,242 | sign-extend after a narrow load (L3 did the unsigned half only) |
-By value class the spill story is starker than the per-mnemonic split: **kit
-stores ~78 K vs tcc ~26 K (3× the spilling)**; loads are closer (107 K vs 92 K).
-
-Where **kit already beats tcc** (do not touch — these are structural wins, mostly
-kit's fused `cmp;b.cond` vs tcc's `cset;cbnz`, plus folded offsets):
+Where **kit already beats tcc** (do not touch — structural wins):
| mnemonic | excess | note |
|---|--:|---|
-| `ldr` | −28,365 | tcc materializes more addresses; kit folds offsets into displacement |
-| `add` | −18,422 | "" |
+| `add` | −21,709 | kit folds offsets into the displacement |
+| `ldr` | −19,360 | far slots now fold into `[sp,#scaled]`; tcc materializes more addresses |
| `cset` | −17,226 | tcc lowers every compare to a materialized bool |
| `cbnz` | −10,817 | "" |
| `cmp` | −9,421 | kit fuses compare into the branch |
-| `movn` | −6,720 | tcc's far-local offset build (see §3) |
-| `nop` | −5,834 | kit's prologue pad is *smaller* than tcc's (§3) |
+| **`movn`** | **−6,720** | **far-local addressing: kit now emits 0 (positive-scaled), tcc still builds the offset** |
+| `nop` | −5,852 | kit's prologue pad is smaller |
| `stp` | −3,722 | tcc pairs more aggressively |
-**Concentration.** The excess lives in the giant functions. Top 12 by size hold
-~80 K insns; `_sqlite3VdbeExec` alone is **22,991** (5% of all text), of which
-**5,162 (22%) are pure far-local address arithmetic** (2,757 `sub xN,x29` +
-2,405 `movk`). Frame-size distribution: kit mean **112 B**, max 3,520, 181
-functions > 256 B; tcc mean **53 B**, max 3,392, 58 functions > 256 B.
-
----
-
-## 3. The two structural realities
-
-**(a) The frame-addressing tax** (`sub` +27 K, `movk` +4 K, and part of
-`stur`/`ldur`). Both kit and tcc home fixed locals at **negative x29 offsets** —
-this is *not* an addressing-mode difference. aarch64 reaches negative offsets
-only with unscaled `stur`/`ldur` (±256). Past that, an address must be built. The
-difference is *how*:
-
-- **kit** emits `sub x17, x29, #off` **per access** (+ a `movk` when off > 4095),
- then `stur`/`ldur [x17]`. Cited at `aa_fp_off_slot = -off` (negative), the
- single-pass TOP-RECORD layout (`src/arch/aa64/native.c:112`).
-- **tcc** falls through `arm64_ldrx` (`arm64-gen.c:317`) to **register-indexed**
- addressing: `movn/movz x30, #off ; ldr/str [x29, x30]`, reusing x30. One `movn`
- covers any frame < 64 KB, and the index reg is shared across nearby accesses.
-
-In `VdbeExec`: kit 2,757 `sub` + 2,405 `movk`; tcc **66** `sub`, 1,668 indexed
-accesses, 1,719 `movn`. kit pays ~2–3 insns per far access where tcc pays ~1, and
-**kit makes more far accesses** because it touches the frame more (reality (b)).
-
-**(b) kit computes-into-a-register-then-moves/spills; tcc computes-into-place**
-(`mov` +70 K, stores +52 K). 8,511 movs just capture a return value out of x0;
-~19 K cluster around `bl` (arg setup `mov mov bl` / retval `bl mov`); the rest are
-materialization chains (`ldur;mov`) and value-stack copies. The value-stack
-residency work (§6) moved partway — it traded frame round-trips for register
-moves (net fewer insns, but `mov` count *rose*). Closing the rest is
-destination-driven (lazy) value placement — which is the *absence* of eager
-relocation, **not** an added pass, so it is compile-time-neutral-to-faster (§8).
-The cost is correctness/invasiveness, not throughput. Treated as a separate
-track (§4, Lever 4; mechanism + phased plan in §8).
-
-**Non-levers (ruled out empirically):**
-- *Prologue nop pad.* kit reserves a `b <body>; nop×3` patch region per function;
- tcc reserves `ARM64_FUNC_STACK_SETUP_SLOTS` nops too, patched in `gfunc_epilog`
- (`arm64-gen.c:1417`). tcc has **more** nops (13,757 vs kit 7,923). Not a
- differentiator.
-- *Branch lowering.* kit's fused `cmp;b.cond` already beats tcc's `cset;cbnz`
- (the −17 K `cset`, −10 K `cbnz`, −9 K `cmp` deficits). Leave it.
-
----
-
-## 4. Ranked open levers
-
-| # | Lever | Est. saving | Risk | Compile-speed | §|
-|---|---|--:|---|---|---|
-| **1** | **Fix B — positive-offset -O0 frame addressing** | **~33 K (−7.3%) → ~1.24× tcc** | med (alloca layout, CFI, base-reg) | neutral/faster | §5 |
-| 2 | Epilogue slim (drop the x16 dance) | ~2,633 (−0.6%) | low | neutral | §6 |
-| 3 | Load-with-extend (`ldrsb/ldrsh` vs `ldrb;sxtb`) | ~1,300 | low-med | neutral | §7 |
-| 4 | Compute-into-place (`mov`/spill; lazy dest placement) | up to ~50–70 K | high (correctness/invasiveness) | neutral/faster | §8 |
-
-Levers 1–3 are smaller-and-faster with bounded risk; Lever 4 is the deep,
-high-count frontier — high *correctness* risk and invasiveness, but
-compile-time-neutral-to-faster (it removes work, §8), so it is *not* gated by the
-compile-speed budget. Re-profile after each landing — the phase shares shift.
-
----
-
-## 5. Lever 1 — positive-offset -O0 frame addressing (Fix B)
-
-**The idea:** address fixed locals as **positive *scaled* offsets** from a stable
-frame-base anchor. A scaled `ldr/str [base, #+ofs]` reaches **32,760 B** for
-64-bit (4095×8), covering any realistic frame in **one instruction, zero address
-building** — strictly better than both kit's current `sub`-per-access *and* tcc's
-`movn`+indexed scheme. This deletes essentially all 29,379 `sub xN,x29` and the
-~4,000 frame-offset `movk`.
-
-### 5.1 The infrastructure already exists
-
-The aa64 backend already has a positive-offset layout and a deferred-patch list:
-
-- **BOTTOM-RECORD layout** (`fp_at_bottom`): the saved pair sits at the *bottom*,
- `fp = sp`, slots stack **above** at positive offsets, `aa_fp_off_slot(a,off) =
- frame_size − off` (`src/arch/aa64/native.c:291`, layout diagram :120-136). It is
- currently **gated to -O1 known-frame** small frames with `out_stack == 0`
- (`native.c:1835`), because the positive offset `N − off` needs `frame_size` (N),
- which the single-pass path does not know before the body (`native.c:217`).
-- **`AAPatch` deferred-patch list** (`native.c:185-202`, resolved in
- `aa_func_end`): the single-pass path **already defers the prologue `sub sp`** and
- patches it once `frame_size` is final (`native.c:72-82`). The same mechanism
- extends to slot offsets.
-
-So Fix B is **not a new subsystem** — it is *extending BOTTOM-RECORD to the -O0
-single-pass path* by deferring the per-slot offset, reusing `AAPatch` and the
-deferred-prologue machinery.
-
-### 5.2 The single-pass mechanic
-
-1. Emit each slot access as a scaled `str/ldr [base, #placeholder]` (imm12) and
- record an `AAPatch{insn_pos, slot_off, scale}`.
-2. In `aa_func_end`, once `frame_size` (N) is final, patch `imm12 = (N − slot_off)
- >> scale`. Because positive scaled imm12 reaches 32,760 B, the placeholder is
- **always one word** — no 1-word→2-word expansion (the old worry, which was a
- *negative*-offset artifact, is gone).
-3. The base anchor is set once in the prologue and never clobbered. Two viable
- anchors (decide during implementation):
- - **`fp` at the bottom** (true BOTTOM-RECORD): `mov x29, sp` *after* the
- deferred frame `sub`. fp becomes the stable bottom anchor. CFA becomes
- `fp + N` (frame-size-dependent — patch the CFI in `aa_func_end`, see §5.4).
- - **A dedicated callee-saved base reg `xFB`**: keep fp at the top (top-record
- CFA `fp+16`, unwind unchanged) and add `mov xFB, sp` after the frame `sub`.
- Costs one reserved callee-saved register + its save/restore, but leaves
- unwind/CFI untouched. Note the single-pass path currently *uses no
- callee-saved registers* (`native.c:77`, enforced in `nd_scratch_acquire`) —
- reserving `xFB` relaxes that invariant.
-
-The setup insn lands in the existing reserved prologue region, so **only
-functions that overflow ±256 pay it** — small functions keep today's
-`ldur [x29,#-n]` and emit nothing extra.
-
-### 5.3 alloca / VLA — why this is *not* a fallback problem
-
-How **tcc** handles alloca (`arm64-gen.c:2305-2334`): `gen_vla_alloc` does
-`sub sp, sp, xN` at runtime; `gen_vla_sp_save/restore` spill/reload sp to a frame
-slot at scope boundaries. **Fixed locals stay x29-relative; sp floats.** x29 is a
-stable anchor that alloca cannot disturb.
-
-Fix B's base anchor (fp-at-bottom or `xFB`) **plays exactly tcc's x29 role** — it
-is set once and is unaffected when alloca lowers sp. The VLA memory is sp-relative
-(below the anchor); fixed-local addressing through the anchor is untouched. So
-alloca is **safe by construction, with no `CG_FUNC_HAS_ALLOCA` flag and no
-fallback path** — that worry was specific to a *bare-sp-relative* scheme.
-
-Caveat to settle: today's BOTTOM-RECORD excludes `has_alloca` *and* requires
-`out_stack == 0` (`native.c:1835`). The -O0 variant must support `out_stack > 0`
-(outgoing stack args coexist with positive slot addressing) and alloca. Cleanest
-first cut: choose the layout **at `aa_func_end`**, where `has_alloca` is finally
-known — patch slots positive (Fix B) when no alloca was seen, or fall back to the
-current negative top-record `sub` form when it was. The deferred-patch list makes
-this a per-function decision at resolve time, not an up-front guess.
-
-### 5.4 Open design points
-
-- **CFI.** fp-at-bottom makes CFA frame-size-dependent (`fp + N`); emit/patch the
- CFA rule in `aa_func_end`. The `xFB` variant avoids this entirely (fp stays at
- top). Gate on `test-dwarf`/`test-debug`.
-- **Sub-word reach.** Scaled imm12 reaches only 4095×size — 16 KB for 4-byte,
- 8 KB for 2-byte, 4 KB for 1-byte. Frames past those for *narrow* accesses keep
- the indexed/`sub` fallback. Rare; the 64-bit/32-bit common case covers nearly
- everything.
-- **Index registers (`[base, xN, lsl #s]`) and q-regs.** Verify the scaled forms
- for 128-bit (reaches 65,520 B) and the few indexed loads.
-- **`frame_size_final` plumbing.** It is "set in `aa_func_begin_known_frame` …
- unread on the single-pass path" (`native.c:214-217`). The single-pass path must
- now compute and store it before the patch loop in `aa_func_end`.
-
-### 5.5 Expected payoff & gate
-
-**−29,379 `sub` − ~4,000 frame `movk` ≈ −33 K insns (−7.3% text) → ~1.24× tcc**,
-concentrated in the giant functions (`VdbeExec` alone −5,162, ~22% of the
-function). Setup cost is ~1–3 insns only in big-frame functions, vastly
-out-weighed. Compile speed is **neutral-to-faster** (bounded O(far-accesses) patch
-list; ~33 K fewer instructions to emit).
-
-**Gate: run-correctness + determinism** (it changes emitted bytes): compile
-sqlite twice and `cmp` the objects (self-identical), `make test-toy test-parse-ok
-test-parse-err test-smoke-x64 test-smoke-rv64 test-dwarf test-debug`, the sqlite
-e2e (`84|2`), and a clang-differential probe over alloca/VLA + large-frame +
-&local-address-taken + outgoing-stack-arg functions.
-
----
-
-## 6. Lever 2 — epilogue slim (quick win)
-
-kit emits a **3-insn** epilogue using x16 as scratch:
-
-```
-add x16, x29, #0 ; copy fp to scratch
-ldp x29, x30, [x16] ; restore fp/lr
-add sp, x16, #16 ; sp = fp + 16
-ret
-```
-
-tcc emits **2 insns** (`arm64-gen.c:1668`):
-
-```
-mov sp, x29 ; sp = fp (recovers from any alloca too)
-ldp x29, x30, [sp], #16 ; restore fp/lr, post-increment sp
-ret
-```
-
-`mov sp, x29` then a post-indexed `ldp [sp],#16` needs no scratch and is correct
-under alloca (it restores sp from fp). **−2,633 insns (−0.6%).** Byte-safe modulo
-the intended change; investigate why the x16 form was chosen (likely incidental).
-Gate: run-correctness + determinism + the alloca probe above.
-
----
+**fp-vs-sp is resolved.** Both compilers home locals at **fp-relative** offsets
+(`[x29]`: kit 124 K mem-ops, tcc 87 K — tcc is *not* sp-relative). The old gap
+was how each reached a *far* local (past `stur`'s ±256): kit built the address
+(`sub xN,x29,#off; ldur`), tcc used register-indexed `movn x30,#off; ldr
+[x29,x30]`. L1 moved far slots to **positive-scaled off sp** (`ldr/str [sp,#+ofs]`,
+one insn, no address build) — so kit now does far addressing *better* than tcc
+(`movn` 0 vs 6,720; 28 K `[sp]` accesses are the migrated slots).
-## 7. Lever 3 — load-with-extend (quick win)
+**The remaining gap is register pressure, not addressing.** kit emits ~72 K
+stores vs tcc's ~25 K (3×) and 102 K loads vs 87 K: it keeps fewer live values in
+registers across statements than tcc's value stack, so it stores-and-reloads the
+frame far more. That excess (`mov` arg-setup + `stur`/`ldur`/`str`) is the
+compute-into-place / value-residency frontier (§4).
-kit emits **0** `ldrsb`/`ldrsh`; tcc emits 1,313. kit lowers a signed narrow load
-as `ldrb; sxtb` (two insns) where one `ldrsb` suffices. This was flagged blocked
-in `[[o0-copy-convert-coalescing]]` (needs a load-with-extend rider on the
-`MemAccess`, not a register rename). **~−1,300 insns.** Gate: run-correctness +
-determinism; signed/unsigned char/short load corpus.
+**Concentration.** The excess still clusters in the giant functions
+(`_sqlite3VdbeExec` etc.); per-function spill counts dominate there.
---
-## 8. Lever 4 — compute-into-place (the `mov`/spill frontier)
-
-The biggest *raw* counts — `mov` +70 K and stores +52 K — are reality (b): kit
-computes into a register, then moves/spills to place. This is the largest lever,
-and the one most worth understanding correctly, because the obvious framing
-("add a register allocator / a cleanup pass") is wrong and would indeed cost
-compile time. **The correct framing is the opposite: compute-into-place is the
-*absence* of eager relocation, not extra analysis — it emits fewer instructions
-and runs no extra pass. tcc is faster *because* of it.** So this lever is
-compile-time-neutral-to-*positive*; its real cost is correctness/invasiveness,
-not throughput.
-
-### 8.1 How tcc does it — lazy location + demand-driven registers
-
-tcc's `SValue` separates *what a value is* from *where it lives* (`tcc.h:479`). A
-value-stack entry stays unmaterialized — a constant, a memory ref
-(`VT_LOCAL|VT_LVAL`), CPU flags (`VT_CMP`), or a register — until something
-*forces* it. When forced, `gv(rc)` materializes it into a register **of the
-class the consumer asks for** (`tccgen.c:1844`); `get_reg(rc)` picks that
-destination *at the point of use* (`tccgen.c:1487`); the result lands there and
-`vtop->r` is updated to point at it. There is no IR, CFG, or second pass — the
-"analysis" is the recursion shape: the parent passes down the register it wants.
-Laziness is therefore *cheaper*, not costlier.
-
-### 8.2 The three mov classes — tcc vs kit (with emit sites)
-
-| class | excess | tcc | kit today |
-|---|--:|---|---|
-| **retval capture** | ~8,511 | result SValue records `r = x0`, emits nothing; next forcing use materializes it (`tccgen.c:6334`) | pre-allocates a cache reg, emits `mov cachereg,x0` after `bl` (`native_direct_target.c:1952,2010`) |
-| **arg setup** | ~19,000 | `gv(RC_R(n))` materializes each arg **directly into x0–x7** at the call (`arm64-gen.c:1187`) | args already sit in general cache regs; `native_arg_shuffle` moves them into x0–x7 |
-| **materialize chains** (`ldur;mov`) | ~18,000 | `gv` loads straight into the consumer's demanded reg | `nd_materialize_operand` loads into an *arbitrary* scratch, then `mov`s to the consumer's reg (`native_direct_target.c:832,954`) |
-
-### 8.3 Why kit pays — three missing handoffs, not missing analysis
-
-kit's NDT cache is *already* a lazy-location model, and its bookkeeping is
-already cheap (`nd_cache_alloc` O(≤8), `nd_touch` O(1), flush O(ncached)≤8). The
-movs come from three gaps:
-
-1. **ABI registers are excluded from the cache.** x0–x7 are non-allocable
- (`aa64/native.c:3843`), so a result can't *stay* in x0 and an arg can't be
- *produced into* x_k — both route through a general cache reg + a mov.
-2. **No destination-hint surface.** A producer always writes to a
- cache-allocator-chosen register; no consumer can say "produce this in
- register R." The only operand hint today is `OPK_FLAG_KILL` (`cgir.h:210`),
- which transfers ownership of a *dead source* but does not request a
- *destination*.
-3. **kit materializes args eagerly, before the call.** By the time `nd_call`
- runs, args are CGLocals already in cache regs → they must be shuffled. tcc
- materializes args *at* the call, while still unmaterialized, straight into
- arg regs.
-
-### 8.4 Phased plan — cheapest-first, each compile-neutral-or-faster
-
-Each phase adds a *preferred-register field*, not a pass; every mov removed is
-also one fewer instruction to emit (faster compile). Uniform gate:
-**run-correctness + determinism** (compile twice + `cmp`; full
-toy/parse/smoke/dwarf/debug; sqlite e2e `84|2`; clang-differential probes around
-calls/varargs/sret). The one real correctness surface is the **eviction
-discipline for ABI-register holders**.
-
-- **Phase 1 — adopt x0 as the result holder (~8,511 movs, the clean win).**
- Record "this result local *is in* x0" as a short-lived cache entry instead of
- allocating a cache reg + emitting `mov cachereg,x0`. The next call already
- evicts caller-saved cache regs — reuse that path to spill/move x0 only if it
- is still live at the next clobber. For `x = f()`, `f()+1`, `g(f())` (result
- consumed before the next call) it is **zero movs**. Self-contained, reuses the
- flush-at-call machinery, exactly tcc's "retval stays in x0." Correctness rule:
- evict the x0-holder before *any* op that clobbers x0 (calls, and x0-using
- lowering like div/varargs).
-- **Phase 2 — a destination-hint field on the operand (~18 K materialize
- chains).** Add an optional "preferred destination register" beside
- `OPK_FLAG_KILL`; when the consumer's wanted reg is known and free,
- `nd_materialize_operand` loads straight into it. The cache already supports
- "write into reg R" — only the hint needs to flow. This is the general form of
- tcc's `gv(rc)`.
-- **Phase 3 — args into arg-regs (~19 K, the structural one).** Thread the hint
- from the call *back* to the arg producers (the frontend knows, when generating
- arg-expr k, that its result feeds arg slot k — pass an abstract "destined for
- arg k" the arch resolves to x_k), or keep args unmaterialized until `nd_call`.
- Non-conflicting args (the majority) then produce directly into x_k;
- `native_arg_shuffle` still resolves the rare cross-arg cycles (tcc handles the
- same case by spilling conflicts to stack first). Biggest payoff, most
- plumbing — do last.
-
-Prior partial work this builds on: value-stack residency (register-resident
-return/args/result), lazy transient homes, lazy-dup + copy/convert coalescing
-(`[[o0-value-stack-residency]]`, `[[o0-copy-convert-coalescing]]`).
-
-History of landed codesize work (context for §1's 1.33×), newest last:
-
-- temp-slot reuse `3516b3d3`, addr-result caching `f75874cb`, subscript fusion
- `442fdc55`, transient-liveness + lazy-dup `ecbd00dc`/`73e128a2` — took the
- object 3.45×→1.57× tcc (`[[o0-codesize-vs-tcc]]`).
-- value-stack residency (register-resident return/args/result + spill-only-live-
- set), lazy transient frame homes, copy+convert coalescing — took emitted insns
- ~539 K→455 K and the ratio 1.57×→1.33× (`PERF-TCC-SLIM.md` §10c/§10d,
- `[[o0-value-stack-residency]]`, `[[o0-lazy-transient-homes]]`).
+## 4. Remaining levers (ranked)
+
+| # | Lever | Est. saving | Risk | §|
+|---|---|--:|---|---|
+| **1** | Args materialized into arg regs (L4 Phase 3) | ~30 K `mov` | high (frontend/CG-API + arg-shuffle correctness) | §4.1 |
+| **2** | Reduce frame spilling / extend value residency | up to ~50 K (stores+reloads) | high (register-allocation reach) | §4.2 |
+| **3** | byte/half far slots → positive-scaled (extend L1) | ~1,940 `sub` + some `movk` | med (overflow fallback) | §4.3 |
+| **4** | Signed load-with-extend (`ldrsb`/`ldrsh`) | ~1,242 `sxtb`/`sxth` | med (frontend widening load) | §4.4 |
+| **5** | rv64 L4-P1 (flip `ndt_result_reg_stable`) | rv64 only | low (on-target check) | §4.5 |
+
+Levers 1–2 are the bulk of what's left and both reduce to one structural change:
+**demand-driven value placement** — keep a value where its consumer wants it,
+materialized at the point of use, instead of producing into a cache-allocator
+register and relocating. tcc gets this from its lazy `SValue`/`gv(rc)` model
+(`tcc.h:479`, `tccgen.c:1844`); kit's NDT cache is already a lazy-location model
+but is missing the destination-hint handoff (below). Re-profile after each
+landing — the highest lever moves as work is removed.
+
+### 4.1 Lever 1 — args into arg-regs (L4 Phase 3)
+
+The `mov` excess is now dominated by **arg setup**: ~30 K `mov xK, reg` moving an
+argument from a general cache register into x0–x7 at the call. The arg's producer
+ran *before* the call, so it wrote a cache-allocator register; `native_arg_shuffle`
+then moves it into the ABI arg register. There is **no backend-only fix** — a
+cached arg costs exactly one `mov`, and forcing it through the home instead
+(load-into-xK) trades the `mov` for a spill+reload (worse).
+
+The fix is to **produce the arg directly in xK**, which needs the call's
+arg-register target threaded *back* to the arg-expression producer. The frontend
+knows, when generating arg-expr *k* of `f(a0,a1,…)`, that its result feeds arg
+slot *k*; the CG API has no way to express "destined for arg *k*" today. Add that
+hint (an abstract arg-slot the arch resolves to xK), have the NDT allocate xK as
+the producer's destination when free, and let `native_arg_shuffle` resolve the
+rare cross-arg cycles (tcc spills conflicts to the stack first). Biggest single
+remaining win; most plumbing (frontend + CG API + backend); highest correctness
+surface (arg marshalling). The general destination-hint mechanism it introduces
+also subsumes the residual materialize/value-stack `mov`s.
+
+### 4.2 Lever 2 — reduce frame spilling / extend value residency
+
+The deepest structural gap: kit spills ~3× as much as tcc (§3). The NDT register
+cache flushes the live-across set at every call and holds a bounded set across
+statements; tcc's value stack keeps more operands unmaterialized and threads them
+through. Closing this is the path to ~1.0× parity, and it overlaps Lever 1 (both
+are "place the value where it's used, don't relocate"). It is the highest-reach
+change to the -O0 register model and should follow Lever 1's destination-hint
+surface. Measure the spill sites in the giant functions first to scope it.
+
+### 4.3 Lever 3 — byte/half far slots (extend L1)
+
+L1 moved only **4/8-byte** far slots to positive-scaled `[sp,#ofs]`. **1/2-byte**
+far slots (~1,940 `sub xN,x29` + `movk`, see §3) still build the address, because
+scaled byte/half reach (4 KB/8 KB) is below the largest frame (`VdbeExec`
+6,544 B) so a one-word patch is not *guaranteed* to fit. Extend the
+`AA_PATCH_SLOT` resolver to byte/half with an **overflow fallback** (when the
+scaled offset exceeds imm12 for that size, keep the `sub`/indexed form) and lift
+the `sz ∈ {2,3}` gate in `aa_emit_mem`. Self-contained aa64 change; the fallback
+is the only new code. (Also in this bucket but low-value: the ~4 K
+address-of-local `sub xN,x29` from `aa_load_addr`, which would become `add
+xN,sp,#ofs` — count-neutral for the common ≤4095 frame, only saving the `movk`
+in big frames.)
+
+### 4.4 Lever 4 — signed load-with-extend
+
+A signed narrow load is still `ldrb/ldrh ; sxtb/sxth` (two insns) where `ldrsb`/
+`ldrsh` would do one. L3 captured only the *unsigned* half: a zero-extending load
+already fills the register, so the following `uxtb`/`uxth` is provably redundant
+and is elided. The signed case can't be done the same way — **kit's CG integer
+types are sign-agnostic** (signedness lives in the convert op, not the type), so
+`nd_load` can't know the value feeds a sign-extend and must emit the
+zero-extending `ldrb`. The fix is a **frontend widening signed load**: emit the
+load at the promoted width with a signedness rider on `MemAccess`, so `aa_emit_mem`
+picks `ldrsb`/`ldrsh` and no separate convert is generated. Frontend + CG-API +
+encoder; ~1,242 insns.
+
+### 4.5 Lever 5 — rv64 L4-P1
+
+riscv64's `a0` is clobbered only at call boundaries (div/rem/mul take explicit
+operands), so result-in-result-register caching is sound there — flip
+`ndt_result_reg_stable` for rv64 after an **on-target div-after-result check**
+(build the p1-x64 probe for rv64 and run it through the emulator / a rv64 host).
+Small, low-risk, rv64-only win.
---
-## 9. Gate & measurement methodology
+## 5. Gate & measurement methodology
- **Measure on RELEASE** (`make bin RELEASE=1`); the ASan default inverts costs.
- **Size metric = `.text` machine code**, not object bytes (format-skewed). Use
- `kit size` / the §1 histogram diff. The instruction metric is deterministic to
- ±0.2 M on a fixed binary — small deltas are real.
+ `kit size` / the §1 histogram diff. The instruction metric is deterministic on
+ a fixed binary — small deltas are real.
- **Codesize changes alter emitted bytes**, so the gate is **run-correctness +
- determinism** (compile twice + `cmp`; full toy/parse/smoke/dwarf/debug; sqlite
- e2e `84|2`; clang-differential probes), *not* the byte-identity gate used for
+ determinism** (compile sqlite twice + `cmp` for self-identity; `make test-toy
+ test-parse-ok test-parse-err test-smoke-x64 test-smoke-rv64 test-dwarf
+ test-debug`; the sqlite ecosystem e2e at O0+O1, golden + vs-clang; and
+ clang-differential probes), *not* the byte-identity gate used for
compile-speed-only refactors.
+- **Verify every shared-NDT change on x64 and rv64, not just aa64** — the L4-P1
+ x86-64 RAX miscompile (now gated) is the cautionary tale: a shared change can
+ be a win on the reference arch and a silent miscompile elsewhere. Keep a
+ clang-differential probe per arch for any register-residency change.
- Re-profile after each landing and re-rank §4 — the highest lever moves as work
is removed.