Files
LithosAnanake/FABRIC-3.5.md
T
Claude 3a4e5cff07 FABRIC-3.5.md §XLIII: B3 settled -- the target drains its own queue at its own checkpoint
The last genuine design question in the reshuffle, and once again the
mechanism already exists, built for the structurally identical problem
and corrected twice in the field.

Delivery must execute the payload inside the target, but §XXXII retired
the pump so kernel-Hermes may not VM-EXEC into anyone. The payload must
wait somewhere the target consumes under its own power, at a moment when
doing so is safe -- and safe is the hard part, since interpreting
mid-word, mid-unwind or nested inside someone else's dispatch is the
reentrancy class this codebase has been bitten by repeatedly.

vm_core.c already has all of it: sk_vm_at_outermost_interpret() as the
predicate, a per-word checkpoint in execute_colon_word() gated on it, a
defer-don't-lose discipline for the nested case whose own comment
explains that a nested word does not own the stack it is running on, and
placement before the error and unwind checks so it only acts when the VM
is in a clean resumable state. That is the exact predicate, placement and
deferral semantics delivery needs, because the switcher had to answer the
same question.

Ruling: kernel-Hermes enqueues and publishes the fact, dispatching
nothing; the target drains its own queue in its own context at its own
outermost checkpoint. One published fact, two independent consumers --
the switcher for eligibility, the VM itself for drain -- and neither
dispatches into anyone.

Notes this is §XX's proven pattern generalized: the pump's defect was
never draining but draining from outside, and Hera was special-cased into
safety. Retire the pump and every VM does what she already does, so the
special case disappears by becoming universal. Also notes recursive drain
is prevented for free, since interpreting increments the same depth
counter that defines the boundary.

Names three constraints rather than leaving them to be discovered:
INPUT_BUFFER_SIZE is 1025 so a payload above 1024 bytes cannot be
interpreted in one drain, starvation is real but is the switcher's
existing risk rather than a new one, and draining one message per
checkpoint rather than the whole queue keeps the work bounded.

FABRIC-3.6.md: B3 cleared, B4 added for the payload bound.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VkM1zHGvBerLF6aqkHPweP
2026-09-19 12:18:46 +00:00

269 KiB
Raw Blame History

FABRIC-3.5.md — the Tripod/kernel reshuffle

Status: REOPENED 2026-09-19, by direct instruction ("reopen 3.5 and write it up as a gap analysis section"), to add §XXXI — a gap-analysis sweep of the whole FABRIC set. The design phase remains closed; §XXXI adds findings, not new design.

Prior status, kept rather than overwritten: DESIGN PHASE CLOSED as of 2026-09-19, by direct instruction ("close the document for now"). Not yet archival. That close stood for the duration of the sweep and its substance is unchanged — every design question was and remains ruled. The reopen is recorded rather than the close deleted, per this series' own rule against silently rewriting a prior state.

What is closed, and what is deliberately not. Every design question this document opened is ruled — see §XXX.7. No code has been written and no code is authorized. The execution sequence (§XXVI.6, as amended by §XXX.7) is untouched: the surgical strip, the build, the Isabelle/HOL pass, the documentation sweep, the SBOM, the merge to master, and the v2.1.0 tag all remain ahead.

This is therefore a narrower close than §XXVI.5 specified, and the difference is stated rather than glossed. §XXVI.5 ruled that this document closes when the tag exists, in FABRIC-2.md's CLOSED/ARCHIVAL form, naming the tag it closed at. The tag does not exist, so claiming that close would be a label running ahead of the real state — the exact error FABRIC-3.md §I.2 corrected when it rolled LITHOS_VERSION back, and the same standard §XXIX applied to the LTS question. The archival close specified by §XXVI.5 still stands and still happens at v2.1.0. This header is the provisional one the instruction's own "for now" asks for.

Nothing is carried forward and no successor is created. As §XXVI.5 establishes, this document is not part of the FABRIC-0 → -1 → -2 → -3 chain, so closing it hands nothing to a successor. FABRIC-3.md remains open, living and authoritative for its own topic (bare metal boot); nothing here supersedes it. The live artifact from this document is its punch list (§XXVI.6 / §XXX.7) — that, not this header, is what the build works from.

Open by design, not by omission: items 10–17 of §XXVI.6 are execution, not decisions. Three things recorded here are explicitly outside this reshuffle and still need their own authorization — the .claude/CLAUDE.md/MANIFEST.md documentation reconciliation (§XXVI.1), the src/*.c.bak hygiene question (§XXII.5), and the stray refs/heads/v2.0.1 branch and PR #1 (§XXVI.4).


Topic and original framing, as opened. Standalone working document, opened 2026-09-18 by direct instruction ("we write this in FABRIC-3.5.md as its own document, it's this important"). Topic: relocating Hermes into the kernel, reconstituting the Tripod as Hera/Artemis/Hestia, generalizing BIRTH, and adopting one fleet-wide failure/recovery ladder.

Why 3.5 and not 4, and why not a section of 3. FABRIC-4.md is explicitly the lower-discipline scratchpad for theory-stage ideas "caught early... often missing a stated 'why' on purpose." This is not that: it is a ratified architectural direction with a stated shape, and it needs the full discipline every numbered FABRIC document carries. But it is also not bare-metal boot, which is FABRIC-3.md's declared topic, and it is not a successor to FABRIC-3.md — FABRIC-3.md remains open, living, and authoritative for its own topic; nothing here supersedes it, and this document does not trigger the close-and-carry-forward discipline the FABRIC-0 → -1 → -2 → -3 chain used (that chain triggers on closing a document). It sits at 3.5 because it is a peer in rigor and an outsider in topic.

How to use this document. Same discipline as the rest of the series: write the decision and its reasoning down before building, close items with a dated note citing real evidence, never silently drop a stale claim. New findings and decisions for the reshuffle get added here, not to FABRIC-3.md.

Scope discipline, stated up front: this is a reorganization, not an invention. Per the instruction opening it — "most of the underlying logic (Hermes routing, Artemis block I/O, Hera lifecycle/ACL, Console framebuffer) already exists and is expected to be relocated and rewired rather than rewritten." Any item in this document that turns out to need a genuinely new algorithm is by that fact out of scope and gets flagged, not quietly built.

No code is authorized by this document. Design only. Per Captain Bob's Law (.claude/CLAUDE.md): never apply a fix not explicitly requested.

Provenance convention used below. This series distinguishes what was read from what was recalled, so each claim says which. "Traced 2026-09-18" = the file was read directly while writing this document. "Per §X" = the fact is reported by an existing FABRIC section that did its own end-to-end read; it is cited, not re-verified here. "Open" = not decided. Several items below are deliberately marked for re-verification before any code is written, because a reorganization that trusts a stale map moves the wrong things.


I. The current shape, traced rather than recalled

Before moving anything, what is actually there today. All of §I traced 2026-09-18 against the working tree at commit e56974e, except where marked.

I.1 — Hermes is a FORTH capsule VM, and the routing layer is FORTH, not C

capsules/hermes/init.4th is short and load-bearing. Its block 5116 says it outright:

Hermes owns the one real, canonical COMMON-CH. Every other VM subscribes into it via VM-EXEC (see Artemis's and Hera's own init.4th); Hera always exists first, so Hermes adds her here rather than Hera subscribing to a VM that doesn't exist yet at Hera's own birth.

It then does S" common:messaging.4th" EXEC, MSG-CD-INIT, 1 MY-CH-ID !, and adds members 1 and 0 to COMMON-CH. That is the whole of Hermes's own identity: it is the VM that happens to hold the canonical channel.

The mechanism it holds is capsules/common/messaging.4th (504 lines) — generic per-VM messaging vocabulary that every participating VM loads its own copy of. This matters more than it first appears. Block 5005 allocates the arenas with CREATE ... ALLOT:

CREATE MSG-ARENA MSG-MAX MSG-CELLS * CELLS ALLOT
CREATE CH-ARENA  CH-MAX  CH-CELLS  * CELLS ALLOT
CREATE MBR-ARENA MBR-MAX MBR-CELLS * CELLS ALLOT

CREATE/ALLOT carves out of the dictionary of whichever VM is interpreting. So today there is no single routing fabric — there are N per-VM copies of the same vocabulary, and Hermes is distinguished only by convention (it holds COMMON-CH; it is routing table index 1). The arbitration is distributed and conventional, not centralized and structural.

I.2 — Hermes is a hard, by-name dependency of its peers

CORRECTED 2026-09-18 by §XIII.1 — this section undercounts. It names two call sites; the full repo-wide inventory is 14 by-name FORTH references across 5 files, plus 3 C-side sites and a DoE CSV schema site. The two below are real, are among the only three live ones, and remain the sharpest edge — but read §XIII.1 for the actual extent. Left in place rather than rewritten, per this series' own "never silently drop a stale claim" rule.

Artemis does not talk to an abstract routing layer. It talks to a VM called Hermes, by name, via VM-EXEC. Two live call sites in capsules/artemis/init.4th:

  • line ~138: S" 2 ENQUEUE-READY" S" Hermes" VM-EXEC
  • line ~501: S" 2 COMMON-CH @ CH-ADD-MBR" S" Hermes" VM-EXEC

Both execute FORTH text inside Hermes's own dictionary. Either one breaks the moment Hermes is not a VM with a dictionary. This is the single most concrete consequence of the move and §III.3 is about it.

I.3 — The routing table pins 0/1/2 and says so

messaging.4th block 5006, verbatim comment:

VM routing table. §XIX: 16 slots, not 8 — see FABRIC-3.md. Hera/Hermes/Artemis stay 0/1/2 (artemis:init.4th depends on it); identities get NEW slots 3–10, never renumbered.

VM-NAMES-INIT registers Hera 0, Hermes 1, Artemis 2, then identities from 3 up via DOE-IDX>MSG-IDX ( doe-idx -- msg-idx ) 2 +. "Never renumbered" is an existing, written commitment, and the reshuffle walks straight into it (§IV.2).

I.4 — Message types in use, and the free numbers

Traced across messaging.4th: SPAWN-EVENT 1, PAUSE-EVENT 2, RESUME-EVENT 3, KILL-EVENT 4, CONSOLE-CMD-EVENT 7, ELEVATE-REQUEST 8, BLK-ATTACH-EVENT 9. Reserved status sentinels at the top of the range: MSG-NACKED 253, MSG-DELIVERED 255.

5 and 6 are unallocated gaps; 10 is the next in sequence. Relevant to §VII, which needs a number for SOS. Also relevant: per §XX, SPAWN-EVENT was grepped for every consumer across the tree and has none — "an unwired placeholder, not working infrastructure." Do not model SOS on it as though it were a working precedent.

I.5 — BIRTH is already VM-agnostic in C; Hera's monopoly is registration-time

Per §XX, which read mama_word_birth (mama_forth_words.c:229, BIRTH's C implementation) in full:

It is genuinely VM-agnostic — it operates on whichever vm calls it, with no check that the caller is Hera specifically. The only Hera-specific logic anywhere in it is the reverse guard preventing Hera from re-birthing herself. capsule_birth_baby() treats vm->stadium_vm_id generically as "whoever is birthing this VM." So "any VM could theoretically birth another" is already true in the code... Hera's practical monopoly on BIRTH is a registration-time fact (only register_mama_forth_words() registers it), not a check inside BIRTH itself.

This is the single most important pre-existing fact in this document. §VI is not a redesign of BIRTH; it is a change to which VMs get it registered, plus an inheritance rule.

I.6 — Hestia already has a real bind mechanism, built and live-verified

Per §XXXII.2, closed 2026-09-16: CONSOLE-ATTACH ( name-c name-u -- ok? ) exists, is a plain unconditional primitive (deliberately not an identity capability bit — that was tried and rejected on two counts), and was verified end to end on amd64: UNATTENDED-BIRTH a VM named bob → CONSOLE-ATTACH bob → S" bob" USE → typing 5 6 + . at the new console relayed to and executed on the target identity VM, printing [zuse@bob~user] 11 ok>.

Also live: capsule_console_birth() (minimal relay proxy, bare username) paired to capsule_runcap_birth() (the real identity, <username>~user) by name convention plus one VM-NAME-REG executed inside the console VM's own dictionary — confirmed per-console-VM, not a global singleton, so after-the-fact pairing already works structurally.

§V is therefore mostly a promotion of an existing, proven mechanism to an architectural role — not new construction.

I.7 — The headless-until-login invariant

Per §VIII.1 (ratified) and §XXXII.2: g_wirebind_attached_username is what sk_console_identity_present() reads (capsule_wirebind.c:321-327), and that function is the live "no thumbdrive, no prompt" gate. §XXXII.2 states the rule as an explicit invariant for any new birth path: never set that global, never touch the console-pairing mechanism, or you bring up a console with no human present and reopen a gate that was deliberately closed. §V and §VI both have to carry this invariant forward.

I.8 — The idle pump, and why Hera is special-cased in it

Per §XX: repl.c's idle loop walks the VM registry every idle beat, VM-EXECing "MSG-TICK" into every other live VM to drain its queue, skipping Hera — because Hera is the pump and self-targeting VM-EXEC would hit a known reentrancy class. Hera instead gets a direct MSG-TICK word dispatch in her own context, guarded by a fresh vm_find_word + acl_allow check every tick rather than a cached flag.

So there is already a kernel-resident component driving the messaging layer's clock. The pump is in repl.c, in C, today. §III is in part an admission that the arbiter's center of gravity is already partly there.


II. What the reshuffle is, in one paragraph

Hermes stops being a peer VM on the Stadium floor and becomes a kernel-resident arbiter between the floor and the HAL. It stops being a client of ACL-checked, Hermes-routed messaging and becomes the routing and arbitration layer itself. The Tripod — which was Hera, Hermes, Artemis — is reconstituted as Hera, Artemis, Hestia, with Artemis keeping its existing storage/block-arbitration role and Hestia taking the vacated third slot. Hestia's own role widens from owning the drawing fabric to being the bind point where users and agent VMs (the GPIO VM of FABRIC-4.md §2, future networking VMs) attach. BIRTH becomes a general primitive any VM with standing may invoke, with authority bounded by inheritance rather than by a separate check. And every VM in the fleet — current and future — adopts one escalation ladder on failure: recover gently, escalate to a from-scratch birth, then fail fast and total.


III. Hermes moves into the kernel

III.1 — Decided (Captain Bob, 2026-09-18)

Hermes is no longer a Tripod VM. It becomes kernel-resident, sitting between the Stadium floor and the HAL. It no longer participates in Hermes-routed, ACL-checked messaging the way the other VMs do — it is the routing/arbitration layer now, not a client of it.

III.2 — What that actually changes, structurally

The self-reference in "Hermes cannot be a client of Hermes" is the whole point, and it has a concrete reading against §I.1: today Hermes holds COMMON-CH inside a FORTH dictionary that is itself subject to the same ACL and Stadium-heat economy as every other VM's. An arbiter whose own arbitration state can be evicted, cooled, or ACL-denied by the mechanism it arbitrates is a circular dependency that happens not to have bitten yet. Moving it into the kernel breaks the circle by construction.

Note this is also what makes §VII's SOS exception (Hermes cannot route a message announcing its own routing failure) not a special case but a direct corollary — see §VIII.

III.3 — The by-name VM-EXEC dependency has to go somewhere. Open.

SUPERSEDED 2026-09-18 by §XIV.1/§XIV.2 — this question dissolved. The three options below all assume the existing FORTH callers must be preserved. Captain Bob ruled the FORTH messaging layer legacy, so §XIII.1's inventory becomes input to a dead-FORTH cleanup rather than a migration constraint. Left in place, not rewritten.

§I.2's two Artemis call sites are the sharp edge. S" ..." S" Hermes" VM-EXEC requires a target VM with a dictionary. Three shapes are visible; none is chosen here:

  1. Kernel primitives replace the by-name calls. ENQUEUE-READY and CH-ADD-MBR become registered C words callable directly by any VM, and Artemis simply calls them. Closest to "relocated, not rewritten"; largest registration-surface change.
  2. A vestigial Hermes name that the kernel answers. VM-EXEC to Hermes is intercepted and serviced by the kernel arbiter, leaving caller text untouched. Smallest diff, but it preserves a fiction and this project has a standing distaste for exactly that (§XXXII.2 Q3's rejection of an unreachable gate on doctrine grounds, not just mechanics).
  3. Artemis subscribes through a new kernel-side registration call at birth, removing the peer-to-peer step entirely.

Decide this on paper before touching code. Whichever wins, the two Artemis lines and the STARTUP-BANNER/LOG-INFO" scaffolding in capsules/hermes/init.4th are the concrete edit set, plus whatever else a full grep for S" Hermes" VM-EXEC turns up — that grep has not been run exhaustively for this document and must be, repo-wide, before scoping.

III.4 — The per-VM-arena question. Open, and the biggest unknown here.

RESOLVED 2026-09-18 by §XIV.1/§XIV.2. Answered in the "full centralization" direction, but without the migration cost this section feared: with the FORTH legacy there is no live state to port. This was named the root dependency throughout §X; it is closed. Left in place, not rewritten.

Per §I.1 the arenas are per-VM dictionary allocations. If arbitration centralizes into the kernel, does MSG-ARENA/CH-ARENA/MBR-ARENA centralize with it, or does each VM keep its own outbox/inbox with only the routing centralized? These are very different amounts of work and very different risk:

  • Routing-only centralization keeps messaging.4th largely intact, keeps the Stadium heat economy's "sender pays" property (per §XX: STADIUM-RES-PULL/-PUSH operate on the calling VM's own reservoir), and is genuinely a reshuffle.
  • Full arena centralization moves FORTH dictionary state into kernel C, changes every VM's dict_hash, and needs the three-arch parity story re-established from scratch.

Not decided. Flagged here because the executive framing ("relocated and rewired rather than rewritten") holds comfortably for the first and is strained by the second.

III.5 — Parity consequence, stated so it is not discovered late

Any change to what register_*_words() registers, or to what init.4th loads, changes VM dictionary contents and therefore dict_hash. §XX's own verification standard applies: the property that matters is identical dict_hash across amd64/aarch64/riscv64, not an unchanging absolute value. Expect the absolute hashes to move; require the cross-arch identity to hold. Acceptance is the three-arch QEMU boot to zuse)ok> with zero UNKNOWN WORD faults, per .claude/CLAUDE.md's non-negotiable criteria — there is no other test.


IV. The Tripod reconstituted: Hera, Artemis, Hestia

IV.1 — Decided (Captain Bob, 2026-09-18)

Tripod = Hera, Artemis, Hestia. Artemis keeps its existing role (storage/block arbitration) unchanged. Hestia takes the slot Hermes vacates.

IV.2 — The "never renumbered" collision. Open, needs a ruling.

DISSOLVED 2026-09-18 by §XIV.1/§XIV.2 — via option 3 below, exactly as this section predicted. The FORTH routing table is legacy, so the slot-numbering question has no subject. Option 3's own caveat ("settle §III.4 first") was the correct call. Left in place, not rewritten.

§I.3's comment commits to Hera/Hermes/Artemis at 0/1/2 and to identities at 3–10 never being renumbered. Hermes leaving index 1 forces a choice, and the existing comment forecloses the laziest option:

  1. Hestia takes index 1. Tidy, preserves the "Tripod occupies 0/1/2" shape, and requires no identity renumbering — but it silently redefines what slot 1 means, and artemis:init.4th's dependence on the numbering is on 2, not 1, so it may be survivable.
  2. Index 1 is retired; Hestia takes a fresh slot. Honest, costs a slot out of 16, and leaves a permanent hole that needs a comment explaining itself forever.
  3. The routing table stops being a FORTH array at all, because §III moved routing into the kernel — in which case this question dissolves into §III.4's arena question and should not be answered separately.

Not decided. Note option 3 makes 1 and 2 moot, so settle §III.4 first. This ordering matters: answering IV.2 before III.4 risks ratifying a table that the kernel move deletes.


V. Hestia's role expands: the bind point

V.1 — Decided (Captain Bob, 2026-09-18)

Beyond owning the drawing fabric, Hestia becomes the bind point for users and agent VMs — GPIO VM, future networking VMs, and whatever follows. The attach point that was previously implicit or undecided now lives here, explicitly.

V.2 — This is mostly promotion of an existing mechanism, not new construction

Per §I.6, CONSOLE-ATTACH already resolves a target's liveness by name before birthing a console (so a typo refuses cleanly with no orphaned VM — verified live), and console/user pairing already works after the fact. The reshuffle's contribution is declaring that this is the fleet's attach point, and then asking the question the current mechanism does not yet answer: it was built for human at a console attaching to an identity VM. An agent VM — a GPIO VM per FABRIC-4.md §2, pinned and born at boot with no human anywhere near it — is a different shape wearing the same word.

V.3 — Genuinely open

  1. Does an agent VM bind through CONSOLE-ATTACH or through a sibling word? sk_repl_dispatch_line()'s pairing check reconstructs the target as console_get_vm_name() + "~user" (per §XXXII.2, and the cause of a real bug caught live when a console was named anything else). A GPIO VM is not a ~user identity. Either the naming convention widens or agent binding takes its own path.
  2. Does binding an agent VM touch g_wirebind_attached_username? It must not — §I.7. State this as an explicit invariant in whatever implements it, exactly as §XXXII.2 required of unattended birth.
  3. Is Hestia's bind role ACL-gated, and if so where? Per .claude/CLAUDE.md's hard rule and §XXXII.2 Q3's correction: policy belongs in ACL.4th, never in C. The shape is ' CONSOLE-ATTACH ACL-PIN (or a sibling word's equivalent) in ACL.4th, not a C-side capability check — and specifically not a vm_identity_has_cap() gate, which §XXXII.2 proved unreachable because identity.installed is 0 for Hera/Hermes/Artemis and for every console-proxy VM.
  4. Does Hestia-as-bind-point survive Hestia being a Tripod member? A Tripod VM is pinned and born at boot (session_register()/session_set_pinned(), per FABRIC-2.md §H.12 step 4; "pinned sessions never leave," §H.1 — cited from .claude/CLAUDE.md's and §XX's references, not re-read for this document; verify before relying on it). Whether the drawing-fabric owner and the bind-point arbiter are the same VM or two roles that happen to share a name is not settled here.

VI. BIRTH as a general primitive

VI.1 — Decided (Captain Bob, 2026-09-18)

BIRTH is a general primitive, not Hera-exclusive. Any VM with standing can invoke it. A birthed VM inherits its ACL/DNA from the birthing VM. A VM can never birth something with more authority than it itself holds — enforced by inheritance, not by a separate authorization check.

VI.2 — Why this is small in C and large in policy

Per §I.5 the C implementation is already VM-agnostic; the monopoly is which registration function hands out the word. So the mechanical change is registration, and the design work is entirely in the inheritance rule.

The rule as stated is elegant precisely because it needs no gate: if a child's authority is derived from the parent's, "cannot exceed the parent" is not a check that can be bypassed, forgotten, or misconfigured — it is a property of how the child is constructed. This is the same doctrine §XXXII.2 Q3 landed on from the opposite direction (a bespoke gate that can never open is worse than no gate), and the same doctrine ZUSE-ELIGIBILITY-ADD already carries.

VI.3 — "ACL/DNA" needs a precise referent. Open, and this is the real work.

The system has a word-level ACL: four DictEntry fields (acl_ttl, acl_allow, acl_mode, acl_pinned), a C primitive ACL-INHERIT (traced 2026-09-18 in capsules/ACL.4th's own primitive list, line 6), and the semantics "pin is one-way; inheritance clears pin, copies mode" (per .claude/CLAUDE.md). What §VI.1 describes is VM-level inheritance — a different axis. Two things must be settled before any code:

  1. Is VM-level DNA just "the child's dictionary is built from the parent's, with per-word state carried by the existing ACL-INHERIT"? If yes, this is genuinely a reshuffle and the existing primitive does the work. If no, something new is being invented and that contradicts this document's own scope discipline — flag it rather than build it.
  2. Hard constraint: .claude/CLAUDE.md states "Never add acl_* fields to DictEntry beyond the four already present." Any DNA design that wants a fifth per-word field is ruled out at the door. If VM-level authority needs state, it belongs on the VM, not on every dictionary entry.

VI.4 — "With standing" is undefined. Open.

§VI.1 says any VM with standing. Standing is not currently a concept in this codebase — traced: no such notion appears in the ACL vocabulary or the Stadium words. Candidate readings: (a) standing = simply having BIRTH registered, which collapses it into §VI.2's registration question and makes the phrase decorative; (b) standing = a Stadium-economy property (sufficient reservoir to pay for a child, consistent with "sender pays"); (c) standing = an ACL.4th policy predicate. Not decided. (a) is the minimal reading and the one most consistent with "do not over-engineer"; (b) is the one most consistent with the rest of the fleet's physics.

VI.5 — Invariants that must survive

  • The reverse guard preventing Hera from re-birthing herself (per §I.5) still has to hold, and now has to hold for every birther with respect to itself.
  • capsule_birth_baby() must stay the generic path — per §XXXII.2 it is already what every (p) capsule uses across four call sites, and capsule_runcap_birth() was explicitly established as the wrong tool for build-time-sourced births.
  • Unattended/agent births must not set g_wirebind_attached_username (§I.7).
  • ' BIRTH must not enter capsules/ACL.4th — per .claude/CLAUDE.md, ACL.4th is shared and portable, BIRTH is kernel-only, and the file deliberately omits it (comment at ~line 64). Generalizing BIRTH does not change this; pin it in a kernel-specific capsule.

VII. One fleet-wide failure/recovery ladder

VII.1 — Decided (Captain Bob, 2026-09-18)

A single escalation shape, applied consistently across the fleet — current VMs and future ones (GPIO, networking, whatever comes):

  1. Recover gently first. For a VM with continuity to preserve (Hera being the type case), this means warm restart — resume with prior state intact.
  2. Escalate if gentle recovery fails. Fall back to a from-scratch BIRTH — no continuity, rebuild fleet-state/trust from nothing.
  3. Fail brutally once genuinely exhausted. No lingering, no partial states. Once recovery options are spent the failure is fast and total, not a slow degrade.

The value here is uniformity: one shape for Hera, Hermes, Hestia and every future VM, instead of bespoke handling per component.

VII.2 — SOS as a standard message type

Decided: SOS is a standard message type any VM can emit — not specific to any one component. It applies to non-Tripod VMs and to two of the three Tripod VMs (Hera and Hestia). Hermes is the exception, for a structural reason — §VIII.

Grounding and open points:

  • It takes a number from §I.4's space. 5 and 6 are unallocated gaps; 10 is next in sequence. Picking a gap vs. appending is a small call but should be made deliberately, not by whoever types first.
  • Do not model it on SPAWN-EVENT. Per §XX that constant has zero consumers anywhere in the tree — it is an unwired placeholder. A message type with no dispatcher is exactly the failure mode SOS cannot afford.
  • Open: who consumes an SOS, and what do they do with it? The executive framing says SOS is emitted; it does not say who acts. Given §IX (no VM assumes another's authority), the consumer set is constrained but not specified.

VII.3 — The existing precedent this should be reconciled with

WITHDRAWN 2026-09-18 by §XV.2 — this section's premise is wrong. It assumed the ladder needs a detection mechanism. Death is compudynamic: no detector, no declarer, nothing to build. The closing claim below ("a recovery ladder with no detection story only ever triggers on failures loud enough to notice by accident") does not hold. The wellness-check idea stands on its own merits, unrelated to this ladder. Left in place, not rewritten.

Per §XX, closed as captured-for-later: Captain Bob's self-healing idea — "all VMs participating in a wellness check at their nearest participating wellness center" — a distributed liveness/failure-detection concept "deliberately not built as part of this fix."

That idea and this ladder are the same problem approached from two ends: wellness checks are how a failure gets detected; the ladder is what happens after. Neither document has reconciled them. Flagged as open, and as the natural next thing to settle, because a recovery ladder with no detection story only ever triggers on failures loud enough to notice by accident.

VII.4 — Open: what "warm restart" concretely means

Warm restart implies a state boundary — what is "prior state" for a VM, and where does it live across the restart? Candidate anchors traced or cited: the Stadium reservoir/heat state, the VM's dictionary, its message arenas (§III.4 — if those centralize into the kernel, warm restart gets easier, which is an argument to settle §III.4 first), its identity struct. Not designed here. Note also that "fail brutally" must not mean "leave the Stadium's conservation invariant K broken" — per FABRIC-3.md §XVIII that invariant is continuously verified, so a brutal death still has to return what it held. Cited from §XVIII's heading, not re-read for this document; verify before relying on it.


VIII. Hermes's exception: the sinking semaphore

VIII.1 — Decided (Captain Bob, 2026-09-18)

Hermes cannot emit SOS, because Hermes is the message arbiter and it cannot route a message announcing its own routing failure. Instead of emitting SOS, Hermes raises a semaphore signaling that it is sinking, and holds it until it goes down. This gives the rest of the fleet a window to shut down cleanly rather than going dark with no warning.

VIII.2 — Why this is a corollary, not a special case

This is the direct consequence of §III.2. The instant Hermes stops being a client of its own routing layer, any failure-announcement it might send has no transport — the transport is the thing that failed. A semaphore is the right shape precisely because it is not a message: it does not require the failed subsystem to work, and a held-until-death signal degrades correctly (release = gone) rather than requiring a successful final transmission.

Worth stating plainly because it reframes the whole design: the fleet's failure signalling is two mechanisms, not one, and which one a VM uses is determined by whether it is above or below the routing layer. Everything on the Stadium floor sends SOS. The arbiter beneath it raises a semaphore. That is a clean rule with no exceptions list to maintain.

VIII.3 — Open

FULLY SETTLED — §XVI (2026-09-18) and §XXIII (2026-09-19). The first bullet, where the semaphore lives, is answered by §XXIII: a kernel-resident single-writer scalar read through a registered primitive, and a one-way latch rather than a counting semaphore. The note below is kept as written.

PARTLY SETTLED 2026-09-18 by §XVI. The second bullet (what a clean-shutdown routine does) is answered: forced blk_flush(0), then BYE; Hera's is suicide via a permanent arch_halt() loop. The third (is the window bounded) answers itself — the window ends when Hera halts, a bound by construction rather than by timer (§XVI.4). The first bullet, where the semaphore mechanically lives, remains open. Left in place, not rewritten.

  • Where does the semaphore live, mechanically? It must be readable by VMs whose messaging is dying. Kernel-resident, below the routing layer, is the only coherent answer, but the concrete form is unspecified.
  • What does a VM's "clean shutdown routine" actually do? §IX makes this fleet-wide behavior, so it needs a single definition — and it must satisfy §VII.4's note about the conservation invariant.
  • Is the window bounded? "Holds it until it goes down" gives no deadline. If Hermes sinks slowly, do VMs shut down on the signal alone or wait for release?

IX. Birther failure: no provisional handoff

IX.1 — Decided (Captain Bob, 2026-09-18)

If a VM responsible for birthing/lifecycle decisions (total Hera failure is the type case) exhausts its own ladder — warm restart, then from-scratch BIRTH — and still cannot recover, there is no provisional handoff of its authority to another VM. No VM assumes partial or stolen authority on another's behalf. Instead, the rest of the fleet runs its own shutdown routines — the same clean-shutdown behavior §VIII's sinking semaphore triggers, applied fleet-wide.

IX.2 — Why this is consistent rather than merely austere

It is the same invariant as §VI, read from the other side. §VI says a VM can never birth something with more authority than it holds. §IX says a VM can never acquire authority it was not born with, even when the holder is gone and the authority is going begging. Together: authority only ever flows down the birth graph, never sideways and never up. No VM ends up holding authority it wasn't born with, and no VM limps in a half-failed state — every path terminates in either full recovery or fast, clean termination.

That is a strong, checkable property, and it is worth writing down as the thing to defend when some future failure mode makes a provisional handoff look temporarily attractive.

IX.3 — Open

PREMISE REJECTED 2026-09-18 by §XV.2. The first bullet asks the wrong question. Nobody declares a VM dead — it dies a compudynamic death. The "circular by construction" difficulty was an artifact of assuming death is a judgment someone renders rather than a physical outcome of the physics. Left in place, not rewritten.

  • What has standing to declare a birther dead? Circular by construction: the fleet must conclude Hera is gone without Hera participating. Ties directly to §VII.3's undesigned detection story.
  • Does fleet shutdown mean the kernel halts, or that the floor empties and the kernel survives? Materially different outcomes, not specified.

X. Sequencing — dependencies between the open questions

These do not decompose into independent work items; several answers foreclose others. The dependency order that fell out of writing this up:

  1. §III.4 (arenas: routing-only vs. full centralization) is the root. It determines whether §IV.2's routing-table question exists at all, and materially changes §VII.4's warm restart.
  2. §III.3 (the by-name VM-EXEC dependency) — needs a repo-wide S" Hermes" VM-EXEC grep first, which has not been done.
  3. §IV.2 (slot numbering) — only if §III.4 leaves a FORTH routing table standing.
  4. §VI.3 (what "ACL/DNA" refers to) — independent of the above; gates all of §VI.
  5. §VII.3 (reconciling the ladder with the wellness-check idea) — gates §VII.2's consumer question and §IX.3's declare-dead question, which are the same question twice.
  6. §V.3 (agent-VM binding) — needs §VI settled, since an agent VM is birthed before it is bound.

XI. Punch list — design only, no code authorized

Nothing here is started. Nothing here authorizes an edit. Numbered for discussion order, not execution order.

  1. ✅ CLOSED 2026-09-18 (§XIII.1). Repo-wide by-name Hermes inventory: 14 FORTH sites across 5 files, 3 C-side sites, 1 DoE CSV schema site. Only 3 are live. §I.2 corrected.
  2. ✅ CLOSED 2026-09-18 (§XIII.4). All three citations verified against source; §V.3.4 and §VII.4 may now be relied on.
  3. ⬜ Settle §III.4: routing-only vs. full arena centralization. Root dependency — do first.
  4. ⬜ Settle §III.3 given 3.
  5. ⬜ Settle §IV.2 if and only if 3 leaves a FORTH routing table standing.
  6. ⬜ Settle §VI.3: precise referent for "ACL/DNA", against the four-field DictEntry constraint.
  7. ⬜ Settle §VI.4: what "standing" means, or rule it decorative.
  8. ⬜ Reconcile §VII.3: the ladder vs. the wellness-check detection idea.
  9. ⬜ Assign SOS a message-type number deliberately (§VII.2), with a named consumer.
  10. ⬜ Specify the sinking semaphore's mechanism and the clean-shutdown routine (§VIII.3).
  11. ⬜ Answer §IX.3: what declares a birther dead, and what fleet shutdown means concretely.
  12. ⬜ NEW 2026-09-18 (§XIII.5). Rule on whether the Hestia Tripod leg is a new singleton, distinct from today's plural per-attach console proxies. Gates §IV.1 and §IV.2 — sits above the slot-numbering question, not beside it.
  13. ⬜ NEW 2026-09-18 (§XIII.2), not part of this reshuffle. Authorize a capsules/MANIFEST.md correction pass for two false claims: block 4055's "immutable ABI" (FABRIC-2.md declared it stale and it was never corrected) and block 2049's stated contents. Reported, not fixed, per Captain Bob's Law.

Acceptance for anything that eventually comes out of this list, per .claude/CLAUDE.md's non-negotiable criteria: three-arch QEMU boot (amd64, aarch64, riscv64, one at a time, in the foreground, clean before qemu), all reaching zuse)ok> with zero UNKNOWN WORD faults, logs present under logs/, and dict_hash identical across all three architectures. There is no other test.


XII. Explicitly not decided anywhere in this document

Collected so nothing here is mistaken for settled. Updated 2026-09-18 after §XV's rulings. Most of this list is now closed — see §XIV.2 and §XV.6 for what resolved, dissolved, or had its premise rejected.

Still genuinely open — updated again 2026-09-18 after §XVII closed item 12. Nothing structural remains; what follows are small local decisions, one placement question, and one build-time check.

  • Item 9 — SETTLED 2026-09-19 by §XXI: actionable to the sender's birther, advisory to peers; payload descriptive, never executable; number 10.
  • §VIII.3 first bullet — SETTLED 2026-09-19 by §XXIII: a single-writer kernel scalar in kernel-Hermes's scope, read through a registered primitive; a one-way latch, not a counting semaphore. This was the last open design question.
  • Item 16 (§XV.4) — an implementation check rather than a decision: every message-holding teardown path must reach STADIUM-EVICT, verifiable via fleet_conserved.
  • Item 17 — SETTLED 2026-09-19 by §XXIV: a separate word; BYE left alone. The reshuffle therefore changes zero live registered words.
  • Item 18 — SETTLED 2026-09-19 by §XXVII: the empty floor is unreachable (Hera is pinned, unkillable, and every path by which she can cease already halts or reboots the machine). Build nothing; §XVI.7's proposed shape is withdrawn.

§XIII.5's proposed Hestia resolution shape remains analysis, not a ruling. §XIV.4's and §XIV.5's proposals were both superseded by §XV.4 and §XV.3 respectively.

What is decided, all by Captain Bob on 2026-09-18 and recorded in §III.1, §IV.1, §V.1, §VI.1, §VII.1, §VII.2, §VIII.1 and §IX.1: Hermes goes into the kernel and becomes the arbiter rather than a client; the Tripod becomes Hera/Artemis/Hestia; Hestia becomes the bind point; BIRTH is general with authority bounded by inheritance; the fleet shares one gentle→from-scratch→brutal ladder; SOS is a standard message type with Hermes excepted via a sinking semaphore; and a failed birther's authority is never provisionally handed off.


XIII. Punch-list items 1 and 2 worked — CLOSED 2026-09-18, with three corrections to this document's own §I

Worked per the series' standing discipline: items 1 and 2 were the two pure-investigation entries on §XI, carrying no design commitment, so they were executable without a ruling. All findings below traced directly against the working tree at e56974e on 2026-09-18. No code written, no design question answered, nothing in the tree modified.

XIII.1 — Item 1 CLOSED: the by-name Hermes inventory. §I.2 undercounted by an order of magnitude.

Correction to §I.2, which is wrong as written. It claimed the by-name dependency was Artemis's "two live call sites" and framed Artemis as the dependent. The real inventory is 13 by-name FORTH references across 5 files, plus 3 independent C-side sites. §I.2's two sites are real and still the sharpest edge, but they are not the extent of it, and the executive framing ("relocated and rewired") has to cover all of it. Full inventory:

# Site Call Live?
1 common/msg.4th:7 S" MSG-ACK-LAST" S" Hermes" VM-EXEC dead — §XIII.2
2 common/msg.4th:10 S" MSG-NACK-LAST" S" Hermes" VM-EXEC dead — §XIII.2
3 artemis/init.4th:138 S" 2 ENQUEUE-READY" S" Hermes" VM-EXEC live
4 artemis/init.4th:501 S" 2 COMMON-CH @ CH-ADD-MBR" S" Hermes" VM-EXEC live
5 process.4th:12 S" 1 EVENT-EMIT" S" Hermes" VM-EXEC (SPAWN) dead — §XIII.2
6 process.4th:15 S" 2 EVENT-EMIT" S" Hermes" VM-EXEC (PAUSE) dead
7 process.4th:21 S" 3 EVENT-EMIT" S" Hermes" VM-EXEC (RESUME) dead
8 process.4th:26 S" 4 EVENT-EMIT" S" Hermes" VM-EXEC (KILL-VM) dead
9 doe-campaign.4th:6 S" Hermes" BIRTH not auto-run
10 doe-campaign.4th:7 S" LOAD-DOE" S" Hermes" VM-EXEC not auto-run
11 doe-campaign.4th:17 S" DOE-WORK" S" Hermes" VM-EXEC not auto-run
12 doe-campaign.4th:28 S" DOE-WORK" S" Hermes" VM-EXEC not auto-run
13 doe-campaign.4th:54 S" 1959 1 EXEC-DOE" S" Hermes" VM-EXEC not auto-run
14 messaging.4th:77 S" Hermes" 1 VM-NAME-REG live (§I.3)

The good news is real: only 3 of the 14 are live production paths (#3, #4, #14). Seven are dead code (§XIII.2), four are in a campaign orchestrator docs/working/architecture/ DOE-LIBRARY-HOWTO-20260819.md itself records as "Not auto-run anywhere." This materially lowers the estimated cost of §III.3 — but it must be re-confirmed rather than trusted, since "not auto-run" is a claim about today's boot path, not a guarantee nothing invokes it.

XIII.2 — Seven of those sites are dead code, and MANIFEST.md carries a claim FABRIC-2 already declared stale

common/msg.4th is loaded by nothing. Traced: S" common:msg.4th" EXEC appears exactly once in the tree — inside common/msg.4th's own header comment, as usage documentation. No capsule EXECs it. capsules/init.4th (block 2049) loads ACL.4th, block-acl.4th, zuse-eligibility.4th, lib.4th, fabric.4th, font.4th, common:messaging.4th — and nothing else.

This is already known and already written down. FABRIC-2.md:2770-2773 states it outright: the HERMES-ACK/HERMES-NACK wrappers "are now obsolete (every VM has its own local MSG-ACK-LAST/MSG-NACK-LAST — no VM-EXEC indirection needed) but the file itself was left in place, unloaded, rather than deleted unprompted; capsules/MANIFEST.md's 'immutable ABI' claim for block 4055 is now stale." messaging.4th:411 carries the same note from the other side: "common:msg.4th's HERMES-ACK/NACK indirection is retired."

capsules/MANIFEST.md was never corrected. Line 317 still reads: "Immutable: this is the cross-VM ACK/NACK ABI. Every messaging VM (Hera, Hermes, Artemis) loads this at birth. Changing the block or the word names breaks the Hermes delivery protocol." That is false on every clause. A second MANIFEST claim also fails against the source: line 52 describes block 2049 as loading "compudynamics, VM-INIT, lib, common:msg, fleet-k, process; BIRTHs Artemis + Hermes" — init.4th block 2049 loads none of common:msg/process and contains no BIRTH at all (Hermes and Artemis are birthed from C — §XIII.3).

process.4th is likewise EXEC'd nowhere, making sites #5–#8 dead with it.

Reported, not fixed, per Captain Bob's Law ("if you identify a bug, report it; do not fix it unless the user says to"). Flagged here because a reshuffle that reads MANIFEST.md as current will conclude the ACK/NACK path is a live immutable Hermes ABI and scope around a constraint that stopped existing in FABRIC-2.md's era. Recommend a MANIFEST.md correction pass as its own authorized item; it is not part of this reshuffle.

XIII.3 — The Tripod is hardcoded as a name-triple in C, in three independent places

Not previously recorded in this document. Each is an edit site the moment the Tripod's membership changes (§IV.1):

  1. capsule_birth.c:793-796 — is_fleet_foundation is literally vm_name_prefix_eq_nocase(capsule_name, "Hera") || ... "Hermes" || ... "Artemis". It gates StadiumPatronHeader setup and the session_register()/session_set_pinned() pinning calls. This is the Tripod, encoded as a string test. Swapping Hermes for Hestia is a one-line change here — and that single line is what makes a VM pinned.
  2. kernel_main.c:864-874 — vm_interpret(mama, "S\" Hermes\" BIRTH") plus a capsule_vm_find_by_name_nocase("Hermes", ...) liveness check and console banner. The comment cites FABRIC-2.md D.7 (birth-by-message-only, 2026-08-28): the Tripod legs "must be alive session-less so a later thumbdrive-attach flow has a running Hermes/Artemis to message." Note the stated rationale for Hermes being born early is precisely that other flows need it available — which a kernel-resident arbiter satisfies trivially and permanently. §III is consistent with D.7's intent, not in tension with it.
  3. kernel_main.c:1007-1009 — sk_vm_switch_signal_register(hermes_entry.vm_id), the preemptive context-switch registration from FABRIC-3.md §XXVIII.

A fourth site is not code but schema, and is the expensive one: doe_log.c. The per-tick DoE CSV hardcodes the Tripod across six columns — 16/17/18 (hera_heat_q48, hermes_heat_q48, artemis_heat_q48, populated by doe_log_heat_by_name("Hera"/"Hermes"/ "Artemis") at lines 206-208) and 23/24/25 (switch_*_readiness), with column 22 documented as "0=Hera, 1=Hermes, 2=Artemis by current registration order." Changing Tripod membership changes the DoE CSV schema, which bears directly on comparability with every campaign already run (§XXXIV/§XXXV's segmented per-ISA design in FABRIC-3.md is mid-flight). Flagged as a real cost, not a blocker, and not something to resolve by quietly renaming a column.

XIII.4 — Item 2 CLOSED: all three flagged citations verified against source

§XI item 2 flagged three citations this document took from headings and .claude/CLAUDE.md rather than from the source. All three check out; §V.3.4 and §VII.4 may now be relied on.

  • FABRIC-2.md §H.1 — real, at line 3699, "Session = Stadium patron, admission restated." The pinning claim is at line 3713: pinned sessions are exempt from "the normal departure path (heat decay / COOL): they never leave, permanently." Confirmed as §V.3.4 used it.
  • FABRIC-2.md §H.12 — real, at line 4057, "Implementation punch list (2026-09-03)." Independently corroborated from the code side: capsule_birth.c:785-788 cites "§H.12 step 4" by name for the session_register()/session_set_pinned() soft-fail convention.
  • FABRIC-3.md §XVIII — confirmed. K is vm_physics_fleet_heat_sum() over all live VMs, which the reservoir-transfer accounting holds exactly at Q48_ONE, surfaced as the fleet_k_q48/fleet_conserved CSV columns. §VII.4's warning stands as written: a brutal death must still return what it held, or it breaks a continuously-verified invariant.

XIII.5 — New finding, and the most consequential of this pass: Hestia today is plural, ephemeral, and capsule-less

This was not known to §IV or §V when they were written, and it changes what §IV.1 is asking for. Traced in capsule_console.c:

  • Hestia has no capsule. capsules/ contains hermes/ and artemis/ directories but no console/. A console VM's entire personality is a 3-line C string literal, CONSOLE_IDENTITY_SRC (capsule_console.c:28-31): Block 4997, S" common:messaging.4th" EXEC, MSG-CD-INIT. That is all of it.
  • Hestia is deliberately not on the routing table. The source comment is explicit: "No COMMON-CH subscription: a console's own traffic is direct 1:1 with its paired user VM (CONSOLE-CMD-EVENT), not broadcast, so there's no need to resolve an index in Hermes's own routing table for it."
  • Hestia is plural and per-attach. capsule_console_birth(const char *console_name, ...) is called from three sites (capsule_wirebind.c:245, mama_forth_words.c:1591 and :2015) and mints a fresh heap-built single-entry capsule directory each time. There are as many console VMs as there are attachments.

Consequence. Hera, Hermes and Artemis are singular, pinned, born-at-boot, capsule-backed, routing-table-indexed. Hestia today is none of those five things. So §IV.1's "Hestia takes the vacated third slot" is not a relocation of an existing pinned VM — as stated it would create a Hestia that does not currently exist. That is worth naming plainly, because it is the one place this document's "reorganization, not invention" scope discipline is genuinely strained.

CORRECTED 2026-09-18 by §XVII — the paragraph above compares the wrong object. The singular pinned Hestia does already exist: it is the HAL drawing fabric (hal/console.c, framebuffer.c, vt100.c, font_8x16.c), already singular, permanent and kernel-resident — it simply has no VM face and no owner. The per-attach proxies this section measured against are what binding produces, not what Hestia is. The scope discipline is therefore not strained: the leg is an existing singleton given an owner, and the only genuinely new artifact is a capsule personality. The resolution shape proposed below was right; this justification for it was wrong. RATIFIED — see §XVII.2. Left in place, not rewritten.

The shape that resolves it without inventing anything — offered as analysis, not ratified, Captain Bob's call: distinguish Hestia (singular, pinned, capsule-backed, the Tripod leg — owns the drawing fabric and is the bind point, per §V.1) from console proxies (plural, ephemeral, per-attach — exactly what capsule_console_birth() mints today, unchanged). The Tripod leg is new; the proxies are untouched. This reading makes §V.1's "bind point" precise: the leg is what you bind to, the proxy is what binding produces.

XIII.6 — §V and §VI are load-bearing for each other, which neither section noticed

If Hestia is the bind point (§V.1), it is Hestia that must mint console proxies — and minting a VM is BIRTH. Today all three capsule_console_birth() call sites run in Hera's or the caller's context. So "Hestia is the bind point" cannot be implemented while BIRTH remains Hera-exclusive; it requires §VI.1's generalization. They are one change, not two.

Mechanically this already works, which strengthens both: two of the three call sites pass vm->stadium_vm_id — whichever VM invoked — rather than a hardcoded Hera (capsule_wirebind.c:245 passes mama_vm->stadium_vm_id because that path genuinely is Hera's). That matches §I.5's finding that capsule_birth_baby() treats stadium_vm_id generically as "whoever is birthing this VM." Parentage is already generic; only registration is not.

This also supplies the first concrete answer to §VI.4's open "what is standing": Hestia needs BIRTH to do its declared job, so it has standing by role. That is evidence for reading (a) (standing = having BIRTH registered), not a ruling.

XIII.7 — Effect on §X's sequencing

Nothing found here dislodges §III.4 as the root dependency. Two adjustments:

  • §III.3 is cheaper than §I.2 implied (3 live sites, not a broad web) and can be scoped as soon as §III.4 lands.
  • §IV.1 needs a prior ruling that §IV.2 does not cover: is the Tripod's Hestia leg a new singleton distinct from today's proxies (§XIII.5)? That question sits above the slot numbering, not beside it. Added to the punch list as item 12.

XIII.8 — Punch-list status after this pass

  • ✅ Item 1 CLOSED — inventory complete (§XIII.1), 14 FORTH sites + 3 C sites + 1 CSV schema site; §I.2 corrected.
  • ✅ Item 2 CLOSED — all three citations verified against source (§XIII.4).
  • ⬜ Item 12, NEW — rule on §XIII.5: is the Hestia Tripod leg a new singleton, distinct from today's per-attach proxies? Gates §IV.1 and §IV.2.
  • ⬜ Item 13, NEW, not part of this reshuffle — authorize a capsules/MANIFEST.md correction pass for the two false claims in §XIII.2 (block 4055 "immutable ABI"; block 2049 contents). Reported, not fixed.

Items 3–11 unchanged and still open.


XIV. Two rulings that collapse the root dependency (Captain Bob, 2026-09-18)

XIV.1 — RULING: the existing FORTH messaging layer is legacy. Kernel Hermes is greenfield.

Captain Bob, 2026-09-18, verbatim: "as far as Hermes the VM and Hermes the kernel component, forget all the forth existing and call it legacy for now. We'll do a code cleanup for all the dead FORTH anyway."

So kernel-Hermes is not a migration of capsules/common/messaging.4th. It is a new C implementation; the FORTH messaging layer becomes legacy on arrival and is retired by a separate dead-FORTH cleanup pass. This reverses the framing §III.3 and §III.4 were built on — both assumed the existing FORTH had to be carried across.

XIV.2 — What this collapses

Four of this document's open questions dissolve rather than get answered:

  • §III.4 (the root dependency) — RESOLVED by fiat, in the "full centralization" direction, but without the migration cost that made it frightening. The kernel owns message and channel state; messaging.4th's per-VM CREATE/ALLOT arenas are legacy, not something to centralize. §III.4 feared a painful port of live FORTH state; there is no port.
  • §III.3 (the by-name VM-EXEC dependency) — dissolved. The three-option framing (kernel primitives / vestigial name / birth-time registration) was about preserving callers. With the FORTH legacy, §XIII.1's 14-site inventory stops being a migration constraint and becomes an input to the cleanup pass. §XIII.1's inventory is still the right list — its purpose changed, not its content.
  • §IV.2 (routing-table slot numbering) — dissolved, exactly as §IV.2's own option 3 predicted: "the routing table stops being a FORTH array at all... in which case this question dissolves into §III.4's arena question." It did. VM-NAMES-INIT's 0/1/2 pinning and the "never renumbered" commitment are legacy artifacts, not constraints on the new design.
  • §XI item 13 (the MANIFEST.md corrections) — absorbed. Bob's "we'll do a code cleanup for all the dead FORTH anyway" covers §XIII.2's findings. Still worth doing deliberately; no longer a separate ask.

§XIII.1, §XIII.2 and §XIII.3 keep their value — the C-side sites (§XIII.3) are not legacy and remain real edit sites, and the dead-FORTH inventory now feeds the cleanup.

XIV.3 — RULING: Artemis is the persistence path. And it needs no layering inversion.

Captain Bob, 2026-09-18: "For persistence though, use Artemis as much as possible. We're good with that part for sure." Ratified.

A layering question this raises, traced 2026-09-18 and answered cleanly. Artemis is a VM on the Stadium floor; kernel-Hermes sits below the floor. "Persist via Artemis" reads at first like an inversion — the kernel arbiter calling up into a floor VM. It is not, because storage is already two layers, not one:

  • The block subsystem is kernel/shared C, below the floor: src/block_subsystem.c, src/blkio_*.c, src/starkernel/virtio/virtio_blk.c. capsule_loader.c:98-100 calls blk_subsys_init() and blk_subsys_add_raw_device() directly, with no Artemis involved at all.
  • Artemis is the storage policy arbiter on the floor: it owns attach/registration (blk_subsys_attach_device() via the BLK-ATTACH primitive — repl.c:525-531 states it outright, "Artemis's own domain now, not Hera's") and the on-disk Artemis format.

So the ruling is satisfiable without inversion, by keeping the two straight: floor-level and fleet persistence goes through Artemis, its domain and its format; anything kernel-Hermes itself needs uses the same block subsystem Artemis is built on — never a new one, and never by messaging Artemis.

Worth recording that the kernel already calls up into Artemis by name today (repl.c:539-541, S" %llu HERA-BLK-ATTACH-REQ" S" Artemis" VM-EXEC), with a comment explaining it is a direct VM-EXEC rather than a real message because Hera can't use her own MSG-SEND there. That precedent exists; this section is about not depending on it.

XIV.4 — The one real hazard in the persistence ruling, and the shape that avoids it

Kernel-Hermes must not depend on Artemis for the persistence it needs during its own failure. This is the same structural trap as §VIII: if Hermes must write something at death and reaching Artemis requires routing, the dependency closes a circle on exactly the mechanism that is broken. §VIII solved the signalling case with a semaphore; the persistence case needs the same discipline rather than a second special case.

Recommended shape — analysis, not a ruling, Captain Bob's call: kernel-Hermes holds only reconstructible state. If the arbiter's routing state can be rebuilt from the live VM registry (which capsule_birth.c already maintains as the authority on who exists), then:

  • Hermes has nothing that must be persisted, so the Artemis dependency never arises.
  • §VII.4's "what does warm restart mean for Hermes" becomes trivial — rebuild from the registry, no saved state to reconcile.
  • Losing in-flight messages when the arbiter dies is consistent with §VII.1 step 3's "no lingering, no partial states," and with §VIII's clean-shutdown window.

This costs message durability across an arbiter death. Flagged as the real trade, and the one to rule on: if in-flight messages must survive Hermes dying, Hermes needs durable state, and that durable state cannot route through Artemis at death-time.

XIV.5 — Still open, and now more urgent, not less: the scheduler firewall

Raised 2026-09-18 and not yet ruled on. Recorded here because §XIV.1 changes its timing.

The fleet already has a scheduler: §XXVIII built timer-driven preemption, and MSG-SEND (messaging.4th block 5021) already calls SWITCH-MARK-WORK → sk_vm_switch_signal_mark_work(), setting has_work on the target's switch slot; sk_vm_switch_signal_tick() switches when readiness >= SK_SWITCH_READINESS_THRESHOLD and has_work. Message arrival is already a scheduling input. This directly contradicts VM-PHYSICS-DYNAMIC-FLEET-DESIGN-20260705.md's "Explicitly not wanted: a VM scheduler. Nothing gets built that decides whose turn it is to execute" — a live tension in the project that predates this document and is not this document's to resolve, but is this document's to not make worse.

Why §XIV.1 sharpens it. Today that coupling runs FORTH MSG-SEND → C primitive → switch slot. With the FORTH legacy, kernel-Hermes becomes the caller of sk_vm_switch_signal_mark_work() directly — eligibility marking moves inside the arbiter. That is the consolidation to guard against, and it happens by default unless the boundary is stated up front. Greenfield is the right time to state it; retrofitting it later is how this becomes a conventional scheduler by accident.

Proposed invariant, for ruling: kernel-Hermes carries, resolves and ACL-checks. It publishes facts — "this VM has mail" — and never reads readiness, never orders traffic by priority, and never decides turn order. switch.c remains the sole owner of "who runs next."

The argument for it is empirical and from this project's own history: §XXVIII.2 is a record of switch-storms caused by two sources of truth disagreeing about who was running (vm_log_attributed_vm() vs. the real trampoline state) — QEMU pinned near 100%, serial log frozen solid. This codebase punishes duplicated authority with hangs. Two deciders would be worse than one scheduler.

XIV.6 — Punch list after these rulings

  • ✅ Item 3 (§III.4, the root) — RESOLVED by §XIV.1. Kernel owns messaging; FORTH is legacy.
  • ✅ Item 4 (§III.3) — DISSOLVED by §XIV.1; §XIII.1's inventory becomes cleanup input.
  • ✅ Item 5 (§IV.2) — DISSOLVED by §XIV.1, as §IV.2 option 3 predicted.
  • ✅ Item 13 — ABSORBED into the dead-FORTH cleanup pass.
  • ⬜ Item 14, NEW — rule on §XIV.5's scheduler firewall invariant. Raised, not ruled.
  • ⬜ Item 15, NEW — rule on §XIV.4: does message durability survive an arbiter death? If yes, Hermes needs durable state that cannot route through Artemis at death-time.
  • ⬜ Items 6, 7, 8, 9, 10, 11, 12 unchanged and still open.

Note the shape of what remains. With the root dependency resolved, the surviving open items are no longer about mechanism — they are about authority: who may birth (6, 7), who decides a VM is dead (8, 11), who owns turn order (14), and what survives a death (15). That is a better class of question to be left with, and §IX's "authority only ever flows down the birth graph" is the principle most of them should be tested against.


XV. The authority questions, answered (Captain Bob, 2026-09-18)

§XIV.6 observed that the surviving open items were no longer about mechanism but about authority — who may birth, who declares a VM dead, who owns turn order, what survives a death. All four are answered here, and three of the four are answered by rejecting the question's premise rather than by naming an owner. That pattern is the finding.

XV.1 — RULING: any VM may birth a near-clone of itself, carrying its own ACL DNA

Captain Bob, verbatim: "Any VM can birth another VM that is almost a clone of itself with its own ACL DNA whatever ya wanna call it."

This closes §VI.3 and §VI.4 together:

  • §VI.3 ("ACL/DNA" needs a precise referent) — ANSWERED: inheritance by copy, from the parent. The child is almost a clone of its birther: it starts from the parent's own state and carries its own DNA — a copy it owns and may diverge, not a shared reference to the parent's. This is the reading §VI.3 hoped for: it uses the existing word-level ACL-INHERIT semantics (pin cleared, mode copied) rather than inventing a VM-level construct, so it stays inside the "reorganization, not invention" scope, and it does not need a fifth acl_* field on DictEntry (the constraint §VI.3 flagged).
  • §VI.4 ("standing" is undefined) — ANSWERED: standing is decorative. "Any VM can birth" confirms reading (a) — standing means nothing more than having BIRTH registered. Per §I.5 the C implementation is already VM-agnostic, so this is a registration change, exactly as §VI.2 predicted. No Stadium-economy gate, no ACL.4th predicate. Consistent with .claude/CLAUDE.md's "do not over-engineer — if the user says 'BIRTH is a primitive', that is the complete specification."

Why the authority bound still holds without a check. A near-clone cannot exceed its parent because it is built from the parent. §VI.1's "enforced by inheritance, not by a separate authorization check" is therefore literal, not aspirational — there is no gate to bypass because there is no gate.

XV.2 — RULING: nobody declares a VM dead. It dies a compudynamic death.

Captain Bob, verbatim: "Nobody declares a VM dead, it died a compudynamic death."

This rejects the premise of §IX.3 and of §VII.3's detection gap, and both were wrong to ask what they asked. §IX.3 asked "what has standing to declare a birther dead?" and called it "circular by construction." The circularity was an artifact of assuming death is a judgment someone renders. It is not: death is a physical outcome of the compudynamics — heat exhausts, the reservoir empties, the patron leaves the Stadium floor. There is no detector, no quorum, no supervisor, and nothing to build.

This retires the largest scheduler-shaped risk in the whole design. §VII's ladder appeared to need a supervisor — some component watching liveness and deciding to escalate — and §XIV.5 warned that a supervisor with a timer is a scheduler's twin. With death compudynamic, that component does not exist and is not needed. The ladder is self-applied while a VM is alive; once it is dead it is dead, which is exactly §VII.1 step 3's "fast and total, not a slow degrade." No actor, no handoff (§IX.1 already forbade one), no partial states.

Consequence for §VII.3. The "wellness check at the nearest participating wellness center" idea recorded in FABRIC-3.md §XX does not need to be reconciled with this ladder as a detection mechanism, because no detection mechanism is required. It remains an independent idea on its own merits. §VII.3's framing ("a ladder with no detection story only fires on failures loud enough to notice by accident") was built on the same wrong premise and is withdrawn.

XV.3 — RULING: nobody owns turn order. Turn order is compudynamic.

Captain Bob, verbatim: "Nobody owns the turn order. The turn order is compudynamic."

This answers §XIV.5 (item 14) with a stronger invariant than the one proposed there. §XIV.5 suggested a firewall — "switch.c remains the sole owner of who runs next." That named an owner, and naming an owner is what a conventional scheduler is. The correct invariant names none:

Turn order is an emergent output of the compudynamics, not a decision any component makes. Kernel-Hermes publishes facts into that system — "this VM has mail" — and computes no ordering. Neither does anything else.

This is consistent with the project's own standing position, independently of this document: FABRIC-3.md §XIX records the design direction as "fix the gap without a scheduler," with the reasoning that "a fixed round-robin across identities would just be a different hardcoded policy — not in the spirit of" the design. §XV.3 is that same position restated for the reshuffle.

One honest observation, offered as a flag rather than an objection. Traced 2026-09-18: sk_vm_switch_signal_tick() gates on readiness >= SK_SWITCH_READINESS_THRESHOLD && has_work, and SK_SWITCH_READINESS_THRESHOLD is 50u, a hardcoded constant (capsule_vm_switch_signal.c:52). That is a fixed policy sitting in the middle of a mechanism this section calls compudynamic — the same category §XIX rejected as "a different hardcoded policy."

The layering does reconcile: the compudynamic turn-attractor (§XIX/§XXI) decides what work is assigned, and the CPU-level switch signal is the mechanical follower that moves the processor. Turn order in the meaningful sense is the attractor, and it is compudynamic as ruled. But the constant 50 is a real seam, and it is the exact place where a hardcoded policy could quietly start deciding turns. Flagged for later, not proposed as work here — it predates this reshuffle and is FABRIC-3.md §XXVIII's own territory. Noted so it is not discovered later and mistaken for something this document introduced. (FABRIC-4.md §1's fixed-then-adaptive rate graduation is the precedent for how such a constant earns its way to adaptive: measure against a fixed baseline first, self-tune only after.)

XV.4 — §XIV.4 reframed: the question was persistence; the answer is conservation

§XIV.4 asked whether in-flight messages must survive an arbiter's death, and warned that answering "yes" would force Hermes to hold durable state it could not safely write through Artemis at death-time. Captain Bob, on that item: "What survives the death of what? I don't understand that last part." Fair — it was posed as a persistence question, which is the wrong frame and the reason it read as opaque.

The right frame, given §XV.2. A message is not only content; it is heat. Traced 2026-09-18 in capsules/common/messaging.4th: MSG-ALLOC pulls Q.SLOT from the caller's own reservoir before admitting a message (rolling back via STADIUM-RES-PUSH on refusal, line 127), and MSG-FREE-NODE releases it with STADIUM-EVICT (line 135-136). An arbiter dying while holding N messages is holding N × Q.SLOT of fleet heat.

So the question is not "must the payloads be saved." It is:

Does a dying arbiter return the heat its undelivered messages were holding?

  • Payloads may vanish. Consistent with §VII.1 step 3's "no lingering, no partial states."
  • The heat may not. K = vm_physics_fleet_heat_sum() is held at Q48_ONE and continuously verified (FABRIC-3.md §XVIII, fleet_k_q48/fleet_conserved). Heat that dies with the arbiter is heat that breaks a live invariant.

The answer to "what survives the death of what" is therefore: nothing survives except K.

This dissolves §XIV.4's hazard rather than resolving it. There is nothing to persist, so there is no death-time write, so there is no circular dependency on Artemis, so the stateless-arbiter shape §XIV.4 recommended is simply correct rather than a trade-off. And since death is compudynamic (§XV.2), a VM's death is already a Stadium eviction event — returning heat is not a special case bolted onto death, it is what dying consists of. §XIV.3's "Artemis for persistence" ruling stands untouched and applies to fleet/floor persistence; the arbiter simply never needed any.

The one thing to verify when this is eventually built (not a design question, an implementation check): that every message-holding structure's teardown path actually reaches STADIUM-EVICT, so a mass teardown returns heat rather than leaking it. fleet_conserved already makes that directly testable — a leak shows up as K ≠ Q48_ONE, not as a silent error. Added as punch item 16.

XV.5 — What these four rulings have in common

Three of the four answers reject the question rather than answer it. There is no death declarer, no turn-order owner, and nothing to persist — and in each case the thing that seemed to need an authority turned out to be an outcome of the physics instead. The one question that was answered directly (§XV.1, who may birth) was answered with "anyone," which also declines to create an authority.

That is the actual defence against becoming a conventional scheduler, and it is stronger than the firewall §XIV.5 proposed. A firewall constrains a component that owns something. These rulings mean there is no owner to constrain. The design's protection is structural, not procedural: authority only ever flows down the birth graph (§IX.2), and everything else — death, turn order, conservation — is physics rather than policy.

XV.6 — Punch list after these rulings

  • ✅ Item 6 (§VI.3, "ACL/DNA" referent) — ANSWERED (§XV.1): inheritance by copy; existing ACL-INHERIT semantics suffice; no fifth DictEntry field.
  • ✅ Item 7 (§VI.4, "standing") — ANSWERED (§XV.1): decorative; reading (a) confirmed.
  • ✅ Item 8 (§VII.3, ladder vs. detection) — WITHDRAWN (§XV.2): premise wrong, no detection mechanism required.
  • ✅ Item 11 (§IX.3, who declares a birther dead) — PREMISE REJECTED (§XV.2): nobody; compudynamic death.
  • ✅ Item 14 (§XIV.5, scheduler firewall) — ANSWERED (§XV.3), with a stronger invariant than the one proposed: no owner, not a named owner.
  • ✅ Item 15 (§XIV.4, message durability) — DISSOLVED (§XV.4): reframed from persistence to conservation; nothing survives but K.
  • ⬜ Item 16, NEW (§XV.4) — implementation check, not a design question: verify every message-holding teardown path reaches STADIUM-EVICT so mass teardown returns heat. Testable directly via fleet_conserved.
  • ⬜ Item 9 (§VII.2 — SOS message-type number and its consumer) — still open. Note §XV.2 narrows it: with no death-declarer, an SOS consumer cannot be a supervisor, so the question is now specifically what a peer does on receipt.
  • ⬜ Item 10 (§VIII.3 — the sinking semaphore's mechanism, the clean-shutdown routine's definition, and whether the window is bounded) — still open, and now the largest remaining design item.
  • ⬜ Item 12 (§XIII.5 — is the Hestia Tripod leg a new singleton, distinct from today's per-attach proxies?) — still open, and still gates §IV.

Remaining open: items 9, 10, 12, 16. Item 10 is the substantive one; 9 is downstream of it, 12 is independent, and 16 is a build-time check rather than a decision.


XVI. Item 10 settled: clean shutdown is forced flush + BYE; Hera's is suicide (Captain Bob, 2026-09-18)

Captain Bob, verbatim: "A clean shutdown is a VM termination where flush is forced and the VM says BYE. Hera is the special case that should just halt the processor. Call it suicide I guess."

This settles §VIII.3's second open point (what a clean-shutdown routine concretely does) and, as a consequence, its third (is the window bounded). Traced 2026-09-18: the ruling is almost entirely existing mechanism, with exactly one behavioural change to a word that already exists.

XVI.1 — Every piece of this already exists

  • Forced flush — blk_flush(0). block_subsystem.c:988-990; the source comment states it directly: "0 is the 'flush all' sentinel." Already called that way at block_subsystem.c:849, and already the Stadium's own write-back action for STADIUM_BEHAVIOUR_MIGRATE (stadium.c:336, :373). Nothing new to build.
  • A child VM's BYE — system_word_bye() (src/word_source/system_words.c:143-147): sets vm->halted = 1, signalling the REPL to stop. That is already precisely "a VM termination." Confirmed by mama_forth_words.c:2262-2263's own doc comment: "In child VMs this word is never registered; children use the standard system_word_bye which sets vm->halted and returns to the parent's REPL."
  • Halting the processor — arch_halt(), declared at include/starkernel/arch.h:64 and implemented on all three architectures (amd64/arch.c:207, aarch64/arch.c:126, riscv64/arch.c:170). Three-arch parity already holds, so the non-negotiable acceptance criteria are satisfiable without new per-arch work.

So a child VM's clean shutdown is blk_flush(0) then BYE — two existing calls in sequence. This sits comfortably inside the "reorganization, not invention" scope discipline.

XVI.2 — The one real change: Hera's BYE does not currently halt. It cold-restarts.

This is the finding of this section, and it is a behavioural change to a live, registered word — not a gap to fill. mama_word_bye() (mama_forth_words.c:2265-2271) is Hera's own BYE, and its doc comment says outright: "Hera-only: reap all children then cold-restart the machine." The body is exactly two actions:

capsule_vm_kill_all_nonmama();   /* "BYE: reaping children" */
arch_cold_reset();               /* "BYE: cold restart"     */

Against Bob's ruling:

  • capsule_vm_kill_all_nonmama() already matches the "fleet shuts down" half. (The contract is documented downstream too — include/starkernel/capsule_birth.h:339 describes it as called "from Hera's BYE immediately before arch_cold_reset() to reap all children.")
  • arch_cold_reset() does not match. The ruling is halt, not restart. A cold reset reboots the machine; suicide stops it. These are opposite outcomes, and today's code does the wrong one.

Implementation shape, traced: arch_cold_reset() is __attribute__((noreturn)) (arch.h:67), but arch_halt() is not — it halts only until the next interrupt and is called once per idle iteration in normal use (amd64/arch.c:200-205's own comment). A permanent halt is therefore arch_disable_interrupts() followed by a for(;;) arch_halt(); loop — which is exactly the pattern amd64's own arch_cold_reset() already uses as its unreachable fallback (for (;;) __asm__ volatile ("cli; hlt");, arch.c:218) and the same shape the kernel panic path uses. No new primitive; an existing pattern applied deliberately.

Flagged for a conscious decision, not resolved here. .claude/CLAUDE.md carries a hard rule: "Never modify a registered, tested word to 'fix' it." This is not a fix — it is a ruled behavioural change — which is a different thing, and the rule's own reasoning ("the problem is almost certainly in the caller") does not apply. But the change is real and should be made knowingly rather than discovered in a diff. Whether Hera's suicide is spelled as a changed BYE or as a separate word leaving BYE's cold-restart intact is an implementation choice, and is not decided here. Worth noting that cold-restart-on-BYE is plausibly wanted behaviour for an interactive operator typing BYE at Hera's REPL, which is an argument for two words rather than one.

XVI.3 — A shared-source constraint on where the flush goes

system_word_bye() lives in src/word_source/system_words.c — vendored/shared source that must compile and behave correctly in the hosted build as well as the kernel (.claude/CLAUDE.md: gate kernel-only code with #ifdef __STARKERNEL__). Forcing a blk_flush(0) inside system_word_bye() itself would change hosted-build behaviour for a word that has nothing to do with this design.

So the forced flush belongs in the kernel-side shutdown path that calls BYE, not inside BYE. Stated here as a constraint on implementation so the obvious-looking edit is not made in the obvious-looking place.

XVI.4 — §VIII.3's third question answers itself: Hera's halt is the bound

§VIII.3 asked whether the shutdown window is bounded, noting that "holds it until it goes down" gives no deadline. It is bounded, and by construction rather than by a timer: the window ends when Hera halts the processor. Children flush and say BYE; Hera reaps whatever remains and stops the machine. There is no deadline to choose, no timeout to tune, and — consistent with §XV.3 — nothing that has to decide when time is up. The sequence terminates because its last step is physically terminal.

This also satisfies §VII.1 step 3's "no lingering, no partial states" literally: after arch_halt() in a disabled-interrupt loop there is no state left to linger in.

XVI.5 — Item 9 is now half-answered

§VII.2 left two questions on SOS: which message-type number, and "who consumes an SOS, and what do they do with it?" The second is answered for the two fleet-fatal cases: on Hermes's sinking semaphore (§VIII.1) and on total birther failure (§IX.1), the response is now concrete — forced flush, BYE, and Hera's suicide as the terminal step.

Still genuinely open, and narrower than before: what a peer does on receiving an ordinary VM's SOS. A single non-Tripod VM in trouble is presumably not fleet-fatal, so "everyone shuts down" cannot be the universal answer — but §XV.2 rules out the obvious alternative, since with no death-declarer an SOS consumer cannot be a supervisor that decides another VM's fate. The remaining question is therefore specifically: is a peer's SOS actionable at all, or is it purely advisory/diagnostic? Plus the trivial number allocation (§I.4: 5 and 6 are unallocated gaps, 10 is next in sequence).

XVI.6 — Punch list after this ruling

  • ✅ Item 10 (§VIII.3) — SETTLED (§XVI). Clean shutdown = blk_flush(0) + BYE; Hera = suicide via a permanent arch_halt() loop. Window bounded by Hera's halt (§XVI.4). All mechanism exists on all three architectures.
  • 🔶 Item 9 (§VII.2) — HALF-ANSWERED (§XVI.5). Consumer action defined for the two fleet-fatal cases. Open: whether a peer SOS is actionable or advisory, plus the number.
  • ⬜ Item 12 (§XIII.5) — Hestia Tripod leg: new singleton or not? Unchanged; still gates §IV.
  • ⬜ Item 16 (§XV.4) — implementation check: teardown paths must reach STADIUM-EVICT.
  • ⬜ Item 17, NEW (§XVI.2) — decide whether Hera's suicide replaces BYE's cold-restart or becomes a separate word. Small, but it changes a live registered word either way.

XVI.7 — One residual edge, raised not answered

If Hera is already dead, nobody performs the halt. §XV.2 makes death compudynamic and §IX.1 forbids any VM assuming her authority, so on total Hera failure there is no actor left to run arch_halt(). The machine would be a live kernel with an empty Stadium floor — which is precisely the "lingering, partial state" §VII.1 step 3 rules out.

The shape that resolves it without violating §IX: an empty floor is a kernel-observable condition, not a VM's decision. The kernel noticing "no live VMs remain" and halting is not a VM inheriting Hera's authority — it is the machine having nothing left to run. That keeps authority flowing only down the birth graph (§IX.2) while still terminating.

Offered as analysis, not a ruling — Captain Bob's call. Added as punch item 18.


XVII. Item 12 CLOSED — "Console" already exists as a singleton; it just has no VM face

Taken as closing item 12 (§XIII.5's Hestia-singleton question), the largest remaining design item and the one gating §IV.

§XIII.5 framed this wrongly, and the error is worth naming before the answer. It compared the Tripod legs against capsule_console_birth()'s per-attach proxies, found Hestia "plural, ephemeral and capsule-less," and concluded that a singular pinned Hestia "does not currently exist" — so §IV.1 would be creating one, straining the reorganization-not-invention discipline. That comparison picked the wrong object. Traced 2026-09-18: there are three distinct things called "console" in this kernel, not two.

XVII.1 — The three consoles, traced

  1. The HAL console — the drawing fabric itself. src/starkernel/hal/console.c (14 KB), framebuffer.c (13 KB), vt100.c (43 KB), font_8x16.c (27 KB) — ~98 KB of kernel C. Already singular. Already permanent. Already kernel-resident. It is not a VM and has no owner in the fleet sense.
  2. The console proxy VMs — capsule_console_birth(), plural, ephemeral, minted per attach from a 3-line C string literal (§XIII.5).
  3. The active-console-VM name — a single global, console_set_vm_name()/ console_get_vm_name() (include/starkernel/console.h:190-191), swapped by every dispatch in the "switch, do work, switch back" pattern.

XVII.2 — RULING: the Hestia Tripod leg is (1), given a VM face. Not a promoted proxy.

The singular, permanent Hestia this design wants already exists in substance — it is the drawing fabric. What it lacks is an owner: today the fabric is an ownerless kernel singleton that any VM reaches into through a global. So §IV.1's "Hestia takes the vacated third slot" is not the invention §XIII.5 feared. It is giving an existing singleton a VM face.

That resolves the tension cleanly and without inventing anything:

  • The Hestia leg is singular, pinned and born-at-boot — because the fabric it owns already is all three.
  • The proxies are untouched. They stay plural and ephemeral, which is correct: they are what binding produces, not what Hestia is (§XIII.5's own phrasing, now properly grounded).
  • §V.1's "bind point" falls out rather than being asserted. The VM that owns the fabric is necessarily what a user or agent VM binds to.

§XIII.5's proposed shape was right; its justification was wrong. It proposed exactly this leg/proxy split but called the leg new. The leg is not new — only its VM face is.

XVII.3 — The evidence that the fabric actually wants an owner

This is not merely tidy. The ownerless global in (3) has already caused a real, live bug, documented in the kernel's own header (console.h:193-203): console_get_vm_name() alone is not safe for save-then-restore, because it returns a pointer into the single internal buffer that console_set_vm_name() copies into — so an intervening call in the normal "switch, do work, switch back" pattern overwrites the very bytes the saved pointer points at before the restore runs. Found live 2026-08-28; the restore silently no-opped. console_save_vm_name() exists solely to work around it.

That is the signature of shared mutable state with no owner. Giving the fabric an owning VM is a structural fix for the class, not just a reorganization — and it is an argument for Bob's reshuffle that this document had not previously identified.

XVII.4 — What this commits to, concretely

Consequences of the ruling, so they are not discovered late:

  • Hestia needs a capsule personality. capsules/hestia/init.4th — a new directory beside hermes/ and artemis/. This is the one genuinely new artifact, and it is a capsule, not a mechanism.
  • is_fleet_foundation changes membership (capsule_birth.c:793-796): the name-prefix triple becomes Hera / Artemis / Hestia. That single line is what makes a VM pinned (§XIII.3), so it is also what makes Hestia a Tripod leg.
  • kernel_main.c:865's S" Hermes" BIRTH becomes Hestia's birth, and the switch-signal registration at :1007 follows it.
  • The DoE CSV schema changes — doe_log.c's six Tripod-named columns (§XIII.3). Still the expensive consequence, still not to be resolved by quietly renaming a column.
  • Hestia must have BIRTH registered (§XIII.6): binding mints a proxy, and minting is BIRTH. Already settled in principle by §XV.1's "any VM can birth."
  • Hestia becomes the owner of the console_set_vm_name() discipline — the fix in §XVII.3 is available once there is an owner, but is not in this reshuffle's scope. Flagged, not scheduled.

XVII.5 — §IV is unblocked, and §XIII.5's marker corrected

With item 12 closed, §IV.1 (Tripod = Hera/Artemis/Hestia) stands as ratified with no outstanding objection, and §IV.2 remains dissolved per §XIV.2. §XIII.5's "this is the one place the scope discipline is genuinely strained" no longer holds and is corrected there.

XVII.6 — Punch list after this ruling

  • ✅ Item 12 — CLOSED (§XVII). The Hestia leg is the existing drawing-fabric singleton given a VM face; proxies unchanged. §IV unblocked.
  • 🔶 Item 9 — half-answered (§XVI.5): open is whether an ordinary peer's SOS is actionable or advisory, plus the number allocation.
  • ⬜ §VIII.3 first bullet — where the sinking semaphore mechanically lives.
  • ⬜ Item 16 (§XV.4) — implementation check: teardown paths must reach STADIUM-EVICT.
  • ⬜ Item 17 (§XVI.2) — does Hera's suicide replace BYE's cold-restart or become a separate word?
  • ⬜ Item 18 (§XVI.7) — if Hera is already dead, nobody performs the halt.

Nothing structural remains open. Items 9, 17 and 18 are small, local decisions; §VIII.3's first bullet is a placement question; item 16 is a build-time check. Every load-bearing architectural question this document opened — what Hermes becomes, what the Tripod is, what BIRTH means, who owns death, who owns turn order, what survives, what shutdown is — is now ruled.


XVIII. The Hestia component, designed in full (2026-09-18)

§XVII settled what the Hestia leg is. This section designs the component. Everything below traced 2026-09-18 against e56974e; decisions are marked, proposals are marked, and open questions are marked.

XVIII.1 — Hestia is three layers, and only the middle one is new

Layer What Where it lives today Change
0 — the fabric Raw output: framebuffer, VT100, glyph raster, UART hal/console.c, framebuffer.c, vt100.c, font_8x16.c (~98 KB C) None. Stays kernel HAL.
1 — Hestia the VM Tripod leg. Owns fabric policy, is the bind point Does not exist New — and it is a capsule, not a mechanism.
2 — console proxies Per-attach relay VMs, one per bound user/agent capsule_console_birth() None, except their birther becomes Hestia.

The design is almost entirely layer 1, and layer 1 is mostly a relocation of vocabulary that already exists in the wrong dictionary (§XVIII.3).

XVIII.2 — The constraint that shapes everything: the HAL must never know a VM exists

This is not a preference, it is a defended invariant with a documented rationale (include/starkernel/console.h:220-224, verbatim):

console.c is a clean HAL module with no dependency on capsule/WIREBIND logic (zuse_session, capsule_wirebind_attached_username()) — pulling either in directly here would be a real layering violation, not just a style preference.

So Hestia-the-VM may not be implemented by making console.c VM-aware. Ownership flows one way only: Hestia reaches down into the HAL; the HAL never reaches up.

The sanctioned pattern for the cases where the HAL does need something VM-shaped already exists and should be reused rather than reinvented: console_set_user_prefix_provider() (console.h:231) — a callback that lets repl.c supply the "user" half of the [user@VMName] prompt "without console.c knowing anything about VMs, sessions, or WIREBIND." Any further HAL→Hestia coupling this design needs takes that shape: a registered callback, never an include.

Consequence, stated plainly: Hestia's "ownership" of the fabric is by convention, vocabulary and registration — it is the VM that holds the drawing words — not by any enforcement inside console.c. That is weaker than it sounds and is exactly right: it keeps the HAL reusable and the layering intact.

XVIII.3 — Hestia's dictionary: 1,051 lines of Hestia vocabulary currently live in Hera

The most concrete finding in this section. capsules/fabric.4th's own header reads:

fabric.4th — Console drawing-fabric coordinate machinery. FABRIC-0.md item 4.3.3. 45-degree cavalier orthographic projection... Raw pixel write (PLOT/FB-WIDTH/FB-HEIGHT) is C; this capsule is the FORTH-side policy on top of it.

It is named Hestia's fabric, it is the FORTH-side fabric policy — and it loads into Hera's dictionary (capsules/init.4th block 2049: S" fabric.4th" EXEC, S" font.4th" EXEC).

Capsule Lines Loads into today Should load into
fabric.4th 319 Hera (init.4th b2049) Hestia
font.4th 733 Hera (init.4th b2049) Hestia
turtle.4th 99 sdk.4th (opt-in, not boot) unchanged — it is SDK, not fabric

Plus the C primitives underneath: PLOT, FB-WIDTH, FB-HEIGHT (src/word_source/framebuffer_words.c:62-64). Their registration moves to Hestia's word table; the implementations do not change.

DECIDED (proposal — Captain Bob's call, but this is the direct consequence of §XVII.2): fabric.4th and font.4th move from init.4th to capsules/hestia/init.4th. This is a pure relocation of 1,051 lines — the single most obviously-correct edit in the whole reshuffle, and the clearest evidence that the fabric wants an owner: it already has the vocabulary, just attached to the wrong VM.

Parity consequence (§III.5): Hera's dict_hash will shrink, Hestia's is new. Per §XX's standard, the property that matters is cross-architecture identity, not an unchanging absolute.

XVIII.4 — Hestia's capsule personality

capsules/hestia/init.4th — a new directory beside hermes/ and artemis/, and the one genuinely new artifact this reshuffle creates.

Contents, in load order:

  1. S" common:messaging.4th" EXEC + MSG-CD-INIT — Hestia is a Tripod leg and a full messaging participant. Note this differs from the proxies, which deliberately skip COMMON-CH (capsule_console.c:23-26: "a console's own traffic is direct 1:1 with its paired user VM"). The leg subscribes; the proxies still do not.
  2. S" fabric.4th" EXEC and S" font.4th" EXEC — per §XVIII.3.
  3. Hestia's own bind vocabulary (§XVIII.5).
  4. A WELCOME/banner word, matching hermes/init.4th's and artemis/init.4th's shape.

Block allocation is a real to-do, not a formality. mkcapsule's actual constraints, per §XXXII.2's correction: block range [2048, 5120) and a hard 16-content-line-per-block cap (not the 1024-byte framing .claude/CLAUDE.md still describes — that doc error is outstanding). Block 4997 is already taken by the proxy's own C string literal (capsule_console.c:28-31), so Hestia's range must avoid it. Allocate against capsule-reserved.txt and verify with mkcapsule --lint capsules/ (§XXIV's collision machinery covers this; it is blind to content, only to numbers).

XVIII.5 — The bind point: flow and vocabulary

Binding, end to end, after the reshuffle:

  human attaches thumbdrive        agent VM needs a console
            │                                │
            ▼                                ▼
      WIREBIND path                    CONSOLE-ATTACH
            │                                │
            └──────────► Hestia ◄───────────┘
                           │
                    BIRTH a proxy          (§XV.1 makes this legal:
                           │                any VM may birth)
                           ▼
                 proxy paired to target
                 (VM-NAME-REG, unchanged)
  • CONSOLE-ATTACH ( name-c name-u -- ok? ) already exists and is live-verified (§I.6, closed 2026-09-16). It resolves the target's liveness before birthing, so a typo refuses cleanly with no orphaned VM. Keep it exactly as is. The change is who runs it: today the three capsule_console_birth() call sites run in Hera's or the caller's context (capsule_wirebind.c:245, mama_forth_words.c:1591, :2015). Under this design Hestia is the birther.
  • This is why Hestia needs BIRTH registered (§XIII.6) — binding mints a proxy, and minting is BIRTH. Already legal per §XV.1.
  • Mechanically this already works: two of the three call sites pass vm->stadium_vm_id — whichever VM invoked — so parentage is already generic (§I.5, §XIII.6). Only registration changes.

OPEN — agent-VM binding and the ~user convention (§V.3.1, still unresolved). The pairing check in sk_repl_dispatch_line() reconstructs the target as console_get_vm_name() + "~user" — a convention built for human at a console → identity VM. A GPIO VM (FABRIC-4.md §2) is not a ~user identity. §XXXII.2 records that a console named anything else "silently falls back to direct interpretation with no error," caught live. Two shapes, neither chosen: (a) widen the convention so the suffix is a property of the binding, not hardcoded; (b) give agent binding its own word alongside CONSOLE-ATTACH. (a) is preferable on no-bespoke-gate grounds but touches a live, bug-prone path; flagged rather than decided.

XVIII.6 — Headless-until-login survives, and the invariant that makes it survive

Hestia the leg is born at boot. A console session is not. These must not be conflated, or §VIII.1's ratified headless-until-login policy reopens — the exact failure §XXXII.2 warned about for unattended birth.

INVARIANT, to be stated in whatever code implements Hestia:

Hestia's birth at boot must not set g_wirebind_attached_username, must not cause sk_console_identity_present() (repl.c:122) to report an identity, and must not mint a proxy. Hestia owns the fabric from boot; it presents nothing until something binds.

This is the same invariant §XXXII.2 imposed on unattended identity birth, applied to a second path. It is cheap to hold — Hestia births no proxy until asked — but it is exactly the kind of thing that gets violated by a well-meaning "initialise the console at startup" line.

XVIII.7 — What Hestia does not own

Stated because a component that owns "the bind point" attracts responsibilities that belong elsewhere, and this document has already had to defend against exactly that (§XV.3):

  • Not turn order. Nothing owns it (§XV.3).
  • Not routing. That is kernel-Hermes (§III), which sits below Hestia.
  • Not the lifecycle of anything but its own proxies. Hestia births proxies; it does not birth, kill, or supervise identity VMs. Authority flows down the birth graph only (§IX.2).
  • Not the prompt's user half. That stays repl.c's, supplied to the HAL via the existing provider callback (§XVIII.2). Moving it into Hestia would be the layering violation console.h explicitly names.

XVIII.8 — Hestia's own failure behaviour

Hestia is a Stadium-floor VM above the routing layer, so §VIII.2's rule applies without an exception: Hestia emits SOS; only Hermes uses the semaphore. Its clean shutdown is §XVI's: forced blk_flush(0), then BYE.

A pleasant property worth recording: Hestia's death degrades gracefully by construction. Because layer 0 is HAL C (§XVIII.1) and ownership is by convention rather than enforcement (§XVIII.2), a dead Hestia leaves the framebuffer, VT100 and UART fully functional — the fabric simply becomes ownerless again, which is precisely today's status quo. The kernel can still print. This is the opposite of Hermes, whose death takes the transport with it. So the Tripod's three legs have three genuinely different failure characters: Artemis's death costs storage arbitration, Hestia's costs only fabric policy, Hermes's costs the transport itself — which is why only Hermes needed the semaphore.

XVIII.9 — Punch list for Hestia

Design items, in dependency order. No code authorized.

  1. ⬜ Ratify §XVIII.3's relocation: fabric.4th + font.4th move from init.4th to capsules/hestia/init.4th; PLOT/FB-WIDTH/FB-HEIGHT registration moves to Hestia.
  2. ⬜ Allocate Hestia's block range against capsule-reserved.txt, avoiding 4997; verify with mkcapsule --lint (§XVIII.4).
  3. ✅ SETTLED 2026-09-19 by §XX — neither option. The pairing is already recorded at bind time; the ~user reconstruction is a redundant derivation and is removed. Agent binding then needs no new mechanism.
  4. ⬜ Confirm Hestia in the is_fleet_foundation triple (capsule_birth.c:793-796) and its birth at kernel_main.c:865, replacing Hermes's (§XVII.4).
  5. ⬜ Handle the doe_log.c CSV schema change (§XIII.3) — still the expensive consequence.
  6. ⬜ State §XVIII.6's headless invariant in the implementation.
  7. ⬜ Not in this reshuffle: the console_set_vm_name() ownerless-global fix (§XVII.3) is now available once Hestia exists, but remains out of scope. Flagged, not scheduled.

XIX. RULING: the third Tripod leg is named Hestia, not Console (Captain Bob, 2026-09-19)

The Tripod is: Hera, Artemis, Hestia. Applied throughout this document — 99 occurrences renamed; three verbatim quotations deliberately left saying "Console" (§XIX.4).

XIX.1 — The naming convention, stated so it stops being re-litigated

antiprosopos (Gk. representative/delegate) was considered first and rejected, on Captain Bob's reasoning, which is the correct reasoning and is recorded here because it generalizes:

"Now this is just a name of a Greek mythos god, it doesn't have to carry full/perfect [fidelity], it's only a contextual display... After all, Artemis isn't exactly the perfect name for a memory manager/storage manager either."

That is the convention, and it had never been written down: a Greek deity name, evocative rather than literal. Checked against what exists — Hera (queen/mother → the Mama VM) is a close fit, Hermes (messenger → messaging) is the closest, and Artemis (huntress, wilderness → block storage and freemap arbitration) is a loose fit that has never caused anyone a problem. Fidelity was never the standard.

antiprosopos failed on being the wrong kind of word: a common noun straining for literal accuracy, which is the opposite of how the other three work. Hestia — goddess of the hearth — is a deity, is evocative (the hearth is the fixed centre of the household, where everyone gathers), and carries a pleasing incidental: Hestia is the one who never leaves. That maps onto the pinned-session model (FABRIC-2.md §H.1, "they never leave, permanently") without having been chosen for it.

Two costs that antiprosopos carried are retired by this choice, not merely tolerated:

  • Pattern. antiprosopos is Greek vocabulary, not a deity — it half-joined the set. Hestia joins it fully.
  • Typo surface. Twelve characters of unusual spelling, in a path where a VM-name mismatch fails silently (§XIX.5). Hestia is six, comparable to Hermes and Artemis.

XIX.2 — Why renaming at all is structural, not cosmetic

§XVII.1 found three distinct things in this kernel all called "console" — the HAL fabric, the VM, and the per-attach proxies — and that collision is exactly what led §XIII.5 to the wrong conclusion about whether a singular Console existed. Naming the VM separately resolves it by construction:

Layer Name after this ruling
0 — the drawing fabric (HAL C) the console — console.c, unchanged
1 — the Tripod leg, owns the fabric Hestia
2 — per-attach relay VMs console proxies — unchanged

"Hestia owns console.c" is unambiguous in a way "Console owns console.c" cannot be. This document has already paid once for that ambiguity; the rename is what stops it recurring.

XIX.3 — What is NOT renamed

The ruling names a VM. It does not rename the subsystem:

  • The HAL stays console.c/console.h and every console_*() function. Renaming them would be churn across ~98 KB of C for no gain, and would re-create the confusion by making the HAL sound like the VM.
  • capsule_console_birth() / CONSOLE_IDENTITY_SRC stay — they mint console proxies, still called console proxies.
  • CONSOLE-ATTACH stays. .claude/CLAUDE.md's "never modify a registered, tested word" applies with more force to renaming one, and the word remains accurate: it attaches a console. Recommended, not ruled.
  • CONSOLE-CMD-EVENT and sk_console_identity_present() stay — both are layer 0/2 business (a console session), not Hestia's (§XVIII.6).

Net code-visible rename surface: the VM's registry name, its capsule directory (capsules/hestia/init.4th), and the places the Tripod is enumerated (§XIX.6).

XIX.4 — Three quotations deliberately still say "Console"

Quoted source is not silently edited to match a later decision:

  1. Line 26 — the opening brief's own wording ("Hera lifecycle/ACL, Console framebuffer").
  2. §XVIII.3 — capsules/fabric.4th's verbatim header, "Console drawing-fabric coordinate machinery." That is what the file says today; it changes when the file is edited, and that edit belongs to the §XVIII.3 relocation, not to this ruling.
  3. §XVII's heading — a meta-reference to the name itself, which is that section's subject.

XIX.5 — ~user is a separate question, unaffected

antiprosopos was first raised as a replacement for the ~user registry-name suffix. That remains punch item 3 (§XVIII.5), and the Hestia ruling does not touch it. Recorded so the two never get conflated:

Traced 2026-09-18 — as a suffix it does not fit the buffers. VM_NAME_MAX is 64 and USER_IDENTITY_USERNAME_MAX 32, so the name fits in principle, but every construction buffer is sized + 8, tuned exactly to "~user" (5 chars + NUL = 6): capsule_wirebind.c:222 (=40), mama_forth_words.c:1855/:1991 and repl.c:1243 (=72), each guarding on a hardcoded + 6. A 14-byte suffix would not overflow them — it would make them refuse, silently failing to bind usernames over 26 characters and registry names over 58. Six memcpy sites and five guards hardcode that 6, and FABRIC-3.md §XXVI was itself a truncation fix.

And it would not answer item 3 regardless: a GPIO VM is neither a user nor a human's representative. The structural fix is making the suffix a property of the binding; settle that first and land any suffix rename inside it, since the same eleven sites are touched either way.

XIX.6 — Edit surface

All were already edit sites for the Tripod membership change (§XVII.4); the rename changes the string, not the count:

  • capsule_birth.c:793-796 — is_fleet_foundation's prefix triple becomes Hera / Artemis / Hestia (matched by vm_name_prefix_eq_nocase, as today).
  • kernel_main.c:865 — S" Hestia" BIRTH; the liveness check and the switch-signal registration at :1007 follow.
  • capsules/hestia/init.4th — the new capsule directory (§XVIII.4).
  • doe_log.c — CSV columns become hestia_heat_q48 and switch_hestia_readiness. Still the expensive consequence (§XIII.3), but hestia is the same width as hermes, so the schema churn is a rename rather than a reflow.

XIX.7 — One hazard, recorded and generalized

A VM-name mismatch in this system fails quietly. §XXXII.2 records it caught live on the first attempt: a console named anything other than the expected form "silently falls back to direct interpretation with no error" — typing 5 6 + . printed a direct 11 instead of a relayed one, with nothing reported.

Hestia is short and ordinary enough that this is no longer an argument about the name. But the class remains: if the §XVIII.5 binding work touches that path anyway, making an unresolved or mismatched name say so would retire the hazard for every VM name, not just this one. Raised as punch item 19, not scheduled.


XX. Item 3 SETTLED: delete the suffix from the mechanism — neither widen it nor add a word

§XVIII.5 posed agent-VM binding as a choice between (a) widening the ~user convention and (b) a sibling word for agent binding. Both options are wrong, because both accept a premise the code does not support. Traced 2026-09-18/19: the pairing is already recorded explicitly, and the ~user reconstruction is a redundant second derivation of information the mechanism already holds.

XX.1 — The relay does not route by name. It routes by a slot recorded at bind time.

sk_repl_dispatch_line() (repl.c:1225-1265) does two separate things, and only one of them involves the suffix:

  1. The guard builds paired_name = console_get_vm_name() + "~user" (repl.c:1243-1247) and checks capsule_vm_find_by_name(paired_name, ...) is VM_STATE_LIVE.
  2. The actual relay then sends CONSOLE-CMD-EVENT 0 3 S" <line>" 0 MSG-SEND — to index 3, with this comment verbatim: "to-index 3: the fixed convention this console's own VM-NAME-REG entry for its paired user VM uses (set once at pairing time — see the pairing word)."

PAIR-TEST's own header says the same from the other side (mama_forth_words.c:1542-1543): "Registers the pairing in the console's own VM-name routing table at the fixed index (3) that relay uses."

So the target is already stored, per-proxy, at bind time. The suffix string is used only to answer "is my partner alive?" — a question the stored pairing can answer directly, without deriving anything.

XX.2 — This is the documented cause of the known silent-failure bug

CONSOLE-ATTACH's own doc comment (mama_forth_words.c:1931-1940) states it outright:

sk_repl_dispatch_line()'s own pairing check (repl.c) reconstructs the target as console_get_vm_name() + "~user"... A console named anything other than the identity's own base name reconstructs the wrong target string and silently falls back to direct interpretation — no error, just quietly never relays.

CONSOLE-ATTACH already carries a workaround for this: it "takes exactly one name, not an independently chosen console name," collapsing two arguments into one so the reconstruction cannot disagree. That is a constraint imposed by the derivation, not by the problem.

XX.3 — RULING (proposed): the bind records its target; nothing re-derives it

Replace the derivation with a read of what bind time already stored.

  • sk_repl_dispatch_line()'s guard asks "does this console have a recorded pairing, and is that target live?" instead of "does <name>~user exist?"
  • CONSOLE-ATTACH resolves the target as given, rather than appending ~user (mama_forth_words.c:2000).
  • ~user survives as a naming convention, not as a mechanism. It still disambiguates rajames the proxy from rajames~user the identity in the prompt — the whole point of the [user@VMName] form (§XXV follow-up, 2026-09-13). It simply stops being load-bearing.

Four things fall out, none of which needed designing:

  1. Agent-VM binding stops being a special case. A GPIO VM named gpio binds by the same unchanged path as bob~user, because nothing appends anything to its name. There is no agent-binding mechanism to build.
  2. CONSOLE-ATTACH's one-name constraint can relax if ever wanted — it exists only to prevent a reconstruction that no longer happens. Not proposed as work; noted so the constraint is not later mistaken for a requirement.
  3. The silent-fallback class is dissolved rather than fixed. With no string to mismatch, the states are "a pairing is recorded" or "none is" — determinate and reportable. This retires punch item 19 (§XIX.7's "make mismatch loud") by removing what could mismatch, which is strictly better than making it noisy.
  4. The suffix-rename question becomes purely cosmetic (§XIX.5). Once lookup no longer derives the suffix, the remaining sites only ever construct a name.

XX.4 — The six memcpy sites split cleanly in two

This is what makes the change small. Traced 2026-09-18:

Site Role Under this ruling
capsule_wirebind.c:229 names the identity VM at birth stays — naming
mama_forth_words.c:1864 (UNATTENDED-BIRTH) names the registry entry stays — naming
mama_forth_words.c:1579 (PAIR-TEST) names both halves; diagnostic word stays — naming
repl.c:1247 re-derives the partner for a liveness guard goes
mama_forth_words.c:2000 (CONSOLE-ATTACH) re-derives the target goes

Only the two derivation sites change. The naming sites are untouched, which is why ~user can remain a convention without remaining a mechanism.

XX.5 — What "binding" means for an agent VM, stated because §V.1 conflated two things

§V.1 called Hestia "the bind point for users and agent VMs." Worth separating, since the two need different things and only one of them is binding:

  • Reachability — being addressable by other VMs — is kernel-Hermes's job (§III), not Hestia's. A GPIO VM is reachable because routing exists, not because it bound to anything.
  • Fabric access — a VM that wants to draw — is vocabulary and ACL, not binding: it needs Hestia's drawing words (§XVIII.3), gated the ACL.4th way if gated at all.
  • Binding proper — acquiring a console proxy so a human can type at a target — is what CONSOLE-ATTACH does, and after §XX.3 it works for any live VM by name.

So an agent VM does not routinely bind at all. It binds only when a human wants to attach to it for inspection or control — and that is the ordinary console path, unchanged. The "agent-binding problem" was an artifact of the hardcoded suffix; removing the suffix removes the problem rather than solving it.

XX.6 — Punch list

  • ✅ Item 3 — SETTLED (§XX.3): the bind records its target; nothing re-derives it. Agent binding needs no new mechanism.
  • ✅ Item 19 — RETIRED (§XX.3.3): dissolved by removing the mismatch, not by making it loud.
  • 🔶 Item 9 — unchanged: is an ordinary peer's SOS actionable or advisory, plus its number.
  • ⬜ §VIII.3 first bullet — where the sinking semaphore mechanically lives.
  • ⬜ Items 16, 17, 18 — unchanged (teardown/STADIUM-EVICT check; BYE vs. a separate suicide word; the empty-floor halt).

Caveat on scope, per §XIV.1. repl.c's relay and the routing-slot-3 convention are part of the FORTH messaging layer Captain Bob ruled legacy. This ruling states the principle — the pairing is recorded, never derived — which kernel-Hermes must carry forward. It is not a proposal to patch the legacy path in place.


XXI. Item 9 SETTLED: SOS travels up the birth graph. It is actionable there and advisory everywhere else.

§VII.2 left two questions: who consumes an SOS and what do they do with it, and which number. The first turns out not to need a new rule — §IX.2's "authority only ever flows down the birth graph" already answers it, and answering it supplies something this document had been missing without noticing.

XXI.1 — RULING: actionability is a property of the relationship, not of the message

  • To the sender's birther — actionable. The parent may re-birth. This is the parent exercising its own birth authority (§XV.1: any VM may birth), not inheriting the child's. §IX.2 is untouched: nothing flows sideways or up.
  • To every other VM — advisory. A peer may adjust its own behaviour (stop sending to a failing VM, note it, shed work that depended on it). A peer may not act on the sender's behalf, re-birth it, or reclaim anything of its own.

One rule, no exceptions list: an SOS is actionable exactly where the birth graph already grants authority.

XXI.2 — Why this does not smuggle in a supervisor

Pressure-tested against the doctrine and the rulings, because "parent watches child and restarts it" is precisely the shape §XIV.5 warned about:

  • Message-driven, not polling. The parent acts on receipt, during its own tick. That is the doctrine's own sanctioned form — VM-PHYSICS-DYNAMIC-FLEET-DESIGN-20260705.md: "a VM's own internal state (e.g. Hermes's message queue) determining whether it does anything." No timer, no scan, no watch loop.
  • Optional, never obligatory. A parent that ignores an SOS is behaving correctly. Nothing polices it.
  • No declarer. The sender reports its own trouble. Nobody judges anyone else's liveness, so §XV.2 holds exactly as ruled.
  • No authority transfer. The parent uses authority it already had, on a VM it already birthed.

A supervisor is a component that must watch and may decide another's fate. This is neither.

XXI.3 — What this actually closes: the ladder's missing actor

§VII.1's ladder has three rungs, and until now only the first and third had an obvious actor. With SOS settled, each rung has exactly one:

Rung Actor
1. Recover gently (warm restart) The VM itself, while still alive — self-applied (§XV.2)
2. Escalate to a from-scratch BIRTH The birther, on receipt of the SOS
3. Fail brutally Nobody. Compudynamic death (§XV.2)

And this is why §IX.1's fleet-shutdown rule is not a special case. Hera has no birther — capsule_birth.c:178 sets her name at the root, and the Tripod legs are birthed by her (kernel_main.c:865). So for Hera, rung 2 has no actor, and the ladder falls straight through to rung 3. §IX.1's "no provisional handoff → the fleet runs its own shutdown" is not an extra rule bolted on for Hera; it is what the ladder already does when you have no parent.

The ladder and §IX are the same rule seen from two ends: rung 2 exists if and only if you have a birther.

XXI.4 — An SOS must not carry an executable payload

A real hazard specific to this system, worth stating before anything is built. In the existing mechanism a message's payload is FORTH text executed by the receiver — messaging.4th block 5041 says so directly of BLK-ATTACH-EVENT: "Payload is FORTH text, executed by the receiver, same as every other message here."

An SOS whose payload were executed by the birther would let a failing VM drive its own parent — precisely inverting the authority direction §XXI.1 exists to preserve, and doing it at the moment the sender is least trustworthy.

So: an SOS payload is identifying and descriptive only — who is failing, and what for — never instructions. The birther decides what to do; the sender only reports. This is a deliberate departure from the "payload is executed text" convention, and the reason is the whole point of the message.

XXI.5 — SOS is best-effort, and that is not a gap

A VM that dies instantly emits nothing. There is no SOS, nobody notices, and per §XV.2 nobody declares anything — it is simply gone.

This is correct, not a hole to plug. Covering it would require something watching for absence, which is a detector, which §XV.2 rules out. SOS covers exactly the case where a VM can see its own failure coming; the silent case is covered by nothing, deliberately. (FABRIC-3.md §XX's wellness-check idea remains the separate, optional thing that could address it — on its own merits, not as a dependency of this design. §VII.3 was withdrawn for assuming otherwise.)

XXI.6 — RULING: SOS takes 10. Do not reuse the gaps.

Verified 2026-09-19: message types in use are 1, 2, 3, 4, 7, 8, 9; 5 and 6 are genuinely unallocated — a repo-wide search for either as a message-type constant returns nothing (the sole 6 is CH-CELLS, unrelated). Status sentinels sit at the top: MSG-NACKED 253, MSG-DELIVERED 255.

Take 10, the next in sequence, rather than filling a gap. The reasoning is legibility across the transition: during the dead-FORTH cleanup (§XIV.1) legacy constants and the new kernel enumeration coexist, and logs will span both eras. Monotonic allocation means a number appearing in any log means exactly one thing, forever. A reused gap would make an old log ambiguous for no gain. The gaps stay gaps; a documented hole is more honest than a silent reuse.

Scope note per §XIV.1: SOS properly belongs to kernel-Hermes's own message-type enumeration, since the FORTH layer is legacy. The legacy numbering is not binding on the new enum — but it should be honoured as a floor, for the same log-legibility reason.

XXI.7 — Named consumers, so this does not become another SPAWN-EVENT

§I.4 recorded the warning: SPAWN-EVENT was grepped across the whole tree and has zero consumers — "an unwired placeholder, not working infrastructure." A message type with no dispatcher is exactly what SOS cannot be.

So, stated explicitly: SOS's consumer is the sender's birther, whose handler is the rung-2 re-birth decision (§XXI.3), plus any peer that chooses to shed dependence on the sender (advisory, optional, and correct to omit entirely in a first implementation). If no birther is live, the SOS is simply never consumed — a determinate outcome, not a silent gap, because the birth graph says plainly whether a consumer exists.

XXI.8 — Punch list

  • ✅ Item 9 — SETTLED (§XXI). Actionable to the birther, advisory to peers; payload descriptive, never executable; number 10; consumer named.
  • ⬜ §VIII.3 first bullet — where the sinking semaphore mechanically lives. Now the last open design question in this document.
  • ⬜ Items 16, 17, 18 — the STADIUM-EVICT teardown check; BYE vs. a separate suicide word; the empty-floor halt when Hera is already gone.

XXII. Execution method: the surgical strip is punch item 1, and it is iterative (Captain Bob, 2026-09-19)

Captain Bob, verbatim: "As far as any of the existing forth we anticipate will no longer be used. That's probably going to get totally stripped out before we code. The existing gaps we need to fill in will reveal themselves as we code/test iterate and dead FORTH code falls out at the same time. That's going to be the first very surgical item on the punchlist."

So the method is: strip the anticipated-dead FORTH first, then code; further dead FORTH falls out as coding and testing reveal the real gaps, and gets stripped in the same loop. This becomes punch item 1, ahead of everything else in §XVIII.9 and §XX.

XXII.1 — Two categories, and only one of them can be stripped first

"Strip before we code" must not be read as "delete the working system before the replacement exists." The legacy FORTH splits cleanly, and the split is the ordering:

  • Category A — dead today, independent of this reshuffle. Nothing loads or invokes it now. Strippable immediately, before a line of new code, with no dependency on kernel-Hermes existing. This is the genuine first item.
  • Category B — dead on arrival of kernel-Hermes. capsules/hermes/init.4th, most of capsules/common/messaging.4th, the routing table, the slot-3 pairing convention (§XX.1). These are load-bearing until the replacement boots. Stripping them first would remove a working messaging layer with nothing behind it. They fall out during the iterate loop, as Captain Bob describes — not before it.

§XIII.1's fourteen-site inventory is the map for both; §XIII.2 already identified the Category A members it found.

XXII.2 — The trap: reachability here cannot be established by grep, in either direction

Recorded because this document nearly walked into it while building the strip inventory. A first pass counted textual references per capsule and produced a list with init-l8-diverse.4th, init-l8-omni.4th, init-l8-stable.4th, init-l8-temporal.4th, init-l8-transition.4th, init-l8-volatile.4th, sdk.4th and others at zero references — i.e. apparently dead.

They are not dead. They are DoE experiment infrastructure, and a second pass including experiments/, tools/ and docs/ found 18–19 references each for the init-l8-* family and 9 for sdk.4th. The first pass searched only capsules/ and src/.

The failure is structural, not a careless grep:

  • It under-reports. Capsules are birthed by name from runtime strings — S" <name>" BIRTH, campaign scripts, EXEC of a name assembled at run time. Static reference counting cannot see any of that.
  • It over-reports. A capsule named for a common word returns noise: process scores 180 hits across the tree, essentially all prose. (process.4th's actual deadness was established the other way — by confirming nothing EXECs it — not by counting.)

So no strip list in this document is authoritative, including any list this document produces. Reachability must be established per capsule by all three routes: loaded from a boot path, invoked from experiments//tooling, or present in the capsule directory mkcapsule bakes into the image.

XXII.3 — The acceptance asymmetry, which is what makes "surgical" necessary

.claude/CLAUDE.md is unambiguous that the three-architecture QEMU boot is the only acceptance test there is. For this particular task it is necessary but not sufficient, and the gap is specific:

  • A boot exercises the boot path. Deleting something Category A and boot-reachable fails loudly and immediately — good.
  • A boot does not exercise DoE/experiment capsules. Deleting one passes all three architectures cleanly and surfaces months later, when a campaign is run.

That asymmetry is the real hazard of this item. And the stakes are not only code: .claude/CLAUDE.md records that the logs/ artifacts "are audit artifacts — they are committed to the repo. Do not delete them," and that the DoE campaign report is patent support material. Experiment infrastructure is part of a measurement apparatus with an evidentiary role.

Rule, proposed: treat every experiments/-reachable and DoE-related capsule as live by default. Removing one is its own decision with its own justification, never a side effect of a messaging cleanup.

XXII.4 — How the loop should run

Consistent with this series' own discipline and with the acceptance criteria:

  1. Strip Category A only, in a commit separate from any new code, so a single file can be restored without unpicking a feature. Nothing else in the same commit.
  2. Boot all three architectures after the strip, before writing anything new — establishing that the strip alone is clean, rather than discovering later which half of a mixed commit broke something.
  3. Then code, and let the gaps reveal themselves as Captain Bob describes.
  4. Each time coding proves a Category B item dead, strip it in its own commit too, with its own three-arch boot.

Expected and not a defect: every strip changes dict_hash. Per §III.5 and §XX's standard, the property that must hold is cross-architecture identity, never an unchanging absolute value. mkcapsule --lint capsules/ must stay clean throughout, and freed block ranges should be returned to capsule-reserved.txt rather than silently reused (§XVIII.4).

XXII.5 — What rides along

  • The MANIFEST.md corrections (§XIII.2, formerly item 13) belong to this pass: block 4055's "immutable ABI" claim that FABRIC-2.md:2773 already declared stale, and block 2049's wrong contents list. Correct them as the files they describe are stripped, so the manifest never describes a file that no longer exists.
  • SPAWN-EVENT (§I.4: zero consumers, confirmed twice) is Category A by definition.
  • src/*.c.bak — .claude/CLAUDE.md flags vm.c.bak, doe_metrics.c.bak and inference_engine.c.bak as tracked-but-stale repo hygiene debt, to "report it if it comes up; don't delete unprompted." It has now come up. Not part of this reshuffle, but the natural companion pass — flagged, not scheduled, and still requiring its own authorization.

XXII.6 — Punch list, re-ordered

1. ⬜ SURGICAL STRIP — Category A dead FORTH. Establish reachability per §XXII.2's three routes (never by grep alone), strip in isolated commits, three-arch boot each time, mkcapsule --lint clean, MANIFEST corrected alongside, freed blocks returned. Precedes all other items.

Then, unchanged in content but now downstream of it:

  1. ⬜ Hestia: relocate fabric.4th + font.4th; move PLOT/FB-WIDTH/FB-HEIGHT registration (§XVIII.9.1).
  2. ⬜ Hestia's block range against capsule-reserved.txt, avoiding 4997 (§XVIII.9.2).
  3. ⬜ Hestia into is_fleet_foundation; kernel_main.c:865 birth; switch-signal registration (§XVIII.9.4, §XIX.6).
  4. ⬜ doe_log.c CSV schema change (§XIII.3) — still the expensive consequence.
  5. ⬜ State §XVIII.6's headless invariant in the implementation.
  6. ⬜ §VIII.3 first bullet — where the sinking semaphore mechanically lives. The last open design question.
  7. ⬜ Item 16 — every message-holding teardown path reaches STADIUM-EVICT; verify via fleet_conserved (§XV.4).
  8. ⬜ Item 17 — Hera's suicide: replace BYE's cold-restart, or a separate word (§XVI.2).
  9. ⬜ Item 18 — the empty-floor halt when Hera is already gone (§XVI.7).
  10. ⬜ Category B strips, each as coding proves the item dead (§XXII.1, §XXII.4.4).

XXIII. §VIII.3 SETTLED: the sinking latch is a kernel-resident scalar, read through a registered primitive

The last open design question in this document. §VIII.1 ruled what the signal is and why it is not a message; this settles where it mechanically lives and who reads it.

XXIII.1 — What the mechanism has to satisfy

Collected from the rulings, because together they constrain the answer almost completely:

  1. Readable by a VM whose messaging is dying (§VIII.2) — so it cannot itself route.
  2. Below the routing layer (§III.2) — kernel-resident, in kernel-Hermes's own scope.
  3. Must not require Hermes to still be working — the whole point is that the transport is the thing that failed.
  4. Must not make the kernel a supervisor (§XV.3, §XIV.5) — the kernel may publish a fact; it may not decide another VM's fate.
  5. Must be actionable by a VM that has no thread of its own — per FABRIC-3.md §XXVIII there is no per-VM native stack; every VM runs on the one shared kernel C stack and only executes when dispatched.

XXIII.2 — RULING: a single-writer kernel scalar, exposed as a registered primitive

The latch is a plain scalar in kernel-Hermes's own translation unit, written only by Hermes, read by every VM through a primitive registered into its dictionary — e.g. SINKING? ( -- flag ).

This invents nothing. The registration path is the one §XX of FABRIC-3.md already established and fixed: register_child_vm_words() hands the eight STADIUM-* primitives to every child VM, and Hera's two registration sites mirror them so her dictionary is a proper superset. A ninth kernel-state reader is that same pattern, not a new mechanism — and it satisfies (1) and (3) exactly, because reading a scalar needs no queue, no channel, no routing table and no working arbiter.

Deliberately a plain unconditional primitive, not gated. Same doctrine CONSOLE-ATTACH and ZUSE-ELIGIBILITY-ADD already carry (§XXXII.2 Q3): restricting who may call it, if that is ever wanted, is ' SINKING? ACL-PIN in ACL.4th, never a bespoke C check. Reading a fact is not an authority.

Keep it a bare scalar, not a field in a larger struct. Robustness, not style: the reader must get a coherent answer while the writer's own subsystem is failing. A standalone word-sized value cannot be observed mid-update; a field inside a structure being torn down can.

XXIII.3 — It is a one-way latch, not a counting semaphore — and that is what makes it safe

"Semaphore" was the word in the original framing; the accurate word is latch, and the distinction matters both for naming and for correctness.

It is monotonic. Per §VII.1, a VM attempts its own gentle recovery before anything is announced. So Hermes raising the latch means recovery has already failed and it is committed to going down — there is no path back to "not sinking." Raise once; never lower.

That gives a concurrency story with nothing in it:

  • Single writer (Hermes, mainline), many readers, one irreversible transition. A reader either sees the old value or the new one, and both are valid — seeing "not sinking" one moment before the raise is indistinguishable from having read a moment earlier, which is fine.
  • This is the discipline this codebase has already sanctioned twice, and reuses on purpose. heartbeat.c:33-34 states it verbatim: "Single writer (mainline, via vm_tick()'s Loop #7 site), single reader (the ISR's re-arm call) — no lock needed, per §21.1's finding." FABRIC-3.md §XXVIII Stage 3 then reused that same sanctioned pattern for a second variable rather than "inventing new locking or overturning the ruling itself." This is the third such variable, and it takes the same route for the same reason.

Naming it a semaphore in code would imply counting and blocking semantics it does not have and must not acquire. Call it what it is.

XXIII.4 — Who reads it, and when: the kernel delivers the fact, the VM performs the act

The subtle part, and the one constraint (5) forces. A VM cannot poll. With no per-VM thread (§XXVIII), a VM only runs when dispatched — so "every VM watches the latch" is not implementable as stated.

So the check belongs at the dispatch point, and the split of responsibility is the whole design:

On giving a VM its turn, the kernel checks the latch. If raised, that VM runs its own clean-shutdown routine — blk_flush(0) then BYE (§XVI) — in its own context, instead of its normal work.

  • The kernel publishes a fact and delivers it. It does not shut anyone down, does not decide who dies, does not order anything. §XV.3 and §XIV.5 hold.
  • Each VM performs its own shutdown, which is exactly what §IX.1 already ruled for fleet-wide shutdown ("the rest of the fleet runs its own shutdown routines") and §XVI defined the content of.
  • No timer, no watcher, no supervisor. A VM acts when it next runs, which is the doctrine's own sanctioned "a VM's own internal state determining whether it does anything."

And the window closes by itself. Per §XVI.4, Hera's suicide bounds it: children flush and BYE as they are dispatched, Hera reaps what remains and halts the processor. §VIII.3's "is the window bounded?" needs no deadline because the last step is physically terminal — there is nothing to time out.

XXIII.5 — The two-mechanism rule, now complete

§VIII.2 predicted the shape; with §XXI and §XXIII both settled it closes cleanly, with no exceptions list to maintain:

Sender is… Signal Why
On the Stadium floor (Hera, Artemis, Hestia, identity VMs, agent VMs) SOS, a routed message (§XXI) Routing is available to them; it travels up the birth graph to the one VM with authority to act
The arbiter beneath the floor (kernel-Hermes) the sinking latch (§XXIII) It cannot route a message announcing that routing has failed

Which mechanism a VM uses is determined entirely by which side of the routing layer it sits on. Nothing is special-cased, and there is no list of exceptions to keep in sync — a property worth defending if a future component is ever tempted to want "just a small" third signal.

XXIII.6 — Punch list: the design phase is complete

Every design question this document opened is now ruled. What remains is build sequencing and three small local decisions:

  1. ⬜ SURGICAL STRIP — Category A (§XXII.6). Precedes everything.
  2. ⬜ Hestia: relocate fabric.4th + font.4th; move PLOT/FB-WIDTH/FB-HEIGHT (§XVIII.9.1).
  3. ⬜ Hestia's block range, avoiding 4997 (§XVIII.9.2).
  4. ⬜ Hestia into is_fleet_foundation; birth at kernel_main.c:865; switch registration (§XIX.6).
  5. ⬜ doe_log.c CSV schema (§XIII.3) — the expensive consequence.
  6. ⬜ §XVIII.6's headless invariant stated in the implementation.
  7. ⬜ Item 16 — teardown paths reach STADIUM-EVICT; verify via fleet_conserved (§XV.4).
  8. ⬜ Item 17 — Hera's suicide: replace BYE's cold-restart, or a separate word (§XVI.2).
  9. ⬜ Item 18 — the empty-floor halt when Hera is already gone (§XVI.7).
  10. ⬜ Category B strips, each as coding proves the item dead (§XXII.4).

Items 8 and 9 are the only two carrying an unmade decision; the rest are execution. Nothing on this list is authorized by this document — per Captain Bob's Law, no code without an explicit instruction.


XXIV. Item 17 SETTLED: Hera's suicide is a separate word; BYE is left alone (Captain Bob, 2026-09-19)

Captain Bob, 2026-09-19: "Separate word for Hera's suicide, leave BYE alone."

XXIV.1 — What this buys, beyond settling the question

mama_word_bye() (mama_forth_words.c:2265-2271) keeps its current behaviour exactly: capsule_vm_kill_all_nonmama() then arch_cold_reset() — reap children, cold-restart the machine. That is the right thing for an operator typing BYE at Hera's REPL (§XVI.2 flagged this as the argument for two words), and it stays untouched.

The larger consequence: this reshuffle now modifies the behaviour of zero live registered words. §XVI.2 flagged Hera's BYE as the one place the design proposed changing one, and noted that .claude/CLAUDE.md's hard rule — "Never modify a registered, tested word to 'fix' it" — did not strictly apply to a ruled change but that the change should be made knowingly. It is now not made at all. Every word this design touches is either new or unchanged.

XXIV.2 — The two words differ only in their terminal action

They share their first half. Per §XVI.1 and §XXIII.4 everything already exists:

BYE (unchanged) the suicide word (new)
Reap remaining children capsule_vm_kill_all_nonmama() same
Terminal action arch_cold_reset() — reboot arch_disable_interrupts() + for(;;) arch_halt(); — stop
Registered on Hera only Hera only

The permanent-halt idiom is not new either — it is the pattern amd64's own arch_cold_reset() already uses as its unreachable fallback (for (;;) __asm__ volatile ("cli; hlt");, arch.c:218) and the same shape the kernel panic path uses. arch_halt() is declared at arch.h:64 and implemented on all three architectures (amd64:207, aarch64:126, riscv64:170), so three-arch parity holds with no per-arch work.

Registration: Hera-only, the same way mama_word_bye is registered only in register_mama_forth_words() and never by register_child_vm_words(). A child VM must not have it — §IX.1's no-handoff rule means no other VM may stop the machine on Hera's behalf.

Checked 2026-09-19 — HALT, SCUTTLE, SINK, DIE, EXPIRE, SUICIDE and GO-DOWN are all unregistered; none collides with an existing word.

Recommendation: SCUTTLE. To scuttle is to deliberately sink the vessel you command — which is exactly what this word does, performed by the one VM with the standing to do it. It also rhymes with the imagery already in the design without colliding with it: Hermes sinks (§VIII.1, a thing that happens to it), Hera scuttles (a thing she does).

HALT is the obvious alternative and the one caution worth stating: it reads naturally and matches arch_halt(), but vm->halted already exists and means something quite different — a single VM stopping, not the machine. Reusing the stem invites exactly the ambiguity §XIX renamed a Tripod leg to avoid.

Captain Bob's call; this section does not decide it.

XXIV.4 — Where its ACL pin goes: the project's own docs give three different answers

The new word is kernel-only and privileged, exactly like BIRTH — so it needs the same ACL treatment, and traced 2026-09-19, there is no single agreed answer in this repo:

  1. .claude/CLAUDE.md's hard rule: "ACL policy belongs in ACL.4th, never in C. No policy logic in kernel_main.c, no vm_find_word + field assignment for pinning."
  2. .claude/CLAUDE.md, later: "' BIRTH in shared capsules breaks the hosted build — BIRTH is kernel-only. Pin it in a kernel-specific capsule, not in ACL.4th."
  3. capsules/ACL.4th's own comment (block 4005): "BIRTH/CAPSULE-BIRTH are omitted: kernel-only, not in hosted VM. Pinned in C (kernel_main.c) after capsule load instead, so this file stays host-portable."

The live code does (3), and (3) is what (1) forbids. kernel_main.c:771-782 is literally vm_find_word(mama_vm_ptr, "BIRTH", 5) followed by birth->acl_mode = ACL_MODE_STRICT; birth->acl_pinned = 1; — a vm_find_word + field assignment for pinning, in kernel_main.c, printing "ACL: BIRTH pinned STRICT". It works and it is deliberate; it simply contradicts the stated rule.

Recommendation: follow the live precedent — pin the suicide word in kernel_main.c beside BIRTH and CAPSULE-BIRTH — on the grounds that matching working code beats matching a rule the working code already breaks, and that a lone exception is worse than a consistent one. BYE's own pin stays where it is, in ACL.4th:70 (['] BYE ACL-STRICT ['] BYE ACL-PIN), which is correct because BYE is not kernel-only.

Reported, not fixed (Captain Bob's Law): .claude/CLAUDE.md carries two statements that disagree with each other and with the code. Added as punch item 20 — a documentation reconciliation, outside this reshuffle, and not something to resolve by quietly editing one of the three.

XXIV.5 — Punch list

  • ✅ Item 17 — SETTLED (§XXIV): separate word, BYE untouched. Reshuffle now changes zero live registered words.
  • ⬜ Item 18 (§XVI.7) — the empty-floor halt when Hera is already gone. The last undecided item in this document.
  • ⬜ Item 20, NEW (§XXIV.4) — reconcile the three-way ACL-pinning contradiction between .claude/CLAUDE.md (twice) and ACL.4th/kernel_main.c. Documentation, not code; outside this reshuffle.
  • Build sequencing otherwise unchanged (§XXIII.6), with the surgical strip still item 1.

XXV. DoE constraint relaxed; a formal re-verification pass closes the sequence (Captain Bob, 2026-09-19)

Captain Bob, verbatim: "Don't worry about the DoE, I'm anticipating a rewrite of its invocative paths. AND at the end redo a lot of Isabelle/HOL."

XXV.1 — RULING: the DoE CSV schema stops being "the expensive consequence"

§XIII.3 named doe_log.c's six Tripod-hardcoded columns as the costly part of changing Tripod membership, and §XVII.4/§XIX.6 carried it forward as such. With the DoE's invocative paths slated for rewrite, that cost is absorbed rather than paid. Two items collapse:

  • Punch item 5 (the CSV schema change) is no longer a constraint on the reshuffle. The columns become hestia_heat_q48/switch_hestia_readiness as part of the DoE rewrite, not as a reluctant edit to a schema that had to stay comparable.
  • §XIII.1's sites 9–13 — doe-campaign.4th's five S" Hermes" VM-EXEC calls — fold into the same rewrite. They were already classed "not auto-run"; they are now explicitly someone else's problem, in a good way.

One thing this does not relax, stated to avoid over-reading it. §XXII.3's strip-safety rule stands on its own footing: the hazard there was a naive grep deleting init-l8-*, the workloads and sdk.4th by accident, passing all three architectures cleanly and surfacing months later. Rewriting how the apparatus is invoked is not the same as discarding the apparatus, and its evidentiary role is unchanged — .claude/CLAUDE.md still records the logs/ artifacts as committed audit artifacts and the campaign report as patent support material. Experiment capsules remain live-by-default during the strip.

XXV.2 — RULING: Isabelle/HOL re-verification is the closing item

A formal re-proof pass runs at the end, after the build work. That gives the sequence a deliberate symmetry: the surgical strip opens it (§XXII.6 item 1), formal re-verification closes it.

Correction to .claude/CLAUDE.md while establishing scope. It states proof/ "contains 23 Isabelle/HOL theory files (same composition as the standalone StarForth repo — 18 core VM/ word-category theories + 5 ACL theories)." Traced 2026-09-19: there are 52 .thy files — 5 ACL_* and 47 StarForth_* — and all 52 are listed in proof/ROOT, so none is orphaned. proof/COVERAGE.md says 52 and is correct; CLAUDE.md is stale by more than a factor of two. Added to item 20's documentation reconciliation. Build invocation is isabelle build -D proof/ directly; neither Makefile defines a target (that part of CLAUDE.md is accurate).

XXV.3 — The standard this pass must meet is already written, and it is not "the proofs build"

proof/COVERAGE.md states the goal in Captain Bob's own framing:

Prove the StarForth VM system as close to bare metal as possible under Isabelle/HOL, and — just as importantly — identify precisely what cannot be proven and why. A clean pass/fail isn't the deliverable; the boundary between "formally verified" and "not, for this specific reason" is.

So the deliverable of the closing pass is a restated boundary, not a green build. A suite that still compiles while quietly covering less than it did would satisfy the build and fail the actual standard.

XXV.4 — The honest cost: this reshuffle moves messaging out of the provable region

Not previously stated anywhere in this document, and it should be, before it is discovered. COVERAGE.md defines what the suite can reach: words "that operate purely on modelled per-VM state (stacks, dictionary, and the ~40 scalar fields this suite has added to an abstract vm_state record)" — and, explicitly, where an implementation "reaches outside that model (raw pointers, file-scope statics, TIB/stdio, an unmodelled subsystem), the theory says so explicitly rather than silently modelling something else."

Today's messaging layer is FORTH over per-VM dictionary state — inside the model. Kernel-Hermes makes it C with file-scope statics — outside it. Likewise the sinking latch (§XXIII.2) is by design a kernel-resident scalar.

So the reshuffle is very likely a net reduction in formal coverage of the messaging area. That may well be the right trade — §III.2's circular-dependency argument for moving it is strong — but it is a real cost, and by this project's own standard the pass must name it rather than let the boundary quietly move.

Partially offsetting, and cheap: the one-way latch is close to the easiest thing in this design to prove. Single writer, one irreversible transition, both observable values valid (§XXIII.3) — and StarForth_Mutex.thy, StarForth_Concurrent.thy and StarForth_Transition.thy already exist as the homes for exactly that reasoning.

XXV.5 — What the pass will actually touch, mapped

Theory Why this reshuffle touches it
ACL_No_Escalation.thy The central one — see below.
ACL_Inherit_Clears_Pin.thy §XV.1's "almost a clone with its own ACL DNA" is inheritance semantics
ACL_Pin_Monotone.thy Same monotonicity shape as the §XXIII.3 latch
StarForth_System_Words.thy BYE lives here; §XXIV adds a Hera-only sibling beside it
StarForth_Framebuffer_Words.thy PLOT/FB-WIDTH/FB-HEIGHT registration moves to Hestia (§XVIII.3)
StarForth_Lifecycle_Words_Hosted.thy Birth/lifecycle adjacency
StarForth_Mutex / _Concurrent / _Transition The sinking latch's safety argument

The observation worth carrying into the pass: ACL_No_Escalation.thy already proves a no-escalation property at the word level. §VI.1's central safety claim — a VM can never birth something with more authority than it itself holds — is the same theorem one level up, at VM granularity. And §IX.2's "authority only ever flows down the birth graph, never sideways and never up" is a graph invariant of exactly the kind this suite is good at.

The reshuffle's most important safety property is therefore already half-proven, at a different granularity. Lifting it is a better-defined task than it would be from a standing start — and per §VI.1 the property is structural (enforced by inheritance, no gate to bypass), which is precisely the kind of claim that survives formalization.

XXV.6 — Punch list

  • ✅ Item 5 (DoE CSV schema) — RELAXED (§XXV.1). Absorbed by the DoE invocative-path rewrite; no longer a constraint.
  • ⬜ Item 21, NEW — the closing Isabelle/HOL pass (§XXV.2–XXV.5). Runs last. Deliverable is a restated boundary per §XXV.3, explicitly including §XXV.4's coverage loss, not a green build.
  • ⬜ Item 20 grows (§XXV.2): CLAUDE.md's "23 theory files" joins the ACL-pinning contradiction in the documentation reconciliation.
  • ⬜ Item 18 (§XVI.7) — the empty-floor halt. Still the last undecided design item.
  • Sequence otherwise unchanged (§XXIII.6): surgical strip first, build, formal re-verification last.

XXVI. The close protocol: doc sweep → SBOM → master HEAD → tag → close this document (Captain Bob, 2026-09-19)

Captain Bob, verbatim: "and documentation sweep, SBOM + spdx before we make it masters head and tag it. That's when FABRIC-3.5.md is closed."

So the tail of the sequence is fixed, and this document now has a defined close condition:

  1. Build work (§XXIII.6), ending with the Isabelle/HOL pass (§XXV.2).
  2. Documentation sweep (§XXVI.1).
  3. SBOM + SPDX regeneration (§XXVI.2).
  4. Merge to master HEAD (§XXVI.4).
  5. Tag (§XXVI.3).
  6. Close FABRIC-3.5.md (§XXVI.5).

XXVI.1 — The documentation sweep, inventoried

Item 20 has been accumulating known-stale claims throughout this document. Consolidated, so the sweep is a checklist rather than a hunt. Every entry below was traced, not recalled:

.claude/CLAUDE.md — four separate errors found during this design pass:

  • "23 Isabelle/HOL theory files (18 core + 5 ACL)" → 52, all in proof/ROOT (§XXV.2).
  • LITHOS_VERSION "currently 2.0.1" → 2.0.0 (Makefile.starkernel:83); FABRIC-3.md §I.2 corrected it down on 2026-09-04 and CLAUDE.md never followed.
  • The ACL-pinning three-way contradiction (§XXIV.4) — its own hard rule vs. its own later advice vs. ACL.4th/kernel_main.c's live behaviour.
  • mkcapsule's "1024-byte-per-block" framing → the real constraints are block range [2048, 5120) and a 16-content-line cap (§XVIII.4, per §XXXII.2's own correction, still outstanding).
  • Plus, after this reshuffle: the Tripod described as "(Hera the Mama VM, two Hermes instances, Artemis)" becomes Hera / Artemis / Hestia, and the superseded-subsystem note must add FABRIC-3.5.md.

capsules/MANIFEST.md (§XIII.2) — block 4055's "immutable ABI" claim that FABRIC-2.md:2773 already declared stale, and block 2049's wrong contents list. Correct as the files they describe are stripped (§XXII.5), so the manifest never describes a file that no longer exists.

Superseded subsystem docs — .claude/TRIPOD.md, HERMES.md, ARTEMIS.md, CONSOLE.md already carry superseded headers, but after this reshuffle they describe a fleet that no longer exists. Same for docs/03-architecture/tripod/{TRIPOD,HERMES,ARTEMIS,README}.md, which .claude/CLAUDE.md still lists as current for Tripod. Decide deliberately: update, or mark archival and point at FABRIC-3.5.md. Not a silent edit either way.

Also in scope: capsules/README.md (lists hermes/init.4th as a live personality), ROADMAP.md / docs/lithosananke/ROADMAP.md, CHANGELOG.md, README.md, docs/lithosananke/SYSTEM_ARCHITECTURE.md and M7.1.md, and proof/COVERAGE.md (after the §XXV pass — and per §XXV.3 its boundary statement matters more than its file count). capsules/BLOCK_MAP.md is a generated artifact; regenerate, don't hand-edit.

XXVI.2 — SBOM: a real target, and currently stale

make sbom (Makefile:740-752) runs syft: syft dir:. -o spdx-json=sbom.spdx.json -o spdx=sbom.spdx. Reproducible, not hand-maintained — so this step is "run the target and commit the output," provided syft is available.

Two things to check rather than assume:

  • The committed SBOM was generated 2026-07-24 (sbom.spdx's own Created: field) — roughly two months stale before this reshuffle adds or removes a single file.
  • DocumentName: StarForth, not LithosAnanke — plausibly an artifact of the monorepo split, since syft derives it from the directory name. Worth a look during the sweep; a bill of materials naming the wrong product is exactly the kind of thing a licensee or patent reviewer reads literally.
  • The target lives in the hosted Makefile, not Makefile.starkernel.

XXVI.3 — The tag: version semantics here are real, and 2.0.1 is not available

Tags are sacred ground (.claude/CLAUDE.md) — "each tag has logs attached proving its state." And this project has already corrected itself once on exactly this, so the policy must be read before choosing, not after.

Per FABRIC-3.md §I.2 (2026-09-04), Makefile.starkernel's own versioning policy is semantic, not sequential:

  • v2.0.0 = the QEMU release line, even-major/LTS.
  • v2.0.1 = the SER5 hardware-track line (RDRAND backend + thumbdrive image goal).

§I.2 rolled LITHOS_VERSION back from 2.0.1 to 2.0.0 precisely because "claiming 2.0.1 implies hardware-track progress that was never actually verified on real hardware." That verification is still FABRIC-3.md's open topic. This reshuffle is QEMU-only work.

So 2.0.1 must not be used for this, and the version for the reshuffle is a genuine decision — Captain Bob's, not this document's. The shape of it: a structural Tripod change on the QEMU line reads as a minor bump (v2.1.0) rather than a patch, and VERSION (the embedded engine, currently 3.1.0) likely moves too, since BIRTH registration, the new Hera-only word and the messaging layer all change. Both are hand-edited in Makefile.starkernel — the bump-z/bump-y targets were removed 2026-08-15 as non-functional and must not be resurrected from memory.

XXVI.4 — Unexpected remote state, to investigate before the merge — not to clean up

Traced 2026-09-19 against origin. Reported, not acted on, per the standing rule about unfamiliar state:

  • refs/heads/v2.0.1 still exists at fcba5282. FABRIC-3.md §I.2 item 5 records that branch as deleted "local and origin" on 2026-09-04, having first confirmed it a strict ancestor, and concludes "master is the repo's only branch from here on." It is not. Either the deletion did not take or the branch was recreated. Worth resolving before a tag is cut, since the whole point of §I.2 was that the v2.0.1 name carries a claim.
  • refs/pull/1/head exists at bcf41e5 — this document's own §XIII commit. A pull request was opened against this branch at some point; this session did not create one, and .claude/CLAUDE.md's subversion-like workflow says "no pull requests unless explicitly requested." Check what it is before merging, in case the intended path is through it.

XXVI.5 — How this document closes

FABRIC-3.5.md closes when the tag exists. The close is a dated header in the form FABRIC-2.md uses — "Status: CLOSED/ARCHIVAL as of <date>" — stating what the document produced and naming the tag it closed at.

Two things the close must get right, both established in this document's own opening:

  1. It does not trigger the carry-forward chain. The FABRIC-0 → -1 → -2 → -3 discipline triggers on closing a document and hands its open items to a successor. This document is not in that chain — it is a standalone topic document — so closing it creates no successor and hands nothing forward.
  2. FABRIC-3.md remains open, living and authoritative for its own topic (bare metal boot). Nothing here supersedes it, and closing this document does not close that one.

Any item still open at close time is either carried into FABRIC-3.md explicitly (if it has become bare-metal-boot work) or recorded in the close header as deliberately unfinished — never silently dropped.

XXVI.6 — Punch list, complete

Design phase: complete. One item (18) still carries an unmade decision; the rest is execution.

# Item §
1 Surgical strip — Category A. Precedes everything §XXII.6
2 Hestia: relocate fabric.4th + font.4th; move PLOT/FB-* registration §XVIII.9
3 Hestia's block range vs. capsule-reserved.txt, avoiding 4997 §XVIII.4
4 Hestia into is_fleet_foundation; birth at kernel_main.c:865; switch registration §XIX.6
5 §XVIII.6's headless invariant stated in the implementation §XVIII.6
6 Kernel-Hermes; the sinking latch; SOS type 10 §III, §XXI, §XXIII
7 BIRTH generalization + ACL/DNA inheritance §VI, §XV.1
8 The Hera-only suicide word (SCUTTLE recommended); pin beside BIRTH §XXIV
9 item 18 — CLOSED §XXVII: empty floor unreachable, build nothing §XXVII
10 Teardown paths reach STADIUM-EVICT; verify via fleet_conserved §XV.4
11 Category B strips, each as coding proves the item dead §XXII.4
12 Isabelle/HOL pass — deliverable is the restated boundary §XXV
13 Documentation sweep §XXVI.1
14 make sbom; check Created: and DocumentName §XXVI.2
15 Resolve the stray v2.0.1 branch and PR #1 §XXVI.4
16 Version RULED 2.1.0 (§XXVIII); merge to master, tag v2.1.0. Engine VERSION still open §XXVIII
17 Close this document §XXVI.5

No code is authorized by this document. Per Captain Bob's Law, nothing on this list starts without an explicit instruction.


XXVII. Item 18 SETTLED: the empty floor is unreachable — build nothing

§XVI.7 raised it and §XVI.7's own proposed shape (the kernel observing an empty floor and halting) was offered as analysis. Tested against the code 2026-09-19, the premise is false: the state cannot occur. The ruling is therefore to build nothing, and the value of the item is the invariants it surfaced.

XXVII.1 — The premise, tested

An empty Stadium floor requires Hera to be gone. Every other VM is her descendant, and capsule_vm_kill_all_nonmama() spares her by construction. So item 18 reduces to: can Hera cease to exist while the machine keeps running?

Five paths, all traced:

Path Result Evidence
Compudynamic death (heat decay / COOL) Impossible — she is pinned stadium.c:453 if (header->flags & STADIUM_FLAG_PIN) return -1;; :559 if (h->flags & STADIUM_FLAG_PIN) continue;. Hera is pinned via session_set_pinned() (§H.10/§H.12 step 4)
KILL Impossible — refused mama_word_kill()'s own doc: "Destroy a named VM unconditionally. Hera cannot be killed."
Reaped with the children Impossible — spared by name capsule_vm_kill_all_nonmama()
BYE Machine reboots arch_cold_reset() (§XXIV.2) — not a lingering state
Scuttle (§XXIV) / panic Machine halts Both end in a permanent arch_halt() loop

Every path by which Hera can cease already terminates the machine. There is no route to a live kernel over an empty floor.

Note the pinning result generalises: all three Tripod legs are pinned by the same is_fleet_foundation line (§XIII.3), so no Tripod leg can die a compudynamic death. §XV.2's compudynamic death applies to the unpinned population — identity VMs, agent VMs — which is exactly where it was always aimed.

XXVII.2 — RULING: build nothing

.claude/CLAUDE.md is direct about this class: "Don't add error handling, fallbacks, or validation for scenarios that can't happen." An empty-floor halt would be precisely that — defensive code for an unreachable state, carrying its own maintenance and its own risk of firing wrongly. §XVI.7's proposed shape is withdrawn, not adopted.

A supporting argument, recorded but not load-bearing (the ruling stands on §XXVII.1 alone): even were the state reachable it would be provably terminal rather than a judgement call, since BIRTH is a word a VM must call — zero VMs means zero possible births, so nothing could ever run again. That is a dead end, not a decision. But it is moot.

XXVII.3 — The panic path already is the empty-floor halt

sk_hal_panic() (hal/hal.c:490-506) prints System halted. and executes:

while (1) {
    arch_halt();
}

That is character-for-character the idiom §XXIV.2 specified for Hera's suicide word. The two converge on the same terminal state by different triggers, which is a pleasing consistency rather than a duplication: a scuttle is a panic Hera chose. Whatever kills her involuntarily already routes here; what she does deliberately looks the same from outside. Nothing new is needed at either end.

XXVII.4 — What this ruling depends on: three invariants to defend

The ruling is conditional on facts that a future change could break silently. State them as invariants, since that is the durable part of this item:

  1. Hera stays pinned. If she is ever admitted unpinned, or the is_fleet_foundation triple (capsule_birth.c:793-796) stops covering her, she becomes evictable and the empty floor becomes reachable. Note this triple is edited by this very reshuffle (§XIX.6, Hermes out, Hestia in) — so the change most likely to break this is one already on the punch list.
  2. KILL keeps refusing Hera, and any new teardown path spares her the way capsule_vm_kill_all_nonmama() does. §XXVIII's already-confirmed unguarded-kill UAF shows this area has been got wrong before.
  3. Every involuntary-death path ends at sk_hal_panic() or an equivalent permanent halt — never at a return that leaves the kernel looping over nothing.

If all three hold, the empty floor cannot occur and no code is needed. If any is broken, this ruling must be revisited rather than patched around.

XXVII.5 — Design phase complete

Every design question this document opened is now ruled. Item 18 was the last, and it closes by dissolution — the fourth to do so, after §XIV.2's four, §XV.2's two and §XX's one.

That pattern is worth naming as the document closes its design phase: of the questions that dissolved rather than resolved, every one turned out to rest on a premise the code did not support — a supervisor that was not needed, a declarer that could not exist, a suffix that was never load-bearing, a state that cannot occur. The design got simpler each time it was checked against the tree rather than reasoned about in the abstract. That is the single most useful habit to carry into the build.

Punch list (§XXVI.6) stands as written, with item 9 now closed and one decision outstanding: the version number (§XXVI.3). No code is authorized.


XXVIII. RULING: this release is 2.1.0 (Captain Bob, 2026-09-19)

LITHOS_VERSION becomes 2.1.0, and the tag cut at §XXVI's close is v2.1.0.

XXVIII.1 — Checked against the real policy, not §I.2's paraphrase

§XXVI.3 worked from FABRIC-3.md §I.2's summary. The authoritative statement is the roadmap table in Makefile.starkernel itself (lines ~74–82), which is more specific:

# Roadmap (per docs/lithosananke/ROADMAP.md "Release Versioning Policy" and
# FABRIC-2.md §G — X.0.0 = QEMU release, X.5.0 = hardware bare-metal release):
#   v1.0.x — serial-only production (released)
#   v1.5.x — framebuffer VT100 terminal/console milestone (released)
#   v2.0.0 — QEMU release (even major = LTS): three-arch QEMU story complete
#   v2.0.1 — SER5 hardware-track line: RDRAND backend + generic thumbdrive image goal
#   v2.2.0 — amd64 bare-metal (Beelink SER5) — see ROADMAP "Board-by-board rollout"
#   v2.5.0 — hardware bare-metal release: real per-arch RNG + real-board boot

2.1.0 is unallocated — the ladder currently steps 2.0.0 → 2.0.1 → 2.2.0 → 2.5.0. So the ruling takes a genuine gap rather than overloading an existing rung, and it lands correctly:

  • Below v2.2.0 (amd64 bare-metal) and v2.5.0 (hardware bare-metal release) — right, because this reshuffle is QEMU-only work and must claim no hardware progress. That was §I.2's entire concern when it rolled the version back.
  • Does not disturb v2.0.1's meaning as the SER5 hardware-track line. That name keeps what it means; nothing is reused.
  • Consistent with X.0.0 = QEMU / X.5.0 = hardware bare-metal — a minor bump inside the 2.x QEMU-side band, before the hardware rungs.

A structural change to the Tripod, a new kernel-resident arbiter and a generalized BIRTH are comfortably a minor-version event rather than a patch.

XXVIII.2 — The roadmap table must gain a 2.1.0 line, in three places

The table above is not decoration — it is how the next person reads the ladder. Left alone it will step 2.0.1 → 2.2.0 with a released v2.1.0 tag sitting in a gap the table does not mention. Add the line wherever the table lives:

  1. Makefile.starkernel (the copy above, beside the LITHOS_VERSION edit itself).
  2. docs/lithosananke/ROADMAP.md, "Release Versioning Policy" — the comment names it as the upstream source.
  3. FABRIC-2.md §G is the origin of the scheme but is CLOSED/ARCHIVAL — cite it, do not edit it. Amending a closed document to match a later decision is exactly what this series' discipline forbids.

XXVIII.3 — A trap the documentation sweep must not walk into, with precedent

Makefile.starkernel is not a .md file, and this project has already been bitten by exactly that. FABRIC-3.md §I.2 records it plainly: a FABRIC-series rename looked complete but had silently skipped Makefile.starkernel, Kconfig.kernel, scripts/bleach_zuse_img.sh, four proof/*.thy files and src/starkernel/arch/amd64/isr.S —

the original sweep's file-list only matched --include=*.md/*.c/*.h/*.4th, which silently skipped every file without one of those four extensions. Found by re-grepping with the extensions excluded instead of included.

So item 13's sweep (§XXVI.1) must grep by exclusion, not inclusion, and must expect version/policy/citation text in Makefiles, Kconfig.*, scripts/*.sh, proof/*.thy, .S sources and .gitea/workflows/. The technique §I.2 used to fix it — ordered placeholder substitution, one pass per file — is the recorded-working method and should be reused rather than reinvented.

And the same exclusion that §I.2 made deliberately still applies: .claude/settings.local.json's historical command log and ClaudeEXPORT/ are frozen audit records — rewriting them would falsify an audit trail, not fix a stale citation. Leave them.

XXVIII.4 — Still unruled: VERSION, the embedded engine string

LITHOS_VERSION and VERSION are independently tracked (.claude/CLAUDE.md, and confirmed at Makefile.starkernel:83-84). 2.1.0 settles the kernel version; the engine version VERSION ?= 3.1.0 is not settled by it.

It plausibly moves — this reshuffle changes BIRTH's registration, adds a Hera-only word, relocates the framebuffer primitives' registration and replaces the messaging layer, all of which are engine-side. But it is a separate string with its own meaning, and .claude/CLAUDE.md warns it "does not auto-sync with the standalone StarForth repo's own version." Captain Bob's call, at tag time.

Both are hand-edited. The bump-z/bump-y targets were removed 2026-08-15 as non-functional (they referenced fields that never existed in the generated include/version.h) and must not be resurrected from memory.

XXVIII.5 — Punch list

  • ✅ Version — RULED: LITHOS_VERSION = 2.1.0, tag v2.1.0 (§XXVIII).
  • ⬜ Roadmap-table line added in two live places, FABRIC-2.md §G cited not edited (§XXVIII.2). Folds into item 13.
  • ⬜ Item 13's sweep greps by exclusion, per §XXVIII.3's precedent.
  • ✅ VERSION (engine) — CLOSED as a rule (§XXX.6), applied at tag time against the real diff rather than guessed now.

Every design question is ruled and the release number is set. What remains is execution, one engine-version call, and the close.


XXIX. Versioning policy closed; the LTS designation deferred with written criteria (2026-09-19)

Captain Bob proposed a revised scheme — odd major = LTS, even major = working line; minor = release; patch = working builds; CI build number appended once CI is exposed — and asked directly whether this release deserves an LTS number as "the very first fully production ready version," inviting disagreement and noting the policy should be closed either way.

Policy: closed here. LTS designation: deferred, with the criteria written down rather than left to judgement. The reasoning below is the disagreement, stated plainly because it was asked for.

XXIX.1 — Why not LTS at this tag: five findings, not an opinion

  1. §XIV is open, and was reproduced live after being filed. "Concurrent WIREBIND attach detection gap — found live, NOT fixed, flagged for later" (2026-09-11). §XXXI then caught it again on the post-Stage-4 rerun: near-simultaneous thumbdrive attach meant 6 of 9 identities silently never triggered WIREBIND: <id> attached at all, with QMP query-block confirming every device was genuinely present. That is the identity-attach path — how humans get into the system — failing silently and reproducibly. It is not a corner case, and "silently" is the part that disqualifies it.
  2. It has never booted on real hardware. That is FABRIC-3.md's entire declared topic, and the roadmap reserves v2.2.0 (amd64 bare-metal) and v2.5.0 (hardware release) for it. "Production ready" invites the question production on what?
  3. ACL Phase 8 is open — PKI / thumbdrive, Ed25519 challenge-response, user minting by Zuse: "⬜ This is the open item — pick up here next" (.claude/CLAUDE.md). That is the last mile of the security story, and the security story is much of the production claim.
  4. At tag time this reshuffle is brand-new code. Kernel-Hermes is greenfield C (§XIV.1), Hestia is a new VM, BIRTH is newly general, the latch is new. LTS normally means soaked — years of support promised on something proven. Declaring it at the moment of a structural rewrite inverts the word. The reshuffle is an excellent reason to tag; it is a poor reason to call the result long-term-stable.
  5. §XXV.4: formal coverage likely decreases at this tag, as messaging moves from modelled FORTH to file-scope C statics. For the audience an LTS label is aimed at — licensees, SSRN, patent support — the proof story is load-bearing, and this is the wrong moment to claim maximum stability while the verified boundary contracts.

And the precedent is this project's own. FABRIC-3.md §I.2 rolled LITHOS_VERSION back from 2.0.1 because the label "got ahead of the real state." "First fully production ready" is a far larger claim than 2.0.1 was. The same instinct that caught that one should catch this.

XXIX.2 — The constructive form: make LTS a test, not a judgement

This document has twice turned a judgement into a structural fact and got a better answer both times (§XV.2's death, §XXVII's empty floor). Same move here. Proposed LTS criteria, to be ratified as part of the policy:

A tag may take an odd (LTS) major only when all hold:

  • Real-hardware boot demonstrated on at least the amd64 reference board, with logs committed as audit artifacts (the existing acceptance standard, extended off QEMU).
  • No known-reproducible silent-failure defect open — §XIV's class specifically: a failure that produces no error and is caught only by inspection.
  • ACL Phase 8 closed, or its absence explicitly scoped out of the LTS claim in writing.
  • The proof/ boundary restated and not contracted relative to the prior LTS (§XXV.3).
  • A soak: the architecture unchanged for at least one full release cycle before the LTS tag — LTS follows stability, never announces it.

Then the first tag meeting them takes the odd major, and nobody has to decide whether it feels ready.

So: 2.1.0 stands as ruled (§XXVIII) — even major, working line, correct under both the old and new schemes. LTS is not refused, only sequenced.

XXIX.3 — Three conflicts in the new scheme, to resolve while closing it

  1. The minor position is already taken, and by a live plan. The current policy uses minor for QEMU-vs-hardware (X.0.0 = QEMU release, X.5.0 = hardware bare-metal release), and v2.5.0 is already on the roadmap. The new scheme wants minor = "release version." Both readings cannot hold. Adopting the new scheme retires X.5.0 semantics and must say so explicitly, or v2.5.0 becomes ambiguous the day it is cut.
  2. Odd/even major collides with "major = breaking change." Binding the major to LTS-ness means a breaking change inside the working line cannot be expressed without also flipping LTS status, and an LTS line cannot take a breaking fix without ceasing to be LTS. Node and the pre-2.6 Linux kernel both used this shape and both paid for it. Worth choosing deliberately: either accept that majors no longer signal breakage, or move LTS-ness out of the version number (a tag suffix, a channel name) and leave the major for compatibility.
  3. 3.x is already in use, by the other string. VERSION ?= 3.1.0 is the embedded engine. If LITHOS_VERSION takes an odd major it lands on 3.0.0 — two different 3.x values in the same generated include/version.h. Not fatal, since they are separate fields, but confusing at precisely the moment a release wants to be legible. Flag before, not after.

XXIX.4 — CI build numbers: use SemVer build metadata, not a fourth field

Specific because it interacts with something already shipped. sbom.spdx / sbom.spdx.json are generated by syft and consumed as SPDX (§XXVI.2), and SPDX/SemVer tooling parses X.Y.Z+build.N correctly while X.Y.Z.N is not valid SemVer — it sorts wrongly or is rejected outright.

Recommendation: 2.1.0+ci.1234. Build metadata is ignored in precedence comparison, which is exactly right for a build number: two builds of the same source rank equal, and the SBOM stays machine-readable for the licensee and patent audiences it exists for.

XXIX.5 — What this section rules and what it leaves open

  • ✅ Policy closed as §XXIX.2–XXIX.4, subject to Captain Bob ratifying the three conflicts in §XXIX.3.
  • ✅ 2.1.0 stands (§XXVIII) — even major, working line.
  • 🔶 LTS deferred, not refused. Criteria written (§XXIX.2); the first tag meeting them takes the odd major.
  • ⬜ Captain Bob to rule on §XXIX.3's three conflicts, and on the engine VERSION string (§XXVIII.4).

The disagreement is narrow and worth stating once more so it is not mistaken for reluctance: the work is real and the tag is deserved. What is not yet earned is the word long-term stable, on code that will be new the day it ships.


XXX. Versioning policy — CLOSED (Captain Bob ratified, 2026-09-19)

§XXIX's recommendations accepted. This section is the closed policy, written as the text to carry into the two live places §XXVIII.2 names. Two items §XXIX left as open alternatives are resolved here rather than left hanging.

XXX.1 — The scheme

Field Meaning
Major Odd = LTS line. Even = working line.
Minor Release within that line
Patch Working builds within a release
+ci.N CI build number, SemVer build metadata (§XXIX.4) — ignored in precedence

XXX.2 — Conflict 1 RESOLVED: X.5.0 semantics are retired

The old encoding (X.0.0 = QEMU release, X.5.0 = hardware bare-metal release) used the minor field to carry a milestone meaning. That is retired — minor now means only "release within the line." Stated explicitly, per §XXIX.3.1, so v2.5.0 cannot be read two ways the day it is cut.

The milestones themselves survive; only their encoding changes. They become roadmap entries rather than arithmetic, which re-expresses the ladder cleanly:

  v1.0.x  — serial-only production (released, historical encoding)
  v1.5.x  — framebuffer VT100 console milestone (released, historical encoding)
  v2.0.0  — QEMU release: three-arch QEMU story complete (released)
  v2.0.1  — SER5 hardware-track line (historical; superseded by this policy)
  v2.1.0  — Tripod/kernel reshuffle: Hermes into the kernel, Tripod = Hera/Artemis/Hestia
  v2.2.0  — amd64 bare-metal bring-up (Beelink SER5)
  v2.x    — further working releases; riscv64 / aarch64 board bring-up
  v3.0.0  — FIRST LTS, when and only when §XXX.5's criteria are met

A pleasing consequence, not a contrivance: the hardware bare-metal release that was going to be v2.5.0 is exactly the point where §XXX.5's first criterion (real-hardware boot) is satisfied. The milestone that used to be encoded in the minor field becomes the natural candidate for the first LTS major. The scheme and the roadmap agree without being forced to.

XXX.3 — Conflict 2 RESOLVED: the line signals breakage, not the major

§XXIX.3.2 offered two options without recommending. Resolved in favour of keeping the major for LTS parity, since that is the scheme's whole point, and relocating the compatibility signal rather than abandoning it:

Breaking changes land on the even (working) line. An odd (LTS) line takes non-breaking fixes only — backports, never redesigns.

So the major no longer means "something broke"; the line you are on tells you whether breakage is possible at all. That is how odd/even schemes are meant to work, and it repairs §XXIX.3.2's objection rather than absorbing it: a breaking change is expressible (bump the minor on the working line), and an LTS line cannot silently acquire one (it is not allowed to).

Corollary worth writing down: an LTS line that needs a breaking fix has failed as an LTS. The fix goes to the working line and the next LTS picks it up. No exceptions, because the exception is the whole failure mode this rule exists to prevent.

XXX.4 — Conflict 3 RESOLVED: the two version strings are independent and may coincide

LITHOS_VERSION reaching 3.0.0 while VERSION (the engine) already sits at 3.1.0 is not a collision to engineer around. .claude/CLAUDE.md already states they do not auto-sync; they are separate fields in the generated include/version.h describing different artifacts — the kernel and the embedded engine.

Ruling: state the independence in the policy text itself, so a reader meeting two 3.x values does not infer a relationship that has never existed. Documentation, not renumbering.

XXX.5 — LTS criteria (ratified)

A tag may take an odd (LTS) major only when all five hold:

  1. Real-hardware boot demonstrated on at least the amd64 reference board, logs committed as audit artifacts — the existing acceptance standard, extended off QEMU.
  2. No known-reproducible silent-failure defect open — §XIV's class specifically: a failure producing no error, caught only by inspection.
  3. ACL Phase 8 closed, or its absence explicitly scoped out of the LTS claim in writing.
  4. The proof/ boundary restated and not contracted relative to the prior LTS (§XXV.3).
  5. A soak — the architecture unchanged for at least one full release cycle. LTS follows stability; it never announces it.

The first tag meeting all five takes the odd major. No judgement call, no debate about whether it feels ready.

XXX.6 — The engine VERSION string: a rule, not a number

§XXVIII.4 left it to tag time. Closed here with a rule instead, because the correct bump is a function of the realised diff and the code is not written yet — picking a number now would be guessing, and guessing is what §I.2's rollback was about.

Engine VERSION bumps at tag time by what actually changed: major — any FORTH-79-visible word semantics change; minor — words added, removed, or their registration relocated; patch — build-only, no dictionary-visible change.

On the current plan this reshuffle is a minor bump (3.1.0 → 3.2.0): BIRTH's registration widens, a Hera-only suicide word is added, PLOT/FB-* registration relocates to Hestia, the messaging layer is replaced — all dictionary-visible, none changing a FORTH-79 word's semantics. Confirm against the real diff at tag time; the rule decides, not this paragraph.

Both strings stay hand-edited in Makefile.starkernel. The bump-z/bump-y targets were removed 2026-08-15 as non-functional and must not be resurrected.

XXX.7 — Punch list

  • ✅ Versioning policy — CLOSED (§XXX). All three §XXIX.3 conflicts resolved; LTS criteria ratified; engine VERSION closed as a rule.
  • ✅ This release is 2.1.0 (§XXVIII), even major, working line.
  • ⬜ Carry §XXX.1–XXX.6 into Makefile.starkernel and docs/lithosananke/ROADMAP.md (§XXVIII.2), replacing the old roadmap table. FABRIC-2.md §G is cited, never edited.
  • ⬜ Engine VERSION applied at tag time per §XXX.6's rule.

Every decision this document required is now made. What remains is execution: the surgical strip, the build, the Isabelle pass, the sweep, the SBOM, the merge, the tag, and the close.


XXXI. Gap-analysis sweep of the FABRIC set (2026-09-19)

Run by direct instruction — "we need a gap analysis sweep and I'd look HARD at the entire fabric document set." The set is 25,493 lines across FABRIC-0 … FABRIC-4 plus this document.

A gap analysis of a series whose central rule is "never silently drop a stale claim" is, in effect, an audit of whether the series kept its own rule. Mostly it did. Where it did not, the failures are structural and in one case operationally urgent.

XXXI.1 — Method, and an honest statement of coverage

Done: mechanical harvesting across all six documents — open/closed checkbox counts, deferral vocabulary (deferred, not in scope, for later, still open, NOT fixed), cross-document reference counts for every long-lived item id — followed by targeted reads of everything the counts flagged.

Not done: an end-to-end read of all 25,493 lines. FABRIC-0.md §1–24 (the theory) and the bulk of FABRIC-1.md's resolved narrative were not read line by line. So this sweep finds structural gaps and orphaned items; it does not certify that every claim in the set is accurate. Stated plainly because §XXII.2 was written about exactly this failure mode, and it would be poor form for the sweep to repeat it.

XXXI.2 — The carry-forward chain held twice and then terminated

The series' defining discipline is that closing a document carries its open items forward. Audited mechanically:

Link Status Evidence
FABRIC-0 → FABRIC-1 Held F0's open item ids are referenced throughout F1 (1.11 ×8, 4.4s ×5, 4.6 ×9, 5.1 ×3)
FABRIC-1 → FABRIC-2 Held, and independently verifiable F2's header claims a "full, non-sampled carry-forward of every open item… 51 items." grep -c '^\s*-\s*\[ \]' FABRIC-1.md returns 51 today. The claim is exactly true
FABRIC-2 → FABRIC-3 Terminated Every long-lived F0 id — 1.11, 4.4s, 4.6, 5.1, 5.2, 5.3 — returns zero references in FABRIC-3.md

FABRIC-3.md contains no checkboxes at all (0 open, 0 closed), having moved to a prose and CLOSED-heading convention. That is a legitimate stylistic choice, but it ends the mechanical audit trail that made the F1 → F2 link provable. After FABRIC-2, "is anything still open?" stops being a grep and becomes a reading exercise — which is how the items in §XXXI.3 and §XXXI.4 went quiet without anyone deciding to drop them.

XXXI.3 — FABRIC-2's close claim is overstated: six design items are open, not hardware-blocked

FABRIC-2.md's close header states it was "closed after §I… was worked through to completion — every item either closed with a dated note or correctly left open pending a physical machine (§I.6)."

It closed with 14 open checkboxes. Eight are the real-hardware USB boot checklist (build the ISO, dd it, boot the machine, capture with no serial log, confirm POST 1012/0/0, reach ok>, document the result) — legitimately FABRIC-3's topic, correctly left open. The claim holds for those.

The other six have no hardware dependency and appear nowhere since:

Item Line Picked up in F3 / F3.5?
§17.4 — framebuffer heat/decay dynamics, undesigned 77, 4366 No (§XXXI.6)
First-touch allocation (identity pubkey → claimed block range) 192 No
Migration state machine — states, transition triggers 228 No
Block-namespace sandboxing / mkcapsule conflict extension 1011 No
ACL EXPIRE — "confirmed genuinely unscoped (2026-08-26)" 1219 No

The one apparent hit for EXPIRE in this document is a false positive — §XXIV.3's name-collision check for the suicide word, unrelated to the ACL item.

None of these is necessarily urgent. All of them are unowned, and the close header says otherwise.

XXXI.4 — FABRIC-0's seven open items, three of which this document re-derived

FABRIC-0.md closed 2026-08-12 with seven open items. All seven remain - [ ] today:

Item Bearing on this reshuffle
4.3 — Console Direct. See below.
1.11 — Dirty-event granularity Blocked on 4.3; unblocks with it
4.4s — (user) prompt segment Plausibly done — FABRIC-3.md §XXV's [user@VMName] work looks like it, never checked off
4.6 — "Artemis last. It works today; it is the thing that cannot be broken" Sequencing advice this reshuffle happens to follow (§IV.1 leaves Artemis unchanged) without citing it
5.1 — Re-run the DoE on the new substrate Overtaken by §XXV.1's DoE rewrite
5.2 — Isabelle/HOL Direct. See below.
5.3 — Shrink the subsystem documents Direct. See below.

Three were re-derived by this document without either side knowing:

  1. 4.3's deferred tail is this reshuffle's Hestia work. Its 2026-08-07 discussion note scopes 4.3.1–4.3.4 and defers, verbatim, "fonts, scrolling, cursor/VT100 semantics, the Hermes message protocol, and Console as a fleet VM under Hera's birth protocol — later 4.3.x items, scoped once this slice is reviewed." That last clause is §XVII and §XVIII. Also worth noting: 4.3 carries a standing instruction — "Captain Bob wants a discussion before any work starts on this item… raise it and wait." Satisfied in substance (Captain Bob directed this work personally), but the linkage was never recorded. The slices themselves are complete: 4.3.1 through 4.3.7f — framebuffer, keyboard on all three architectures, fonts, UTF-8, TTF rasterization — are all [x]. Only the parent stays open, for the tail this document just designed.
  2. 5.2 already specifies the Isabelle target §XXV left generic: "One datatype, one index space, one conservation theorem." §XXV.5 mapped affected theories without citing it — and "one conservation theorem" is almost certainly K, the fleet conservation invariant that §XV.4 independently rediscovered as the only thing that must survive a death. Two documents converged on the same theorem from opposite ends, a year apart in the series, with no cross-reference.
  3. 5.3 already asked for §XXVI.1's documentation work: "ARTEMIS.md, HERMES.md, CONSOLE.md, TRIPOD.md should each reduce to roughly three lines. Any that grows is fighting the design." It also flags that TRIPOD.md's Immediate Goal "currently requires Hera to spawn Hermes and Artemis at boot, which 0.1 undoes" — worth checking, because kernel_main.c:865 does exactly that today (§XIII.3), so either 0.1 was reverted or that note is stale.

XXXI.5 — The urgent finding: FABRIC-3 holds an open item that contradicts §XV.3

This is the one that affects the build, and it should be reconciled before kernel-Hermes is written.

FABRIC-3.md §XXVIII.2 records, and §XXXI restates as still open:

Still open: the MSG-TICK/Stage-3-switch dual-ownership rough edge §XXVIII itself flagged (both mechanisms can independently move control between the same VMs) is unchanged by this pass.

…and in §XXXI's own summary: "the MSG-TICK/Stage-3 dual-ownership rough edge (still open, not observed failing)."

§XV.3 of this document ruled that nothing owns turn order. That ruling is a principle; FABRIC-3 records a live, known, unfixed instance of exactly the dual ownership the principle forbids — two mechanisms independently moving control between the same VMs.

Three consequences:

  • §XIV.5's scheduler-firewall concern was a rediscovery, not a new finding. FABRIC-3 named this seam first. This document should have cited it and did not.
  • Kernel-Hermes lands directly on it. Per §XIV.5, once the FORTH layer is legacy, kernel-Hermes becomes the caller of sk_vm_switch_signal_mark_work() — i.e. the reshuffle moves one of the two owners into the new arbiter. Building on an open dual-ownership defect is how it stops being "not observed failing."
  • "Not observed failing" is not "does not fail." §XXVIII.2's own history is a record of this neighbourhood producing switch storms — QEMU pinned near 100%, serial log frozen solid — when two sources of truth disagreed.

Added as punch item 22, ahead of the kernel-Hermes work rather than alongside it.

XXXI.6 — §17.4 became Hestia's problem and nothing says so

FABRIC-2.md §17.4, still open: "framebuffer utility's internal heat/decay dynamics. Undesigned, blocked on 1.11 specifically… Still open 2026-09-04: 1.11's decision closed, but no dirty-region-tracking mechanism exists yet for this item's physics to attach to — that mechanism, not a further ruling, is the real remaining blocker."

§XVIII.3 relocated the drawing fabric's vocabulary to Hestia and §XVII.2 gave the fabric an owner — without inheriting this open item. The framebuffer now has a VM that owns it and an undesigned heat/decay story attached to nothing, and the two facts live in different documents.

Not a blocker for the reshuffle — §XVIII deliberately scoped Hestia to ownership and binding, not physics. But it is now unambiguously Hestia's, and should be recorded as such rather than left in a closed document. Punch item 23.

XXXI.7 — Three "reported, not scheduled" registries, with inconsistent hygiene

The set carries at least three separate registries of found-but-unfixed defects — FABRIC-0 §25.7 "Reported, not scheduled", FABRIC-1 §C "Reported bugs and dead code, not yet fixed", and scattered prose flags through FABRIC-2/FABRIC-3. They have no consolidated home and no shared convention.

§25.7 is actively misleading as a registry, because resolved entries are sometimes struck through and sometimes not:

  • ~~stadium_admit() never writes stadium_owner[idx]~~ — struck, correctly.
  • "Fleet heat leak… vm_physics_touch() integer-truncation drift" — not struck, but resolved: FABRIC-1.md:468 carries it as - [x] **Fleet heat leak.** Anyone reading §25.7 today would believe the project sits on an unmeasured monotonic leak in a conservation law it makes claims about.
  • bump-z/bump-y broken — not struck, but resolved: removed 2026-08-15 as non-functional (§XXVIII.4).
  • Still genuinely live in §25.7: heartbeat_trust() exported with zero callers; m5_time_trust/m5_variance declared and never used; hotwords_cache_promote()'s NULL write; Kconfig/menuconfig never exercised end to end; and src/*.c.bak tracked in git — which is now reported in three places (§25.7, .claude/CLAUDE.md, and §XXII.5 of this document) and actioned in none.

Punch item 24: consolidate the registries and apply one convention. Reported, not fixed — and deliberately not folded into this reshuffle.

XXXI.8 — What the sweep changes

Nothing in §I–§XXX is invalidated. No ruling is contradicted by a finding here; §XXXI.5 is a conflict between this document's principle and an open defect in the code, not an error in the ruling. Every design decision stands.

New punch items, appended to §XXX.7's list:

  • ⬜ 22 — Reconcile §XV.3 against the MSG-TICK/Stage-3 dual-ownership defect (§XXXI.5). Before kernel-Hermes, not alongside it. The reshuffle moves one of the two owners into the new arbiter.
  • ⬜ 23 — Re-home FABRIC-2 §17.4 (framebuffer heat/decay) as Hestia's (§XXXI.6).
  • ⬜ 24 — Consolidate the three "reported, not scheduled" registries and mark the resolved entries (§XXXI.7).
  • ⬜ 25 — Decide the fate of FABRIC-2's five other orphaned design items (§XXXI.3): first-touch allocation, migration state machine, block-namespace sandboxing, ACL EXPIRE. Adopt, re-home, or explicitly abandon with a reason — but not leave unowned.
  • ⬜ 26 — Reconcile FABRIC-0's seven opens (§XXXI.4): close 4.3 against §XVII/§XVIII, check whether 4.4s is already done, fold 5.2 into item 12's Isabelle pass and 5.3 into item 13's sweep, and confirm the TRIPOD.md/0.1 contradiction.

And one recommendation the sweep argues for on its own evidence: the carry-forward discipline was verifiable only while the documents used checkboxes (§XXXI.2). If the series continues in prose, "what is still open" needs a deliberate mechanism — because two documents have now closed with items nobody chose to drop.


XXXII. Item 22 SETTLED: retire the pump — the switcher becomes the sole mover of control

§XXXI.5 found that FABRIC-3.md carries an open defect — "both mechanisms can independently move control between the same VMs" — which is a live instance of what §XV.3 forbids as a principle. Settled here, and the reshuffle turns out to be the thing that closes it rather than the thing that inherits it.

XXXII.1 — The two mechanisms, and why only one should survive

How it moves control Fate
The MSG-TICK pump (repl.c idle loop) VM-EXECs "MSG-TICK" into a VM — a nested vm_interpret() on the shared kernel C stack Retired
The Stage-3/4 switcher (sk_vm_context_switch()) Saves and restores a per-VM native stack, timer-driven Sole survivor

RULING: kernel-Hermes publishes eligibility; it does not dispatch. It marks a target as having work and the switcher — the one mechanism that moves control — delivers when it next resumes that VM on the VM's own stack. Two control-movement paths collapse into one.

XXXII.2 — Three reasons this is the right direction, not merely a tidy one

  1. The pump is already legacy by §XIV.1. It exists for exactly one purpose: to drain per-VM FORTH message queues by executing MSG-TICK inside each VM's dictionary. Once kernel-Hermes holds the queues, there is nothing for it to drain. Retiring it is not a new decision — it is a consequence of a ruling already made.
  2. The pump is a known scaling wall, already band-aided once. FABRIC-3.md §XXII records it walking the entire live-VM registry every idle beat, VM-EXECing into every VM — "O(N) per tick, forever, with no cap" — caught live as a real bottleneck at just 9 VMs on riscv64, where a campaign that finished in 200–340 s on amd64/aarch64 never finished at all. Fixed by round-robin batching (SK_MSG_PUMP_BATCH), which bounds cost to O(K) at the price of per-VM latency depending on registry position. Its own comment names the endgame: "the real fleet is headed toward hundreds of VMs, at which point this isn't a riscv64 quirk — unconditional O(N) per second is a wall on every architecture." Retiring the pump removes the wall and the latency tradeoff together.
  3. "Not observed failing" is worth very little in this specific subsystem, and FABRIC-3.md says so itself. §XXVIII.1 is a correction to §XXVIII's own "Stage 3 CLOSED" claim: "Zero fault indicators" was true only of what was visible, because "Stage 3 emits nothing per switch by default, so a switch storm and a healthy idle REPL produce an identical serial log. The corruption below was already live in every one of those 'clean' Stage 3 boots; nothing in that verification pass could have shown it." A defect flagged "not observed failing" in a subsystem that is invisible by default is not evidence of health. That is the project's own finding, not an inference.

XXXII.3 — Reconciling with §XV.3: one mover, still no owner

§XV.3 ruled that nothing owns turn order. Naming the switcher the sole mover looks like it names an owner. It does not, and the distinction is the one §XV.3 already relies on:

Moving control is a mechanism. Deciding whose turn it is is a policy. The switcher is the mechanism. The compudynamics remains the policy, and it is still owned by nobody.

So §XV.3 stands unamended. The dual-ownership defect was never about deciding — both the pump and the switcher move control without either claiming to choose who runs next. It is about two physical paths for the same transfer, which is a correctness problem (stack ownership), not an authority problem. Collapsing them resolves it without touching the principle.

And it makes §XXIII.4 implementable. That section ruled the sinking latch is checked "on giving a VM its turn" — which silently assumes there is one place a VM gets its turn. With two movers that check has two homes, or one, and misses. §XXIII.4 already depended on this ruling without saying so.

XXXII.4 — The dependency this rests on, verified

The ruling strands identity VMs unless the switcher covers them. Checked: FABRIC-3.md §XXX — "Stage 4 — WIREBIND-scope extension, built and verified live (2026-09-14/15)," Captain Bob's call to proceed immediately, built in the same one-commit-per-increment discipline as Stages 0–3.

So the switcher already covers WIREBIND identity VMs, not just the Tripod fleet. Had Stage 4 still been the "ratified-decision-only, not implemented" step §XXVIII left it as, this ruling would have been unbuildable and identity VMs would have lost their delivery path entirely. It holds only because Stage 4 landed.

XXXII.5 — Open for the build, not decided here

  1. Where the payload waits. Delivery must still cause the payload to execute inside the target — messaging.4th block 5041: "Payload is FORTH text, executed by the receiver, same as every other message here." Under this ruling the switcher resumes the target on its own stack, so kernel-Hermes must leave the payload somewhere the target consumes on resume. The shape of that hand-off is a build decision.
  2. Hera's special case. Per §XX of FABRIC-3.md, Hera is skipped by the pump and instead gets a direct MSG-TICK word dispatch in her own context, because "she is the pump" and self-targeting VM-EXEC would hit a known reentrancy class. With the pump gone, that special case should go with it — her queue is kernel-held like everyone's. Confirm rather than assume.
  3. The reentrancy guards. g_mama_interpreting and g_idle_pump_active (repl.c) exist specifically because the pump makes nested vm_interpret() calls. They likely retire with it — but carefully, since §XXVIII.2's switch storm lived in this neighbourhood. Remove only what is provably unreachable, per §XXII's own strip discipline.

XXXII.6 — Punch list

  • ✅ Item 22 — SETTLED (§XXXII): retire the pump; the switcher is the sole mover of control; kernel-Hermes publishes eligibility and does not dispatch. Folds into the kernel-Hermes work (item 6) rather than preceding it — §XXXI.5 filed it ahead on the assumption it was a separate defect to fix first; it is instead a property of how kernel-Hermes is built.
  • ⬜ Items 23–26 unchanged (§XXXI.8): re-home §17.4; consolidate the reported-not-scheduled registries; decide FABRIC-2's five orphaned design items; reconcile FABRIC-0's seven opens.

Worth recording as the pattern this makes four of: §XXXI.5 read as a risk the reshuffle would walk into. Examined, it is a defect the reshuffle removes — because the mechanism that creates it is the same FORTH messaging layer already ruled legacy. The gap analysis found a problem that the design had already solved without noticing.


XXXIII. Sizing kernel-Hermes: what actually has to be reimplemented (2026-09-19)

Why this section exists. §XIV.1 ruled the FORTH messaging layer legacy and kernel-Hermes greenfield. That was the right call, but it quietly exited this document's own scope discipline — "a reorganization, not an invention… relocated and rewired rather than rewritten" — and nobody counted what the rewrite actually amounts to. With the build trigger approaching and confidence the stated concern, an unsized centerpiece is the wrong thing to carry into it.

capsules/common/messaging.4th is 504 lines defining 85 words. That number is the reason to look, and it turns out to be badly misleading in the reassuring direction.

XXXIII.1 — The 85 words, by category

Category Words Fate in kernel C
Field accessors (MSG-TYPE@/!, CH-ID@/!, MBR-VM@/!, …) ~34 Vanish — they become struct members
Channel abstraction (CH-*) 28 Mostly unexercised — see §XXXIII.2
Message core (alloc/free, send, deliver, tick, reap, ack/nack, redeliver) ~20 Must survive. The real work.
Member list (MBR-*) 7 Collapses to a small membership list
Events (EVENT-EMIT/WAIT/DRAIN) 3 Dead — see §XXXIII.3
Elevation (ELEVATE-*, SEND-ELEVATE-REQUEST) 5 Live. Must survive.
Status/diagnostics (MSG-USED, CH-USED, MSG-STATUS) 3 Cheap, keep

XXXIII.2 — Finding: the channel abstraction serves exactly one static channel

28 of the 85 words are channel machinery. Traced across capsules/, src/ and experiments/ for callers outside messaging.4th itself:

Word Live callers outside messaging.4th
CH-REQUEST none — only capsules/MANIFEST.md (documentation)
CH-ACCEPT none — documentation only
CH-CONFIRM none — documentation only
CH-CLOSE none — documentation only
CH-MINT-ID none — documentation only
CH-ADD-MBR 3 — the only live channel operation in the system

The entire CH-NEGOTIATING → CH-OPEN → CH-CLOSING handshake has no caller anywhere. The whole live channel lifecycle is: Hermes creates COMMON-CH once at birth, adds itself and Hera (hermes/init.4th:20-21), Artemis adds itself (artemis/init.4th:501). One channel, three members, created at boot and never negotiated, never closed, never reaped.

So CH-ARENA, CH-MAX 16, CH-ALLOC, CH-FREE-NODE, CH-FIND-FREE-SLOT, CH-COOL-ALL, CH-TOTAL-HEAT, CH-REAP-SAFE, per-channel Stadium heat accounting and the three-state machine all exist to support a generality nothing has ever used. Console proxies explicitly opt out of channels entirely (capsule_console.c:23).

Kernel-Hermes does not need a channel subsystem. It needs one broadcast membership list.

XXXIII.3 — Finding: the event words die with process.4th

EVENT-EMIT, EVENT-WAIT, EVENT-DRAIN have exactly one live caller — capsules/process.4th, which §XIII.2 established is EXEC'd nowhere. The only other reference is MANIFEST.md.

They go out with it, along with SPAWN-EVENT (§I.4: grepped tree-wide, zero consumers, "an unwired placeholder"). Do not port the event layer.

XXXIII.4 — What must actually survive

Stripped of the above, kernel-Hermes's real surface is:

  1. Message allocation and release, coupled to Stadium heat. MSG-ALLOC pulls Q.SLOT from the caller's own reservoir and rolls back on refusal; MSG-FREE-NODE returns it via STADIUM-EVICT. This is the hard part, and the one that cannot be approximated — §XV.4 established that K conservation is the only thing that must survive a death, and fleet_conserved verifies it continuously.
  2. Send / deliver / deliver-all / tick / reap / ack / nack / redeliver-nacked — the core protocol, ~9 words of real logic.
  3. One broadcast membership list (replacing 28 words of channel machinery).
  4. The elevation path — SEND-ELEVATE-REQUEST / ELEVATE-GRANT, live in zuse-eligibility.4th (loaded at boot by init.4th) and mama_forth_words.c. Must survive; it is the word-ACL elevation ask carried to Zuse.
  5. Status words, cheap and worth keeping for the same reason §XXVIII.1 gives: a subsystem that emits nothing by default cannot be debugged.

Roughly half the file is accessors that become struct fields, and another third is generality with no caller. The genuinely new C is the heat-coupled allocator plus about nine protocol words. That is a much smaller thing than "reimplement 504 lines of messaging," and it is the first honest estimate this document has had.

XXXIII.5 — The caveat, and the part that is not mine to decide

"Unexercised today" is not "unwanted." The channel abstraction may have been built for a future that has not arrived — per-VM private channels, capability-scoped groups. Deleting it is a decision, not an observation, and it belongs to Captain Bob.

Two honest qualifications on the evidence:

  • §XXII.2's lesson applies in weakened form. Capsules are birthed by name from runtime strings, which defeats grep; FORTH words are not — a word is invoked by its literal text, so this search is far more reliable than the capsule-reachability search that nearly deleted the init-l8-* family. But a human can still type CH-REQUEST at a REPL, and no grep sees that.
  • The negotiation machinery is described in MANIFEST.md as though live. That is the same class of stale-documentation problem §XIII.2 found for block 4055 — another entry for item 13's sweep.

Recommendation: build kernel-Hermes for one broadcast membership and no negotiation, and keep the FORTH channel words in the Category B strip until something asks for them. If per-VM channels are wanted later, they are a clean addition to a working arbiter rather than a speculative port into a new one.

XXXIII.6 — Effect on confidence, stated plainly

This was opened expecting to find the scope larger than advertised. It is smaller, and for a reason that generalizes: the FORTH layer accreted a general mechanism where the system only ever used a specific one. The reshuffle's value is partly that it forces that accounting.

What remains genuinely hard is narrow and now named: the heat-coupled allocator, because K is continuously verified and any drift is visible. That is the piece to build first and prove first — before send/deliver, before delivery hand-off, before the latch. If K holds across alloc/free/evict under the new arbiter, the rest is protocol plumbing.

Punch list: ⬜ Item 27 — rule on §XXXIII.5: does kernel-Hermes implement channel negotiation, or one broadcast membership? ⬜ Item 28 — build and prove the heat-coupled allocator first, verified against fleet_conserved, before any protocol work.


XXXIV. The transition state: how two messaging layers coexist without corrupting K (2026-09-19)

The question nobody asked. §XXII.1 draws the ordering distinction that protects the strip — Category B (hermes/init.4th, most of messaging.4th, the routing table, the slot-3 pairing) is "load-bearing until the replacement boots." That sentence implies a window in which both messaging layers are live, and nothing in this document says how they behave in it.

XXXIV.1 — The real hazard is not control. It is heat.

Item 22 (§XXXII) dealt with two mechanisms moving control. The transition creates the same shape one layer down: two allocators drawing on the same per-VM Stadium reservoir, both counted in K.

The budget is explicit. messaging.4th derives it:

Q.SLOT = (Q.1 - Q.1/3) / (MSG-MAX + CH-MAX - 1)

— each VM's reservoir (Q.1 = Q48_ONE, per stadium.c's "item 4.1's reservoir starts each VM's quota at 1.0"), less COMMON-CH's Q.1/3 floor, split across 32 messages + 15 channel slots = 47 items. That pool is sized for one messaging layer.

Conservation itself is not at risk — K holds as long as each layer honestly pulls and returns, regardless of who holds the heat. What is at risk is the budget: two layers each sized to consume the same reservoir is oversubscription, and the failure mode is allocation refusals under load, not silent drift. fleet_conserved makes it visible rather than quiet, which is the saving grace.

A helpful interaction with §XXXIII: retiring the channel abstraction removes CH-MAX - 1 = 15 of those 47 slots, so the divisor falls from 47 to 32 and per-message heat rises by ~47%. Deleting the unused generality recovers roughly a third of the per-VM message budget — useful headroom precisely during the window when two layers share it.

XXXIV.2 — RULING: partition by message type. One owner per message, never shared.

During the transition, every message type is owned by exactly one layer. A given message is allocated, held, delivered and released by that layer alone. No message crosses.

This is the heat-side analogue of §XXXII's control-side ruling, and it has the same effect: two systems coexist without any shared object. Double-accounting becomes structurally impossible rather than carefully avoided — there is no message for both layers to charge for.

XXXIV.3 — The staging, reusing §XXVIII's own discipline rather than inventing one

FABRIC-3.md §XXVIII faced the structurally identical problem — standing up a second control-transfer mechanism beside a live one — and solved it with "explicit go/no-go gates, not attempted as one pass," each stage inert until the next enables it: Stage 1 allocated per-VM stacks with "nothing executes on them yet"; Stage 2 built the switch primitive "cooperative only, no timer"; Stage 3 turned on the timer. One commit per increment, full three-architecture acceptance before the next begins.

That discipline is the answer here, applied to messaging:

Stage Content Live layers
A Kernel-Hermes structures + heat-coupled allocator. Wired to nothing, draws no heat. Boot is byte-identical; dict_hash unmoved (no FORTH edit) FORTH only
B Prove the allocator (item 28): a test path allocates and frees N messages, fleet_conserved holds across it. Heat drawn and returned within the test FORTH only
C Cut over exactly one message type — the narrowest live one. BLK-ATTACH-EVENT (a single Hera↔Artemis request/ack pair) is the natural candidate Both, disjoint
D Remaining live types cut over one at a time Both, disjoint, shrinking
E Category B strip: hermes/init.4th, messaging.4th, routing table, slot-3 pairing Kernel only

Stage A is the one that makes this safe, and it is the stage most likely to feel like wasted work. It is not: it is what makes every later stage a revert of one commit.

XXXIV.4 — A coupling nobody noticed: the Tripod change cannot land before the cutover

§XIX.6 puts Hestia into is_fleet_foundation (capsule_birth.c:793-796) and replaces Hermes's birth at kernel_main.c:865. But FORTH Hermes must stay alive through Stages C and D — it is still routing every type not yet cut over.

So the two changes are sequenced, not simultaneous, and the intermediate state is explicit:

  • is_fleet_foundation temporarily holds four names — Hera, Artemis, Hestia and Hermes. That line is an OR-chain of prefix comparisons, so a fourth is trivial; what matters is that it is deliberate and temporary, not an oversight for a later reader to "clean up."
  • Both births run at kernel_main.c: Hestia's is added, Hermes's is not yet removed.
  • Hermes's birth and its is_fleet_foundation entry are removed in Stage E, with the strip, not with the Tripod change.

This contradicts nothing in §XIX — it sequences it. Worth writing down because the obvious reading of §XIX.6 ("swap Hermes for Hestia") would break Stages C and D on the first boot.

XXXIV.5 — Cost and rollback, honestly

Rollback is one commit per stage, which is the entire reason for staging. Tags remain the oracle (.claude/CLAUDE.md: "when in doubt about the correct state of any file or branch, look at the tag first").

The cost is acceptance runs. Per .claude/CLAUDE.md the only valid acceptance is all three architectures booted in QEMU, one at a time, in the foreground, clean before qemu. Stages A–E plus the Category B strips is at least eight full three-architecture cycles, each with logs committed. That is the price of the discipline, and it is the same price §XXVIII paid across Stages 0–4. Naming it now so it is a plan rather than a surprise — amd64 under TCG is the slow one, and .claude/CLAUDE.md warns the session must stay engaged through long runs.

XXXIV.6 — What would make me stop and re-plan

Stated as tripwires rather than hopes, since the point of this pass is confidence:

  • Stage B fails — fleet_conserved will not hold across a bare alloc/free cycle under the new allocator. That is the piece §XXXIII named as the only genuinely hard one, and failing it early is the cheap outcome. Do not proceed to C.
  • Stage C shows allocation refusals that the FORTH-only baseline did not. That is the oversubscription of §XXXIV.1 arriving, and the answer is to retire channels first (§XXXIII, recovering ~a third of the budget) rather than to push on.
  • dict_hash diverges across architectures at any stage. Not "changes" — changing is expected and normal (§III.5). Diverging between amd64/aarch64/riscv64 is the signal that something non-deterministic entered the dictionary, and it is the one failure this project's acceptance criteria are specifically built to catch.

XXXIV.7 — Punch list

  • ⬜ Item 29 — adopt §XXXIV.2's partition rule and §XXXIV.3's Stage A–E staging as the kernel-Hermes build plan (replacing item 6's single line).
  • ⬜ Item 30 — sequence the Tripod change per §XXXIV.4: Hestia added alongside Hermes, is_fleet_foundation temporarily four names, Hermes removed only at Stage E.
  • Item 28 (prove the allocator) becomes Stage B and keeps its priority.
  • Item 27 (channels: negotiate or one membership) gains urgency — §XXXIV.1 shows retiring channels directly relieves the transition's budget pressure.

XXXV. Pre-mortem: it is six months on, the reshuffle failed — why? (2026-09-19)

Run at Captain Bob's request before pulling the build trigger. Method: assume failure, work backwards, and ground every cause in this project's own recorded failure record rather than in generic risk. Two of the six are new findings that change the design; one is a hypothesis that weakened when checked, recorded as such.

XXXV.0 — This codebase's dominant failure mode, evidenced

Before the specific causes, the pattern they mostly instantiate. Things here fail without saying anything:

  • §XXVIII.1: "Stage 3 emits nothing per switch by default, so a switch storm and a healthy idle REPL produce an identical serial log. The corruption was already live in every one of those 'clean' Stage 3 boots."
  • §XXXII.2 (FABRIC-3): a console-name mismatch "silently falls back to direct interpretation with no error — just quietly never relays."
  • §XIV/§XXXI: near-simultaneous attach — 6 of 9 identities silently never triggered their attach line.
  • §XX: "Referencing an undefined word during FORTH compilation doesn't raise a hard error, it silently drops the definition being compiled."

Four independent subsystems, one signature. Any pre-mortem for this project that does not put silent failure first is not about this project.

XXXV.1 — Cause 1, most likely: the sinking latch never fired, because nothing could raise it

This is the pre-mortem's principal finding, and it is a design gap, not a risk.

§VIII.1 has Hermes raise a latch when it detects itself sinking, holding it until it goes down, to give the fleet a window. That was conceived when Hermes was a VM on the Stadium floor — a patron with heat, a TTL and a reservoir, capable of degrading.

Kernel-Hermes is none of those things. stadium_admit() is called only from capsule_birth.c:808, on birth. Kernel-Hermes is not born, so it is not a Stadium patron: no heat, no TTL, no eviction, no compudynamic death. Which raises the question §VIII never had to answer:

What does it mean for a kernel-resident C subsystem to sink?

Its real failure modes are:

  • A bug → sk_hal_panic() → while(1){arch_halt();}. Immediate and total. No window at all — the fleet never gets to shut down cleanly, which is precisely what §VIII.1 exists to provide.
  • Resource exhaustion → allocation refusals. Degraded, but not dying, and recoverable. Raising a monotonic "I am sinking" latch here would be wrong (§XXIII.3 made it irreversible).

And §VII's ladder does not apply to it either. Rung 1 is warm restart, rung 2 is a from-scratch BIRTH — but BIRTH births VMs. There is no birthing a kernel subsystem. §XXI.3's table gives every rung an actor for VMs; kernel-Hermes sits outside that table entirely.

So the failure story for the one component that cannot use SOS is undesigned. Six months on, the most likely shape of "it failed" is: the arbiter hit a bug, panicked, halted the machine instantly, and every VM died mid-write with no window — with the latch never set, because there was no state between healthy and gone.

Punch item 31, and it gates §VIII/§XXIII rather than following them. Candidate directions, none chosen: give kernel-Hermes an explicit degraded state it can detect and announce (queue exhaustion, corrupted routing table) and raise the latch there; or accept that its only failure is a panic, and make the panic path itself raise the latch and yield briefly before halting; or conclude §VIII was scoped to a component that no longer exists and retire it. The third is a real possibility and should not be dismissed for tidiness.

XXXV.2 — Cause 2: the fleet hit a ceiling of 16

SK_SWITCH_MAX_SLOTS is 16 (capsule_vm_switch_signal.c:42), with the header noting compaction "after SK_SWITCH_MAX_SLOTS attach/detach cycles."

That was a reasonable bound when the switcher was one of two ways a VM got control. §XXXII made it the only one. So the moment that ruling lands, 16 becomes a hard ceiling on how many VMs can receive control at all — and §XXII's own comment says "the real fleet is headed toward hundreds of VMs."

This is the §XXII pump story repeating one layer over: a bound that is invisible at Tripod scale, fine at 9, and a wall later. §XXII found its wall live, at 9 VMs, on the slowest architecture, mid-campaign. This one is findable now, on paper, before it is built on.

Punch item 32: size the switch-slot table against the intended fleet before §XXXII's ruling is implemented. Not necessarily hard — but it must be a decision, not a default carried forward from when it did not matter.

XXXV.3 — Cause 3, weakened on inspection and recorded honestly

The hypothesis: the reshuffle's tripwires (§XXXIV.6) all depend on fleet_conserved, while §XXV.1 has the DoE's invocative paths being rewritten in the same period — so the instrument that verifies K would be under reconstruction exactly when it is most needed.

Checked, and it is weaker than it looked. vm_physics_conserved() lives in capsule_vm_physics.c:492 — a kernel function, independent of the DoE — and is exposed as a FORTH word at mama_forth_words.c:2099 (VM-CONSERVED?). doe_log.c only reports it, in columns 19/20. The capability survives a DoE rewrite.

The residual is real but smaller: what the DoE provides is the continuous, per-tick record of K across a long run. A word you have to call gives you a point sample. Stage B's proof (§XXXIV.3) should therefore not lean on the DoE — it needs its own deterministic alloc/free/check loop calling VM-CONSERVED? directly. Recorded so Stage B is designed for an instrument that exists rather than one mid-rewrite.

XXXV.4 — Cause 4: Hestia turned out to be ceremonial

§XVIII.2 established that the HAL must never know a VM exists, so Hestia's ownership of the fabric is by vocabulary and registration — not enforcement. §XVIII.2 called that "weaker than it sounds and exactly right."

The six-months-on failure: nothing changed. Other VMs still reach the framebuffer, Hestia owns the fabric only in the documentation, and the reorganization was cosmetic — a rename plus a capsule.

The test is narrow and checkable: after the reshuffle, is PLOT/FB-WIDTH/FB-HEIGHT registered anywhere except Hestia's word table? In particular, register_child_vm_words() must not hand them out the way it hands out the eight STADIUM-* primitives. If it does, ownership is a fiction on day one.

Punch item 33: verify the framebuffer primitives are registered to Hestia alone, and treat any other registration site as a defect rather than a convenience.

XXXV.5 — Cause 5: the strip took something live, and the usual net was down

§XXII.2's near-miss (the init-l8-* family reading as zero-reference) plus §XXII.3's acceptance asymmetry — a bad DoE-capsule strip passes all three architectures cleanly and surfaces months later during a campaign.

What makes this worse than §XXII assumed: §XXV.1 has the DoE's invocative paths being rewritten in the same window. The campaign that would eventually catch a wrong strip is itself in flux, so "we would have noticed at the next campaign" is a weaker assurance than it was when §XXII was written.

Mitigation, no new mechanism needed: do the Category A strip before the DoE rewrite begins, while the existing campaign apparatus still runs, and run one campaign against the stripped tree as the strip's real acceptance — not just a boot.

XXXV.6 — Cause 6: it was declared done, and it wasn't

The project's most repeated failure, by its own record: §XXVIII's "Stage 3 CLOSED" corrected by §XXVIII.1; §XXV corrected by §XXVI ("wrong VM tested"); FABRIC-2's close header overclaiming (§XXXI.3); the v2.0.1 bump rolled back (§I.2).

Already partly defended — §XXIX refused the LTS label and wrote criteria instead, §XXXIV.6 states tripwires as stop conditions. The remaining exposure is the acceptance definition itself: three green boots prove the boot path, and §XXXV.0 says this codebase's failures do not announce themselves on the boot path.

Punch item 34: for each of Stages A–E, name in advance the observation that would prove it worked — not "it booted," but the specific evidence, chosen before the stage runs. §XXVIII.1 is the cautionary case: a stage that emits nothing cannot be verified by reading a log that looks identical either way.

XXXV.7 — What the pre-mortem changes

Two new design gaps (§XXXV.1's undesigned kernel-arbiter failure story, §XXXV.2's slot ceiling), two verification corrections (§XXXV.3's Stage B instrument, §XXXV.6's name-the-evidence rule), one sequencing change (§XXXV.5: strip before the DoE rewrite), and one checkable invariant (§XXXV.4).

  • ⬜ 31 — design kernel-Hermes's own failure story, or retire §VIII. Gates §VIII/§XXIII.
  • ⬜ 32 — size SK_SWITCH_MAX_SLOTS against the intended fleet before §XXXII lands.
  • ⬜ 33 — framebuffer primitives registered to Hestia alone, verified not assumed.
  • ⬜ 34 — per-stage success evidence named in advance, per §XXVIII.1's lesson.
  • ⬜ 35 — Category A strip before the DoE rewrite, with a campaign as its acceptance.

The honest summary: the most likely cause of failure is not the messaging rewrite. That is now sized (§XXXIII), staged (§XXXIV) and instrumented. It is that the one component which cannot call for help has no defined way to fail — and that this codebase's failures are characteristically quiet.


XXXVI. Item 31 SETTLED: kernel-Hermes sinks when its own heat accounting stops adding up

§XXXV.1 found §VIII's latch had no trigger: kernel-Hermes is not a Stadium patron, so it has no heat, no TTL and no compudynamic death, and §VII's ladder does not reach it. Settled here — §VIII survives, with a trigger that is concrete, cheap, and already precedented twice in this codebase.

XXXVI.1 — What can actually go wrong, enumerated

Mode Shape Is it "sinking"?
kmalloc failure for a message Refuse the send, return the heat No — refusal, by design
Budget exhaustion (reservoir empty, Q.SLOT unaffordable) MSG-ALLOC refuses, rolls back No — this is the economy working
Corrupted internal state — queues, routing table, free list, heat bookkeeping Still executing, can no longer route correctly, and may not know Yes. This is the only real one.
Hard bug in its own code sk_hal_panic() → while(1){arch_halt();} No — it is already gone. §XXXVI.4

Only the third mode is a state between healthy and gone, which is precisely what a held-until-death latch is for. The other three are either recoverable or instantaneous.

XXXVI.2 — What §VIII's window is actually worth, which is why it is not retired

§XXXV.1 raised retiring §VIII as a legitimate option. It is not the right call, and the reason is concrete: the window buys a flush.

Per §XVI, clean shutdown is forced blk_flush(0) then BYE. block_subsystem.c carries dirty RAM blocks (g.dirty_ram) that a hard halt loses. Artemis holds block state; identity VMs may have written. The window's value is not messaging continuity — §XV.4 already ruled that payloads may vanish. It is that dirty blocks reach the disk. That is worth a mechanism.

XXXVI.3 — RULING: the sinking condition is a failed self-audit of its own heat accounting

Kernel-Hermes continuously audits its own heat bookkeeping. When the audit fails, it can no longer be trusted to route, and that is sinking: raise the latch, hold it, let the fleet flush and BYE, then stop.

The invariant is internal and needs no reservoir of its own:

  sum(heat held in live messages)  ==  total_pulled - total_returned

If those diverge beyond an epsilon, the arbiter's accounting is corrupt — and an arbiter that has lost track of the heat it holds is exactly an arbiter that will break K for everyone else. It is the earliest honest moment it can know it is failing.

This is not a new mechanism. It is the third instance of a pattern already in the tree:

  • MSG-K (messaging.4th:364) — the FORTH messaging layer already audits its own heat: MSG-TOTAL-HEAT CH-TOTAL-HEAT + STADIUM-RES@ + STADIUM-WORD-HEAT +. The layer being replaced already does this.
  • vm_physics_conserved() (capsule_vm_physics.c:492) — the C shape, verbatim: sum, compare against the expected constant, tolerate VM_PHYSICS_EPSILON_Q48.

One honest adaptation: MSG-K is per-VM — it sums a patron's arena plus its own reservoir. Kernel-Hermes has no reservoir, holding heat pulled from senders'. So the formula is not a straight port; the invariant above is its analogue for a non-patron holder. MSG-K establishes the idea is already this layer's practice, not that the code lifts across.

A property worth naming: §XV.4 ruled that K is the only thing that must survive a death. Under this ruling, a violation of that same conservation is what announces the death. The invariant being protected and the alarm are the same quantity — no second mechanism, nothing to keep in sync.

XXXVI.4 — Panics get no window, deliberately

A panic raises nothing and flushes nothing. sk_hal_panic() prints and halts, and it should stay that way:

  • The corruption that caused the panic may be in the block layer. Flushing from a panic path risks writing corrupt data over good data — worse than losing the dirty blocks.
  • Standard practice in any kernel: do not do complex work in the panic path.

So the accepted loss is stated rather than engineered around: if kernel-Hermes panics, dirty blocks are lost and the fleet gets no window. The self-audit exists precisely to catch the recoverable-but-doomed case before it becomes a panic. It narrows the window of silent loss; it does not close it, and claiming otherwise would be the overclaiming §XXXV.6 warns about.

XXXVI.5 — §VII's ladder is VM-only. Stated, because §XXI.3 implied otherwise

§XXI.3's table gives every rung an actor and reads as universal. It is not. Rung 1 is warm restart, rung 2 is a from-scratch BIRTH — and BIRTH births VMs. There is no birthing a kernel subsystem, and no parent to birth it.

Kernel-Hermes has a two-state life: correct, or stopped. No rungs, no recovery, no re-birth. The ladder applies to Stadium patrons. Kernel-Hermes is not one.

This does not weaken §VII — it bounds it, and the bound was always implicit in "authority flows down the birth graph" (§IX.2): a thing that was never born has no place on that graph.

XXXVI.6 — Is a self-audit a supervisor? No

Worth testing, given how much of this document defends against exactly that (§XIV.5, §XV.3, §XXI.2):

  • It watches only itself — no other component's liveness, no other component's fate.
  • It decides nothing about anyone else. It publishes one fact about its own state; VMs act on their own shutdown routines (§XXIII.4, §IX.1).
  • It is not a timer over the fleet. The natural cadence is per-allocation and per-release — the moments its own bookkeeping changes — not a scan of anything.

A component checking its own invariants is not a supervisor. It is the same thing vm_physics_conserved() already does, one scope down.

XXXVI.7 — Punch list

  • ✅ Item 31 — SETTLED (§XXXVI): §VIII kept; the sinking condition is a failed self-audit of kernel-Hermes's own heat accounting; panics deliberately get no window; §VII's ladder is VM-only.
  • ⬜ Item 36, NEW — choose the audit's epsilon and cadence. VM_PHYSICS_EPSILON_Q48 (5% of Q48_ONE, per §25.7's own note) is the existing precedent but is far too loose for this purpose — it was sized for fleet drift, not for detecting a corrupt ledger. A tighter bound is wanted; a wrong one either cries wolf or never fires.
  • ⬜ Item 37, NEW — Stage B (§XXXIV.3) should prove the audit as well as the allocator. They are the same arithmetic seen from two sides, and proving them together costs nothing extra.
  • Items 32–35 unchanged (§XXXV.7).

XXXVII. Item 36 SETTLED: the audit epsilon is zero — and §XXXVI's invariant needed correcting first

XXXVII.1 — Correction to §XXXVI.3: the invariant as written is wrong

§XXXVI.3 proposed sum(heat held in live messages) == total_pulled - total_returned. Traced 2026-09-19, that cannot hold, and the reason is by design:

: MSG-COOL-ONE ( m -- )
  DUP MSG-HEAT@ Q-DECAY Q.* SWAP MSG-HEAT! ;

MSG-COOL-ONE multiplies a message's heat by Q-DECAY (65208/65536 ≈ 0.99499) and writes it back. The difference is returned to nothing — it is destroyed. So a message allocated at Q.SLOT and cooled a few times holds less than was pulled for it, and the two sides of the invariant diverge without any bug at all. This is Loop #3's decay doing its job.

§XXXVI.3 is corrected here rather than rewritten, per this series' own rule. The ruling it made — that the sinking condition is a failed self-audit of kernel-Hermes's own heat accounting — stands. The arithmetic it proposed does not.

XXXVII.2 — The numbers, which settle the epsilon question on their own

Quantity Value
Q48_ONE 65536 (q48_16.h:67)
VM_PHYSICS_EPSILON_Q48 3277 (capsule_vm_physics.c:57) — 5% of Q48_ONE
Q.SLOT today = (Q.1 − Q.1/3) / (32 + 16 − 1) ≈ 929
Q.SLOT after retiring channels (§XXXIII) = /32 ≈ 1365

The existing epsilon is 3.5 messages' worth of heat at today's Q.SLOT, or 2.4 after §XXXIII. Kernel-Hermes could lose track of three entire messages and the audit would not fire. That is not "too loose" as §XXXVI.7 put it — it is loose enough to hide the exact failure the audit exists to catch.

VM_PHYSICS_EPSILON_Q48 is not wrong; it was sized for a different quantity — fleet heat, which is genuinely noisy. It must not be reused here.

XXXVII.3 — RULING: epsilon is zero, made achievable by ledgering the decay

The corrected invariant:

  sum(heat held in live messages)  ==  total_pulled − total_returned − total_decayed

All four terms are exact integers. Q.SLOT is a compile-time constant, pulls and returns are symmetric, and decay's per-application delta is computable at the moment it is applied — heat_before − heat_after, already in hand. Nothing here is a measurement. There is no noise to tolerate.

Epsilon is zero. Any divergence is a bug, and the audit fires on the first unit.

The one change required is small and is the whole trick: today decay is an invisible overwrite — MSG-HEAT! destroys heat and records nothing. Kernel-Hermes must record what it destroys. One counter, incremented by the delta on every decay application. That converts an unauditable quantity into an exact one, and it is why the epsilon can be zero rather than guessed.

This also answers §XXXVI.7's "wrong either way" worry directly: a guessed epsilon either cries wolf or never fires. A zero epsilon over an exact ledger does neither — it is a correctness check, not a threshold.

XXXVII.4 — Cadence: O(1) counters, never a scan

MSG-TOTAL-HEAT (messaging.4th:267) computes its sum by scanning the whole arena — MSG-MAX 0 DO … LOOP. Fine for 32 slots per VM. As a fleet-wide kernel structure it is §XXII's O(N)-per-beat wall again, the mistake this reshuffle has now identified three times (the pump, SK_SWITCH_MAX_SLOTS, this).

So the audit maintains running counters rather than recomputing a sum: held, pulled, returned, decayed, each updated at the three moments they change — allocate, release, decay. The check is then one comparison of four integers, O(1), and can run on every mutation without a cadence decision at all.

A scan remains useful as a verification of the counters themselves — run it in Stage B and in diagnostics, not on the hot path. That is the same split vm_physics_conserved() already uses: cheap invariant continuously, expensive truth occasionally.

XXXVII.5 — What this does not change

Fleet-level vm_physics_conserved() keeps its 5% band. Different quantity, genuinely noisy, and outside this reshuffle. §XXXIV.6's tripwire still uses it as written — the two epsilons answer different questions and should not be unified.

XXXVII.6 — Flagged, not resolved: decay means fleet heat is not strictly conserved

A finding this section surfaced and deliberately does not settle. If message heat decays and is destroyed, and MSG-FREE-NODE/STADIUM-EVICT return only what remains, then heat pulled from a VM's reservoir does not fully return to it. Fleet heat would drift downward over a long enough run — which is precisely what FABRIC-0.md §25.7 warned of for a different cause (vm_physics_touch() share truncation, since fixed per FABRIC-1.md:468): "a long enough run would trip VM-CONSERVED?. Nobody has measured the rate."

Limit of this trace, stated honestly: STADIUM-EVICT's return semantics were not read end to end for this section, so whether the decayed portion is genuinely lost or accounted elsewhere is not established here. It may be intended — decay is the physics — in which case "conservation" means "conserved modulo decay" and the 5% band is doing exactly that job.

Punch item 38: establish whether decayed message heat returns to the reservoir. It bears on §XV.4's "nothing survives but K" and on §XXXIV.6's tripwire, and it is cheap to settle by reading one function. Reported, not fixed.

XXXVII.7 — Punch list

  • ✅ Item 36 — SETTLED (§XXXVII): epsilon zero; achieved by ledgering decay as an explicit counter; VM_PHYSICS_EPSILON_Q48 explicitly not reused. §XXXVI.3's arithmetic corrected in place.
  • ⬜ Item 38, NEW — does decayed message heat return to the reservoir? (§XXXVII.6)
  • ⬜ Item 37 unchanged — Stage B proves the audit alongside the allocator, and should use the scan to verify the counters (§XXXVII.4).
  • Items 32–35 unchanged.

XXXVIII. Item 38 SETTLED: decayed heat does not return — and a reaped message returns nothing at all

§XXXVII.6 flagged this as unestablished and named the one function to read. Read now, end to end. The answer is worse than §XXXVII.6 estimated, and the cause is a design conflation rather than a bug.

XXXVIII.1 — The trace

Four facts, each from source:

  1. Message heat is Stadium cell heat. MSG-HEAT@/! are not local fields — messaging.4th:222-223: : MSG-HEAT@ ( m -- q ) MSG-STADIUM-CELL@ STADIUM-HEAT@ ;. So MSG-COOL-ONE decays header->heat in the Stadium cell itself.
  2. Eviction returns whatever remains, and the code knows exactly why (stadium.c:459-466): "the departing patron's remaining heat must flow back to its owner's reservoir before the cell returns to the free list, or every reap leaks heat and Σ(resident) + reservoir drifts below Q48_ONE." → stadium_quotas[slot].reservoir += header->heat;
  3. Reaping fires exactly when heat reaches zero (messaging.4th:302-305): MSG-REAP → MSG-SCAN @ MSG-HEAT@ 0 = IF … MSG-FREE-NODE.
  4. Therefore, at the moment a naturally-aged message is evicted, header->heat is 0.

A message that ages out returns reservoir += 0. The entire Q.SLOT pulled to send it is destroyed.

This is not the partial loss §XXXVII.6 anticipated. The eviction path's heat-return — which exists specifically to stop reaps leaking — is bypassed completely for every aged-out message, because decay has already drained the field to zero before eviction reads it. The guard works; decay walks around it.

XXXVIII.2 — The root cause: one field carrying two incompatible meanings

header->heat is used as both:

  • An activity metric — Loop #3's "quiescent words lose heat over time." Decaying this is correct and is the physics.
  • A reservation currency — Q.SLOT is described in messaging.4th:44 as "per-item admission heat for MSG-SEND/CH-ACCEPT," pulled from a reservoir and expected back.

Decaying a budget token destroys budget. A message sitting in a queue quietly becomes cheaper, and on reap refunds nothing. The two meanings are irreconcilable in one field, and the leak is the visible consequence.

Whether it is intentional is genuinely arguable. "You pay to send; if nobody collects, you forfeit" is a coherent backpressure economy. But nothing in the tree replenishes reservoirs — every VM starts at Q48_ONE and the system only ever conserves or loses — so the forfeit is permanent and the pool is finite. A long-lived VM that sends messages nobody collects eventually cannot send at all.

XXXVIII.3 — A precision boundary, stated rather than glossed

There are two distinct heat accountings in this system, and this finding lands in one of them:

  • Stadium accounting — patron-cell heat plus stadium_quotas[].reservoir, audited by MSG-K (MSG-TOTAL-HEAT + CH-TOTAL-HEAT + STADIUM-RES@ + STADIUM-WORD-HEAT). This is what leaks.
  • VM-physics accounting — vm_physics_fleet_heat_sum() over execution_heat_q48, which is what vm_physics_conserved() checks against Q48_ONE and what FABRIC-3.md §XVIII verified live as K.

Whether the two are coupled was not established here, so this section does not claim the message leak trips VM-CONSERVED?. It may be invisible to it entirely. What it does claim is narrower and certain: Stadium reservoir heat is destroyed, permanently, at Q.SLOT per reaped message. Establishing the coupling is worth doing before §XXXIV.6's tripwire is relied on — punch item 39.

XXXVIII.4 — RECOMMENDATION: separate the reservation from the age

Greenfield C is the moment this costs nothing to fix, because kernel-Hermes is writing the structure anyway (§XXXIII). Give a message two fields instead of one:

Field Semantics Decays? Returned at free?
reservation the Q.SLOT pulled from the sender No In full, always
age drives reaping when it expires Yes n/a

The leak disappears, reaping keeps working (it reads age, not the budget), and the backpressure economy is preserved in its honest form: sending costs you the reservation for as long as the message lives, and you get it back when it dies — rather than being silently taxed for messages nobody collected.

Captain Bob's call, because it is an economic behaviour change, not a repair. Option (a) — keep the forfeit, document it as intentional — remains legitimate, but should then be stated as a designed sink rather than left looking like the leak stadium.c:460 was written to prevent.

XXXVIII.5 — This makes §XXXVII's ledger simpler, not harder

A pleasing consequence. §XXXVII had to correct §XXXVI's invariant because decay broke it, and introduced a decayed counter to restore exactness:

  held == pulled − returned − decayed          (§XXXVII, with decay in the budget)
  held == pulled − returned                    (with reservation separated from age)

Separating the fields restores §XXXVI.3's original invariant exactly, and deletes the counter §XXXVII.3 had to invent. The epsilon stays zero either way — but under this recommendation it is zero over a simpler ledger, with one fewer thing to get wrong.

Both §XXXVI.3 and §XXXVII.3 were right about the destination and wrong about the obstacle. The obstacle was never decay itself; it was decay being applied to the wrong field.

XXXVIII.6 — Punch list

  • ✅ Item 38 — SETTLED (§XXXVIII): no. Decayed heat does not return, and an aged-out message returns nothing at all — the eviction guard is bypassed because decay zeroes the field first.
  • ⬜ Item 39, NEW — establish whether Stadium accounting and VM-physics accounting are coupled, i.e. whether this leak is visible to VM-CONSERVED? (§XXXVIII.3). Do this before relying on §XXXIV.6's tripwire.
  • ⬜ Item 40, NEW — rule on §XXXVIII.4: separate reservation from age, or keep the forfeit and document it as a designed sink.
  • Items 32–35, 37 unchanged.

XXXIX. Item 39 SETTLED: the two accountings are not coupled — and §XXXIV.6's tripwire watches the wrong one

§XXXVIII.3 declined to claim the message-reap leak was visible to VM-CONSERVED?, and filed this to settle before §XXXIV.6's tripwire was relied on. Settled: they are entirely separate systems, and the tripwire as written would have stayed green through exactly the failure it was meant to catch.

XXXIX.1 — The trace: no coupling, in either direction

Check Result
capsule_vm_physics.c references to stadium 0 — and it does not include stadium.h (only starforth_config.h, for the unrelated STADIUM_CAPACITY_TICK)
capsule_vm_physics.c references to reservoir 0
stadium.c references to vm_physics 1 — a comment, not code
stadium_words.c references to vm_physics 0

vm_physics_fleet_heat_sum() sums n->physics.execution_heat_q48 over the physics node list. Every writer of that field is internal: initialise (is_root ? Q48_ONE : 0), transfer between two VMs, redistribute on death. Nothing reads or writes a Stadium reservoir or patron cell. Message allocation cannot move it.

XXXIX.2 — They are not even the same shape of invariant

StadiumVMQuota.reservoir's own doc comment (stadium.c:105-109):

"Invariant: Σ(resident patron heat) + reservoir == Q48_ONE, checked the same way vm_physics_conserved() checks the fleet sum."

…and stadium.h:484 confirms "conservation is per-VM."

Scope Quantity Target
Stadium Per-VM patron-cell heat + that VM's reservoir Q48_ONE each
VM-physics ("K") Fleet-wide Σ execution_heat_q48 over live VMs Q48_ONE total

Two invariants, two scopes, two disjoint data sets — sharing only the Q48.16 format, the constant Q48_ONE, and the word "heat." That shared vocabulary is what made them look like one system, and is why §XV.4 and §XXXIV.6 conflated them.

XXXIX.3 — Only one of them has a checker

vm_physics_conserved() is the only *_conserved() function in the kernel. Confirmed by exhaustive grep: the sole other match is mama_word_vm_conserved(), its FORTH wrapper.

The Stadium side has:

  • a documented invariant (the comment above),
  • a diagnostic print — stadium_words.c:328, "Stadium conservation: resident_sum=" — a human-readable dump, not a boolean,
  • MSG-K (messaging.4th:364), a FORTH word covering only the messaging slice, which someone must call.

There is no automated, continuous verification of Stadium conservation anywhere. Which means §XXXVIII's leak — Q.SLOT destroyed per reaped message — is invisible to every automated check in the system. It is not that the checks disagree; it is that nothing is looking.

XXXIX.4 — CORRECTION to §XXXIV.6: the tripwire monitors a quantity the failure cannot move

§XXXIV.6 states: "Stage B fails — fleet_conserved will not hold across a bare alloc/free cycle under the new allocator."

That is wrong, and wrong in the dangerous direction. fleet_conserved reports vm_physics_conserved(), which sums execution_heat_q48. Message allocation never touches that field. A kernel-Hermes allocator that leaked every reservation it took would leave fleet_conserved reporting a serene 1.

This is precisely §XXXV.0's signature — a check that passes while the thing it is supposed to guard is broken — and I built one into the plan. Item 39 existed to catch it before it mattered, and did.

Corrected tripwire for Stage B:

Stage B verifies (a) kernel-Hermes's own ledger (§XXXVII.3, epsilon zero) and (b) the Stadium per-VM invariant for the sending VM — Σ(resident patron heat) + reservoir == Q48_ONE — before and after the alloc/free cycle. fleet_conserved is not evidence for this stage and should not be cited as such.

Since §XXXIX.3 shows no boolean Stadium checker exists, Stage B either calls MSG-K, or reads stadium_words.c:328's dump, or — better, and cheap — adds the missing stadium_conserved() as its first act. That function is a five-line analogue of vm_physics_conserved() and the system has wanted one since the invariant was written down.

XXXIX.5 — Consequences for two earlier rulings

§XV.4 ("nothing survives except K") needs its K qualified. The heat a dying arbiter holds is Stadium heat. The K that FABRIC-3.md §XVIII verified live, and that §XV.4 cited, is VM-physics K. The arbiter's heat was never in the verified quantity. The ruling's substance stands — a dying arbiter must return what it holds — but it must say Stadium conservation, not K.

§XXXVI/§XXXVII's self-audit is promoted from safety net to sole instrument. With no continuous Stadium checker, kernel-Hermes auditing its own ledger would be the only automated guard on the messaging pool's conservation. That raises the value of §XXXVII.3's zero epsilon considerably — and argues for §XXXIX.4's stadium_conserved() as a second, independent check rather than relying on the arbiter to police itself alone.

XXXIX.6 — A naming hazard worth fixing in the sweep

Two conserved quantities, both Q48.16, both targeting Q48_ONE, both called "heat," differing only in scope — and the FABRIC series calls one of them "K" without qualification. Every future reader will conflate them exactly as §XV.4 and §XXXIV.6 did.

Added to item 13's documentation sweep: wherever "K" or "conservation" appears, say which — fleet K (vm_physics) or quota conservation (Stadium, per-VM).

XXXIX.7 — Punch list

  • ✅ Item 39 — SETTLED (§XXXIX): not coupled. Disjoint data, disjoint code, different invariant scopes.
  • ✅ §XXXIV.6 corrected (§XXXIX.4): Stage B verifies the ledger and Stadium per-VM conservation; fleet_conserved is not evidence for it.
  • ⬜ Item 41, NEW — add stadium_conserved(), the missing boolean analogue of vm_physics_conserved(). Cheap, and the invariant has been documented-but-unverified since it was written.
  • ⬜ Item 42, NEW — qualify "K"/"conservation" throughout the docs (§XXXIX.6). Folds into item 13.
  • ⬜ Item 40 unchanged — separate reservation from age (§XXXVIII.4).

XL. Item 40 SETTLED: preserve the consumption model, ledger the sink — and a correction to §XXXVIII's framing

§XXXVIII.4 recommended separating reservation from age to remove what it called a leak. Further evidence found while settling this says the behaviour is a deliberate model, not a defect — and the recommendation is revised accordingly.

XL.1 — The authors knew. It is documented at the reap site.

messaging.4th, block 5025, immediately above MSG-REAP:

( K reap only fires at heat=0: freed K contribution is 0. ) ( Force-reap not yet implemented. If added: explicit ) ( K redistribution will be required here. )

That is not an oversight; it is a design note. It states outright that reaping returns zero, and anticipates that a force-reap — freeing a message before it has cooled — would need "explicit K redistribution," i.e. handing the remaining heat back. The zero-return at natural reap is intended: the heat is modelled as consumed over the message's life, not held as a refundable deposit.

XL.2 — CORRECTION to §XXXVIII

§XXXVIII called it "a real, structural leak" and framed the two-meanings conflation as an error. That framing was wrong. The trace in §XXXVIII stands exactly — decay drains the field, reap returns nothing, the eviction guard is bypassed — but the interpretation does not. It is a consumption economy, documented where it happens, and §XXXVIII did not find that comment.

What survives from §XXXVIII: the mechanism, the arithmetic, and the observation that nothing replenishes. What is withdrawn: the word "leak," and the §XXXVIII.4 recommendation to change the economics as part of this reshuffle.

XL.3 — Two further roles for message heat, both consistent with the model

  • MSG-NACK-LAST halves it (messaging.4th:432): DUP MSG-HEAT@ 2 / OVER MSG-HEAT!. A rejected message ages twice as fast — a penalty expressed in the same currency. This only makes sense if heat is a time-to-live, reinforcing §XL.1.
  • Delivery order does not use it. MSG-DELIVER-ALL scans the arena in slot order and delivers anything not already delivered or nacked — no heat comparison anywhere. So heat plays no scheduling role, which is consistent with §XV.3 (nothing owns turn order) and means changing it cannot perturb delivery sequencing.

XL.4 — RULING: keep the model; make the sink explicit

Kernel-Hermes preserves the consumption economy exactly — heat decays, natural reap returns nothing — and additionally records what it consumes.

The Stadium per-VM invariant then reads:

  Σ(resident patron heat) + reservoir + consumed  ==  Q48_ONE

This is what makes item 41's stadium_conserved() possible at all. Without a consumed term, a checker over a system with a designed sink reports non-conservation as normal operation — which forces a tolerance that hides real bugs, or makes the checker useless. With it, the invariant is exact and the epsilon stays zero.

A pleasing arc worth recording: §XXXVII.3 invented a decayed counter to keep the ledger exact. §XXXVIII.5 proposed deleting it by separating the fields. It turns out to be the thing that makes Stadium conservation checkable — the counter was right, and for a better reason than the one it was introduced for.

XL.5 — Why not separate the fields now, despite it being cleaner

§XXXVIII.4's design is still the better architecture. It is rejected for this reshuffle on three grounds, none of them about elegance:

  1. Scope. This document's own discipline is "a reorganization, not an invention." Changing the fleet's economic model is an invention, and it would be smuggled in under a relocation.
  2. Measurement comparability. Every DoE campaign and the patent-supporting figures were taken under the current economics. §XXV.1 relieves some of that by rewriting the DoE's invocative paths, but it does not make a behavioural change to the economy free.
  3. The economics are not understood well enough to change safely. §XL.3 found two additional roles for message heat while settling a question about a third. That is a subsystem still yielding surprises, and changing its semantics mid-reshuffle is how a relocation becomes an outage.

XL.6 — The question that is real and is deliberately not answered here

Nothing replenishes a Stadium reservoir. Each VM's quota starts at Q48_ONE and thereafter only moves within that VM or is consumed. So a long-lived VM's capacity to send declines monotonically, and the decline accelerates with fleet size — batched pumping (§XXII's SK_MSG_PUMP_BATCH) means longer queue residency, more decay per message, more consumption.

This is the fourth instance of this reshuffle's recurring pattern: latent at Tripod scale, visible at fleet scale — after the pump's O(N) walk, SK_SWITCH_MAX_SLOTS, and MSG-TOTAL-HEAT's arena scan.

Whether a reservoir should ever replenish is a genuine open design question about the compudynamics, not a reshuffle decision. Punch item 43, explicitly outside this work, and properly FABRIC-3.md's or a successor's topic. Recorded here so it is not lost — and so that whoever meets a VM that has quietly stopped being able to send knows where it was first named.

XL.7 — Punch list

  • ✅ Item 40 — SETTLED (§XL): preserve the consumption model; kernel-Hermes ledgers consumed; §XXXVIII's "leak" framing withdrawn and its recommendation revised.
  • ✅ Item 41 clarified — stadium_conserved() checks Σ patron + reservoir + consumed == Q48_ONE, exactly, epsilon zero.
  • ✅ §XXXVII.3's counter vindicated — kept, and now load-bearing for §XL.4.
  • ⬜ Item 43, NEW, outside this reshuffle — should a Stadium reservoir ever replenish? (§XL.6)
  • Items 32–35, 37, 42 unchanged.

XLI. The build punchlist: phases and testable tasks

Asked: is it now possible to break this into phases and small testable tasks? Answer: for Phases 0–2, yes. For Phase 3, no — and the blockers are three named things, not general caution. Each task below carries its own check, which also discharges item 34 ("name the evidence before the stage runs").

Standing rules for every task: one task per commit; three-architecture QEMU acceptance (clean qemu, amd64/aarch64/riscv64, foreground, one at a time) before the next; logs committed; dict_hash identical across architectures (changing is expected, diverging is the stop condition, §XXXIV.6). No task starts without explicit authorization.

XLI.0 — Phase 0: preparation, no behaviour change

# Task Check
0.1 Establish Category A reachability per §XXII.2's three routes — boot path, tooling, baked capsule directory. Never by grep alone. A written list with the route that proves each entry dead
0.2 Strip common/msg.4th 3-arch boot; mkcapsule --lint clean
0.3 Strip process.4th (takes EVENT-EMIT/WAIT/DRAIN with it, §XXXIII.3) 3-arch boot; lint clean
0.4 Strip SPAWN-EVENT 3-arch boot; lint clean
0.5 Correct MANIFEST.md blocks 4055 and 2049 as their files are stripped (§XXII.5) Manifest describes no file that no longer exists
0.6 Return freed block ranges to capsule-reserved.txt lint clean
0.7 Add stadium_conserved() (item 41) — Σ patron + reservoir + consumed == Q48_ONE, epsilon zero, the five-line analogue of vm_physics_conserved() Returns true on a clean boot, all three arches
0.8 Verify PLOT/FB-WIDTH/FB-HEIGHT are registered nowhere but the table Hestia will own (item 33) Read-only audit; a second registration site is a defect to report

Phase 0 gate: all three architectures boot to zuse)ok>, stadium_conserved() true, no UNKNOWN WORD.

XLI.1 — Phase 1: Hestia, with messaging untouched

# Task Check
1.1 Allocate Hestia's block range against capsule-reserved.txt, avoiding 4997 mkcapsule --lint clean
1.2 Create capsules/hestia/init.4th — messaging load, MSG-CD-INIT, banner. No fabric yet Boots; Hestia absent from fleet (not yet birthed)
1.3 Add Hestia to is_fleet_foundation — four names, Hermes retained (§XXXIV.4) 3-arch boot; Hermes still live
1.4 Birth Hestia at kernel_main.c, alongside Hermes's existing birth Registry shows both; dict_hash identical across arches
1.5 Register Hestia for switch signals (:1007 pattern) Boot clean; no switch storm (§XXVIII.2's failure shape)
1.6 Move fabric.4th from init.4th to hestia/init.4th Hera's dict shrinks, Hestia's grows; cross-arch identity holds
1.7 Move font.4th likewise Same
1.8 Move PLOT/FB-WIDTH/FB-HEIGHT registration to Hestia's table only A non-Hestia VM calling PLOT gets UNKNOWN WORD — verify positively
1.9 Assert §XVIII.6's headless invariant in code: Hestia's birth sets no g_wirebind_attached_username, mints no proxy Boot headless with no thumbdrive; no prompt appears

Phase 1 gate: Tripod is Hera/Artemis/Hestia plus Hermes; drawing works from Hestia and only Hestia; headless policy intact.

XLI.2 — Phase 2: the allocator and its audit (Stage A/B, inert)

This is the phase that matters most and the one §XXXIII named as the only genuinely hard part.

# Task Check
2.1 Kernel-Hermes message/membership structures. Wired to nothing; draws no heat Boot byte-identical; dict_hash unmoved (no FORTH edit)
2.2 Heat-coupled allocate: pull Q.SLOT from the caller's reservoir, roll back on refusal Unit path: N allocs against a VM with known reservoir; refusal at the right count
2.3 Release: return remaining heat via the eviction path Reservoir restored exactly for an undecayed message
2.4 The four counters — held, pulled, returned, consumed (§XL.4) Each increments at exactly one site
2.5 Decay: apply, and record the delta into consumed (§XXXVII.3) consumed grows by exactly heat_before − heat_after
2.6 Self-audit: held == pulled − returned − consumed, epsilon zero (§XXXVII.3) Holds across the cycle; deliberately corrupt a counter → audit fires on the first unit
2.7 Stage B proof (§XXXIV.3, corrected by §XXXIX.4): alloc/free cycle verifying (a) the ledger and (b) stadium_conserved() before and after. fleet_conserved is not evidence here Both true before and after, all three arches
2.8 Scan-based cross-check of the counters, for diagnostics only, off the hot path (§XXXVII.4) Scan agrees with counters

Phase 2 gate — and the project's real go/no-go: if 2.7 fails, stop and re-plan. Do not proceed to Phase 3. §XXXIII said if K holds across alloc/free/evict under the new arbiter, the rest is protocol plumbing; this is where that is established or disproved, cheaply, with nothing else built on top.

XLI.3 — Phase 3: cutover — NOT YET BUILDABLE

Three specific blockers, each a ruling or a design, not a task:

  1. Item 27 — channels: negotiation or one broadcast membership? §XXXIII.5 recommends one membership and keeping the FORTH channel words until something asks for them. Unruled, and it changes what Phase 3 builds.
  2. Item 32 — SK_SWITCH_MAX_SLOTS is 16. §XXXII made the switcher the sole mover of control, so 16 becomes a fleet ceiling. Whether this is a constant bump or a table redesign changes Phase 3's shape, and cannot be estimated until it is decided.
  3. §XXXII.5.1 — the delivery hand-off is undesigned. Delivery must cause the payload to execute inside the target (messaging.4th block 5041). Under §XXXII the switcher resumes the target on its own stack, so kernel-Hermes must leave the payload somewhere the target consumes on resume. That mechanism does not exist on paper. It is the last genuine design question in the reshuffle, and it is on the critical path.

What Phase 3 will look like once unblocked (shape only, per §XXXIV.3): cut over BLK-ATTACH-EVENT alone — one request/ack pair, the narrowest live type — under §XXXIV.2's partition rule, one message type owned by exactly one layer; then the remaining types one at a time; gate each on stadium_conserved() plus the ledger, and on no new allocation refusals against the FORTH-only baseline (§XXXIV.6).

XLI.4 — Phase 4: Category B strip

Only after every live type is cut over. hermes/init.4th, the routing table, the slot-3 pairing convention, then messaging.4th itself. Hermes leaves is_fleet_foundation and kernel_main.c here — not earlier (§XXXIV.4). One commit per item; 3-arch boot each.

XLI.5 — Phase 5: close-out

Isabelle/HOL pass, deliverable being the restated boundary including §XXV.4's coverage loss, not a green build (§XXV.3) · documentation sweep, grepping by exclusion (§XXVIII.3), with CLAUDE.md's four errors, the TRIPOD.md/0.1 contradiction, and item 42's K-qualification · make sbom, checking Created: and DocumentName · LITHOS_VERSION = 2.1.0, engine VERSION per §XXX.6's rule · resolve the stray v2.0.1 branch and PR #1 · merge to master · tag v2.1.0 · archival close of this document.

XLI.6 — The honest assessment

Buildable now: Phases 0, 1, 2 — 25 tasks, each with a check, each one commit. That is roughly half the work and includes the hard part (Phase 2), which is deliberate: the riskiest piece is reachable early and fails cheaply.

Not buildable: Phase 3, pending two rulings and one design. Phases 4 and 5 are shape-complete but depend on 3.

A caution worth stating. Phases 0–2 can be executed without answering the Phase 3 blockers, and that is a feature — but it means arriving at a working, audited, inert allocator with the cutover still undesigned. Better to settle §XXXII.5.1's hand-off before Phase 2 finishes, so Phase 3 starts with momentum rather than a pause. It is the last real design question, and it is the natural next iteration.


XLII. The punchlist moves to FABRIC-3.6.md (2026-09-19)

Captain Bob: "If the punchlist is sufficiently long, maybe it goes to a 3.6 document to work from and annotate as we work." Agreed, and not on preference — this series has a stated rule and this document has just met it.

XLII.1 — The threshold is the series' own, and it was crossed while writing §XLI

FABRIC-1.md's close header gives the rule verbatim:

"At ~4,400 lines and 51 open items scattered among hundreds of resolved ones, continuing to append here made the still-open work hard to find — the same reasoning FABRIC-0.md itself was closed for at 7,595 lines."

This document stands at 4,452 lines with 40+ open punch items scattered among hundreds of settled rulings. Same threshold, same shape, same reason. Annotating §XLI's 25 tasks here, commit by commit, would bury the design record this document exists to be.

XLII.2 — The split, and why it is stated so firmly

FABRIC-3.5.md (this) FABRIC-3.6.md
Holds The design record — rulings, reasoning, evidence The work — tasks, results, dates, commits
State Design phase closed; archival close at v2.1.0 (§XXVI.5) Living until the work is done
On conflict Authoritative Defers, and records the discrepancy

Design rulings are never restated in FABRIC-3.6.md, only cited. That is deliberate: §XXXI found this series' documents losing track of each other, and duplication is how two documents begin to disagree. If executing a task proves a ruling wrong, FABRIC-3.6.md records the finding and this document is amended — never a silent divergence.

XLII.3 — The split repairs the defect §XXXI.2 found, in the successor

§XXXI.2's central finding was that the carry-forward discipline was mechanically auditable (grep -c '\- \[ \]') through FABRIC-2.md, and broke at FABRIC-3.md, which has no checkboxes at all — after which "is anything still open?" stopped being a grep and became a reading exercise. That is how §XXXI.3's six design items went quiet without anyone deciding to drop them.

FABRIC-3.6.md restores the checkbox convention. Every task, every blocker, every close-out step is a - [ ]. The gap analysis found a structural defect in how this series tracks work, and the next document in it does not repeat that defect — which is a better outcome than the findings list alone.

XLII.4 — What did not move

§XLI stays here, as the source §XXXVI.6 instantiates: the phase structure, the reasoning for Phase 2's early placement, and the three Phase 3 blockers with their justification. What moved is the checklist, not the argument.

FABRIC-3.6.md also carries an explicit out-of-scope section — item 43, §XXXI's items 23–26, and src/*.c.bak — so that work recorded here as deliberately excluded is not quietly absorbed by an executing session. §XXXI.3's six orphans are exactly what happens when that is left implicit.

XLII.5 — Punch list

  • ✅ FABRIC-3.6.md opened — execution log, checkbox convention, Phases 0–5, findings log, out-of-scope carry.
  • ⬜ B1/B2/B3 (items 27, 32, §XXXII.5.1) — the Phase 3 blockers, tracked in both documents. §XXXII.5.1, the delivery hand-off, is the last genuine design question in the reshuffle.
  • Design work continues here; execution is tracked there.

XLIII. B3 SETTLED: the target drains its own queue, at its own outermost checkpoint

The last genuine design question in the reshuffle (§XXXII.5.1, Phase 3 blocker B3). Settled, and once again the mechanism already exists — built for the structurally identical problem and corrected twice in the field.

XLIII.1 — What the hand-off has to satisfy

Delivery must cause the payload to execute inside the target — messaging.4th block 5041: "Payload is FORTH text, executed by the receiver, same as every other message here." But §XXXII retired the pump, so kernel-Hermes may not VM-EXEC into anyone. The payload must therefore wait somewhere the target consumes under its own power, at a moment when doing so is safe.

"Safe" is the hard part: interpreting a payload while the target is mid-word, mid-unwind, or running nested inside someone else's dispatch is the reentrancy class this codebase has been bitten by repeatedly (§XX, §XXVIII.2).

XLIII.2 — The safe-boundary machinery is already there

Traced 2026-09-19 in vm_core.c:

  • sk_vm_at_outermost_interpret() — return g_vm_interpret_depth <= 1; (:1164), with g_vm_interpret_depth incremented and decremented by vm_interpret() itself (:1173, :1206).
  • A per-word checkpoint in execute_colon_word() (:946), gated on that predicate.
  • A defer-don't-lose discipline, in the checkpoint's own words (:925-935): if a word is "executing nested inside a VM-EXEC/VM-CALL dispatch… vm here does not own the physical stack it's currently running on," so acting would corrupt the enclosing VM's call chain. It is therefore "left pending… so the next checkpoint at the outermost frame picks it up instead of losing it."
  • And it runs before the error/abort/exit-colon checks, "deliberately, so a switch never happens mid-unwind" — acting only when the word left the VM "in a clean, resumable state."

That is the exact predicate, the exact placement, and the exact deferral semantics delivery needs. It exists because the switcher needed to answer the same question — when is it safe to act on this VM? — and it has already been corrected twice in the field (2026-09-13 ×2, 2026-09-14).

XLIII.3 — RULING

Kernel-Hermes enqueues the payload onto the target's own pending queue and publishes the fact. It dispatches nothing. The target drains its own queue, in its own context, at its own outermost interpret checkpoint — the same checkpoint, the same sk_vm_at_outermost_interpret() gate, and the same defer-if-nested discipline the switcher already uses.

One published fact, two independent consumers — which is precisely the shape §XV.3 and §XXXII established:

Consumer Uses the fact for
The switcher eligibility — whether to resume this VM (§XXXII)
The target VM itself drain — whether to interpret a pending payload at its next safe boundary

Neither dispatches into anyone. Kernel-Hermes publishes; the switcher moves control; the VM does its own work. Three roles, no overlap, no second owner of anything.

XLIII.4 — This is §XX's proven-safe pattern, generalized to the whole fleet

FABRIC-3.md §XX already established the safe shape, for Hera specifically. With the pump skipping her, she drains her own queue via "a plain word dispatch in her own dictionary, not a VM-EXEC dispatch into anyone else's input buffer — no reentrancy risk."

The pump's defect was never draining; it was draining from outside. Hera was special-cased into safety. Retire the pump (§XXXII) and every VM does what Hera already does — the special case disappears by becoming universal. That is a relocation of a proven pattern, not an invention, and it keeps this document's scope discipline intact at the last design question.

XLIII.5 — Recursive drain is prevented for free

Interpreting a payload calls vm_interpret(), which increments g_vm_interpret_depth — so any checkpoint reached during the drain sees depth > 1 and will not drain again. The same counter that defines the safe boundary also makes the drain non-reentrant, with no flag, no lock and no new state.

XLIII.6 — Three constraints, named rather than discovered later

  1. INPUT_BUFFER_SIZE is 1025 — 1024 content bytes plus NUL, and .claude/CLAUDE.md calls this non-negotiable because vm_interpret() is the shared dispatch path for both REPL lines and LOADed block content. A payload above 1024 bytes cannot be interpreted in one drain. Today's MSG-PADDR/MSG-PLEN are out-of-line and carry no such bound. Kernel- Hermes needs either a payload bound at send time or a chunking story — decide at build time, do not discover it at Stage C.
  2. Starvation is possible and bounded. A VM in a tight loop never returns to outermost depth, so never drains. The switcher can still preempt it, but preemption does not create a boundary — the VM must still reach one. This is the switcher's existing risk, not a new one, and it should be recorded as shared rather than solved twice.
  3. Drain one message per checkpoint, not all. Draining the whole queue at one boundary runs unbounded work in a single checkpoint — the §XXII wall pattern in miniature. One per checkpoint is bounded and fair. Recommended, not ruled.

XLIII.7 — Punch list

  • ✅ B3 — SETTLED (§XLIII): the target drains its own queue at its own outermost checkpoint; kernel-Hermes publishes and never dispatches.
  • ⬜ Item 44, NEW — payload bound or chunking, against INPUT_BUFFER_SIZE 1025 (§XLIII.6.1). Decide before Phase 3.
  • ⬜ B1 (item 27) and B2 (item 32) remain — both are rulings, not designs.

Phase 3's design blocker is cleared. What remains before it is buildable are two decisions, neither of which requires further investigation: channels (one membership or negotiation), and the switch-slot ceiling. The reshuffle has no undesigned mechanism left.