Commit Graph
42 Commits
Author SHA1 Message Date
Robert Allan JamesandClaude Sonnet 5 3e201c82a5 Stage E: Category B strip -- remove Hermes, messaging.4th, old routing (FABRIC-3.6.md task 4.1-4.4)
All FORTH-owned message types were cut over to kernel-Hermes in Phase 3
(tasks 3.8-3.10). This removes the now-dead FORTH messaging layer and
the Hermes VM itself: capsules/common/messaging.4th, capsules/hermes/init.4th,
the slot-3 VM-NAME-REG pairing convention, and every load-site/birth-site
reference across capsules/init.4th, artemis/init.4th, hestia/init.4th,
doe-campaign.4th (Artemis-only now), capsule_console.c, capsule_mint.c,
capsule_wirebind.c, capsule_birth.c, and kernel_main.c.

Verified on all three architectures: clean boot, mkcapsule --lint clean
(36 files, 0 violations), zero UNKNOWN WORD, identical dict_hash across
amd64/aarch64/riscv64 for every VM, and a full mint -> WIREBIND-attach ->
USE -> relay round-trip exercising the two highest-risk edits
(capsule_console.c/capsule_mint.c).

Found, not fixed: deleting messaging.4th removes SEND-ELEVATE-REQUEST,
which was the only caller of KH-ELEVATE-SEND and the only path to
ELEVATE-GRANT (zuse-eligibility.4th, still loaded at boot) -- Phase 8
PKI's own elevation entrypoint. Needs a decision before Phase 5 close-out.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-22 18:29:51 -04:00
Robert Allan JamesandClaude Sonnet 5 cab9b5a31f Stage D: ELEVATE-REQUEST real cutover -- FABRIC-3.6.md task 3.10
Reachability verified live before writing any code, per this
project's own standing rule (grep cannot establish reachability
alone): FIND SEND-ELEVATE-REQUEST / FIND ELEVATE-GRANT / FIND
CH-REQUEST all resolve on a live Hera boot, though grep across
capsules/experiments/docs found zero callers of SEND-ELEVATE-REQUEST
-- a real, complete, directly-callable entrypoint (H.5/H.8's own
design) with no current automatic trigger, not dead code.

Correction to a prior finding, made in the course of this check: task
3.8's write-up claimed "Hera's own pre-existing inability to load
common:messaging.4th" -- false. capsules/init.4th (Hera's own
MAMA_INIT capsule) loads it directly, and SEND-ELEVATE-REQUEST lives
and works in her dictionary right now. Task 3.8's own actual scope is
unaffected by this correction.

Cutover: SEND-ELEVATE-REQUEST (messaging.4th) no longer ends in
CH-REQUEST's COMMON-CH/MSG-SEND path; it now calls KH-ELEVATE-SEND
(repl.c), a new C word wrapping sk_hermes_send_one(), registered
unconditionally for every VM. from/to are derived from the calling VM
and sk_get_mama_vm() directly in C, never taken from the stack -- a
real correctness improvement over CH-REQUEST's own initiator-only
gate, which only existed because a caller COULD pass the wrong from
value; deriving it in C makes that spoof structurally impossible.
SK_HERMES_MSG_TYPE_ELEVATE_REQUEST reuses ELEVATE-REQUEST's own value
(8), same partition-rule reasoning as tasks 3.8/3.9. Delivery is
unchanged task 3.4 machinery. No new static-buffer lifetime caveat --
the payload-aliasing fix landed before this task started.

CH-REQUEST (messaging.4th) is now dead code, its one real caller just
removed -- found, not fixed, per Captain Bob's Law.

Verified live on all three architectures: 0 0 0 0 S" DUP"
SEND-ELEVATE-REQUEST (deliberately-invalid pubkey, so ELEVATE-GRANT
correctly refuses -- the check is the pipeline running, not a grant
succeeding) fires the evidence line and completes cleanly, DUP
unaffected afterward. Zero UNKNOWN WORD, dict_hash identical across
all three architectures (changed uniformly from prior runs -- one new
word registered -- not diverged, matching SXXXIV.6's own rule).

This closes task 3.11 (Phase 3 gate): all of messaging.4th's live
FORTH-owned message types (BLK-ATTACH-EVENT, CONSOLE-CMD-EVENT,
ELEVATE-REQUEST) are now real kernel-Hermes cutovers. Phase 4 may
begin.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-22 17:00:11 -04:00
Robert Allan JamesandClaude Sonnet 5 5c07745c18 Fix SkHermesMessage payload-aliasing defect (task 3.8 findings log)
sk_hermes_send_one() -- the single funnel every sender, including
sk_hermes_publish(), already goes through -- used to store the
caller's own payload_addr pointer as-is. Two sends before either
drains meant both messages pointed at the same caller-owned buffer,
whichever send wrote last silently winning: real, confirmed live (a
second identity thumbdrive attached at boot alongside Zuse's own left
its WIREBIND pairing silently never happening).

SkHermesMessage gains an inline payload_buf[SK_HERMES_CHUNK_MAX_
PAYLOAD] field; sk_hermes_send_one() now memcpy()s the caller's
payload into it and points payload_addr at that copy instead. No
sender or reader call site needed to change -- every existing reader
already only ever reads through payload_addr, which still points at
valid bytes of the same length, now message-owned. g_kh_blk_attach_buf/
g_kh_console_cmd_buf (repl.c, tasks 3.8/3.9) no longer need to survive
past their own send call; their doc comments, which had claimed the
old aliasing shape was benign, are corrected.

Verified the original bug is actually gone: reproduced the exact
original scenario (a real, sequentially-minted rajames identity
attached at boot alongside Zuse's own, two simultaneous BLK-ATTACH-
EVENT sends in one idle-loop pass) -- WIREBIND pairing, USE, and the
Stage D relay all now work where WIREBIND previously silently failed.
Standard three-ISA acceptance also clean: zero UNKNOWN WORD, dict_hash
identical across all three and matching every prior acceptance run in
this document.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-22 14:16:29 -04:00
Robert Allan JamesandClaude Sonnet 5 3975dc61cd Stage D: CONSOLE-CMD-EVENT real cutover -- FABRIC-3.6.md task 3.9
Send-side cutover: sk_repl_dispatch_line()'s FORTH-string
"CONSOLE-CMD-EVENT 0 3 S\" ...\" 0 MSG-SEND" interpret is replaced with
a direct sk_hermes_send_one() call (SK_HERMES_MSG_TYPE_CONSOLE_CMD,
kernel_hermes.h, deliberately reusing CONSOLE-CMD-EVENT's own value 7,
same partition-rule reasoning as task 3.8's BLK-ATTACH-EVENT cutover).

Real finding along the way: kernel-Hermes's drain only ever runs as a
side effect of vm_interpret() being called on the target VM. Task
3.8's target (Hera) is always being interpreted via the interactive
REPL loop; task 3.9's target is a WIREBIND identity's own ~user VM, a
passive receiver nothing else drives. The existing idle-loop pump only
ticked VMs with the old FORTH MSG-TICK word ACL-allowed -- a VM minted
with the STD79-lockdown personality never has it, so the pump silently
skipped it forever and queued messages never delivered. Fixed by
adding an unconditional, direct sk_hermes_drain_checkpoint() call in
the same pump loop, independent of the MSG-TICK gate.

Verified live on all three architectures: WIREBIND-attach a real
identity, USE into it, type a plain console line, confirm the Stage D
evidence line and correct relayed execution result. Zero UNKNOWN WORD,
dict_hash identical across all three ISAs and matching task 3.8's own
baseline. The FABRIC-3.md-documented USE/BINDSTEP crash did not
reproduce in any of these live sessions (recorded as a finding, not
chased further).

Depends on the zuse_root_pubkey_known fix already landed in ac4d431 --
without it, WIREBIND cannot attach any identity on a fresh boot at
all, which blocked this task's own verification until found and fixed
separately.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-22 13:34:20 -04:00
Robert Allan JamesandClaude Sonnet 5 bbfd9103f4 Stage C: cut over BLK-ATTACH-EVENT alone -- FABRIC-3.6.md task 3.8
The reply leg (Artemis -> Hera ack) that used to flow through
common:messaging.4th's MSG-SEND/MSG-TICK now goes through
kernel-Hermes's sk_hermes_send_one()/sk_hermes_drain_checkpoint()
instead -- FORTH Hermes never sees a BLK-ATTACH-EVENT message again
(SXXXIV.2's partition rule). The request leg was never real FORTH
messaging traffic to begin with (a direct VM-EXEC, no type tag,
forced by Hera's own inability to load common:messaging.4th), so it
is untouched.

New KH-BLK-ATTACH-SEND (repl.c) wraps sk_hermes_send_one(), reached
from capsules/artemis/init.4th's HERA-BLK-ATTACH-REQ. Delivery reuses
task 3.4's already-wired sk_hermes_drain_checkpoint(); BLK-ATTACH-ACK
itself is unchanged, just reached by a different layer.
SK_HERMES_MSG_TYPE_BLK_ATTACH deliberately reuses BLK-ATTACH-EVENT's
own value (9) to document this as a cutover of the same message, not
a new one.

Two real bugs found on the way, both recorded in FABRIC-3.6.md's
findings log:
- A popped FORTH CREATE-buffer address was raw-cast to a host pointer
  instead of going through vm_ptr() -- silently read all-zero memory,
  no crash, no error, just a message that arrived and did nothing.
  Fixed; the rule and its exception (repl.c's own dev-addr is
  legitimately a raw pointer, formatted that way by its own pushing
  code) are written up for the next FORTH-facing C word.
- A separate, genuine hang on the very first live exercise of this
  path, never reproduced across ten subsequent boots. Reported, not
  chased -- not blocking, per the task's own check being otherwise
  fully satisfied.

Also found live: log_message() is invisible in this build's actual
serial-log capture at every level -- settled on a single
console_println in the real drain target instead, one line per real
USB attach, not a hot-path.

Final acceptance (logs/20260922-105501, -105758, -110304, disk images
reset before each): dict_hash identical across all three
architectures for every VM, zero UNKNOWN WORD, mkcapsule --lint
clean, real ledger+stadium_conserved(Artemis)=true evidence on every
boot.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-22 11:08:26 -04:00
Robert Allan JamesandClaude Sonnet 5 f37aa0fb17 Channel-open policy hook -- FABRIC-3.6.md task 3.7
Added HERMES-CHANNEL-OPEN? ( req-hi req-lo -- allow? ) at
capsules/ACL.4th block 4008 (default: approve everything) -- the one
word policy authors edit. sk_hermes_channel_open_policy(VM*, VMUuid)
(kernel_hermes.h/.c) is the C-side query that calls it via plain
word-dispatch against the target VM's own dictionary/stack, never
vm_interpret() (avoids task 3.4's input-buffer cursor hazard entirely)
and never decides the answer itself. Fails closed: no policy word,
a policy error, or stack underflow all deny, matching CLAUDE.md's
posture that absence of policy must never mean "always allow."

Two real bugs found and fixed before this was called done:
missing current_executing_entry assignment before calling the word's
func pointer (colon words silently no-op without it, vm_core.c:730 --
no crash, just a wrong answer); and a second FAIL with debug
instrumentation still in place whose precise cause isn't
reconstructable, since no intermediate commit exists for that attempt.

Self-test proves the task's check four ways against the same
unchanged C function: default approve, live redefinition to deny
(zero C change), restore, and a VM with no ACL.4th loaded at all
(fail closed). A fifth check wires the result into task 3.6's
sk_hermes_channel_respond() end to end: a denied policy produces a
NACK and no channel, ledger/stadium_conserved() holding throughout.

Scope, per Captain Bob's ruling: closes with the query built and
proven; sk_hermes_channel_respond() still takes a caller-supplied
approved bool rather than calling the policy internally. Wiring a
real channel-open call site to only this query is deferred to
whichever later task first needs a live decision.

dict_hash identical across amd64/aarch64/riscv64 for every VM, zero
UNKNOWN WORD, mkcapsule --lint clean (38 files, 0 violations).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-22 09:08:48 -04:00
Robert Allan JamesandClaude Sonnet 5 205a49ecd0 ACK/NACK and private-channel negotiation -- FABRIC-3.6.md task 3.6
Extracted sk_hermes_send_one() from sk_hermes_publish()'s own
per-subscriber body -- one code path for both point-to-point and
fan-out delivery, so the ledger can never diverge between them.
Point-to-point addressing turned out to be load-bearing, not
incidental: sk_hermes_publish()'s fan-out sets msg->to to whichever
member it is iterating, so a negotiation message "published" to the
common channel would spuriously reach every member, not just the real
target (checked with advisor() before building the naive version).
"Over the common channel" means every VM is reachable from birth (task
3.2), not that the exchange itself fans out -- messaging.4th's own
CH-REQUEST carried an explicit `to` for the same reason.

sk_hermes_channel_request/respond/close build the mechanics: request ->
grant (creates a private channel, subscribes both parties, sends
CH_GRANT + one ACK) or NACK ("a deny is a NACK", SXLV.1 -- no separate
type); close authorized by membership alone. The grant/deny decision is
a plain caller-supplied `approved` bool -- task 3.7 replaces the call
site that produces it with a real ACL.4th query, not this signature.

Self-test covers the task's own three checks plus a sibling advisor()
flagged: an approved respond() whose channel creation itself fails
(table exhausted) must still fall through to NACK, not a silent false
grant or half-open channel -- verified by exhausting the whole channel
table and confirming the fallback.

Bug found and fixed before this was called done: the first draft
dropped a message via pending_pop() alone, without releasing it first,
leaking its Stadium heat and failing the self-test's own ledger
baseline check (logs/20260922-065946/amd64/, kept as audit trail).
Fixed and re-verified PASS on all three architectures.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-22 07:25:51 -04:00
Robert Allan JamesandClaude Sonnet 5 8a2ee0fdad Payload bound and chunking -- FABRIC-3.6.md task 3.5
sk_hermes_publish() now enforces SK_HERMES_CHUNK_MAX_PAYLOAD (1024) on
every message's payload_len uniformly, chunked or not -- closing the
gap task 3.3 explicitly parked. A chunk carrier is
[SkHermesChunkHeader][content slice], slice capped at
1024 - sizeof(header) rather than 1024 itself, so every message on the
wire satisfies the same one-block bound vm_interpret()'s own drain
limit already requires -- a future chunk-aware drain never has to
special-case a carrier that can't be handed to vm_interpret() as-is.

Deliberately no chunking-sender API: building one would need
kernel-Hermes to own chunk-buffer memory with a real lifetime it has no
way to track (kept alive until every subscriber drains it). Sending is
a loop pattern a caller writes with sk_hermes_chunk_count() +
sk_hermes_publish(), demonstrated by this task's own self-test.
sk_hermes_reassemble() is pure and memory-agnostic: validates msg_id
agreement, exact seq coverage, and per-chunk slice sizes before a
single memcpy, with the total length computed once and checked against
the caller's buffer once -- never order-dependent on which chunk
happens to overflow.

Verified live on all three architectures: a 1024-byte payload as one
message, a 1025-byte send refused outright with the ledger untouched,
and a 3000-byte payload split into 3 chunks, drained, and reassembled
byte-exact against the original. dict_hash unmoved and identical
across architectures.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-22 06:51:48 -04:00
Robert Allan JamesandClaude Sonnet 5 2b1ba031a5 Drain at the outermost checkpoint -- FABRIC-3.6.md task 3.4
sk_hermes_drain_checkpoint() interprets one queued payload per checkpoint
(ruled: one message per checkpoint), reusing sk_vm_at_outermost_interpret()
and placed before the switch-signal block in vm_core.c's existing
cooperative checkpoint (sk_vm_context_switch() doesn't return until
switched back to, so drain must come first or it silently never runs on
a switching checkpoint).

Amends FABRIC-3.5.md SXLIII.5, caught by advisor() before writing the
naive version: "recursive drain is prevented for free" via
g_vm_interpret_depth is true but only for same-message re-drain -- it
doesn't cover the separate same-VM reentrancy hazard FABRIC-3.md SXX
already named for Hera specifically (VMCallState saves rsp/exit_colon/
ecw_nesting only, never input_buffer/input_length/input_pos). Draining
calls vm_interpret() on the same vm whose own vm_interpret() call is
still paused mid-word at the checkpoint; without saving and restoring
the cursor by hand, the enclosing REPL line or LOAD block would be
silently truncated. sk_hermes_drain_checkpoint() snapshots and restores
input_buffer/input_length/input_pos/mode/error/abort_requested around
the call. Not a divergence from the ruling -- cursor preservation is the
implementer's own obligation inside the ruled mechanism.

Gated behind a system-wide pending-total counter so the common
no-message-in-flight case costs one integer read per word dispatch, not
a stadium_max_vm_count()-sized queue scan (also flagged by advisor() as
a real hot-path cost, not deferred).

Verified live on all three architectures: a self-test publishes a real
payload to Hermes, proves the depth gate via VM-EXEC-ing an existing
harmless colon word into Hermes (genuine nested vm_interpret(), depth 2,
must not drain), then drains directly from genuinely-outermost context
and confirms exactly one clean drain. dict_hash unmoved and identical
across architectures.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-22 01:27:11 -04:00
Robert Allan JamesandClaude Sonnet 5 1f6343bc03 Publish path, no dispatch -- FABRIC-3.6.md task 3.3
sk_hermes_publish() allocates one SkHermesMessage per channel member
(heat-cost ruling 2026-09-21: one message per subscriber, funded by
the publisher's own reservoir) and enqueues each onto a new
per-subscriber SkHermesPendingQueue -- found-or-created lazily by
vm_id, sized from stadium_max_vm_count() like the channel/switch
tables. Best-effort across subscribers: a failed allocation or full
queue skips and rolls back just that one subscriber, not the whole
publish -- the natural reading of "ledger and stadium_conserved() hold
across N publishes to M subscribers" (the task's own check), not a
separate ruling.

Dispatches nothing -- sk_hermes_pending_count()/peek()/pop() are the
read/drain primitives task 3.4's real checkpoint-driven drain will
build on; this task's own self-test uses them directly since no
checkpoint hook exists yet.

Verified live on all three architectures: pending-queue table sized
50/202/50 slots (tracking the channel table's own per-arch sizing), a
synthetic publish self-test (2 publishes to 3 subscribers) confirms
exact per-subscriber delivery counts, and ledger/stadium_conserved()
invariants hold both mid-publish and after manually draining every
queue back to baseline.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-22 00:49:22 -04:00
Robert Allan JamesandClaude Sonnet 5 a4afdfa591 Channel table + common channel, inert (B1) -- FABRIC-3.6.md task 3.2
Adds SkHermesChannel: a channel is an index into a boot-time,
stadium_max_vm_count()-sized table (same sizing pattern task 3.1
established for the switch table -- no separate numeric rule was ruled
for this table, so task 3.1's bound is extended directly, flagged as
such rather than restated as a new ruling). No name field, mirroring
messaging.4th's own nameless CH-ARENA.

The common channel (index 0) is created at boot and permanent. Hera
subscribes explicitly in kernel_main.c (she is the one VM never born
through capsule_birth_baby()); every other VM -- Tripod fleet and
future WIREBIND identities alike -- subscribes inside
capsule_birth_baby() itself, the single choke point every other birth
already passes through.

Inert: no publish, no dispatch, no ACK/NACK, no ACL hook (tasks 3.3,
3.6, 3.7). Verified live on all three architectures: channel table
sized to 50/202/50 slots (matching switch-signal's own per-arch
sizing), common-channel fleet self-test confirms all four Tripod
members are members, and a synthetic create/subscribe/unsubscribe/
destroy round-trip against a private topic passes, including refusing
to destroy the common channel.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-22 00:25:13 -04:00
Robert Allan JamesandClaude Sonnet 5 feace42397 Add diagnostic scan cross-check of the Hermes counters -- FABRIC-3.6.md task 2.8
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-21 08:46:07 -04:00
Robert Allan JamesandClaude Sonnet 5 2a2bf6eb35 Stage B proof: add per-VM consumed term to stadium_conserved -- FABRIC-3.6.md task 2.7
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-20 21:31:45 -04:00
Robert Allan JamesandClaude Sonnet 5 313ffc89e6 Add exact-equality Hermes ledger self-audit -- FABRIC-3.6.md task 2.6
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-20 21:02:27 -04:00
Robert Allan JamesandClaude Sonnet 5 d11eb2e5db Add sk_hermes_decay(), ledgering decay into consumed -- FABRIC-3.6.md task 2.5
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-20 18:48:46 -04:00
Robert Allan JamesandClaude Sonnet 5 493d410028 Add the four ledger counters (held/pulled/returned/consumed) -- FABRIC-3.6.md task 2.4
FABRIC-3.5.md SXL.4's ledger: held == pulled - returned - consumed,
epsilon zero. Added sk_hermes_ledger() (an accessor, not a mutator)
plus four static counters in kernel_hermes.c.

Each counter has exactly one increment/decrement site: held/pulled
both move at sk_hermes_alloc()'s single success path, after every
refusal branch has already returned; held/returned both move at
sk_hermes_release()'s single success path. consumed is declared and
always reads 0 -- its one increment site doesn't exist yet, and won't
until task 2.5 gives decay something to record.

sk_hermes_release() now reads the Stadium cell's live header.heat
immediately before calling stadium_evict(), rather than assuming the
original pulled amount -- stadium_evict() zeroes the header as part of
freeing the cell and its own return value is a success code, not the
credited amount, so this is the only point the true remaining heat is
available. Today this always equals the original Q.SLOT pull; once
task 2.5's decay exists, this is what keeps returned correct without
touching this function again.

Self-test (kernel_main.c) extended: snapshots the ledger before
running so it checks its own deltas, verifies held/pulled grow by
exactly got_n * Q_SLOT on allocation with returned/consumed untouched,
then verifies held returns to its starting value and returned grows by
the same amount on release, and checks the audit invariant itself as a
bonus (task 2.6 formalizes this properly).

Three-arch boot clean: amd64/aarch64/riscv64 all reach [zuse@Hera] ok>,
zero UNKNOWN WORD, dict_hash unmoved. All three print PASS with
identical final ledger: held=0 pulled=65536 returned=65536 consumed=0.
No compiler warnings.

Authorized by Captain Bob ("keep going with rhe 6.5 document").

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-20 05:59:09 -04:00
Robert Allan JamesandClaude Sonnet 5 2c1dc1753a Correct sk_hermes_alloc() to admit a real Stadium patron; add sk_hermes_release() -- FABRIC-3.6.md tasks 2.2 (amended) + 2.3
Real finding, caught before building release on a foundation that
couldn't support it: task 2.2's first cut of sk_hermes_alloc() pulled
reservoir heat but never admitted a real Stadium-floor patron -- just a
local in_use flag. FABRIC-3.5.md SXXXIII.4 item 1 says MSG-FREE-NODE
returns heat "via STADIUM-EVICT", which only means something if
allocation admitted something. SXL.4's own invariant, Sigma(resident
patron heat) + reservoir + consumed == Q48_ONE, cannot balance if held
heat is invisible to every term while held. Flagged to Captain Bob
before proceeding; authorized to correct 2.2 in the same pass as
building 2.3 on top of the fix.

sk_hermes_alloc() now calls stadium_admit() with heat = the pulled
amount, behaviour = STADIUM_BEHAVIOUR_DELIVER (matching messaging.4th's
own SB-DELIVER STADIUM-ADMIT exactly), and identity = the message's own
slot index (matching the FORTH precedent -- caught live in the first
boot of this fix that omitting this made stadium_dispatch()'s existing
DELIVER diagnostic print msg_idx=0 for every message instead of a
distinct value). The returned cell index is stored in the message's
own stadium_cell field. Stadium-floor refusal (independent of reservoir
affordability) rolls back the pull the same way the other refusal
paths already do.

sk_hermes_release() -- the function task 2.3 actually asks for -- calls
stadium_evict() on that cell, which itself returns the departing
patron's remaining heat to its owning VM's reservoir, matching
MSG-FREE-NODE's exact shape. Release does not touch the reservoir
directly.

Self-test (kernel_main.c) extended: keeps every allocated message's
pointer, allocates to exhaustion as before, releases all of them, and
checks the reservoir returns to precisely its starting value.
"Undecayed" is true by construction (no decay/TTL logic exists yet,
task 2.5) -- exact restoration, not approximate.

Three-arch boot clean: amd64/aarch64/riscv64 all reach [zuse@Hera] ok>,
zero UNKNOWN WORD, dict_hash unmoved from task 2.2. All three print
identical PASS arithmetic: reservoir0=65536, reservoir_after_alloc=0,
reservoir_final=65536. No compiler warnings.

stadium_dispatch()'s DELIVER-case console output (one line per
eviction) is pre-existing instrumentation, not new -- confirmed real
and load-bearing per stadium.c's own comment, verbose but expected.

Authorized by Captain Bob ("yes").

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-20 04:20:14 -04:00
Robert Allan JamesandClaude Sonnet 5 9d129fdb1c Add sk_hermes_alloc(), the heat-coupled allocator -- FABRIC-3.6.md task 2.2, item 28
The piece FABRIC-3.5.md SXXXIII.6 calls "what remains genuinely hard,"
built and proven first per its own recommendation. Added
src/starkernel/vm/kernel_hermes.c (wired into Makefile.starkernel's
LOADER_EXTRA_SRCS -- this repo lists vm/*.c files explicitly, no glob)
and sk_hermes_alloc()'s declaration in kernel_hermes.h.

Checks stadium_reservoir_peek(vm_id) >= SK_HERMES_Q_SLOT before
touching the reservoir at all -- refusal this way needs no rollback,
since nothing was pulled -- with an explicit rollback path
(stadium_reservoir_push) kept defensively for the pull-then-short case,
though nothing in this single-core kernel is expected to reach it.

SK_HERMES_Q_SLOT = Q48_ONE / SK_HERMES_MSG_MAX (2048), deliberately
simpler than messaging.4th's own formula, which reserves a Q.1/3 floor
for COMMON-CH's own Stadium heat -- kernel-Hermes has no such object
(SXXXIII.4/SXXXIII.5's flat membership list carries no heat of its
own), so there is nothing left for that floor to protect.

Self-test in kernel_main.c, same diagnostic-only synthetic-VM pattern
as the existing Stadium quota grant self-test (lo=3, distinct from
that test's lo=1): reads back the actual granted reservoir rather than
assuming a number, derives expected_n from it, allocates to refusal,
and checks the refusal lands at exactly expected_n, the reservoir
doesn't move on the refused attempt (rollback proven, not assumed),
and the final reservoir is exactly reservoir0 minus got_n times
Q_SLOT.

Three-arch boot clean: amd64/aarch64/riscv64 all reach [zuse@Hera] ok>,
zero UNKNOWN WORD, dict_hash unmoved from task 2.1 (pure C, no FORTH
touched). All three print identical self-test arithmetic: reservoir0=
65536 Q_SLOT=2048 expected_n=32 got_n=32 reservoir_after=0. No compiler
warnings.

Noted, not fixed: Q_SLOT's divisor and SK_HERMES_MSG_MAX are the same
32, so reservoir and arena exhaustion land at exactly the same count by
construction -- this test can't distinguish which refusal reason
fired, only that refusal is correct and rolls back correctly.

Deliberately not evidence for stadium_conserved(): allocating alone
(no release yet, task 2.3) leaves pulled heat held off the Stadium
floor, so the two-term check would correctly read false right now if
run mid-hold. That's expected, not a bug -- Stage B (task 2.7) is
defined as "before and after the alloc/free cycle," not "continuously
during." This task's self-test checks reservoir arithmetic directly
instead.

Authorized by Captain Bob ("Yes continue").

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-19 21:00:20 -04:00
Robert Allan JamesandClaude Sonnet 5 f10fa7ae83 Add kernel-Hermes message/membership structures -- FABRIC-3.6.md task 2.1, Phase 2 begins
Phase 2, task 2.1 only: type definitions, wired to nothing, drawing no
heat -- no allocator, no protocol logic, no registration anywhere.
FABRIC-3.5.md SXXII.4: Phase 2 structures come first and prove nothing
until the allocator is built on top (task 2.2 onward, each its own
commit).

Added include/starkernel/vm/kernel_hermes.h:

SkHermesMessage -- field-for-field mirror of messaging.4th's live
9-cell MSG-* layout (type/from/to/payload addr+len/Stadium cell
index/seq/channel/orig-type), per SXXXIII.4 item 1 ("roughly half the
file is accessors that become struct fields"), plus an explicit
in_use flag for task 2.2's allocator. Deliberately no separate heat
field: per SXL.4, a message's heat IS the Stadium cell it occupies,
not a value copied alongside it -- one source of truth for the
conservation invariant stadium_conserved() (task 0.7) checks.

SkHermesMembership -- one flat broadcast membership list, SXXXIII.4/
SXXXIII.5's recommended replacement for messaging.4th's 28-word channel
abstraction (traced to exactly one live caller, CH-ADD-MBR). Item 27
(negotiation vs. broadcast, Phase 3 blocker B1) is not answered by this
structure and isn't meant to be -- a flat list is correct either way.

Genuinely wired to nothing: no .c file, no Makefile change, no include
from any compiled source. Syntax-checked standalone (gcc -std=c99
-Wall -Wextra -Werror -fsyntax-only) before touching the real build.

Boot byte-identical to task 1.9's baseline on amd64 (same dict_hash
triple, zero UNKNOWN WORD). Did not repeat aarch64/riscv64 -- the file
compiles into no object on any architecture, so there is no mechanism
by which it could diverge.

Authorized by Captain Bob ("yes").

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-19 20:48:24 -04:00
Robert Allan JamesandClaude Sonnet 5 380f0a09c9 Add stadium_conserved() -- FABRIC-3.6.md task 0.7, item 41
Boolean analogue of vm_physics_conserved(), for the Stadium per-VM
quota invariant rather than fleet-wide execution heat
(FABRIC-3.5.md SXXXIX.4). int stadium_conserved(VMUuid vm_id), in
src/starkernel/vm/stadium.c alongside stadium_resident_sum()/
stadium_reservoir_peek() that it's built from, declared in
include/starkernel/vm/stadium.h.

Implements the two-term form -- resident_sum(vm_id) +
reservoir_peek(vm_id) == Q48_ONE -- not the three-term form SXL.4
rules for the eventual system. That ruling's `consumed` term is a
Phase 2 kernel-Hermes ledger deliverable that doesn't exist yet:
nothing draws on any VM's Stadium quota today (task 2.2 is literally
where that wiring gets built), so consumed is honestly zero right now.
Folding it in as a placeholder would be inventing Phase 2 state ahead
of it existing -- the doc comment says so explicitly, so whoever
builds Phase 2's ledger extends this function rather than working
around it.

Wired into the existing per-VM boot diagnostic
(stadium_words_print_boot_diagnostics(), kernel_main.c:810, Hera
only -- the sole existing call site) rather than adding a new one,
printing CONSERVED/DRIFTED the same shape vm_physics_status() already
uses.

Three-arch boot clean: amd64/aarch64/riscv64 all reach [zuse@Hera] ok>,
zero UNKNOWN WORD, all print "Stadium conservation: CONSERVED" with
identical resident_sum=47641 reservoir=17895 sum=65536=Q48_ONE. No
compiler warnings on either edited file (forced recompile checked).

Authorized by Captain Bob ("Yes.").

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-19 13:56:14 -04:00
Robert Allan JamesandClaude Sonnet 5 66beae7fd4 Stage 3 follow-on: message-arrival eligibility hook + trampoline-blind switch-storm fix (FABRIC-3.md §XXVIII.2)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
Implements the message-arrival eligibility signal FABRIC-3.md §XXVIII.1 left
open (has_work per-slot flag, set via new SWITCH-MARK-WORK primitive from
MSG-SEND) so an idle VM never becomes a switch target purely by waiting out
the readiness threshold.

Also root-causes and fixes a second, independent switch-storm: the tick's
"who is current" check used vm_log_attributed_vm(), which can't see a VM
parked in switch.c's own raw trampoline. Replaced with a dedicated
g_switch_current_vm tracked by the switch mechanism itself, and moved
target-slot eligibility reset to the switch decision point instead of
relying on ISR polling to observe a window that can be only a few
instructions wide.

Verified live on all 3 architectures: clean boot to zuse@Hera] ok>, live
cross-VM message dispatch, and (since a quiet log looks identical to a
livelocked storm once the DoE probe is gone) confirmed genuine REPL
liveness via QMP send-key + screendump on aarch64/riscv64, not log
inspection alone.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K
2026-09-14 17:39:20 -04:00
Robert Allan JamesandClaude Sonnet 5 f790d0995e Stage 2: cooperative VM context switch primitive, proven on all 3 arches (FABRIC-3.md §XXVIII)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
Third stage of the preemptive context-switching plan. The real
save/restore switch mechanism now exists -- the first time anything has
ever executed on a VM's own native stack (Stage 1 allocated them,
unused).

New sk_vm_switch_to() (switch.S, one per arch) is an ordinary function
call, not an interrupt -- so unlike Stage 0's trap frame, the ABI already
covers every caller-saved register; only the callee-saved set needs
explicit save/restore (amd64: rbx/rbp/r12-r15, no FP at all since SysV
has no callee-saved XMM; aarch64: x19-x28/x29/x30 + d8-d15; riscv64:
s0-s11/ra + fs0-fs11, FS-gated like Stage 0 but read once and reused for
both halves within one call, since FS is genuine global CPU state, not
part of what's switched). A sibling sk_vm_switch_prime() in the same file
builds the synthetic first-entry frame, kept in assembly so the layout
can never drift out of sync with sk_vm_switch_to() itself.

New switch.c/switch.h: sk_vm_context_switch(from, to) handles first-entry
priming vs. resuming a parked context, and updates registry state (new
VM_STATE_SWITCHED_OUT, distinct from VM_STATE_STOPPED -- STOPPED means no
live frame, this means the opposite). sk_vm_switch_entry() is the minimal
permanent trampoline every freshly-entered VM lands in: no production
behavior defined yet, so it just yields straight back to whoever switched
to it, forever.

Closes the confirmed unguarded-KILL UAF found during planning:
capsule_vm_kill(), mama_word_kill(), and capsule_vm_kill_all_nonmama()
all now refuse (or silently leak rather than free, on the cold-restart
path where arch_cold_reset() wipes everything immediately after anyway)
tearing down a switched-out VM. Side effect found, not built on purpose:
the existing MSG-TICK idle-pump already filters on VM_STATE_LIVE, so it
automatically stopped dispatching into a switched-out VM with zero
changes needed there.

Verified via a temporary SWITCH-TEST probe (boot-triggered, since nothing
can type interactively into a foreground-only QEMU session) that
round-tripped a sentinel through 5 real Hera<->Hermes switches on all 3
architectures: 5/5 rounds, 0 failures, clean continuation to ok>. Probe
fully reverted after capture; kernel_main.c shows zero diff.

Also: Makefile.starkernel's LOADER_EXTRA_SRCS/LOADER_ASM needed the new
files added explicitly (this project's "loader" PE binary is the full
running kernel, not a thin bootstrap stage), and aarch64's switch.S
needed the same #ifndef _WIN32 guard around .hidden that isr.S already
carries (aarch64's loader assembles via clang targeting a PE/COFF
target with no .hidden equivalent) -- caught by a build failure, fixed.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016UNhH1mhi52i6Qihh7ZV5S
2026-09-13 15:25:39 -04:00
Robert Allan JamesandClaude Sonnet 5 57ac3fc304 Stage 1: per-VM native stacks, allocated but not yet executed on (FABRIC-3.md §XXVIII)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
Second stage of the preemptive context-switching plan. Every VM (Hera,
every capsule_birth_baby()-born VM including WIREBIND identities) now
gets its own dedicated 2 MiB native C stack at birth -- but nothing runs
on it yet, that's Stage 2. Pure allocation-machinery proof.

Design correction made before writing code: the plan called for cloning
sk_vm_arena_alloc()'s guard-page pattern, but that pattern turns out to
be Mama-only -- host_services.c's kernel_alloc() gives every baby VM a
plain kmalloc() block for its dictionary arena, not a real guarded PMM
allocation. Stacks get the real treatment instead (new
sk_vm_native_stack_alloc()/_free() in arena.c): independent
pmm_alloc_contiguous() + guard pages for every VM without exception, no
singleton, no kmalloc fallback -- a stack overflow is exactly the
failure mode guard pages exist for, and a corrupted stack could corrupt
whatever saved context Stage 2 trusts.

2 MiB size matches this project's own established kernel-stack
convention (g_kernel_stack/g_rpi5_native_stack), not a guess -- that one
shared 2 MiB stack today already carries all VMs' combined nested
VM-EXEC recursion.

Three new VM struct fields, freed in vm_cleanup() alongside the existing
call_stack free. Allocation failure is non-fatal to birth.

All 3 architectures re-verified clean boot to ok>, no native-stack
allocation failures for any Tripod-fleet VM.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016UNhH1mhi52i6Qihh7ZV5S
2026-09-13 14:30:46 -04:00
Robert Allan JamesandClaude Sonnet 5 e51a8d229e Fix stadium_grant_quota() donor floor; rerun std79 DoE clean, 81/81 (FABRIC-3.md §XVII)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
capsule_birth.c hardcoded every new VM's initial Stadium quota grant to split
from Hera specifically. Since a grant always halves whatever the donor
currently has, Hera's own free list converges toward empty after a bounded
number of grants — independent of whether the Stadium as a whole still had
spare capacity, since VMs she'd granted to earlier typically still held
nearly all of their own share untouched. Past that point every subsequent
VM birth's Stadium grant would be silently refused (soft-failed, non-fatal
by existing design), even with plenty of capacity sitting idle elsewhere.

Fixed by adding an O(1)-maintained free_count to StadiumVMQuota (incremented
in stadium_evict(), decremented at both of stadium_admit()'s free-list-pop
sites, set/adjusted in stadium_grant_quota()'s own split — this also let
grant_quota drop its old O(free-list length) counting walk in favor of an
O(1) read) and stadium_best_donor(), an O(live VM count) scan over quota
slots returning whichever in-use VM currently has the most free cells.
capsule_birth.c's birth path now splits from that VM instead of
unconditionally vm_uuid_hera().

Verified with another full rerun of the 3x9x3 std79 DoE campaign from
scratch — same discipline as the prior Stadium fix (any defect repair
reruns the whole DoE from the top) — one continuous boot per architecture,
all 9 identities simultaneously live throughout. 81/81 trials correct, 0
mismatches, DOE-RUN header sequence md5-identical to every prior run.
aarch64 ~280s total (vs ~290s for the O(ncells)-scan fix alone — confirms
no regression). Both known Stadium defects are now closed together on one
clean campaign rerun.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo
2026-09-11 17:20:07 -04:00
Robert Allan JamesandClaude Sonnet 5 b031b802e3 Rename FABRIC series: FABRIC.md->0, FABRIC-2.md->1, FABRIC-3.md->2, FABRIC-4.md unchanged
FABRIC.md -> FABRIC-0.md
FABRIC-2.md -> FABRIC-1.md
FABRIC-3.md -> FABRIC-2.md (the current/living document)
FABRIC-4.md unchanged (new #3 to follow separately)

Every cross-reference repo-wide updated to match, including doc-comment
citations inside kernel source (.c/.h) files -- done via an ordered
placeholder substitution (FABRIC-3.md->placeholder2, FABRIC-2.md->
placeholder1, FABRIC.md->placeholder0, then placeholders resolved to
final names) in a single pass per file to avoid double-shifting
already-renamed references.

One line in capsules/font.4th grew past the 64-char block-format limit
as a side effect of the longer filename; shortened it and reverified
with mkcapsule --lint (34/34 pass) before rebuilding.

Verified 3-arch boot to ok> (amd64/aarch64/riscv64, each in the
foreground) after the fix; logs and DoE CSVs from this session's
verification runs included per this repo's own audit-artifact
convention.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019YcT3H2PQeyujrzjqS3Var
2026-09-04 11:22:51 -04:00
Robert Allan JamesandClaude Opus 5 a621131ef6 §H.12 step 3: session_set_pinned/session_is_pinned pin-authority choke point
session_is_pinned() reads Session.pinned directly (authoritative, no
Stadium re-derivation); session_set_pinned() writes both Session.pinned
and the mirrored STADIUM_FLAG_PIN bit on the session's own patron cell,
keeping Stadium's internal eviction/admission logic (which must stay
self-contained) in sync without it calling back into session.c.

Added Session.stadium_cell (index into stadium_cells()) -- necessary
plumbing not in the original H.2 field list; the choke point can't reach
the right patron header without it. Moved STADIUM_FLAG_PIN from a
stadium.c-private #define to stadium.h (public) so session.c can
reference it without a duplicate definition.

Verified 3-arch boot to ok> (amd64/aarch64/riscv64).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-03 06:10:11 -04:00
Robert Allan JamesandClaude Sonnet 5 c7c9332321 Stadium: real block-patron admission + MIGRATE dispatch (FABRIC-3.md §B)
stadium_admit()'s mass==1 refusal looked like a hard blocker for 1024-byte
blocks, but stadium_word_dispatch()'s real candidate construction proves
Stadium cells carry pure identity/heat/bookkeeping, never the resident's
actual content -- a block patron follows the same shape (identity=LBN,
payload unused), so this was real, scoped work, not a case for stubbing.

New stadium_blocks.h/.c mirror stadium_words.c's admission/cooling shape,
keyed by (quota_slot, lbn) in a fixed-capacity open-addressing hash table
(tombstone deletion) instead of a dense array, since LBN space isn't
densely bounded like word_id. Wired into block_word_block()/buffer()/
update() (block_words.c), __STARKERNEL__-guarded. stadium_dispatch()'s
MIGRATE case now calls blk_flush(lbn) for real instead of printing
"(stub)". Three new Kconfig constants (STADIUM_BLOCK_HEAT_QUANTUM/
STADIUM_BLOCK_COOL_RATE_Q48/STADIUM_BLOCK_TRACK_CAP_MULT) mirror the
word-patron ones, same three-layer wiring.

VM-COOL/DELIVER/EXPIRE stay explicit punch-list items -- VM-COOL
deferred pending the still-iterating Tripod/Zuse/messaging vision,
DELIVER/EXPIRE are their own future subsystem integrations per
FABRIC.md's own "open, not resolved" notes.

Verified clean compile (zero warnings) and clean boot to REPL with
conservation intact (resident_sum + reservoir == Q48_ONE) on all three
architectures (amd64/aarch64/riscv64); BLOCK/BUFFER touches exercised
live from the REPL with no crash; a 22,000-distinct-block flood loop
against an artificially shrunk Stadium ran clean under heavy admission
load. A live MIGRATE console fire was not directly observed this
session (root-caused to a pre-existing reservoir-floor/density-eviction
interaction unrelated to this change, documented in FABRIC-3.md) --
flagged as an honest follow-up, not silently claimed.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CXjAPTEKrgY2Mrk25KoLDn
2026-08-25 23:28:59 -04:00
Robert Allan JamesandClaude Sonnet 5 89d8c08582 stadium: wire STADIUM_CAPACITY_TICK in as a flat threshold, not a scheduler
Closes FABRIC-2.md's last open §12 Q5 question. fleet_heartbeat_tick_count
is fed by every live VM's own vm_tick(), not one VM's, so it was reaching
HEARTBEAT_INFERENCE_FREQUENCY (shared/borrowed from the per-VM inference
gate) several times faster than intended with more than one VM live -
backwards from FABRIC.md §22.4's required ~1000:1 separation.

What's actually gated turned out to be low-stakes: vm_physics_tick()
(capsule_vm_physics.c:397) is a passive statistics refit - re-sorts a
window of past heat-transfer samples and recomputes a median rate
estimate. It doesn't move heat or arbitrate capacity. Firing too often
just meant a noisier statistic recomputed more frequently than planned,
not incorrect behavior.

Considered and explicitly rejected: scaling the threshold by live VM
count at the check site. That's the first brick of a scheduler - reading
fleet state to adjust a rate dynamically - which this project has
deliberately avoided building. Implemented instead: STADIUM_CAPACITY_TICK
(existing Kconfig symbol, defined but never read by any code path) now
gates vm_physics_heartbeat_tick()'s call directly, replacing the borrowed
HEARTBEAT_INFERENCE_FREQUENCY. Default bumped 1000 -> 4000, a flat
constant picked once for Tripod's known 4-VM topology, same kind of
placeholder as every other frequency knob in Kconfig.kernel - not
computed from anything at runtime. Renamed fleet_last_inference_tick ->
fleet_last_capacity_tick to match. Still one clock, one counter
(fleet_heartbeat_tick_count) - just a bigger flat divisor on it.

Three-arch QEMU acceptance: all clean to ok>, identical Stadium
conservation invariant on all three (resident_sum=43691 reservoir=21845
sum=65536). logs/20260815-093425/amd64, logs/20260815-093521/aarch64,
logs/20260815-093641/riscv64.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-15 09:37:58 -04:00
Robert Allan JamesandClaude Sonnet 5 00e657019e stadium: make VM population bound RAM-derived, not a static array of 4
Replaces STADIUM_MAX_VM_COUNT (Kconfig, hardcoded default 4) with a
boot-time computation, mirroring the pattern stadium_boot_init() already
used for the cell pool. New Kconfig STADIUM_VM_MEMORY_PERCENT (default
50): max_vm_count = (kmalloc_get_stats().free_bytes after the cell array
* STADIUM_VM_MEMORY_PERCENT / 100) / VM_MEMORY_SIZE, floored to 1, no
ceiling (population is not knowable in advance - could be 4, could be
4000). stadium_quotas and word_slots (plus stat_promotions/stat_evictions)
are now kmalloc'd to the computed count instead of declared with a macro.
New accessor stadium_max_vm_count() replaces every STADIUM_MAX_VM_COUNT
reference, including capsule_birth.c's birth-refusal gate.

Two things found and fixed along the way:

- The existing cell-pool budget was sourced from pmm_get_stats(), which
  reflects physical pages PMM hasn't handed to any subsystem yet - but
  the actual allocation is kmalloc(), which draws from the separate,
  fixed-size heap kmalloc_init() (M6) already carved out of PMM before
  stadium_boot_init() ever runs. Budgeting against PMM's leftover and
  allocating from the kmalloc heap are two different pools. Both the
  cell budget and the new VM-count budget now source from
  kmalloc_get_stats() instead.

- stadium_owner[] (which VM's quota owns each cell) was uint8_t, capped
  at 255 slots by a compile-time assert tied to the old macro. Widened
  to uint16_t (65535 slots of headroom) with a runtime clamp + log if
  the computed count ever exceeds that, since there's no ceiling anymore.

Three-arch QEMU acceptance: all clean to ok>, computed VM count genuinely
differs by actual available RAM (amd64/riscv64: 50 slots at -m 1024,
aarch64: 101 slots), Stadium conservation invariant identical across all
three (resident_sum=43691 reservoir=21845 sum=65536).
logs/20260815-080526/amd64, logs/20260815-080826/aarch64,
logs/20260815-080952/riscv64.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-15 08:11:21 -04:00
Robert Allan JamesandClaude Sonnet 5 5a28458b21 starkernel: item 4.2 -- Hermes native on the Stadium (complete)
Migrates Hermes's message/channel lifecycle onto the Stadium's unified
heat/capacity economy: MSG-ALLOC/FREE-NODE and CH-ALLOC/FREE-NODE now
route entirely through stadium_admit()/stadium_evict(), replacing the
old local free-list + independent heat-field mechanism. Eight
kernel-only STADIUM-* FORTH primitives (ADMIT, EVICT, RES@, RES-PULL,
RES-PUSH, HEAT@, HEAT!, WORD-HEAT), VM.stadium_vm_id threaded through
all three vm_core.c dispatch sites (replacing item 4.1's hardcoded
vm_uuid_hera()), and the stadium_owner[idx] fix so evict-credit lands
in the VM that actually admitted a patron, not whoever owned cell 0.

This session's own contribution, on top of that pre-existing
implementation: found and fixed two bugs blocking the item's own K≡1.0
conservation self-check (HERMES-K was reading 0, not 65536):

- Q.SLOT admission-heat fix (capsules/hermes/init.4th): MSG-SEND/
  CH-ACCEPT admitted with Q.1 (the entire fleet-wide "1.0" unit) per
  item, a leftover from before the Stadium migration when each
  message/channel had its own unconstrained heat field. Instantly
  drained the shared, finite reservoir.

- Reservoir floor for word-execution admission (stadium_words.c):
  stadium_word_dispatch() (item 4.1) pulls STADIUM_WORD_HEAT_QUANTUM on
  every word dispatch, not just first admission -- exhausts a VM's
  entire reservoir in ~32 dispatches, starving any application-level
  economy sharing that VM's reservoir before it gets a chance to pull
  anything. word_dispatch_pull() now clamps word-execution's own pulls
  to leave a Q48_ONE/3 floor (same fair-share figure COMMON-CH's own
  floor already uses); application-level pulls are unaffected.

- STADIUM-WORD-HEAT primitive + stadium_words_resident_heat(): the
  floor deliberately leaves word-execution residents holding real
  heat, invisible to HERMES-K's original formula (MSG+CH+reservoir,
  no term for word patrons). Adding this term closes K to exactly
  65536 on all three architectures.

Also rules on two open scope questions in FABRIC.md: MBR-ALLOC/
MBR-FREE-NODE stay off the Stadium (membership records have no heat
field, never did -- the acceptance bullet's inclusion of them was a
completeness gesture predating a check of the actual layout), and
records the effort number (12 implementation files, +759/-120 lines).

Verified: all three architectures boot clean, full self-test passes,
Stadium conservation closes exactly (resident_sum + reservoir =
Q48_ONE) at both the C/Stadium level and the FORTH-level HERMES-K
check.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-07 01:49:23 -04:00
Robert Allan JamesandClaude Sonnet 5 2981ada2a5 starkernel: item 4.1a -- quota granting, Hermes's one-time birth grant
Punch list §25 item 4.1a complete.
New prerequisite item, found while scoping 4.2: no quota-granting mechanism
existed at all. Adds stadium_grant_quota(new_vm_id, from_vm_id) -- a
one-time initial grant at birth, distinct from item 1.3's still-unbuilt
recurring capacity-transfer arbitration. Splits the donor's free list evenly
by cell count, reassigns stadium_owner[] for every moved cell, and grants
the new VM a fresh Q48_ONE reservoir (not a split of the donor's -- per-VM
conservation, same pattern as Hera's own boot grant). Wired into every baby
VM's birth in capsule_birth.c.

Verified via a boot-time self-test in kernel_main.c using a synthetic
identity (not the real UUID pool, not a real capsule birth -- item 0.1's
Hera-alone pruning stays intact). All three architectures booted to ok> with
identical output: grant OK, Hera reservoir=0 (already fully committed to
resident words, correctly unchanged), test-vm reservoir=65536 (fresh
Q48_ONE). dict_hash identical across all three and unchanged from item 4.1's
baseline (0x3d4e1daf289da94f) -- confirms no dictionary word was added.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-05 15:12:30 -04:00
Robert Allan JamesandClaude Sonnet 5 3d0b9351bd starkernel: item 4.1 -- hot words onto the Stadium, density-ranked eviction
Punch list §25 item 4.1 complete.
Replaces the round-robin hotwords cache with Stadium density-ranked
admission/eviction on the kernel side, via the §17.7 reservoir mechanism and a
kernel-side word_id -> cell_index map (no DictEntry change, dict_hash
untouched). Adds stadium_birth_hera() to close the cell-0 panic hazard,
STADIUM_WORD_HEAT_QUANTUM/STADIUM_WORD_COOL_RATE_Q48 Kconfig knobs (flagged
untuned), and a stadium_word_forget() FORGET coherence hook to close a
recycled-word_id aliasing gap.

Verified: all five hotwords_cache_* call sites in dictionary_management.c
bypassed under __STARKERNEL__; word dispatch feeds the Stadium at all three
vm_core.c physics_execution_heat_increment() sites; hosted make unaffected;
all three architectures booted to ok> with matching dict_hash
(0x3d4e1daf289da94f) and matching conservation stats (promotions=354
evictions=0, resident_sum=65536 reservoir=0 sum=65536).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-05 13:37:10 -04:00
Robert Allan JamesandClaude Sonnet 5 9b305a5be7 starkernel: item 3.8 -- VM identifiers as UUID/GUID
Punch list §25 item 3.8 complete. Added after starting item 4.1
surfaced the need to thread a vm_id into stadium_admit()'s new quota
parameter; Captain Bob ruled UUID/GUID rather than keeping the
narrower uint32_t.

New VMUuid type (vm_uuid.h/vm_uuid.c): two uint64_t halves, RFC-4122-
shaped for logging. Not real randomness -- checked directly against
QEMU 10.2.1's actual CPU feature set: amd64 RDRAND and riscv64 Zkr are
both real, available features here; aarch64 has no RNG property on any
CPU model including "max" (verified exhaustively via QMP
query-cpu-model-expansion). Captain Bob ruled a uniform fallback
across all three ISAs rather than a per-architecture split.

Fallback is a deterministic PRNG (splitmix64) seeded from the Mama
capsule's content hash, pre-filling a 16-entry FIFO pool at boot and
refilling with another batch of the same stream when exhausted --
exactly the shape requested. Same capsule booted twice produces the
same id sequence, preserving the dict_hash reproducibility this
session has relied on throughout.

Hera keeps a fixed, reserved all-zero id, not drawn from the pool --
capsule_birth.c uses vm_id == 0 as a load-bearing sentinel in three
places (KILL protection x2, fleet heat-fanout parent-chain
terminator), found by reading before writing any code.

Two real sentinel-collision bugs caught before shipping, same class as
STADIUM_CONTAINS_NONE: vm_uuid_none() (all-ones, not all-zero) for
"not yet assigned"/"no VM" placeholders; confirmed item 3.7's quota
table already used an in_use boolean rather than a vm_id sentinel, so
no second collision was actually possible there -- the dead,
never-referenced STADIUM_QUOTA_SLOT_EMPTY macro was removed.

Blast radius larger than first scoped, flagged mid-work rather than
silently absorbed: capsule_vm_physics.c/.h (the fleet heat-transfer
layer item 2.1 modified earlier this session) has its own vm_id-keyed
node table and walks parent_vm_id chains through the same identity
space, so it needed the same change, plus its callers in
mama_forth_words.c and sk_vm_bootstrap.c.

One live FORTH word contract changed, by explicit ruling: CAPSULE-BIRTH
was ( capsule-id -- vm-id ), a single cell -- can't hold 128 bits.
Captain Bob picked pushing two cells ("there is doubles support in the
FORTH std word set anyway"): ( capsule-id -- vm-id-hi vm-id-lo ).
MAMA-VM-ID changed the same way: ( -- 0 0 ).

Verified: full (not standalone-file) kernel rebuild to catch cross-file
breakage given the size of this change -- it surfaced the
capsule_vm_physics.c blast radius a narrower check would have missed.
Three-architecture boot (amd64, aarch64, riscv64), all reaching ok>
with identical dict_hash=0x3d4e1daf289da94f matching the item-3.7
baseline.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-04 19:50:34 -04:00
Robert Allan JamesandClaude Sonnet 5 e55111c2c5 starkernel: item 3.7 -- per-VM free lists (Phase 3 core complete, for real)
Punch list §25 item 3.7 complete. Added to §25.4 after starting item
4.1 surfaced it as an unbuilt prerequisite -- 3.6's earlier "Phase 3
core complete" claim is corrected in this same commit.

StadiumVMQuota table (size STADIUM_MAX_VM_COUNT, linearly searched by
vm_id -- capsule_birth.c's vm_id is monotonic and never reused, so it
cannot index a table directly, and a 4-entry scan costs nothing). New
per-cell stadium_owner byte array records which quota a cell belongs
to, needed so eviction returns a freed cell to the correct VM's list
and so eviction search stays scoped to the evicting VM's own residents
(quota isolation).

Free-list linkage reuses each cell's `link` field as a next-free
pointer while unresident -- link is documented only as generic "index
into the Stadium, not a pointer," so this is a repurposing, not a
header change. Does not answer the separate, still-open question of
which field carries a multi-cell patron's first continuation-cell
index; item 3.5's mass != 1 refusal stands exactly as it was.

Boot-time: every cell chained into one list in ascending index order,
granted whole to vm_id 0 (Hera), the only VM that exists. Ascending
order preserves item 3.6's "Hera is patron zero" invariant once real
birth-wiring lands.

stadium_admit()'s signature changed to take vm_id -- a change to code
shipped in item 3.5, amended there. Pops the calling VM's free-list
head first (O(1)); only falls back to a same-VM-scoped eviction search
if empty.

Caught a real bug before the boot run: the header zero-fill on
eviction (and the initial free-list build) both left contains == 0,
but 0 is Hera's valid index -- the same collision item 3.1's
STADIUM_CONTAINS_NONE fix addressed, recurring at a new site. Fixed by
explicitly setting contains = STADIUM_CONTAINS_NONE at both free-list
sites.

Explicitly out of scope, reported not invented: granting quota to any
VM other than Hera is capacity arbitration (item 1.3 left "how much
moves per transfer" open). stadium_owner is set once at boot and never
rewritten, so quota_slot_for_vm() refuses every vm_id != 0 permanently
until item 4.2 adds the grant path and owner-array writes.

Verified: three-architecture boot (amd64, aarch64, riscv64), all
reaching ok> with identical dict_hash=0x3d4e1daf289da94f matching the
item-3.6 baseline.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-04 18:09:34 -04:00
Robert Allan JamesandClaude Sonnet 5 72487e7fff starkernel: item 3.6 -- Hera as patron zero, pinned (Phase 3 core complete)
Punch list §25 item 3.6 complete. Phase 3 (§25.4) core is now done:
items 3.1-3.6 all closed.

stadium_evict() now panics via sk_hal_panic() if a resident cell 0
(Hera, patron zero by construction of §6's boot order) is ever
selected for eviction. Placement is deliberate: the check runs before
the pin/contains refusal checks, not after -- if it ran after, a
wrongly-cleared pin would let the ordinary refusal path quietly return
-1 instead of ever reaching the panic, defeating the point of a check
that's supposed to be independent of pin holding.

Per §20.5 #3's explicit wording, not implemented as a filter:
stadium_admit()'s least-dense search is unchanged, still relying on
the general pin skip from item 3.5. Adding a second filter there would
have done exactly what that section warns against ("filtering hides
the bug, asserting reports it").

The panic path is, and will remain, unexercised by the acceptance
mechanism: sk_hal_panic() halts the machine, and triggering it
deliberately is incompatible with the three-arch boot being this
project's sole acceptance test. Correctness rests on the placement
argument, not a test -- same honesty precedent as items 3.4 and 3.5's
other unexercised paths.

Verified: three-architecture boot (amd64, aarch64, riscv64), all
reaching ok> with identical dict_hash=0x3d4e1daf289da94f matching the
item-3.5 baseline.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-04 17:36:23 -04:00
Robert Allan JamesandClaude Sonnet 5 f8a50561b0 starkernel: item 3.5 -- admission and eviction
Punch list §25 item 3.5 complete.

stadium_admit(candidate) places into an unused cell if one exists (no
comparison needed), otherwise finds the least-dense resident -- skipping
pinned and contains-gated patrons, which are never eviction candidates
-- and evicts it only if the candidate is strictly denser, per §19.3.
stadium_evict(cell_index) dispatches the departing patron's behaviour
before clearing its slot, per §17.2.

Caught a real bug before it ran: the first draft used contains == 0 to
mean "holds nothing," but cell index 0 is a valid index (Hera, item
3.6). Fixed with a proper sentinel, STADIUM_CONTAINS_NONE (UINT32_MAX).

A second-pass review found mass was not accounted for: both functions
handled exactly one cell regardless of the candidate's stated mass,
which leaks cells on eviction of any mass > 1 patron and breaks
capacity conservation. Fixed by refusing any candidate with mass != 1
-- multi-cell patrons need the per-VM free lists item 3.2 already
deferred (§22.3), not built here.

Documented, not fixed: the discriminator bitmap can't distinguish free
from continuation cells, so the free-cell scan reads continuation-cell
payload bytes under the header layout -- latent since nothing creates
continuation cells yet, and the mass != 1 refusal keeps it provably
latent. Superseded by the free list when it exists.

Unexercised at runtime: nothing calls either function yet (no real
patron kind is wired to the Stadium). No self-test added -- filling
~74,000+ cells to reach the eviction-on-full branch was judged
impractical, following item 2.2's own precedent for its unexercised
fleet-full path.

Verified: three-architecture boot (amd64, aarch64, riscv64), all
reaching ok> with identical dict_hash=0x3d4e1daf289da94f matching the
item-3.4 baseline.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-04 17:28:38 -04:00
Robert Allan JamesandClaude Sonnet 5 0b47c256fc starkernel: item 3.4 -- density ranking
Punch list §25 item 3.4 complete.

stadium_density(cell_index) reads a header's heat and mass and returns
heat / mass -- a division on demand from fields already stored in the
cell, matching §19.3's "read, not computed by a scheduler" literally.
Stays valid Q48.16 without a special fixed-point routine, since heat
is already Q48.16 and mass is a plain integer divisor.

mass == 0 and an out-of-range cell_index both return 0 rather than
dividing by zero -- an empty or never-admitted slot has no footprint
to be dense within.

Deliberately not built here, per the item's own wording: finding the
densest or least-dense resident (§19.3's admission/eviction
comparison) is item 3.5's scope, not this one's. Nothing calls
stadium_density() yet either.

Verified: three-architecture boot (amd64, aarch64, riscv64), all
reaching ok> with identical dict_hash=0x3d4e1daf289da94f matching the
item-3.3 baseline.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-04 17:16:46 -04:00
Robert Allan JamesandClaude Sonnet 5 378d688898 starkernel: item 3.3 -- behaviour enumeration and dispatch
Punch list §25 item 3.3 complete.

StadiumBehaviour (stadium.h) enumerates exactly the four tags §18.3
already names -- MIGRATE, DELIVER, EXPIRE, COOL -- mapped from §17.1's
patron table: blocks->MIGRATE, messages->DELIVER, ACLs->EXPIRE, words
and VMs both->COOL. Nothing invented; the tag set and mapping were
already in the document.

stadium_dispatch(cell_index, behaviour) dispatches on the tag only,
never asks what kind of patron departed. Handlers are stubs -- the
real actions belong to subsystems not yet migrated onto the Stadium
(Phase 4). Nothing calls stadium_dispatch() yet; item 3.5 is its first
consumer.

The switch is exhaustive with no default case, making §13's "closed
enumeration, fixed at build time" a compiler-enforced property under
this project's -Wall -Werror rather than just prose. Verified live:
temporarily deleted the COOL case, rebuild failed with
error: enumeration value 'STADIUM_BEHAVIOUR_COOL' not handled in
switch [-Werror=switch], restored it, confirmed clean again.

The header's behaviour field stays uint8_t, not the enum type itself,
since C does not guarantee an enum's underlying type and that field's
offset is load-bearing for item 3.1's validated 64-byte layout.

Verified: three-architecture boot (amd64, aarch64, riscv64), all
reaching ok> with identical dict_hash=0x3d4e1daf289da94f matching the
item-3.2 baseline.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-04 17:06:14 -04:00
Robert Allan JamesandClaude Sonnet 5 eb0fd4fffa starkernel: item 3.2 -- Stadium boot-time allocation
Punch list §25 item 3.2 complete.

stadium_boot_init() (src/starkernel/vm/stadium.c) sizes the global
cell array at boot from a real memory-budget query rather than a
hardcoded count: pmm_get_stats().free_bytes at the point of
allocation, times the new STADIUM_MEMORY_PERCENT Kconfig symbol
(default 1%), rounded down to whole 64-byte cells. Matches §17.6's
position (b) literally. Also allocates the header/continuation
discriminator bitmap item 3.1 declared but did not allocate. Both are
kmalloc'd and explicitly zero-filled (kmalloc does not zero).

Called from kernel_main.c immediately before sk_vm_bootstrap_parity(),
i.e. before any VM exists (§6). Failure is soft -- logs and continues,
does not halt boot -- matching the existing precedent one line below
it (VM bootstrap parity failure does the same).

Added a "Stadium: N cells (M KB)" boot console line at the allocation
site so the acceptance logs are evidence the array was actually
allocated, not just that the kernel still boots -- the same blind spot
item 3.1's uncompiled-header gap exposed.

Verified: three-architecture boot (amd64, aarch64, riscv64), all
reaching ok> with identical dict_hash=0x3d4e1daf289da94f matching the
item-3.1 baseline, and the Stadium boot line confirmed present in all
three serial logs (amd64: 74234 cells/4639 KB, aarch64: 161329
cells/10083 KB, riscv64: 76122 cells/4757 KB).

Not built here, reported per §25.0 rule 3: per-VM free lists (§22.3)
-- granted when Hera assigns quota, not this item's scope.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-04 16:48:13 -04:00
Robert Allan JamesandClaude Sonnet 5 1b2f0677de starkernel: item 3.1 reopened -- two Kconfig symbols items 1.1/1.4 deferred here
Punch list §25 item 3.1 re-closed after reopening.

Items 1.1 and 1.4's resolutions both explicitly named this item as
where their Kconfig symbols would be implemented, but 3.1's own stated
scope never mentioned them, so the first close missed both:

- STADIUM_CONTAINS_DEPTH_MAX (default 5) -- item 1.1's contains-chain
  depth cap. No consumer yet; reap-gating enforcement is item 3.5.
- STADIUM_CAPACITY_TICK (default 1000) -- item 1.4's capacity
  arbitration cadence in virtual ticks. No consumer yet; capacity
  arbitration itself is not on the punch list.

Both added following STADIUM_MAX_VM_COUNT's exact pattern:
Kconfig.kernel entry, Makefile.starkernel kconfig_int +
VM_FEATURE_FLAG_VARS forwarding, starforth_config.h fallback default.
stadium.h now includes starforth_config.h and carries two more
C99-portable compile-time checks proving both symbols are defined and
sane, same discipline as the byte-count checks. Declaration only --
not inventing the consuming logic to close this out early.

Verified: three-architecture boot (amd64, aarch64, riscv64), all
reaching ok> with identical dict_hash=0x3d4e1daf289da94f, re-run after
the reopening.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-04 15:15:59 -04:00
Robert Allan JamesandClaude Sonnet 5 d55ec3241b starkernel: item 3.1 -- the Stadium cell and header
Punch list §25 item 3.1 complete.

Defines StadiumPatronHeader and StadiumContinuationCell in new
include/starkernel/vm/stadium.h, unioned as StadiumCell per §3's
closed two-valued union. src/starkernel/vm/stadium.c added to
Makefile.starkernel's LOADER_EXTRA_SRCS/KERNEL_EXTRA_SRCS so the
header's compile-time size checks are actually compiled, not merely
included by something that never builds.

Discriminator ruled an external side bitmap (Captain Bob), not a
header field -- amended into §3 and §23.3 before this code was
written. Item 3.1 declares the bitmap's purpose/indexing in a comment
only; allocating it is item 3.2's scope.

Both cell shapes counted for real at exactly 64 bytes with zero
compiler-inserted padding (three C99-portable negative-array-size
assertions -- no _Static_assert, this project targets C99). Header
matches §23.3's original 32+32 split unchanged, since the
discriminator moving outside the cell left nothing to compete for that
space. Continuation cell matches item 1.12's 4+60 figure unchanged for
the same reason.

Verified the size assertion is actually live: broke it to 63,
confirmed the build failed with the expected negative-array-size
error, restored it, confirmed a clean compile.

Verified: three-architecture boot (amd64, aarch64, riscv64), all
reaching ok> with identical dict_hash=0x3d4e1daf289da94f matching the
item-2.2 baseline. Confirmed stadium.o present in both obj/loader/vm
and obj/kernel/vm post-build on amd64, closing the gap the item-2.2 WIP
exposed (an uncompiled header proves nothing).

Left open, not fabricated: §23.4 #2 ("does a typical message fit in
one cell") is unanswerable today -- no message patron struct exists
anywhere in this tree yet.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-04 14:55:30 -04:00
Robert Allan James a5ed8c3d87 Initial commit — LithosAnanke kernel 2026-08-01 07:49:56 -04:00