6302dcb50ea3da43a0cc99e0c0eea968eaee603b
42
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
3e201c82a5 |
Stage E: Category B strip -- remove Hermes, messaging.4th, old routing (FABRIC-3.6.md task 4.1-4.4)
All FORTH-owned message types were cut over to kernel-Hermes in Phase 3 (tasks 3.8-3.10). This removes the now-dead FORTH messaging layer and the Hermes VM itself: capsules/common/messaging.4th, capsules/hermes/init.4th, the slot-3 VM-NAME-REG pairing convention, and every load-site/birth-site reference across capsules/init.4th, artemis/init.4th, hestia/init.4th, doe-campaign.4th (Artemis-only now), capsule_console.c, capsule_mint.c, capsule_wirebind.c, capsule_birth.c, and kernel_main.c. Verified on all three architectures: clean boot, mkcapsule --lint clean (36 files, 0 violations), zero UNKNOWN WORD, identical dict_hash across amd64/aarch64/riscv64 for every VM, and a full mint -> WIREBIND-attach -> USE -> relay round-trip exercising the two highest-risk edits (capsule_console.c/capsule_mint.c). Found, not fixed: deleting messaging.4th removes SEND-ELEVATE-REQUEST, which was the only caller of KH-ELEVATE-SEND and the only path to ELEVATE-GRANT (zuse-eligibility.4th, still loaded at boot) -- Phase 8 PKI's own elevation entrypoint. Needs a decision before Phase 5 close-out. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
cab9b5a31f |
Stage D: ELEVATE-REQUEST real cutover -- FABRIC-3.6.md task 3.10
Reachability verified live before writing any code, per this project's own standing rule (grep cannot establish reachability alone): FIND SEND-ELEVATE-REQUEST / FIND ELEVATE-GRANT / FIND CH-REQUEST all resolve on a live Hera boot, though grep across capsules/experiments/docs found zero callers of SEND-ELEVATE-REQUEST -- a real, complete, directly-callable entrypoint (H.5/H.8's own design) with no current automatic trigger, not dead code. Correction to a prior finding, made in the course of this check: task 3.8's write-up claimed "Hera's own pre-existing inability to load common:messaging.4th" -- false. capsules/init.4th (Hera's own MAMA_INIT capsule) loads it directly, and SEND-ELEVATE-REQUEST lives and works in her dictionary right now. Task 3.8's own actual scope is unaffected by this correction. Cutover: SEND-ELEVATE-REQUEST (messaging.4th) no longer ends in CH-REQUEST's COMMON-CH/MSG-SEND path; it now calls KH-ELEVATE-SEND (repl.c), a new C word wrapping sk_hermes_send_one(), registered unconditionally for every VM. from/to are derived from the calling VM and sk_get_mama_vm() directly in C, never taken from the stack -- a real correctness improvement over CH-REQUEST's own initiator-only gate, which only existed because a caller COULD pass the wrong from value; deriving it in C makes that spoof structurally impossible. SK_HERMES_MSG_TYPE_ELEVATE_REQUEST reuses ELEVATE-REQUEST's own value (8), same partition-rule reasoning as tasks 3.8/3.9. Delivery is unchanged task 3.4 machinery. No new static-buffer lifetime caveat -- the payload-aliasing fix landed before this task started. CH-REQUEST (messaging.4th) is now dead code, its one real caller just removed -- found, not fixed, per Captain Bob's Law. Verified live on all three architectures: 0 0 0 0 S" DUP" SEND-ELEVATE-REQUEST (deliberately-invalid pubkey, so ELEVATE-GRANT correctly refuses -- the check is the pipeline running, not a grant succeeding) fires the evidence line and completes cleanly, DUP unaffected afterward. Zero UNKNOWN WORD, dict_hash identical across all three architectures (changed uniformly from prior runs -- one new word registered -- not diverged, matching SXXXIV.6's own rule). This closes task 3.11 (Phase 3 gate): all of messaging.4th's live FORTH-owned message types (BLK-ATTACH-EVENT, CONSOLE-CMD-EVENT, ELEVATE-REQUEST) are now real kernel-Hermes cutovers. Phase 4 may begin. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
5c07745c18 |
Fix SkHermesMessage payload-aliasing defect (task 3.8 findings log)
sk_hermes_send_one() -- the single funnel every sender, including sk_hermes_publish(), already goes through -- used to store the caller's own payload_addr pointer as-is. Two sends before either drains meant both messages pointed at the same caller-owned buffer, whichever send wrote last silently winning: real, confirmed live (a second identity thumbdrive attached at boot alongside Zuse's own left its WIREBIND pairing silently never happening). SkHermesMessage gains an inline payload_buf[SK_HERMES_CHUNK_MAX_ PAYLOAD] field; sk_hermes_send_one() now memcpy()s the caller's payload into it and points payload_addr at that copy instead. No sender or reader call site needed to change -- every existing reader already only ever reads through payload_addr, which still points at valid bytes of the same length, now message-owned. g_kh_blk_attach_buf/ g_kh_console_cmd_buf (repl.c, tasks 3.8/3.9) no longer need to survive past their own send call; their doc comments, which had claimed the old aliasing shape was benign, are corrected. Verified the original bug is actually gone: reproduced the exact original scenario (a real, sequentially-minted rajames identity attached at boot alongside Zuse's own, two simultaneous BLK-ATTACH- EVENT sends in one idle-loop pass) -- WIREBIND pairing, USE, and the Stage D relay all now work where WIREBIND previously silently failed. Standard three-ISA acceptance also clean: zero UNKNOWN WORD, dict_hash identical across all three and matching every prior acceptance run in this document. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
3975dc61cd |
Stage D: CONSOLE-CMD-EVENT real cutover -- FABRIC-3.6.md task 3.9
Send-side cutover: sk_repl_dispatch_line()'s FORTH-string
"CONSOLE-CMD-EVENT 0 3 S\" ...\" 0 MSG-SEND" interpret is replaced with
a direct sk_hermes_send_one() call (SK_HERMES_MSG_TYPE_CONSOLE_CMD,
kernel_hermes.h, deliberately reusing CONSOLE-CMD-EVENT's own value 7,
same partition-rule reasoning as task 3.8's BLK-ATTACH-EVENT cutover).
Real finding along the way: kernel-Hermes's drain only ever runs as a
side effect of vm_interpret() being called on the target VM. Task
3.8's target (Hera) is always being interpreted via the interactive
REPL loop; task 3.9's target is a WIREBIND identity's own ~user VM, a
passive receiver nothing else drives. The existing idle-loop pump only
ticked VMs with the old FORTH MSG-TICK word ACL-allowed -- a VM minted
with the STD79-lockdown personality never has it, so the pump silently
skipped it forever and queued messages never delivered. Fixed by
adding an unconditional, direct sk_hermes_drain_checkpoint() call in
the same pump loop, independent of the MSG-TICK gate.
Verified live on all three architectures: WIREBIND-attach a real
identity, USE into it, type a plain console line, confirm the Stage D
evidence line and correct relayed execution result. Zero UNKNOWN WORD,
dict_hash identical across all three ISAs and matching task 3.8's own
baseline. The FABRIC-3.md-documented USE/BINDSTEP crash did not
reproduce in any of these live sessions (recorded as a finding, not
chased further).
Depends on the zuse_root_pubkey_known fix already landed in
|
||
|
|
bbfd9103f4 |
Stage C: cut over BLK-ATTACH-EVENT alone -- FABRIC-3.6.md task 3.8
The reply leg (Artemis -> Hera ack) that used to flow through common:messaging.4th's MSG-SEND/MSG-TICK now goes through kernel-Hermes's sk_hermes_send_one()/sk_hermes_drain_checkpoint() instead -- FORTH Hermes never sees a BLK-ATTACH-EVENT message again (SXXXIV.2's partition rule). The request leg was never real FORTH messaging traffic to begin with (a direct VM-EXEC, no type tag, forced by Hera's own inability to load common:messaging.4th), so it is untouched. New KH-BLK-ATTACH-SEND (repl.c) wraps sk_hermes_send_one(), reached from capsules/artemis/init.4th's HERA-BLK-ATTACH-REQ. Delivery reuses task 3.4's already-wired sk_hermes_drain_checkpoint(); BLK-ATTACH-ACK itself is unchanged, just reached by a different layer. SK_HERMES_MSG_TYPE_BLK_ATTACH deliberately reuses BLK-ATTACH-EVENT's own value (9) to document this as a cutover of the same message, not a new one. Two real bugs found on the way, both recorded in FABRIC-3.6.md's findings log: - A popped FORTH CREATE-buffer address was raw-cast to a host pointer instead of going through vm_ptr() -- silently read all-zero memory, no crash, no error, just a message that arrived and did nothing. Fixed; the rule and its exception (repl.c's own dev-addr is legitimately a raw pointer, formatted that way by its own pushing code) are written up for the next FORTH-facing C word. - A separate, genuine hang on the very first live exercise of this path, never reproduced across ten subsequent boots. Reported, not chased -- not blocking, per the task's own check being otherwise fully satisfied. Also found live: log_message() is invisible in this build's actual serial-log capture at every level -- settled on a single console_println in the real drain target instead, one line per real USB attach, not a hot-path. Final acceptance (logs/20260922-105501, -105758, -110304, disk images reset before each): dict_hash identical across all three architectures for every VM, zero UNKNOWN WORD, mkcapsule --lint clean, real ledger+stadium_conserved(Artemis)=true evidence on every boot. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
f37aa0fb17 |
Channel-open policy hook -- FABRIC-3.6.md task 3.7
Added HERMES-CHANNEL-OPEN? ( req-hi req-lo -- allow? ) at capsules/ACL.4th block 4008 (default: approve everything) -- the one word policy authors edit. sk_hermes_channel_open_policy(VM*, VMUuid) (kernel_hermes.h/.c) is the C-side query that calls it via plain word-dispatch against the target VM's own dictionary/stack, never vm_interpret() (avoids task 3.4's input-buffer cursor hazard entirely) and never decides the answer itself. Fails closed: no policy word, a policy error, or stack underflow all deny, matching CLAUDE.md's posture that absence of policy must never mean "always allow." Two real bugs found and fixed before this was called done: missing current_executing_entry assignment before calling the word's func pointer (colon words silently no-op without it, vm_core.c:730 -- no crash, just a wrong answer); and a second FAIL with debug instrumentation still in place whose precise cause isn't reconstructable, since no intermediate commit exists for that attempt. Self-test proves the task's check four ways against the same unchanged C function: default approve, live redefinition to deny (zero C change), restore, and a VM with no ACL.4th loaded at all (fail closed). A fifth check wires the result into task 3.6's sk_hermes_channel_respond() end to end: a denied policy produces a NACK and no channel, ledger/stadium_conserved() holding throughout. Scope, per Captain Bob's ruling: closes with the query built and proven; sk_hermes_channel_respond() still takes a caller-supplied approved bool rather than calling the policy internally. Wiring a real channel-open call site to only this query is deferred to whichever later task first needs a live decision. dict_hash identical across amd64/aarch64/riscv64 for every VM, zero UNKNOWN WORD, mkcapsule --lint clean (38 files, 0 violations). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
205a49ecd0 |
ACK/NACK and private-channel negotiation -- FABRIC-3.6.md task 3.6
Extracted sk_hermes_send_one() from sk_hermes_publish()'s own
per-subscriber body -- one code path for both point-to-point and
fan-out delivery, so the ledger can never diverge between them.
Point-to-point addressing turned out to be load-bearing, not
incidental: sk_hermes_publish()'s fan-out sets msg->to to whichever
member it is iterating, so a negotiation message "published" to the
common channel would spuriously reach every member, not just the real
target (checked with advisor() before building the naive version).
"Over the common channel" means every VM is reachable from birth (task
3.2), not that the exchange itself fans out -- messaging.4th's own
CH-REQUEST carried an explicit `to` for the same reason.
sk_hermes_channel_request/respond/close build the mechanics: request ->
grant (creates a private channel, subscribes both parties, sends
CH_GRANT + one ACK) or NACK ("a deny is a NACK", SXLV.1 -- no separate
type); close authorized by membership alone. The grant/deny decision is
a plain caller-supplied `approved` bool -- task 3.7 replaces the call
site that produces it with a real ACL.4th query, not this signature.
Self-test covers the task's own three checks plus a sibling advisor()
flagged: an approved respond() whose channel creation itself fails
(table exhausted) must still fall through to NACK, not a silent false
grant or half-open channel -- verified by exhausting the whole channel
table and confirming the fallback.
Bug found and fixed before this was called done: the first draft
dropped a message via pending_pop() alone, without releasing it first,
leaking its Stadium heat and failing the self-test's own ledger
baseline check (logs/20260922-065946/amd64/, kept as audit trail).
Fixed and re-verified PASS on all three architectures.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
|
||
|
|
8a2ee0fdad |
Payload bound and chunking -- FABRIC-3.6.md task 3.5
sk_hermes_publish() now enforces SK_HERMES_CHUNK_MAX_PAYLOAD (1024) on every message's payload_len uniformly, chunked or not -- closing the gap task 3.3 explicitly parked. A chunk carrier is [SkHermesChunkHeader][content slice], slice capped at 1024 - sizeof(header) rather than 1024 itself, so every message on the wire satisfies the same one-block bound vm_interpret()'s own drain limit already requires -- a future chunk-aware drain never has to special-case a carrier that can't be handed to vm_interpret() as-is. Deliberately no chunking-sender API: building one would need kernel-Hermes to own chunk-buffer memory with a real lifetime it has no way to track (kept alive until every subscriber drains it). Sending is a loop pattern a caller writes with sk_hermes_chunk_count() + sk_hermes_publish(), demonstrated by this task's own self-test. sk_hermes_reassemble() is pure and memory-agnostic: validates msg_id agreement, exact seq coverage, and per-chunk slice sizes before a single memcpy, with the total length computed once and checked against the caller's buffer once -- never order-dependent on which chunk happens to overflow. Verified live on all three architectures: a 1024-byte payload as one message, a 1025-byte send refused outright with the ledger untouched, and a 3000-byte payload split into 3 chunks, drained, and reassembled byte-exact against the original. dict_hash unmoved and identical across architectures. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
2b1ba031a5 |
Drain at the outermost checkpoint -- FABRIC-3.6.md task 3.4
sk_hermes_drain_checkpoint() interprets one queued payload per checkpoint (ruled: one message per checkpoint), reusing sk_vm_at_outermost_interpret() and placed before the switch-signal block in vm_core.c's existing cooperative checkpoint (sk_vm_context_switch() doesn't return until switched back to, so drain must come first or it silently never runs on a switching checkpoint). Amends FABRIC-3.5.md SXLIII.5, caught by advisor() before writing the naive version: "recursive drain is prevented for free" via g_vm_interpret_depth is true but only for same-message re-drain -- it doesn't cover the separate same-VM reentrancy hazard FABRIC-3.md SXX already named for Hera specifically (VMCallState saves rsp/exit_colon/ ecw_nesting only, never input_buffer/input_length/input_pos). Draining calls vm_interpret() on the same vm whose own vm_interpret() call is still paused mid-word at the checkpoint; without saving and restoring the cursor by hand, the enclosing REPL line or LOAD block would be silently truncated. sk_hermes_drain_checkpoint() snapshots and restores input_buffer/input_length/input_pos/mode/error/abort_requested around the call. Not a divergence from the ruling -- cursor preservation is the implementer's own obligation inside the ruled mechanism. Gated behind a system-wide pending-total counter so the common no-message-in-flight case costs one integer read per word dispatch, not a stadium_max_vm_count()-sized queue scan (also flagged by advisor() as a real hot-path cost, not deferred). Verified live on all three architectures: a self-test publishes a real payload to Hermes, proves the depth gate via VM-EXEC-ing an existing harmless colon word into Hermes (genuine nested vm_interpret(), depth 2, must not drain), then drains directly from genuinely-outermost context and confirms exactly one clean drain. dict_hash unmoved and identical across architectures. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
1f6343bc03 |
Publish path, no dispatch -- FABRIC-3.6.md task 3.3
sk_hermes_publish() allocates one SkHermesMessage per channel member (heat-cost ruling 2026-09-21: one message per subscriber, funded by the publisher's own reservoir) and enqueues each onto a new per-subscriber SkHermesPendingQueue -- found-or-created lazily by vm_id, sized from stadium_max_vm_count() like the channel/switch tables. Best-effort across subscribers: a failed allocation or full queue skips and rolls back just that one subscriber, not the whole publish -- the natural reading of "ledger and stadium_conserved() hold across N publishes to M subscribers" (the task's own check), not a separate ruling. Dispatches nothing -- sk_hermes_pending_count()/peek()/pop() are the read/drain primitives task 3.4's real checkpoint-driven drain will build on; this task's own self-test uses them directly since no checkpoint hook exists yet. Verified live on all three architectures: pending-queue table sized 50/202/50 slots (tracking the channel table's own per-arch sizing), a synthetic publish self-test (2 publishes to 3 subscribers) confirms exact per-subscriber delivery counts, and ledger/stadium_conserved() invariants hold both mid-publish and after manually draining every queue back to baseline. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
a4afdfa591 |
Channel table + common channel, inert (B1) -- FABRIC-3.6.md task 3.2
Adds SkHermesChannel: a channel is an index into a boot-time, stadium_max_vm_count()-sized table (same sizing pattern task 3.1 established for the switch table -- no separate numeric rule was ruled for this table, so task 3.1's bound is extended directly, flagged as such rather than restated as a new ruling). No name field, mirroring messaging.4th's own nameless CH-ARENA. The common channel (index 0) is created at boot and permanent. Hera subscribes explicitly in kernel_main.c (she is the one VM never born through capsule_birth_baby()); every other VM -- Tripod fleet and future WIREBIND identities alike -- subscribes inside capsule_birth_baby() itself, the single choke point every other birth already passes through. Inert: no publish, no dispatch, no ACK/NACK, no ACL hook (tasks 3.3, 3.6, 3.7). Verified live on all three architectures: channel table sized to 50/202/50 slots (matching switch-signal's own per-arch sizing), common-channel fleet self-test confirms all four Tripod members are members, and a synthetic create/subscribe/unsubscribe/ destroy round-trip against a private topic passes, including refusing to destroy the common channel. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
feace42397 |
Add diagnostic scan cross-check of the Hermes counters -- FABRIC-3.6.md task 2.8
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
2a2bf6eb35 |
Stage B proof: add per-VM consumed term to stadium_conserved -- FABRIC-3.6.md task 2.7
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
313ffc89e6 |
Add exact-equality Hermes ledger self-audit -- FABRIC-3.6.md task 2.6
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
d11eb2e5db |
Add sk_hermes_decay(), ledgering decay into consumed -- FABRIC-3.6.md task 2.5
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
493d410028 |
Add the four ledger counters (held/pulled/returned/consumed) -- FABRIC-3.6.md task 2.4
FABRIC-3.5.md SXL.4's ledger: held == pulled - returned - consumed,
epsilon zero. Added sk_hermes_ledger() (an accessor, not a mutator)
plus four static counters in kernel_hermes.c.
Each counter has exactly one increment/decrement site: held/pulled
both move at sk_hermes_alloc()'s single success path, after every
refusal branch has already returned; held/returned both move at
sk_hermes_release()'s single success path. consumed is declared and
always reads 0 -- its one increment site doesn't exist yet, and won't
until task 2.5 gives decay something to record.
sk_hermes_release() now reads the Stadium cell's live header.heat
immediately before calling stadium_evict(), rather than assuming the
original pulled amount -- stadium_evict() zeroes the header as part of
freeing the cell and its own return value is a success code, not the
credited amount, so this is the only point the true remaining heat is
available. Today this always equals the original Q.SLOT pull; once
task 2.5's decay exists, this is what keeps returned correct without
touching this function again.
Self-test (kernel_main.c) extended: snapshots the ledger before
running so it checks its own deltas, verifies held/pulled grow by
exactly got_n * Q_SLOT on allocation with returned/consumed untouched,
then verifies held returns to its starting value and returned grows by
the same amount on release, and checks the audit invariant itself as a
bonus (task 2.6 formalizes this properly).
Three-arch boot clean: amd64/aarch64/riscv64 all reach [zuse@Hera] ok>,
zero UNKNOWN WORD, dict_hash unmoved. All three print PASS with
identical final ledger: held=0 pulled=65536 returned=65536 consumed=0.
No compiler warnings.
Authorized by Captain Bob ("keep going with rhe 6.5 document").
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
|
||
|
|
2c1dc1753a |
Correct sk_hermes_alloc() to admit a real Stadium patron; add sk_hermes_release() -- FABRIC-3.6.md tasks 2.2 (amended) + 2.3
Real finding, caught before building release on a foundation that
couldn't support it: task 2.2's first cut of sk_hermes_alloc() pulled
reservoir heat but never admitted a real Stadium-floor patron -- just a
local in_use flag. FABRIC-3.5.md SXXXIII.4 item 1 says MSG-FREE-NODE
returns heat "via STADIUM-EVICT", which only means something if
allocation admitted something. SXL.4's own invariant, Sigma(resident
patron heat) + reservoir + consumed == Q48_ONE, cannot balance if held
heat is invisible to every term while held. Flagged to Captain Bob
before proceeding; authorized to correct 2.2 in the same pass as
building 2.3 on top of the fix.
sk_hermes_alloc() now calls stadium_admit() with heat = the pulled
amount, behaviour = STADIUM_BEHAVIOUR_DELIVER (matching messaging.4th's
own SB-DELIVER STADIUM-ADMIT exactly), and identity = the message's own
slot index (matching the FORTH precedent -- caught live in the first
boot of this fix that omitting this made stadium_dispatch()'s existing
DELIVER diagnostic print msg_idx=0 for every message instead of a
distinct value). The returned cell index is stored in the message's
own stadium_cell field. Stadium-floor refusal (independent of reservoir
affordability) rolls back the pull the same way the other refusal
paths already do.
sk_hermes_release() -- the function task 2.3 actually asks for -- calls
stadium_evict() on that cell, which itself returns the departing
patron's remaining heat to its owning VM's reservoir, matching
MSG-FREE-NODE's exact shape. Release does not touch the reservoir
directly.
Self-test (kernel_main.c) extended: keeps every allocated message's
pointer, allocates to exhaustion as before, releases all of them, and
checks the reservoir returns to precisely its starting value.
"Undecayed" is true by construction (no decay/TTL logic exists yet,
task 2.5) -- exact restoration, not approximate.
Three-arch boot clean: amd64/aarch64/riscv64 all reach [zuse@Hera] ok>,
zero UNKNOWN WORD, dict_hash unmoved from task 2.2. All three print
identical PASS arithmetic: reservoir0=65536, reservoir_after_alloc=0,
reservoir_final=65536. No compiler warnings.
stadium_dispatch()'s DELIVER-case console output (one line per
eviction) is pre-existing instrumentation, not new -- confirmed real
and load-bearing per stadium.c's own comment, verbose but expected.
Authorized by Captain Bob ("yes").
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
|
||
|
|
9d129fdb1c |
Add sk_hermes_alloc(), the heat-coupled allocator -- FABRIC-3.6.md task 2.2, item 28
The piece FABRIC-3.5.md SXXXIII.6 calls "what remains genuinely hard,"
built and proven first per its own recommendation. Added
src/starkernel/vm/kernel_hermes.c (wired into Makefile.starkernel's
LOADER_EXTRA_SRCS -- this repo lists vm/*.c files explicitly, no glob)
and sk_hermes_alloc()'s declaration in kernel_hermes.h.
Checks stadium_reservoir_peek(vm_id) >= SK_HERMES_Q_SLOT before
touching the reservoir at all -- refusal this way needs no rollback,
since nothing was pulled -- with an explicit rollback path
(stadium_reservoir_push) kept defensively for the pull-then-short case,
though nothing in this single-core kernel is expected to reach it.
SK_HERMES_Q_SLOT = Q48_ONE / SK_HERMES_MSG_MAX (2048), deliberately
simpler than messaging.4th's own formula, which reserves a Q.1/3 floor
for COMMON-CH's own Stadium heat -- kernel-Hermes has no such object
(SXXXIII.4/SXXXIII.5's flat membership list carries no heat of its
own), so there is nothing left for that floor to protect.
Self-test in kernel_main.c, same diagnostic-only synthetic-VM pattern
as the existing Stadium quota grant self-test (lo=3, distinct from
that test's lo=1): reads back the actual granted reservoir rather than
assuming a number, derives expected_n from it, allocates to refusal,
and checks the refusal lands at exactly expected_n, the reservoir
doesn't move on the refused attempt (rollback proven, not assumed),
and the final reservoir is exactly reservoir0 minus got_n times
Q_SLOT.
Three-arch boot clean: amd64/aarch64/riscv64 all reach [zuse@Hera] ok>,
zero UNKNOWN WORD, dict_hash unmoved from task 2.1 (pure C, no FORTH
touched). All three print identical self-test arithmetic: reservoir0=
65536 Q_SLOT=2048 expected_n=32 got_n=32 reservoir_after=0. No compiler
warnings.
Noted, not fixed: Q_SLOT's divisor and SK_HERMES_MSG_MAX are the same
32, so reservoir and arena exhaustion land at exactly the same count by
construction -- this test can't distinguish which refusal reason
fired, only that refusal is correct and rolls back correctly.
Deliberately not evidence for stadium_conserved(): allocating alone
(no release yet, task 2.3) leaves pulled heat held off the Stadium
floor, so the two-term check would correctly read false right now if
run mid-hold. That's expected, not a bug -- Stage B (task 2.7) is
defined as "before and after the alloc/free cycle," not "continuously
during." This task's self-test checks reservoir arithmetic directly
instead.
Authorized by Captain Bob ("Yes continue").
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
|
||
|
|
f10fa7ae83 |
Add kernel-Hermes message/membership structures -- FABRIC-3.6.md task 2.1, Phase 2 begins
Phase 2, task 2.1 only: type definitions, wired to nothing, drawing no
heat -- no allocator, no protocol logic, no registration anywhere.
FABRIC-3.5.md SXXII.4: Phase 2 structures come first and prove nothing
until the allocator is built on top (task 2.2 onward, each its own
commit).
Added include/starkernel/vm/kernel_hermes.h:
SkHermesMessage -- field-for-field mirror of messaging.4th's live
9-cell MSG-* layout (type/from/to/payload addr+len/Stadium cell
index/seq/channel/orig-type), per SXXXIII.4 item 1 ("roughly half the
file is accessors that become struct fields"), plus an explicit
in_use flag for task 2.2's allocator. Deliberately no separate heat
field: per SXL.4, a message's heat IS the Stadium cell it occupies,
not a value copied alongside it -- one source of truth for the
conservation invariant stadium_conserved() (task 0.7) checks.
SkHermesMembership -- one flat broadcast membership list, SXXXIII.4/
SXXXIII.5's recommended replacement for messaging.4th's 28-word channel
abstraction (traced to exactly one live caller, CH-ADD-MBR). Item 27
(negotiation vs. broadcast, Phase 3 blocker B1) is not answered by this
structure and isn't meant to be -- a flat list is correct either way.
Genuinely wired to nothing: no .c file, no Makefile change, no include
from any compiled source. Syntax-checked standalone (gcc -std=c99
-Wall -Wextra -Werror -fsyntax-only) before touching the real build.
Boot byte-identical to task 1.9's baseline on amd64 (same dict_hash
triple, zero UNKNOWN WORD). Did not repeat aarch64/riscv64 -- the file
compiles into no object on any architecture, so there is no mechanism
by which it could diverge.
Authorized by Captain Bob ("yes").
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
|
||
|
|
380f0a09c9 |
Add stadium_conserved() -- FABRIC-3.6.md task 0.7, item 41
Boolean analogue of vm_physics_conserved(), for the Stadium per-VM
quota invariant rather than fleet-wide execution heat
(FABRIC-3.5.md SXXXIX.4). int stadium_conserved(VMUuid vm_id), in
src/starkernel/vm/stadium.c alongside stadium_resident_sum()/
stadium_reservoir_peek() that it's built from, declared in
include/starkernel/vm/stadium.h.
Implements the two-term form -- resident_sum(vm_id) +
reservoir_peek(vm_id) == Q48_ONE -- not the three-term form SXL.4
rules for the eventual system. That ruling's `consumed` term is a
Phase 2 kernel-Hermes ledger deliverable that doesn't exist yet:
nothing draws on any VM's Stadium quota today (task 2.2 is literally
where that wiring gets built), so consumed is honestly zero right now.
Folding it in as a placeholder would be inventing Phase 2 state ahead
of it existing -- the doc comment says so explicitly, so whoever
builds Phase 2's ledger extends this function rather than working
around it.
Wired into the existing per-VM boot diagnostic
(stadium_words_print_boot_diagnostics(), kernel_main.c:810, Hera
only -- the sole existing call site) rather than adding a new one,
printing CONSERVED/DRIFTED the same shape vm_physics_status() already
uses.
Three-arch boot clean: amd64/aarch64/riscv64 all reach [zuse@Hera] ok>,
zero UNKNOWN WORD, all print "Stadium conservation: CONSERVED" with
identical resident_sum=47641 reservoir=17895 sum=65536=Q48_ONE. No
compiler warnings on either edited file (forced recompile checked).
Authorized by Captain Bob ("Yes.").
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
|
||
|
|
66beae7fd4 |
Stage 3 follow-on: message-arrival eligibility hook + trampoline-blind switch-storm fix (FABRIC-3.md §XXVIII.2)
Implements the message-arrival eligibility signal FABRIC-3.md §XXVIII.1 left open (has_work per-slot flag, set via new SWITCH-MARK-WORK primitive from MSG-SEND) so an idle VM never becomes a switch target purely by waiting out the readiness threshold. Also root-causes and fixes a second, independent switch-storm: the tick's "who is current" check used vm_log_attributed_vm(), which can't see a VM parked in switch.c's own raw trampoline. Replaced with a dedicated g_switch_current_vm tracked by the switch mechanism itself, and moved target-slot eligibility reset to the switch decision point instead of relying on ISR polling to observe a window that can be only a few instructions wide. Verified live on all 3 architectures: clean boot to zuse@Hera] ok>, live cross-VM message dispatch, and (since a quiet log looks identical to a livelocked storm once the DoE probe is gone) confirmed genuine REPL liveness via QMP send-key + screendump on aarch64/riscv64, not log inspection alone. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K |
||
|
|
f790d0995e |
Stage 2: cooperative VM context switch primitive, proven on all 3 arches (FABRIC-3.md §XXVIII)
Third stage of the preemptive context-switching plan. The real save/restore switch mechanism now exists -- the first time anything has ever executed on a VM's own native stack (Stage 1 allocated them, unused). New sk_vm_switch_to() (switch.S, one per arch) is an ordinary function call, not an interrupt -- so unlike Stage 0's trap frame, the ABI already covers every caller-saved register; only the callee-saved set needs explicit save/restore (amd64: rbx/rbp/r12-r15, no FP at all since SysV has no callee-saved XMM; aarch64: x19-x28/x29/x30 + d8-d15; riscv64: s0-s11/ra + fs0-fs11, FS-gated like Stage 0 but read once and reused for both halves within one call, since FS is genuine global CPU state, not part of what's switched). A sibling sk_vm_switch_prime() in the same file builds the synthetic first-entry frame, kept in assembly so the layout can never drift out of sync with sk_vm_switch_to() itself. New switch.c/switch.h: sk_vm_context_switch(from, to) handles first-entry priming vs. resuming a parked context, and updates registry state (new VM_STATE_SWITCHED_OUT, distinct from VM_STATE_STOPPED -- STOPPED means no live frame, this means the opposite). sk_vm_switch_entry() is the minimal permanent trampoline every freshly-entered VM lands in: no production behavior defined yet, so it just yields straight back to whoever switched to it, forever. Closes the confirmed unguarded-KILL UAF found during planning: capsule_vm_kill(), mama_word_kill(), and capsule_vm_kill_all_nonmama() all now refuse (or silently leak rather than free, on the cold-restart path where arch_cold_reset() wipes everything immediately after anyway) tearing down a switched-out VM. Side effect found, not built on purpose: the existing MSG-TICK idle-pump already filters on VM_STATE_LIVE, so it automatically stopped dispatching into a switched-out VM with zero changes needed there. Verified via a temporary SWITCH-TEST probe (boot-triggered, since nothing can type interactively into a foreground-only QEMU session) that round-tripped a sentinel through 5 real Hera<->Hermes switches on all 3 architectures: 5/5 rounds, 0 failures, clean continuation to ok>. Probe fully reverted after capture; kernel_main.c shows zero diff. Also: Makefile.starkernel's LOADER_EXTRA_SRCS/LOADER_ASM needed the new files added explicitly (this project's "loader" PE binary is the full running kernel, not a thin bootstrap stage), and aarch64's switch.S needed the same #ifndef _WIN32 guard around .hidden that isr.S already carries (aarch64's loader assembles via clang targeting a PE/COFF target with no .hidden equivalent) -- caught by a build failure, fixed. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016UNhH1mhi52i6Qihh7ZV5S |
||
|
|
57ac3fc304 |
Stage 1: per-VM native stacks, allocated but not yet executed on (FABRIC-3.md §XXVIII)
Second stage of the preemptive context-switching plan. Every VM (Hera, every capsule_birth_baby()-born VM including WIREBIND identities) now gets its own dedicated 2 MiB native C stack at birth -- but nothing runs on it yet, that's Stage 2. Pure allocation-machinery proof. Design correction made before writing code: the plan called for cloning sk_vm_arena_alloc()'s guard-page pattern, but that pattern turns out to be Mama-only -- host_services.c's kernel_alloc() gives every baby VM a plain kmalloc() block for its dictionary arena, not a real guarded PMM allocation. Stacks get the real treatment instead (new sk_vm_native_stack_alloc()/_free() in arena.c): independent pmm_alloc_contiguous() + guard pages for every VM without exception, no singleton, no kmalloc fallback -- a stack overflow is exactly the failure mode guard pages exist for, and a corrupted stack could corrupt whatever saved context Stage 2 trusts. 2 MiB size matches this project's own established kernel-stack convention (g_kernel_stack/g_rpi5_native_stack), not a guess -- that one shared 2 MiB stack today already carries all VMs' combined nested VM-EXEC recursion. Three new VM struct fields, freed in vm_cleanup() alongside the existing call_stack free. Allocation failure is non-fatal to birth. All 3 architectures re-verified clean boot to ok>, no native-stack allocation failures for any Tripod-fleet VM. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016UNhH1mhi52i6Qihh7ZV5S |
||
|
|
e51a8d229e |
Fix stadium_grant_quota() donor floor; rerun std79 DoE clean, 81/81 (FABRIC-3.md §XVII)
capsule_birth.c hardcoded every new VM's initial Stadium quota grant to split from Hera specifically. Since a grant always halves whatever the donor currently has, Hera's own free list converges toward empty after a bounded number of grants — independent of whether the Stadium as a whole still had spare capacity, since VMs she'd granted to earlier typically still held nearly all of their own share untouched. Past that point every subsequent VM birth's Stadium grant would be silently refused (soft-failed, non-fatal by existing design), even with plenty of capacity sitting idle elsewhere. Fixed by adding an O(1)-maintained free_count to StadiumVMQuota (incremented in stadium_evict(), decremented at both of stadium_admit()'s free-list-pop sites, set/adjusted in stadium_grant_quota()'s own split — this also let grant_quota drop its old O(free-list length) counting walk in favor of an O(1) read) and stadium_best_donor(), an O(live VM count) scan over quota slots returning whichever in-use VM currently has the most free cells. capsule_birth.c's birth path now splits from that VM instead of unconditionally vm_uuid_hera(). Verified with another full rerun of the 3x9x3 std79 DoE campaign from scratch — same discipline as the prior Stadium fix (any defect repair reruns the whole DoE from the top) — one continuous boot per architecture, all 9 identities simultaneously live throughout. 81/81 trials correct, 0 mismatches, DOE-RUN header sequence md5-identical to every prior run. aarch64 ~280s total (vs ~290s for the O(ncells)-scan fix alone — confirms no regression). Both known Stadium defects are now closed together on one clean campaign rerun. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo |
||
|
|
b031b802e3 |
Rename FABRIC series: FABRIC.md->0, FABRIC-2.md->1, FABRIC-3.md->2, FABRIC-4.md unchanged
FABRIC.md -> FABRIC-0.md FABRIC-2.md -> FABRIC-1.md FABRIC-3.md -> FABRIC-2.md (the current/living document) FABRIC-4.md unchanged (new #3 to follow separately) Every cross-reference repo-wide updated to match, including doc-comment citations inside kernel source (.c/.h) files -- done via an ordered placeholder substitution (FABRIC-3.md->placeholder2, FABRIC-2.md-> placeholder1, FABRIC.md->placeholder0, then placeholders resolved to final names) in a single pass per file to avoid double-shifting already-renamed references. One line in capsules/font.4th grew past the 64-char block-format limit as a side effect of the longer filename; shortened it and reverified with mkcapsule --lint (34/34 pass) before rebuilding. Verified 3-arch boot to ok> (amd64/aarch64/riscv64, each in the foreground) after the fix; logs and DoE CSVs from this session's verification runs included per this repo's own audit-artifact convention. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019YcT3H2PQeyujrzjqS3Var |
||
|
|
a621131ef6 |
§H.12 step 3: session_set_pinned/session_is_pinned pin-authority choke point
session_is_pinned() reads Session.pinned directly (authoritative, no Stadium re-derivation); session_set_pinned() writes both Session.pinned and the mirrored STADIUM_FLAG_PIN bit on the session's own patron cell, keeping Stadium's internal eviction/admission logic (which must stay self-contained) in sync without it calling back into session.c. Added Session.stadium_cell (index into stadium_cells()) -- necessary plumbing not in the original H.2 field list; the choke point can't reach the right patron header without it. Moved STADIUM_FLAG_PIN from a stadium.c-private #define to stadium.h (public) so session.c can reference it without a duplicate definition. Verified 3-arch boot to ok> (amd64/aarch64/riscv64). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
c7c9332321 |
Stadium: real block-patron admission + MIGRATE dispatch (FABRIC-3.md §B)
stadium_admit()'s mass==1 refusal looked like a hard blocker for 1024-byte blocks, but stadium_word_dispatch()'s real candidate construction proves Stadium cells carry pure identity/heat/bookkeeping, never the resident's actual content -- a block patron follows the same shape (identity=LBN, payload unused), so this was real, scoped work, not a case for stubbing. New stadium_blocks.h/.c mirror stadium_words.c's admission/cooling shape, keyed by (quota_slot, lbn) in a fixed-capacity open-addressing hash table (tombstone deletion) instead of a dense array, since LBN space isn't densely bounded like word_id. Wired into block_word_block()/buffer()/ update() (block_words.c), __STARKERNEL__-guarded. stadium_dispatch()'s MIGRATE case now calls blk_flush(lbn) for real instead of printing "(stub)". Three new Kconfig constants (STADIUM_BLOCK_HEAT_QUANTUM/ STADIUM_BLOCK_COOL_RATE_Q48/STADIUM_BLOCK_TRACK_CAP_MULT) mirror the word-patron ones, same three-layer wiring. VM-COOL/DELIVER/EXPIRE stay explicit punch-list items -- VM-COOL deferred pending the still-iterating Tripod/Zuse/messaging vision, DELIVER/EXPIRE are their own future subsystem integrations per FABRIC.md's own "open, not resolved" notes. Verified clean compile (zero warnings) and clean boot to REPL with conservation intact (resident_sum + reservoir == Q48_ONE) on all three architectures (amd64/aarch64/riscv64); BLOCK/BUFFER touches exercised live from the REPL with no crash; a 22,000-distinct-block flood loop against an artificially shrunk Stadium ran clean under heavy admission load. A live MIGRATE console fire was not directly observed this session (root-caused to a pre-existing reservoir-floor/density-eviction interaction unrelated to this change, documented in FABRIC-3.md) -- flagged as an honest follow-up, not silently claimed. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CXjAPTEKrgY2Mrk25KoLDn |
||
|
|
89d8c08582 |
stadium: wire STADIUM_CAPACITY_TICK in as a flat threshold, not a scheduler
Closes FABRIC-2.md's last open §12 Q5 question. fleet_heartbeat_tick_count is fed by every live VM's own vm_tick(), not one VM's, so it was reaching HEARTBEAT_INFERENCE_FREQUENCY (shared/borrowed from the per-VM inference gate) several times faster than intended with more than one VM live - backwards from FABRIC.md §22.4's required ~1000:1 separation. What's actually gated turned out to be low-stakes: vm_physics_tick() (capsule_vm_physics.c:397) is a passive statistics refit - re-sorts a window of past heat-transfer samples and recomputes a median rate estimate. It doesn't move heat or arbitrate capacity. Firing too often just meant a noisier statistic recomputed more frequently than planned, not incorrect behavior. Considered and explicitly rejected: scaling the threshold by live VM count at the check site. That's the first brick of a scheduler - reading fleet state to adjust a rate dynamically - which this project has deliberately avoided building. Implemented instead: STADIUM_CAPACITY_TICK (existing Kconfig symbol, defined but never read by any code path) now gates vm_physics_heartbeat_tick()'s call directly, replacing the borrowed HEARTBEAT_INFERENCE_FREQUENCY. Default bumped 1000 -> 4000, a flat constant picked once for Tripod's known 4-VM topology, same kind of placeholder as every other frequency knob in Kconfig.kernel - not computed from anything at runtime. Renamed fleet_last_inference_tick -> fleet_last_capacity_tick to match. Still one clock, one counter (fleet_heartbeat_tick_count) - just a bigger flat divisor on it. Three-arch QEMU acceptance: all clean to ok>, identical Stadium conservation invariant on all three (resident_sum=43691 reservoir=21845 sum=65536). logs/20260815-093425/amd64, logs/20260815-093521/aarch64, logs/20260815-093641/riscv64. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
00e657019e |
stadium: make VM population bound RAM-derived, not a static array of 4
Replaces STADIUM_MAX_VM_COUNT (Kconfig, hardcoded default 4) with a boot-time computation, mirroring the pattern stadium_boot_init() already used for the cell pool. New Kconfig STADIUM_VM_MEMORY_PERCENT (default 50): max_vm_count = (kmalloc_get_stats().free_bytes after the cell array * STADIUM_VM_MEMORY_PERCENT / 100) / VM_MEMORY_SIZE, floored to 1, no ceiling (population is not knowable in advance - could be 4, could be 4000). stadium_quotas and word_slots (plus stat_promotions/stat_evictions) are now kmalloc'd to the computed count instead of declared with a macro. New accessor stadium_max_vm_count() replaces every STADIUM_MAX_VM_COUNT reference, including capsule_birth.c's birth-refusal gate. Two things found and fixed along the way: - The existing cell-pool budget was sourced from pmm_get_stats(), which reflects physical pages PMM hasn't handed to any subsystem yet - but the actual allocation is kmalloc(), which draws from the separate, fixed-size heap kmalloc_init() (M6) already carved out of PMM before stadium_boot_init() ever runs. Budgeting against PMM's leftover and allocating from the kmalloc heap are two different pools. Both the cell budget and the new VM-count budget now source from kmalloc_get_stats() instead. - stadium_owner[] (which VM's quota owns each cell) was uint8_t, capped at 255 slots by a compile-time assert tied to the old macro. Widened to uint16_t (65535 slots of headroom) with a runtime clamp + log if the computed count ever exceeds that, since there's no ceiling anymore. Three-arch QEMU acceptance: all clean to ok>, computed VM count genuinely differs by actual available RAM (amd64/riscv64: 50 slots at -m 1024, aarch64: 101 slots), Stadium conservation invariant identical across all three (resident_sum=43691 reservoir=21845 sum=65536). logs/20260815-080526/amd64, logs/20260815-080826/aarch64, logs/20260815-080952/riscv64. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
5a28458b21 |
starkernel: item 4.2 -- Hermes native on the Stadium (complete)
Migrates Hermes's message/channel lifecycle onto the Stadium's unified heat/capacity economy: MSG-ALLOC/FREE-NODE and CH-ALLOC/FREE-NODE now route entirely through stadium_admit()/stadium_evict(), replacing the old local free-list + independent heat-field mechanism. Eight kernel-only STADIUM-* FORTH primitives (ADMIT, EVICT, RES@, RES-PULL, RES-PUSH, HEAT@, HEAT!, WORD-HEAT), VM.stadium_vm_id threaded through all three vm_core.c dispatch sites (replacing item 4.1's hardcoded vm_uuid_hera()), and the stadium_owner[idx] fix so evict-credit lands in the VM that actually admitted a patron, not whoever owned cell 0. This session's own contribution, on top of that pre-existing implementation: found and fixed two bugs blocking the item's own K≡1.0 conservation self-check (HERMES-K was reading 0, not 65536): - Q.SLOT admission-heat fix (capsules/hermes/init.4th): MSG-SEND/ CH-ACCEPT admitted with Q.1 (the entire fleet-wide "1.0" unit) per item, a leftover from before the Stadium migration when each message/channel had its own unconstrained heat field. Instantly drained the shared, finite reservoir. - Reservoir floor for word-execution admission (stadium_words.c): stadium_word_dispatch() (item 4.1) pulls STADIUM_WORD_HEAT_QUANTUM on every word dispatch, not just first admission -- exhausts a VM's entire reservoir in ~32 dispatches, starving any application-level economy sharing that VM's reservoir before it gets a chance to pull anything. word_dispatch_pull() now clamps word-execution's own pulls to leave a Q48_ONE/3 floor (same fair-share figure COMMON-CH's own floor already uses); application-level pulls are unaffected. - STADIUM-WORD-HEAT primitive + stadium_words_resident_heat(): the floor deliberately leaves word-execution residents holding real heat, invisible to HERMES-K's original formula (MSG+CH+reservoir, no term for word patrons). Adding this term closes K to exactly 65536 on all three architectures. Also rules on two open scope questions in FABRIC.md: MBR-ALLOC/ MBR-FREE-NODE stay off the Stadium (membership records have no heat field, never did -- the acceptance bullet's inclusion of them was a completeness gesture predating a check of the actual layout), and records the effort number (12 implementation files, +759/-120 lines). Verified: all three architectures boot clean, full self-test passes, Stadium conservation closes exactly (resident_sum + reservoir = Q48_ONE) at both the C/Stadium level and the FORTH-level HERMES-K check. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
2981ada2a5 |
starkernel: item 4.1a -- quota granting, Hermes's one-time birth grant
Punch list §25 item 4.1a complete. New prerequisite item, found while scoping 4.2: no quota-granting mechanism existed at all. Adds stadium_grant_quota(new_vm_id, from_vm_id) -- a one-time initial grant at birth, distinct from item 1.3's still-unbuilt recurring capacity-transfer arbitration. Splits the donor's free list evenly by cell count, reassigns stadium_owner[] for every moved cell, and grants the new VM a fresh Q48_ONE reservoir (not a split of the donor's -- per-VM conservation, same pattern as Hera's own boot grant). Wired into every baby VM's birth in capsule_birth.c. Verified via a boot-time self-test in kernel_main.c using a synthetic identity (not the real UUID pool, not a real capsule birth -- item 0.1's Hera-alone pruning stays intact). All three architectures booted to ok> with identical output: grant OK, Hera reservoir=0 (already fully committed to resident words, correctly unchanged), test-vm reservoir=65536 (fresh Q48_ONE). dict_hash identical across all three and unchanged from item 4.1's baseline (0x3d4e1daf289da94f) -- confirms no dictionary word was added. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
3d0b9351bd |
starkernel: item 4.1 -- hot words onto the Stadium, density-ranked eviction
Punch list §25 item 4.1 complete. Replaces the round-robin hotwords cache with Stadium density-ranked admission/eviction on the kernel side, via the §17.7 reservoir mechanism and a kernel-side word_id -> cell_index map (no DictEntry change, dict_hash untouched). Adds stadium_birth_hera() to close the cell-0 panic hazard, STADIUM_WORD_HEAT_QUANTUM/STADIUM_WORD_COOL_RATE_Q48 Kconfig knobs (flagged untuned), and a stadium_word_forget() FORGET coherence hook to close a recycled-word_id aliasing gap. Verified: all five hotwords_cache_* call sites in dictionary_management.c bypassed under __STARKERNEL__; word dispatch feeds the Stadium at all three vm_core.c physics_execution_heat_increment() sites; hosted make unaffected; all three architectures booted to ok> with matching dict_hash (0x3d4e1daf289da94f) and matching conservation stats (promotions=354 evictions=0, resident_sum=65536 reservoir=0 sum=65536). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
9b305a5be7 |
starkernel: item 3.8 -- VM identifiers as UUID/GUID
Punch list §25 item 3.8 complete. Added after starting item 4.1
surfaced the need to thread a vm_id into stadium_admit()'s new quota
parameter; Captain Bob ruled UUID/GUID rather than keeping the
narrower uint32_t.
New VMUuid type (vm_uuid.h/vm_uuid.c): two uint64_t halves, RFC-4122-
shaped for logging. Not real randomness -- checked directly against
QEMU 10.2.1's actual CPU feature set: amd64 RDRAND and riscv64 Zkr are
both real, available features here; aarch64 has no RNG property on any
CPU model including "max" (verified exhaustively via QMP
query-cpu-model-expansion). Captain Bob ruled a uniform fallback
across all three ISAs rather than a per-architecture split.
Fallback is a deterministic PRNG (splitmix64) seeded from the Mama
capsule's content hash, pre-filling a 16-entry FIFO pool at boot and
refilling with another batch of the same stream when exhausted --
exactly the shape requested. Same capsule booted twice produces the
same id sequence, preserving the dict_hash reproducibility this
session has relied on throughout.
Hera keeps a fixed, reserved all-zero id, not drawn from the pool --
capsule_birth.c uses vm_id == 0 as a load-bearing sentinel in three
places (KILL protection x2, fleet heat-fanout parent-chain
terminator), found by reading before writing any code.
Two real sentinel-collision bugs caught before shipping, same class as
STADIUM_CONTAINS_NONE: vm_uuid_none() (all-ones, not all-zero) for
"not yet assigned"/"no VM" placeholders; confirmed item 3.7's quota
table already used an in_use boolean rather than a vm_id sentinel, so
no second collision was actually possible there -- the dead,
never-referenced STADIUM_QUOTA_SLOT_EMPTY macro was removed.
Blast radius larger than first scoped, flagged mid-work rather than
silently absorbed: capsule_vm_physics.c/.h (the fleet heat-transfer
layer item 2.1 modified earlier this session) has its own vm_id-keyed
node table and walks parent_vm_id chains through the same identity
space, so it needed the same change, plus its callers in
mama_forth_words.c and sk_vm_bootstrap.c.
One live FORTH word contract changed, by explicit ruling: CAPSULE-BIRTH
was ( capsule-id -- vm-id ), a single cell -- can't hold 128 bits.
Captain Bob picked pushing two cells ("there is doubles support in the
FORTH std word set anyway"): ( capsule-id -- vm-id-hi vm-id-lo ).
MAMA-VM-ID changed the same way: ( -- 0 0 ).
Verified: full (not standalone-file) kernel rebuild to catch cross-file
breakage given the size of this change -- it surfaced the
capsule_vm_physics.c blast radius a narrower check would have missed.
Three-architecture boot (amd64, aarch64, riscv64), all reaching ok>
with identical dict_hash=0x3d4e1daf289da94f matching the item-3.7
baseline.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
|
||
|
|
e55111c2c5 |
starkernel: item 3.7 -- per-VM free lists (Phase 3 core complete, for real)
Punch list §25 item 3.7 complete. Added to §25.4 after starting item 4.1 surfaced it as an unbuilt prerequisite -- 3.6's earlier "Phase 3 core complete" claim is corrected in this same commit. StadiumVMQuota table (size STADIUM_MAX_VM_COUNT, linearly searched by vm_id -- capsule_birth.c's vm_id is monotonic and never reused, so it cannot index a table directly, and a 4-entry scan costs nothing). New per-cell stadium_owner byte array records which quota a cell belongs to, needed so eviction returns a freed cell to the correct VM's list and so eviction search stays scoped to the evicting VM's own residents (quota isolation). Free-list linkage reuses each cell's `link` field as a next-free pointer while unresident -- link is documented only as generic "index into the Stadium, not a pointer," so this is a repurposing, not a header change. Does not answer the separate, still-open question of which field carries a multi-cell patron's first continuation-cell index; item 3.5's mass != 1 refusal stands exactly as it was. Boot-time: every cell chained into one list in ascending index order, granted whole to vm_id 0 (Hera), the only VM that exists. Ascending order preserves item 3.6's "Hera is patron zero" invariant once real birth-wiring lands. stadium_admit()'s signature changed to take vm_id -- a change to code shipped in item 3.5, amended there. Pops the calling VM's free-list head first (O(1)); only falls back to a same-VM-scoped eviction search if empty. Caught a real bug before the boot run: the header zero-fill on eviction (and the initial free-list build) both left contains == 0, but 0 is Hera's valid index -- the same collision item 3.1's STADIUM_CONTAINS_NONE fix addressed, recurring at a new site. Fixed by explicitly setting contains = STADIUM_CONTAINS_NONE at both free-list sites. Explicitly out of scope, reported not invented: granting quota to any VM other than Hera is capacity arbitration (item 1.3 left "how much moves per transfer" open). stadium_owner is set once at boot and never rewritten, so quota_slot_for_vm() refuses every vm_id != 0 permanently until item 4.2 adds the grant path and owner-array writes. Verified: three-architecture boot (amd64, aarch64, riscv64), all reaching ok> with identical dict_hash=0x3d4e1daf289da94f matching the item-3.6 baseline. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
72487e7fff |
starkernel: item 3.6 -- Hera as patron zero, pinned (Phase 3 core complete)
Punch list §25 item 3.6 complete. Phase 3 (§25.4) core is now done: items 3.1-3.6 all closed. stadium_evict() now panics via sk_hal_panic() if a resident cell 0 (Hera, patron zero by construction of §6's boot order) is ever selected for eviction. Placement is deliberate: the check runs before the pin/contains refusal checks, not after -- if it ran after, a wrongly-cleared pin would let the ordinary refusal path quietly return -1 instead of ever reaching the panic, defeating the point of a check that's supposed to be independent of pin holding. Per §20.5 #3's explicit wording, not implemented as a filter: stadium_admit()'s least-dense search is unchanged, still relying on the general pin skip from item 3.5. Adding a second filter there would have done exactly what that section warns against ("filtering hides the bug, asserting reports it"). The panic path is, and will remain, unexercised by the acceptance mechanism: sk_hal_panic() halts the machine, and triggering it deliberately is incompatible with the three-arch boot being this project's sole acceptance test. Correctness rests on the placement argument, not a test -- same honesty precedent as items 3.4 and 3.5's other unexercised paths. Verified: three-architecture boot (amd64, aarch64, riscv64), all reaching ok> with identical dict_hash=0x3d4e1daf289da94f matching the item-3.5 baseline. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
f8a50561b0 |
starkernel: item 3.5 -- admission and eviction
Punch list §25 item 3.5 complete. stadium_admit(candidate) places into an unused cell if one exists (no comparison needed), otherwise finds the least-dense resident -- skipping pinned and contains-gated patrons, which are never eviction candidates -- and evicts it only if the candidate is strictly denser, per §19.3. stadium_evict(cell_index) dispatches the departing patron's behaviour before clearing its slot, per §17.2. Caught a real bug before it ran: the first draft used contains == 0 to mean "holds nothing," but cell index 0 is a valid index (Hera, item 3.6). Fixed with a proper sentinel, STADIUM_CONTAINS_NONE (UINT32_MAX). A second-pass review found mass was not accounted for: both functions handled exactly one cell regardless of the candidate's stated mass, which leaks cells on eviction of any mass > 1 patron and breaks capacity conservation. Fixed by refusing any candidate with mass != 1 -- multi-cell patrons need the per-VM free lists item 3.2 already deferred (§22.3), not built here. Documented, not fixed: the discriminator bitmap can't distinguish free from continuation cells, so the free-cell scan reads continuation-cell payload bytes under the header layout -- latent since nothing creates continuation cells yet, and the mass != 1 refusal keeps it provably latent. Superseded by the free list when it exists. Unexercised at runtime: nothing calls either function yet (no real patron kind is wired to the Stadium). No self-test added -- filling ~74,000+ cells to reach the eviction-on-full branch was judged impractical, following item 2.2's own precedent for its unexercised fleet-full path. Verified: three-architecture boot (amd64, aarch64, riscv64), all reaching ok> with identical dict_hash=0x3d4e1daf289da94f matching the item-3.4 baseline. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
0b47c256fc |
starkernel: item 3.4 -- density ranking
Punch list §25 item 3.4 complete. stadium_density(cell_index) reads a header's heat and mass and returns heat / mass -- a division on demand from fields already stored in the cell, matching §19.3's "read, not computed by a scheduler" literally. Stays valid Q48.16 without a special fixed-point routine, since heat is already Q48.16 and mass is a plain integer divisor. mass == 0 and an out-of-range cell_index both return 0 rather than dividing by zero -- an empty or never-admitted slot has no footprint to be dense within. Deliberately not built here, per the item's own wording: finding the densest or least-dense resident (§19.3's admission/eviction comparison) is item 3.5's scope, not this one's. Nothing calls stadium_density() yet either. Verified: three-architecture boot (amd64, aarch64, riscv64), all reaching ok> with identical dict_hash=0x3d4e1daf289da94f matching the item-3.3 baseline. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
378d688898 |
starkernel: item 3.3 -- behaviour enumeration and dispatch
Punch list §25 item 3.3 complete. StadiumBehaviour (stadium.h) enumerates exactly the four tags §18.3 already names -- MIGRATE, DELIVER, EXPIRE, COOL -- mapped from §17.1's patron table: blocks->MIGRATE, messages->DELIVER, ACLs->EXPIRE, words and VMs both->COOL. Nothing invented; the tag set and mapping were already in the document. stadium_dispatch(cell_index, behaviour) dispatches on the tag only, never asks what kind of patron departed. Handlers are stubs -- the real actions belong to subsystems not yet migrated onto the Stadium (Phase 4). Nothing calls stadium_dispatch() yet; item 3.5 is its first consumer. The switch is exhaustive with no default case, making §13's "closed enumeration, fixed at build time" a compiler-enforced property under this project's -Wall -Werror rather than just prose. Verified live: temporarily deleted the COOL case, rebuild failed with error: enumeration value 'STADIUM_BEHAVIOUR_COOL' not handled in switch [-Werror=switch], restored it, confirmed clean again. The header's behaviour field stays uint8_t, not the enum type itself, since C does not guarantee an enum's underlying type and that field's offset is load-bearing for item 3.1's validated 64-byte layout. Verified: three-architecture boot (amd64, aarch64, riscv64), all reaching ok> with identical dict_hash=0x3d4e1daf289da94f matching the item-3.2 baseline. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
eb0fd4fffa |
starkernel: item 3.2 -- Stadium boot-time allocation
Punch list §25 item 3.2 complete. stadium_boot_init() (src/starkernel/vm/stadium.c) sizes the global cell array at boot from a real memory-budget query rather than a hardcoded count: pmm_get_stats().free_bytes at the point of allocation, times the new STADIUM_MEMORY_PERCENT Kconfig symbol (default 1%), rounded down to whole 64-byte cells. Matches §17.6's position (b) literally. Also allocates the header/continuation discriminator bitmap item 3.1 declared but did not allocate. Both are kmalloc'd and explicitly zero-filled (kmalloc does not zero). Called from kernel_main.c immediately before sk_vm_bootstrap_parity(), i.e. before any VM exists (§6). Failure is soft -- logs and continues, does not halt boot -- matching the existing precedent one line below it (VM bootstrap parity failure does the same). Added a "Stadium: N cells (M KB)" boot console line at the allocation site so the acceptance logs are evidence the array was actually allocated, not just that the kernel still boots -- the same blind spot item 3.1's uncompiled-header gap exposed. Verified: three-architecture boot (amd64, aarch64, riscv64), all reaching ok> with identical dict_hash=0x3d4e1daf289da94f matching the item-3.1 baseline, and the Stadium boot line confirmed present in all three serial logs (amd64: 74234 cells/4639 KB, aarch64: 161329 cells/10083 KB, riscv64: 76122 cells/4757 KB). Not built here, reported per §25.0 rule 3: per-VM free lists (§22.3) -- granted when Hera assigns quota, not this item's scope. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
1b2f0677de |
starkernel: item 3.1 reopened -- two Kconfig symbols items 1.1/1.4 deferred here
Punch list §25 item 3.1 re-closed after reopening. Items 1.1 and 1.4's resolutions both explicitly named this item as where their Kconfig symbols would be implemented, but 3.1's own stated scope never mentioned them, so the first close missed both: - STADIUM_CONTAINS_DEPTH_MAX (default 5) -- item 1.1's contains-chain depth cap. No consumer yet; reap-gating enforcement is item 3.5. - STADIUM_CAPACITY_TICK (default 1000) -- item 1.4's capacity arbitration cadence in virtual ticks. No consumer yet; capacity arbitration itself is not on the punch list. Both added following STADIUM_MAX_VM_COUNT's exact pattern: Kconfig.kernel entry, Makefile.starkernel kconfig_int + VM_FEATURE_FLAG_VARS forwarding, starforth_config.h fallback default. stadium.h now includes starforth_config.h and carries two more C99-portable compile-time checks proving both symbols are defined and sane, same discipline as the byte-count checks. Declaration only -- not inventing the consuming logic to close this out early. Verified: three-architecture boot (amd64, aarch64, riscv64), all reaching ok> with identical dict_hash=0x3d4e1daf289da94f, re-run after the reopening. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
d55ec3241b |
starkernel: item 3.1 -- the Stadium cell and header
Punch list §25 item 3.1 complete. Defines StadiumPatronHeader and StadiumContinuationCell in new include/starkernel/vm/stadium.h, unioned as StadiumCell per §3's closed two-valued union. src/starkernel/vm/stadium.c added to Makefile.starkernel's LOADER_EXTRA_SRCS/KERNEL_EXTRA_SRCS so the header's compile-time size checks are actually compiled, not merely included by something that never builds. Discriminator ruled an external side bitmap (Captain Bob), not a header field -- amended into §3 and §23.3 before this code was written. Item 3.1 declares the bitmap's purpose/indexing in a comment only; allocating it is item 3.2's scope. Both cell shapes counted for real at exactly 64 bytes with zero compiler-inserted padding (three C99-portable negative-array-size assertions -- no _Static_assert, this project targets C99). Header matches §23.3's original 32+32 split unchanged, since the discriminator moving outside the cell left nothing to compete for that space. Continuation cell matches item 1.12's 4+60 figure unchanged for the same reason. Verified the size assertion is actually live: broke it to 63, confirmed the build failed with the expected negative-array-size error, restored it, confirmed a clean compile. Verified: three-architecture boot (amd64, aarch64, riscv64), all reaching ok> with identical dict_hash=0x3d4e1daf289da94f matching the item-2.2 baseline. Confirmed stadium.o present in both obj/loader/vm and obj/kernel/vm post-build on amd64, closing the gap the item-2.2 WIP exposed (an uncompiled header proves nothing). Left open, not fabricated: §23.4 #2 ("does a typical message fit in one cell") is unanswerable today -- no message patron struct exists anywhere in this tree yet. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
a5ed8c3d87 | Initial commit — LithosAnanke kernel |