a4ad14aec644c6c12f2e926403593a733e33722d
100
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
a4ad14aec6 |
Phase 5 close-out: Isabelle pass, doc sweep, SBOM, version bump (FABRIC-3.6.md tasks 5.1-5.4)
5.1: Isabelle/HOL pass (52 theories, clean) -- restated the boundary rather than just citing the green build: proof/ scope was already entirely outside this reshuffle's footprint (src/starkernel/, capsules/*.4th), so the boundary is unchanged, not moved. 5.2: Documentation sweep. CLAUDE.md's stale WIP banner and Tripod fleet description updated now that Phases 0-4 have actually landed (Hera/ Artemis/Hestia, no Hermes). MANIFEST.md rides the strip -- hermes/init.4th's block table replaced with a deletion note, init.4th/doe-campaign.4th/ hestia/init.4th entries corrected to match the post-strip live files. Confirmed the TRIPOD.md/0.1 contradiction was already resolved (2026-08-13). Settled the superseded-docs call explicitly: archive as-is, do not rewrite. Fixed experiments/bare_metal/README.md's block-size framing (still said 1024-byte budget; real rule is 64 chars x 16 lines). K-qualification checked clean against the two living documents; full retroactive sweep of the closed archival FABRIC corpus explicitly declined as disproportionate. 5.3: make sbom. Installed syft (user-local, approved). Found and fixed a real Makefile bug while at it -- the sbom target hardcoded --source-name StarForth, so DocumentName was wrong even after regenerating. 5.4: LITHOS_VERSION 2.0.0 -> 2.1.0, engine VERSION 3.1.0 -> 3.2.0 (minor, per the dictionary-visible-only rule). Replaced the stale version-comment block in Makefile.starkernel (had the odd/even LTS rule backwards) and docs/lithosananke/ROADMAP.md's retired versioning-policy section with the ratified ladder. Verified on all three architectures; riscv64's first pass hit a transient virtio_blk timeout during boot-time Zuse genesis mint, reported and confirmed non-reproducing on an immediate clean retry. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
3e201c82a5 |
Stage E: Category B strip -- remove Hermes, messaging.4th, old routing (FABRIC-3.6.md task 4.1-4.4)
All FORTH-owned message types were cut over to kernel-Hermes in Phase 3 (tasks 3.8-3.10). This removes the now-dead FORTH messaging layer and the Hermes VM itself: capsules/common/messaging.4th, capsules/hermes/init.4th, the slot-3 VM-NAME-REG pairing convention, and every load-site/birth-site reference across capsules/init.4th, artemis/init.4th, hestia/init.4th, doe-campaign.4th (Artemis-only now), capsule_console.c, capsule_mint.c, capsule_wirebind.c, capsule_birth.c, and kernel_main.c. Verified on all three architectures: clean boot, mkcapsule --lint clean (36 files, 0 violations), zero UNKNOWN WORD, identical dict_hash across amd64/aarch64/riscv64 for every VM, and a full mint -> WIREBIND-attach -> USE -> relay round-trip exercising the two highest-risk edits (capsule_console.c/capsule_mint.c). Found, not fixed: deleting messaging.4th removes SEND-ELEVATE-REQUEST, which was the only caller of KH-ELEVATE-SEND and the only path to ELEVATE-GRANT (zuse-eligibility.4th, still loaded at boot) -- Phase 8 PKI's own elevation entrypoint. Needs a decision before Phase 5 close-out. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
cab9b5a31f |
Stage D: ELEVATE-REQUEST real cutover -- FABRIC-3.6.md task 3.10
Reachability verified live before writing any code, per this project's own standing rule (grep cannot establish reachability alone): FIND SEND-ELEVATE-REQUEST / FIND ELEVATE-GRANT / FIND CH-REQUEST all resolve on a live Hera boot, though grep across capsules/experiments/docs found zero callers of SEND-ELEVATE-REQUEST -- a real, complete, directly-callable entrypoint (H.5/H.8's own design) with no current automatic trigger, not dead code. Correction to a prior finding, made in the course of this check: task 3.8's write-up claimed "Hera's own pre-existing inability to load common:messaging.4th" -- false. capsules/init.4th (Hera's own MAMA_INIT capsule) loads it directly, and SEND-ELEVATE-REQUEST lives and works in her dictionary right now. Task 3.8's own actual scope is unaffected by this correction. Cutover: SEND-ELEVATE-REQUEST (messaging.4th) no longer ends in CH-REQUEST's COMMON-CH/MSG-SEND path; it now calls KH-ELEVATE-SEND (repl.c), a new C word wrapping sk_hermes_send_one(), registered unconditionally for every VM. from/to are derived from the calling VM and sk_get_mama_vm() directly in C, never taken from the stack -- a real correctness improvement over CH-REQUEST's own initiator-only gate, which only existed because a caller COULD pass the wrong from value; deriving it in C makes that spoof structurally impossible. SK_HERMES_MSG_TYPE_ELEVATE_REQUEST reuses ELEVATE-REQUEST's own value (8), same partition-rule reasoning as tasks 3.8/3.9. Delivery is unchanged task 3.4 machinery. No new static-buffer lifetime caveat -- the payload-aliasing fix landed before this task started. CH-REQUEST (messaging.4th) is now dead code, its one real caller just removed -- found, not fixed, per Captain Bob's Law. Verified live on all three architectures: 0 0 0 0 S" DUP" SEND-ELEVATE-REQUEST (deliberately-invalid pubkey, so ELEVATE-GRANT correctly refuses -- the check is the pipeline running, not a grant succeeding) fires the evidence line and completes cleanly, DUP unaffected afterward. Zero UNKNOWN WORD, dict_hash identical across all three architectures (changed uniformly from prior runs -- one new word registered -- not diverged, matching SXXXIV.6's own rule). This closes task 3.11 (Phase 3 gate): all of messaging.4th's live FORTH-owned message types (BLK-ATTACH-EVENT, CONSOLE-CMD-EVENT, ELEVATE-REQUEST) are now real kernel-Hermes cutovers. Phase 4 may begin. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
d4f8a568ee |
Silence routine diagnostic noise; fix a real CRLF double-submit bug
Two separate fixes, found and closed together after task 3.9/prompt-
format landed (Captain Bob: "finish the work first then we'll do
cleanup before 3.10 begins").
1. Logging noise (confirmed live via QMP screendump): three sources
were cluttering ordinary interactive console sessions.
- INFERENCE "Output validation failed, ignoring results"
(vm_runtime.c, kernel; vm_time.c, hosted mirror) and xhci "CSW
status = FAILED"/"unit not ready -- retrying" (xhci.c) were
already log_message(LOG_WARN/LOG_ERROR, ...) calls, just visible
at the default runtime LOG_WARN level -- downgraded to LOG_INFO,
all three fire routinely and self-resolve (xhci.c's own existing
comment already documents the retry as expected SCSI UNIT
ATTENTION behavior, not a driver defect).
- Stadium: dispatch cell=... (stadium.c's stadium_dispatch()) was a
genuine defect: an unconditional console_puts()/console_println()
sequence with no level gating at all, printing on every single
dispatch. Rewritten through log_message(LOG_DEBUG, ...).
2. CRLF double-submit (found while investigating why the cleaned-up
noise still didn't look like a normal single-VM session):
sk_console_readline() (repl.c) breaks on '\r' OR '\n' as independent
terminators, so a line sent as both bytes submits twice -- the real
line, then an immediate empty-line submit on the second byte, each
printing its own " ok". Pre-dates Stage D entirely, not an async-
relay artifact. Fixed with a single non-blocking peek-and-discard
for the paired byte right where the line terminates.
Verified live on all three architectures for both fixes: zero UNKNOWN
WORD, dict_hash identical to every prior acceptance run in this
document (both changes are display/interaction-only).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
|
||
|
|
5c07745c18 |
Fix SkHermesMessage payload-aliasing defect (task 3.8 findings log)
sk_hermes_send_one() -- the single funnel every sender, including sk_hermes_publish(), already goes through -- used to store the caller's own payload_addr pointer as-is. Two sends before either drains meant both messages pointed at the same caller-owned buffer, whichever send wrote last silently winning: real, confirmed live (a second identity thumbdrive attached at boot alongside Zuse's own left its WIREBIND pairing silently never happening). SkHermesMessage gains an inline payload_buf[SK_HERMES_CHUNK_MAX_ PAYLOAD] field; sk_hermes_send_one() now memcpy()s the caller's payload into it and points payload_addr at that copy instead. No sender or reader call site needed to change -- every existing reader already only ever reads through payload_addr, which still points at valid bytes of the same length, now message-owned. g_kh_blk_attach_buf/ g_kh_console_cmd_buf (repl.c, tasks 3.8/3.9) no longer need to survive past their own send call; their doc comments, which had claimed the old aliasing shape was benign, are corrected. Verified the original bug is actually gone: reproduced the exact original scenario (a real, sequentially-minted rajames identity attached at boot alongside Zuse's own, two simultaneous BLK-ATTACH- EVENT sends in one idle-loop pass) -- WIREBIND pairing, USE, and the Stage D relay all now work where WIREBIND previously silently failed. Standard three-ISA acceptance also clean: zero UNKNOWN WORD, dict_hash identical across all three and matching every prior acceptance run in this document. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
21734b1a52 |
Console prompt: identity/machine both sides once redirected off Hera
sk_console_user_prefix() (repl.c) previously always returned "zuse" (or the WIREBIND-attached username) as the left side of the bracket prefix, regardless of which VM the console was actually pointed at -- "[zuse@rajames]" after USE rajames, always showing the authenticating superuser rather than the active identity. Changed on Captain Bob's direct instruction: once the console is redirected into a WIREBIND identity's own console VM (console_get_vm_name() != "Hera"), show that same name on both sides -- "[rajames@rajames]" -- since WIREBIND births the console VM literally named after the identity, so the identity IS that VM, not a separate label. At the top level (still on Hera, nothing has redirected yet), the original zuse_session/WIREBIND-username logic is unchanged. An earlier, more ambitious attempt (separate identity/machine tracked state across every console_set_vm_name() call site) regressed live to a wrong [zuse@Artemis] prompt and was fully reverted before reaching any acceptance run -- the landed fix needed none of that new state, just this one function. Verified live on all three architectures: [zuse@Hera] at the top level and after a live WIREBIND attach (before USE), [rajames@rajames] after USE rajames, with task 3.9's Stage D relay still firing correctly on top of it. Zero UNKNOWN WORD, dict_hash unaffected (display-only change). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
3975dc61cd |
Stage D: CONSOLE-CMD-EVENT real cutover -- FABRIC-3.6.md task 3.9
Send-side cutover: sk_repl_dispatch_line()'s FORTH-string
"CONSOLE-CMD-EVENT 0 3 S\" ...\" 0 MSG-SEND" interpret is replaced with
a direct sk_hermes_send_one() call (SK_HERMES_MSG_TYPE_CONSOLE_CMD,
kernel_hermes.h, deliberately reusing CONSOLE-CMD-EVENT's own value 7,
same partition-rule reasoning as task 3.8's BLK-ATTACH-EVENT cutover).
Real finding along the way: kernel-Hermes's drain only ever runs as a
side effect of vm_interpret() being called on the target VM. Task
3.8's target (Hera) is always being interpreted via the interactive
REPL loop; task 3.9's target is a WIREBIND identity's own ~user VM, a
passive receiver nothing else drives. The existing idle-loop pump only
ticked VMs with the old FORTH MSG-TICK word ACL-allowed -- a VM minted
with the STD79-lockdown personality never has it, so the pump silently
skipped it forever and queued messages never delivered. Fixed by
adding an unconditional, direct sk_hermes_drain_checkpoint() call in
the same pump loop, independent of the MSG-TICK gate.
Verified live on all three architectures: WIREBIND-attach a real
identity, USE into it, type a plain console line, confirm the Stage D
evidence line and correct relayed execution result. Zero UNKNOWN WORD,
dict_hash identical across all three ISAs and matching task 3.8's own
baseline. The FABRIC-3.md-documented USE/BINDSTEP crash did not
reproduce in any of these live sessions (recorded as a finding, not
chased further).
Depends on the zuse_root_pubkey_known fix already landed in
|
||
|
|
ac4d431aee |
Fix zuse_root_pubkey_known never set after a same-session genesis mint
capsule_zuse_boot_load_root_pubkey() is the only writer of zuse_root_pubkey_known, and it only ever runs once, synchronously, at Artemis's boot-time virtio-blk attach -- before genesis mint has happened on a fresh artemis.img, so it finds no marker yet and leaves the flag 0. Nothing re-triggers it after genesis mint completes. Silent result: capsule_wirebind_try_attach()'s own `if (!mama_vm->zuse_root_pubkey_known) return;` gate then refuses every identity attach for the rest of that boot, with no message at all -- meaning WIREBIND was silently dead on any boot that reset artemis.img fresh, which is exactly what this project's own acceptance convention does before every run. Found live verifying FABRIC-3.6.md task 3.9's console-session acceptance. Fixed at the source: install_and_activate() already has the root pubkey in hand (from genesis mint or an already-Zuse re-attach) -- activate it directly there instead of relying on a disk re-read that may never happen in-session. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
7944abca36 |
Findings: task 3.8 payload-aliasing defect found while starting 3.9
While orienting for task 3.9 (CONSOLE-CMD-EVENT cutover, per Captain Bob's ruling to verify via a real QMP/serial-socket console session), booting with a second real identity drive attached alongside Zuse's own exposed a real defect in the already-closed task 3.8 code: g_kh_blk_attach_buf (repl.c) is one static buffer, and sk_hermes_send_one() stores payload_addr as a caller-owned pointer, not a copy. Two real USB-MSC attaches in one sk_repl_idle() pass send before either drains, so both messages alias the same buffer -- only one identity ever completed. Task 3.8's own write-up claimed this "carries the same single-buffer- reuse shape ATTACH-ACK-BUF itself already had... not a new hazard" -- that claim was wrong and is amended in FABRIC-3.6.md's findings log. FORTH's own MSG-SEND had the identical pointer-aliasing shape but never hit the window: MSG-TICK drained from the same sk_repl_idle() pass that queues attaches. Kernel-Hermes drains at interpret checkpoints, which don't fire during that pass at all -- the cutover changed not just how delivery happens but when, opening a window FORTH's own design never had. Reported, not fixed here, per Captain Bob's Law: this is a defect in closed task 3.8 code, found while scoping a different task. A real fix changes SkHermesMessage's own shape to own its payload bytes rather than reference a caller's pointer -- bigger than a repl.c patch, touches every existing sender, needs its own three-ISA acceptance. Task 3.9's own send-side cutover (kernel_hermes.h, repl.c) is written but deliberately left uncommitted -- it would inherit the identical defect shape if shipped now. logs/20260922-111840/amd64/ is the reproduction. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
bbfd9103f4 |
Stage C: cut over BLK-ATTACH-EVENT alone -- FABRIC-3.6.md task 3.8
The reply leg (Artemis -> Hera ack) that used to flow through common:messaging.4th's MSG-SEND/MSG-TICK now goes through kernel-Hermes's sk_hermes_send_one()/sk_hermes_drain_checkpoint() instead -- FORTH Hermes never sees a BLK-ATTACH-EVENT message again (SXXXIV.2's partition rule). The request leg was never real FORTH messaging traffic to begin with (a direct VM-EXEC, no type tag, forced by Hera's own inability to load common:messaging.4th), so it is untouched. New KH-BLK-ATTACH-SEND (repl.c) wraps sk_hermes_send_one(), reached from capsules/artemis/init.4th's HERA-BLK-ATTACH-REQ. Delivery reuses task 3.4's already-wired sk_hermes_drain_checkpoint(); BLK-ATTACH-ACK itself is unchanged, just reached by a different layer. SK_HERMES_MSG_TYPE_BLK_ATTACH deliberately reuses BLK-ATTACH-EVENT's own value (9) to document this as a cutover of the same message, not a new one. Two real bugs found on the way, both recorded in FABRIC-3.6.md's findings log: - A popped FORTH CREATE-buffer address was raw-cast to a host pointer instead of going through vm_ptr() -- silently read all-zero memory, no crash, no error, just a message that arrived and did nothing. Fixed; the rule and its exception (repl.c's own dev-addr is legitimately a raw pointer, formatted that way by its own pushing code) are written up for the next FORTH-facing C word. - A separate, genuine hang on the very first live exercise of this path, never reproduced across ten subsequent boots. Reported, not chased -- not blocking, per the task's own check being otherwise fully satisfied. Also found live: log_message() is invisible in this build's actual serial-log capture at every level -- settled on a single console_println in the real drain target instead, one line per real USB attach, not a hot-path. Final acceptance (logs/20260922-105501, -105758, -110304, disk images reset before each): dict_hash identical across all three architectures for every VM, zero UNKNOWN WORD, mkcapsule --lint clean, real ledger+stadium_conserved(Artemis)=true evidence on every boot. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
f37aa0fb17 |
Channel-open policy hook -- FABRIC-3.6.md task 3.7
Added HERMES-CHANNEL-OPEN? ( req-hi req-lo -- allow? ) at capsules/ACL.4th block 4008 (default: approve everything) -- the one word policy authors edit. sk_hermes_channel_open_policy(VM*, VMUuid) (kernel_hermes.h/.c) is the C-side query that calls it via plain word-dispatch against the target VM's own dictionary/stack, never vm_interpret() (avoids task 3.4's input-buffer cursor hazard entirely) and never decides the answer itself. Fails closed: no policy word, a policy error, or stack underflow all deny, matching CLAUDE.md's posture that absence of policy must never mean "always allow." Two real bugs found and fixed before this was called done: missing current_executing_entry assignment before calling the word's func pointer (colon words silently no-op without it, vm_core.c:730 -- no crash, just a wrong answer); and a second FAIL with debug instrumentation still in place whose precise cause isn't reconstructable, since no intermediate commit exists for that attempt. Self-test proves the task's check four ways against the same unchanged C function: default approve, live redefinition to deny (zero C change), restore, and a VM with no ACL.4th loaded at all (fail closed). A fifth check wires the result into task 3.6's sk_hermes_channel_respond() end to end: a denied policy produces a NACK and no channel, ledger/stadium_conserved() holding throughout. Scope, per Captain Bob's ruling: closes with the query built and proven; sk_hermes_channel_respond() still takes a caller-supplied approved bool rather than calling the policy internally. Wiring a real channel-open call site to only this query is deferred to whichever later task first needs a live decision. dict_hash identical across amd64/aarch64/riscv64 for every VM, zero UNKNOWN WORD, mkcapsule --lint clean (38 files, 0 violations). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
205a49ecd0 |
ACK/NACK and private-channel negotiation -- FABRIC-3.6.md task 3.6
Extracted sk_hermes_send_one() from sk_hermes_publish()'s own
per-subscriber body -- one code path for both point-to-point and
fan-out delivery, so the ledger can never diverge between them.
Point-to-point addressing turned out to be load-bearing, not
incidental: sk_hermes_publish()'s fan-out sets msg->to to whichever
member it is iterating, so a negotiation message "published" to the
common channel would spuriously reach every member, not just the real
target (checked with advisor() before building the naive version).
"Over the common channel" means every VM is reachable from birth (task
3.2), not that the exchange itself fans out -- messaging.4th's own
CH-REQUEST carried an explicit `to` for the same reason.
sk_hermes_channel_request/respond/close build the mechanics: request ->
grant (creates a private channel, subscribes both parties, sends
CH_GRANT + one ACK) or NACK ("a deny is a NACK", SXLV.1 -- no separate
type); close authorized by membership alone. The grant/deny decision is
a plain caller-supplied `approved` bool -- task 3.7 replaces the call
site that produces it with a real ACL.4th query, not this signature.
Self-test covers the task's own three checks plus a sibling advisor()
flagged: an approved respond() whose channel creation itself fails
(table exhausted) must still fall through to NACK, not a silent false
grant or half-open channel -- verified by exhausting the whole channel
table and confirming the fallback.
Bug found and fixed before this was called done: the first draft
dropped a message via pending_pop() alone, without releasing it first,
leaking its Stadium heat and failing the self-test's own ledger
baseline check (logs/20260922-065946/amd64/, kept as audit trail).
Fixed and re-verified PASS on all three architectures.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
|
||
|
|
8a2ee0fdad |
Payload bound and chunking -- FABRIC-3.6.md task 3.5
sk_hermes_publish() now enforces SK_HERMES_CHUNK_MAX_PAYLOAD (1024) on every message's payload_len uniformly, chunked or not -- closing the gap task 3.3 explicitly parked. A chunk carrier is [SkHermesChunkHeader][content slice], slice capped at 1024 - sizeof(header) rather than 1024 itself, so every message on the wire satisfies the same one-block bound vm_interpret()'s own drain limit already requires -- a future chunk-aware drain never has to special-case a carrier that can't be handed to vm_interpret() as-is. Deliberately no chunking-sender API: building one would need kernel-Hermes to own chunk-buffer memory with a real lifetime it has no way to track (kept alive until every subscriber drains it). Sending is a loop pattern a caller writes with sk_hermes_chunk_count() + sk_hermes_publish(), demonstrated by this task's own self-test. sk_hermes_reassemble() is pure and memory-agnostic: validates msg_id agreement, exact seq coverage, and per-chunk slice sizes before a single memcpy, with the total length computed once and checked against the caller's buffer once -- never order-dependent on which chunk happens to overflow. Verified live on all three architectures: a 1024-byte payload as one message, a 1025-byte send refused outright with the ledger untouched, and a 3000-byte payload split into 3 chunks, drained, and reassembled byte-exact against the original. dict_hash unmoved and identical across architectures. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
2b1ba031a5 |
Drain at the outermost checkpoint -- FABRIC-3.6.md task 3.4
sk_hermes_drain_checkpoint() interprets one queued payload per checkpoint (ruled: one message per checkpoint), reusing sk_vm_at_outermost_interpret() and placed before the switch-signal block in vm_core.c's existing cooperative checkpoint (sk_vm_context_switch() doesn't return until switched back to, so drain must come first or it silently never runs on a switching checkpoint). Amends FABRIC-3.5.md SXLIII.5, caught by advisor() before writing the naive version: "recursive drain is prevented for free" via g_vm_interpret_depth is true but only for same-message re-drain -- it doesn't cover the separate same-VM reentrancy hazard FABRIC-3.md SXX already named for Hera specifically (VMCallState saves rsp/exit_colon/ ecw_nesting only, never input_buffer/input_length/input_pos). Draining calls vm_interpret() on the same vm whose own vm_interpret() call is still paused mid-word at the checkpoint; without saving and restoring the cursor by hand, the enclosing REPL line or LOAD block would be silently truncated. sk_hermes_drain_checkpoint() snapshots and restores input_buffer/input_length/input_pos/mode/error/abort_requested around the call. Not a divergence from the ruling -- cursor preservation is the implementer's own obligation inside the ruled mechanism. Gated behind a system-wide pending-total counter so the common no-message-in-flight case costs one integer read per word dispatch, not a stadium_max_vm_count()-sized queue scan (also flagged by advisor() as a real hot-path cost, not deferred). Verified live on all three architectures: a self-test publishes a real payload to Hermes, proves the depth gate via VM-EXEC-ing an existing harmless colon word into Hermes (genuine nested vm_interpret(), depth 2, must not drain), then drains directly from genuinely-outermost context and confirms exactly one clean drain. dict_hash unmoved and identical across architectures. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
1f6343bc03 |
Publish path, no dispatch -- FABRIC-3.6.md task 3.3
sk_hermes_publish() allocates one SkHermesMessage per channel member (heat-cost ruling 2026-09-21: one message per subscriber, funded by the publisher's own reservoir) and enqueues each onto a new per-subscriber SkHermesPendingQueue -- found-or-created lazily by vm_id, sized from stadium_max_vm_count() like the channel/switch tables. Best-effort across subscribers: a failed allocation or full queue skips and rolls back just that one subscriber, not the whole publish -- the natural reading of "ledger and stadium_conserved() hold across N publishes to M subscribers" (the task's own check), not a separate ruling. Dispatches nothing -- sk_hermes_pending_count()/peek()/pop() are the read/drain primitives task 3.4's real checkpoint-driven drain will build on; this task's own self-test uses them directly since no checkpoint hook exists yet. Verified live on all three architectures: pending-queue table sized 50/202/50 slots (tracking the channel table's own per-arch sizing), a synthetic publish self-test (2 publishes to 3 subscribers) confirms exact per-subscriber delivery counts, and ledger/stadium_conserved() invariants hold both mid-publish and after manually draining every queue back to baseline. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
a4afdfa591 |
Channel table + common channel, inert (B1) -- FABRIC-3.6.md task 3.2
Adds SkHermesChannel: a channel is an index into a boot-time, stadium_max_vm_count()-sized table (same sizing pattern task 3.1 established for the switch table -- no separate numeric rule was ruled for this table, so task 3.1's bound is extended directly, flagged as such rather than restated as a new ruling). No name field, mirroring messaging.4th's own nameless CH-ARENA. The common channel (index 0) is created at boot and permanent. Hera subscribes explicitly in kernel_main.c (she is the one VM never born through capsule_birth_baby()); every other VM -- Tripod fleet and future WIREBIND identities alike -- subscribes inside capsule_birth_baby() itself, the single choke point every other birth already passes through. Inert: no publish, no dispatch, no ACK/NACK, no ACL hook (tasks 3.3, 3.6, 3.7). Verified live on all three architectures: channel table sized to 50/202/50 slots (matching switch-signal's own per-arch sizing), common-channel fleet self-test confirms all four Tripod members are members, and a synthetic create/subscribe/unsubscribe/ destroy round-trip against a private topic passes, including refusing to destroy the common channel. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
c19ef365fe |
Dynamic switch table (B2) -- FABRIC-3.6.md task 3.1
Replaces the fixed SK_SWITCH_MAX_SLOTS=16 compile-time array with a boot-time, RAM-derived allocation via a new sk_vm_switch_signal_boot_init(), kmalloc'd to stadium_max_vm_count() entries -- the same pattern session_boot_init() already established for Stadium-derived sizing. Every switch-signal participant is a Stadium VM, so this reuses that bound directly rather than deriving a separate one. Verified live on all three architectures: switch table sized to 50 slots (amd64), 202 slots (aarch64), 50 slots (riscv64) -- all well past the old fixed cap. All three boot to [zuse@Hera] ok> cleanly; dict_hash for Hermes/Hestia identical across architectures, unmoved from pre-task values. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
66ea4a5e74 |
Record task 3.0 rulings (FABRIC-3.5.md §XLVI); close FABRIC-3.6.md task 3.0
Captain Bob ruled all five §XLV.4 sub-items plus task 3.3's heat-cost design point: ACK on channel-open+delivery only; the channel-open ACL hook is a new word in ACL.4th; switch-table sizing mirrors Stadium's stadium_max_vm_count_val; chunks carry (msg_id, seq, is_last); drain one message per outermost-interpret checkpoint; publish costs one message per subscriber. Tasks 3.1-3.7 may now be written precisely. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
f2f7111dff |
Draft Phase 3 task breakdown (3.0-3.11), awaiting review -- FABRIC-3.6.md
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
303b0c7edf |
Record B1/B2/B4 rulings (FABRIC-3.5.md §XLV); clear Phase 3 blockers in FABRIC-3.6.md
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
feace42397 |
Add diagnostic scan cross-check of the Hermes counters -- FABRIC-3.6.md task 2.8
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
2a2bf6eb35 |
Stage B proof: add per-VM consumed term to stadium_conserved -- FABRIC-3.6.md task 2.7
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
313ffc89e6 |
Add exact-equality Hermes ledger self-audit -- FABRIC-3.6.md task 2.6
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
d11eb2e5db |
Add sk_hermes_decay(), ledgering decay into consumed -- FABRIC-3.6.md task 2.5
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
493d410028 |
Add the four ledger counters (held/pulled/returned/consumed) -- FABRIC-3.6.md task 2.4
FABRIC-3.5.md SXL.4's ledger: held == pulled - returned - consumed,
epsilon zero. Added sk_hermes_ledger() (an accessor, not a mutator)
plus four static counters in kernel_hermes.c.
Each counter has exactly one increment/decrement site: held/pulled
both move at sk_hermes_alloc()'s single success path, after every
refusal branch has already returned; held/returned both move at
sk_hermes_release()'s single success path. consumed is declared and
always reads 0 -- its one increment site doesn't exist yet, and won't
until task 2.5 gives decay something to record.
sk_hermes_release() now reads the Stadium cell's live header.heat
immediately before calling stadium_evict(), rather than assuming the
original pulled amount -- stadium_evict() zeroes the header as part of
freeing the cell and its own return value is a success code, not the
credited amount, so this is the only point the true remaining heat is
available. Today this always equals the original Q.SLOT pull; once
task 2.5's decay exists, this is what keeps returned correct without
touching this function again.
Self-test (kernel_main.c) extended: snapshots the ledger before
running so it checks its own deltas, verifies held/pulled grow by
exactly got_n * Q_SLOT on allocation with returned/consumed untouched,
then verifies held returns to its starting value and returned grows by
the same amount on release, and checks the audit invariant itself as a
bonus (task 2.6 formalizes this properly).
Three-arch boot clean: amd64/aarch64/riscv64 all reach [zuse@Hera] ok>,
zero UNKNOWN WORD, dict_hash unmoved. All three print PASS with
identical final ledger: held=0 pulled=65536 returned=65536 consumed=0.
No compiler warnings.
Authorized by Captain Bob ("keep going with rhe 6.5 document").
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
|
||
|
|
2c1dc1753a |
Correct sk_hermes_alloc() to admit a real Stadium patron; add sk_hermes_release() -- FABRIC-3.6.md tasks 2.2 (amended) + 2.3
Real finding, caught before building release on a foundation that
couldn't support it: task 2.2's first cut of sk_hermes_alloc() pulled
reservoir heat but never admitted a real Stadium-floor patron -- just a
local in_use flag. FABRIC-3.5.md SXXXIII.4 item 1 says MSG-FREE-NODE
returns heat "via STADIUM-EVICT", which only means something if
allocation admitted something. SXL.4's own invariant, Sigma(resident
patron heat) + reservoir + consumed == Q48_ONE, cannot balance if held
heat is invisible to every term while held. Flagged to Captain Bob
before proceeding; authorized to correct 2.2 in the same pass as
building 2.3 on top of the fix.
sk_hermes_alloc() now calls stadium_admit() with heat = the pulled
amount, behaviour = STADIUM_BEHAVIOUR_DELIVER (matching messaging.4th's
own SB-DELIVER STADIUM-ADMIT exactly), and identity = the message's own
slot index (matching the FORTH precedent -- caught live in the first
boot of this fix that omitting this made stadium_dispatch()'s existing
DELIVER diagnostic print msg_idx=0 for every message instead of a
distinct value). The returned cell index is stored in the message's
own stadium_cell field. Stadium-floor refusal (independent of reservoir
affordability) rolls back the pull the same way the other refusal
paths already do.
sk_hermes_release() -- the function task 2.3 actually asks for -- calls
stadium_evict() on that cell, which itself returns the departing
patron's remaining heat to its owning VM's reservoir, matching
MSG-FREE-NODE's exact shape. Release does not touch the reservoir
directly.
Self-test (kernel_main.c) extended: keeps every allocated message's
pointer, allocates to exhaustion as before, releases all of them, and
checks the reservoir returns to precisely its starting value.
"Undecayed" is true by construction (no decay/TTL logic exists yet,
task 2.5) -- exact restoration, not approximate.
Three-arch boot clean: amd64/aarch64/riscv64 all reach [zuse@Hera] ok>,
zero UNKNOWN WORD, dict_hash unmoved from task 2.2. All three print
identical PASS arithmetic: reservoir0=65536, reservoir_after_alloc=0,
reservoir_final=65536. No compiler warnings.
stadium_dispatch()'s DELIVER-case console output (one line per
eviction) is pre-existing instrumentation, not new -- confirmed real
and load-bearing per stadium.c's own comment, verbose but expected.
Authorized by Captain Bob ("yes").
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
|
||
|
|
9d129fdb1c |
Add sk_hermes_alloc(), the heat-coupled allocator -- FABRIC-3.6.md task 2.2, item 28
The piece FABRIC-3.5.md SXXXIII.6 calls "what remains genuinely hard,"
built and proven first per its own recommendation. Added
src/starkernel/vm/kernel_hermes.c (wired into Makefile.starkernel's
LOADER_EXTRA_SRCS -- this repo lists vm/*.c files explicitly, no glob)
and sk_hermes_alloc()'s declaration in kernel_hermes.h.
Checks stadium_reservoir_peek(vm_id) >= SK_HERMES_Q_SLOT before
touching the reservoir at all -- refusal this way needs no rollback,
since nothing was pulled -- with an explicit rollback path
(stadium_reservoir_push) kept defensively for the pull-then-short case,
though nothing in this single-core kernel is expected to reach it.
SK_HERMES_Q_SLOT = Q48_ONE / SK_HERMES_MSG_MAX (2048), deliberately
simpler than messaging.4th's own formula, which reserves a Q.1/3 floor
for COMMON-CH's own Stadium heat -- kernel-Hermes has no such object
(SXXXIII.4/SXXXIII.5's flat membership list carries no heat of its
own), so there is nothing left for that floor to protect.
Self-test in kernel_main.c, same diagnostic-only synthetic-VM pattern
as the existing Stadium quota grant self-test (lo=3, distinct from
that test's lo=1): reads back the actual granted reservoir rather than
assuming a number, derives expected_n from it, allocates to refusal,
and checks the refusal lands at exactly expected_n, the reservoir
doesn't move on the refused attempt (rollback proven, not assumed),
and the final reservoir is exactly reservoir0 minus got_n times
Q_SLOT.
Three-arch boot clean: amd64/aarch64/riscv64 all reach [zuse@Hera] ok>,
zero UNKNOWN WORD, dict_hash unmoved from task 2.1 (pure C, no FORTH
touched). All three print identical self-test arithmetic: reservoir0=
65536 Q_SLOT=2048 expected_n=32 got_n=32 reservoir_after=0. No compiler
warnings.
Noted, not fixed: Q_SLOT's divisor and SK_HERMES_MSG_MAX are the same
32, so reservoir and arena exhaustion land at exactly the same count by
construction -- this test can't distinguish which refusal reason
fired, only that refusal is correct and rolls back correctly.
Deliberately not evidence for stadium_conserved(): allocating alone
(no release yet, task 2.3) leaves pulled heat held off the Stadium
floor, so the two-term check would correctly read false right now if
run mid-hold. That's expected, not a bug -- Stage B (task 2.7) is
defined as "before and after the alloc/free cycle," not "continuously
during." This task's self-test checks reservoir arithmetic directly
instead.
Authorized by Captain Bob ("Yes continue").
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
|
||
|
|
f10fa7ae83 |
Add kernel-Hermes message/membership structures -- FABRIC-3.6.md task 2.1, Phase 2 begins
Phase 2, task 2.1 only: type definitions, wired to nothing, drawing no
heat -- no allocator, no protocol logic, no registration anywhere.
FABRIC-3.5.md SXXII.4: Phase 2 structures come first and prove nothing
until the allocator is built on top (task 2.2 onward, each its own
commit).
Added include/starkernel/vm/kernel_hermes.h:
SkHermesMessage -- field-for-field mirror of messaging.4th's live
9-cell MSG-* layout (type/from/to/payload addr+len/Stadium cell
index/seq/channel/orig-type), per SXXXIII.4 item 1 ("roughly half the
file is accessors that become struct fields"), plus an explicit
in_use flag for task 2.2's allocator. Deliberately no separate heat
field: per SXL.4, a message's heat IS the Stadium cell it occupies,
not a value copied alongside it -- one source of truth for the
conservation invariant stadium_conserved() (task 0.7) checks.
SkHermesMembership -- one flat broadcast membership list, SXXXIII.4/
SXXXIII.5's recommended replacement for messaging.4th's 28-word channel
abstraction (traced to exactly one live caller, CH-ADD-MBR). Item 27
(negotiation vs. broadcast, Phase 3 blocker B1) is not answered by this
structure and isn't meant to be -- a flat list is correct either way.
Genuinely wired to nothing: no .c file, no Makefile change, no include
from any compiled source. Syntax-checked standalone (gcc -std=c99
-Wall -Wextra -Werror -fsyntax-only) before touching the real build.
Boot byte-identical to task 1.9's baseline on amd64 (same dict_hash
triple, zero UNKNOWN WORD). Did not repeat aarch64/riscv64 -- the file
compiles into no object on any architecture, so there is no mechanism
by which it could diverge.
Authorized by Captain Bob ("yes").
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
|
||
|
|
1e75bb8039 |
Assert Hestia's headless invariant -- FABRIC-3.6.md task 1.9, Phase 1 closed
Audited first: neither capsules/hestia/init.4th nor her birth block in
kernel_main.c references g_wirebind_attached_username, CONSOLE-ATTACH,
MINT, or any proxy-minting mechanism -- the invariant already held
structurally, by absence. Stated it explicitly anyway, per the task:
added a comment at Hestia's birth site quoting FABRIC-3.5.md SXVIII.6's
invariant verbatim, warning future edits not to add console/wirebind/
proxy code there without re-reading it first.
Verified live with the actual no-thumbdrive boot
(ARCH=<arch> qemu ZUSEDISK=), not the default. All four VMs born
successfully on all three architectures, zero UNKNOWN WORD, and zero
ok> occurrences anywhere in any of the three full logs -- genuinely
silent, matching sk_repl_headless_wait()'s own documented "no banner,
no prompt, no input surface at all." Watched each log's line count
post-birth for 8-10s to confirm it stayed flat rather than eventually
printing something late.
Confirmed no regression on the standard (with-thumbdrive) path on
amd64: dict_hash identical to task 1.8's baseline. Did not repeat that
check on aarch64/riscv64 -- the change is a comment only, cannot
diverge by compiler, and the headless invariant itself was already
proven identically on all three.
Phase 1 is now fully closed (tasks 1.1-1.9). Tripod is Hera/Artemis/
Hestia plus Hermes (retained through Phase 1-3 per SXXXIV.4); Hestia
owns the drawing fabric exclusively; headless-until-login intact with
Hestia in the fleet. Phase 2 is next.
Authorized by Captain Bob ("yes").
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
|
||
|
|
6c6293cd52 |
Move PLOT/FB-WIDTH/FB-HEIGHT registration to Hestia only -- FABRIC-3.6.md task 1.8
Real fix for task 0.8's finding. register_framebuffer_words() was
called unconditionally from register_forth79_words() (word_registry.c),
itself called unconditionally from vm_init_with_host() -- the generic
per-VM bootstrap every VM goes through, with no way to know a VM's
name at that point.
Removed the unconditional call. Added a name-gated call instead in
capsule_birth_baby() (capsule_birth.c), beside the existing
is_fleet_foundation check -- the one place in the birth path where
capsule_name and the newly-allocated VM* are both in scope together:
Hestia gets register_framebuffer_words(), nobody else does.
Positively verified live, exactly as the task's own check demands:
FB-WIDTH via VM-EXEC returns UNKNOWN WORD in Hermes and Artemis, 1280
in Hestia. Confirmed at the fundamental level too: Hera's own base
PARITY:M7.1a word count dropped from 531 to 528 -- exactly the three
words removed -- identical across all three architectures.
Three-arch boot clean: amd64/aarch64/riscv64 all reach [zuse@Hera] ok>,
zero UNKNOWN WORD during boot. mkcapsule --lint capsules/ clean, 38
files / 0 violations (pure C change, no capsule content touched). No
compiler warnings.
Side effect on the vendored hosted build, expected and not a
regression: capsule_birth.c is kernel-only, so the hosted starforth
binary has no Hestia concept and now never registers these words at
all -- confirmed live. framebuffer_words.c's own top comment already
calls this surface "kernel-only, no-op on hosted builds," so the prior
stub registration was already vestigial. Ran a plain `make` sanity
build per CLAUDE.md's own stated purpose for that target; regenerated
lfs/amd64/starforth included here rather than left stale.
Authorized by Captain Bob ("yes").
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
|
||
|
|
5afb33049f |
Move fabric.4th + font.4th to hestia/init.4th -- FABRIC-3.6.md tasks 1.6+1.7 (merged)
Tasks 1.6 and 1.7 are not independent, and the punchlist's split was
wrong: font.4th calls G-LINE/G-ELLIPSE, which are fabric.4th's own
words. Confirmed live before committing to an approach -- removed only
fabric.4th's EXEC from init.4th, left font.4th's in place, booted
amd64: Hera's boot floods with UNKNOWN WORD: 'G-LINE'/'G-ELLIPSE' the
moment font.4th loads (logs/20260919-171047/amd64/, kept as evidence).
Reverted that partial state, asked Captain Bob how to proceed given
neither task can independently pass its own three-arch-boot check, and
was told to use best practices.
Moved both together, in their original relative order, into a new
capsules/hestia/init.4th block 4988 -- a deliberate, documented
deviation from "one task, one commit," not a bundling of convenience.
Three-arch boot clean: amd64/aarch64/riscv64 all reach [zuse@Hera] ok>,
zero UNKNOWN WORD. Hestia's dict_hash identical across all three
architectures. Verified the shrink/grow live, not just inferred from
hash movement: HERE reads 20008 in Hera, 63040 in Hestia post-move.
CART-PLOT in Hera is UNKNOWN WORD; the identical call routed into
Hestia via VM-EXEC reaches the word and fails on a stack underflow
instead, proof it exists there since an unknown word can't underflow.
mkcapsule --lint capsules/ clean, 38 files / 0 violations. MANIFEST.md
updated: init.4th's block 2049 entry no longer lists fabric.4th/
font.4th; hestia/init.4th's entry gains block 4988.
Authorized by Captain Bob ("Use best practices.").
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
|
||
|
|
487769e18a |
Register Hestia for switch signals -- FABRIC-3.6.md task 1.5
Added a fourth capsule_vm_find_by_name_nocase("Hestia", ...) +
sk_vm_switch_signal_register(...) block in src/starkernel/kernel_main.c,
same shape as the existing Hermes/Artemis blocks, placed after all four
fleet members are confirmed born -- the existing comment on this block
already states why: no critical-section protection during setup, so
registering earlier risks the signal firing mid-birth.
Three-arch boot clean: amd64/aarch64/riscv64 all reach [zuse@Hera] ok>,
zero UNKNOWN WORD, hashes identical to task 1.4's baseline (this is
pure C runtime state, doesn't touch any FORTH dictionary). No compiler
warnings.
Took FABRIC-3.5.md SXXXV.0's "invisible by default" warning literally
rather than trusting a clean boot log alone: SXXVIII.2's own recorded
switch-storm signature is "QEMU pinned near 100% CPU, serial log frozen
solid," not an error message. Confirmed normal wall-clock boot time on
all three (~30s) and, since TCG itself always shows ~100% CPU
regardless of guest workload, watched each serial log's line count at
the idle prompt for 5-10s and confirmed it stopped growing rather than
flooding or silently stalling.
Authorized by Captain Bob ("yes").
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
|
||
|
|
10e2d654e5 |
Birth Hestia in kernel_main.c -- FABRIC-3.6.md task 1.4
Added a birth block immediately after Hermes's own, same shape:
S" Hestia" BIRTH followed by a registry-lookup confirmation. Fleet is
now Hera/Hermes/Artemis/Hestia, four VMs, through Phase 1-3
(FABRIC-3.5.md SXXXIV.4) until Phase 4 retires Hermes.
Three-arch boot clean: amd64/aarch64/riscv64 all reach [zuse@Hera] ok>,
zero UNKNOWN WORD. Registry shows all four (BIRTH: Hermes live, BIRTH:
Hestia live, PARITY:BIRTH for all three non-Hera VMs). Hestia's
dict_hash identical across all three architectures (0x31cab513929eea89).
Hera/Hermes hashes unchanged from task 1.3; Artemis's vm_id shifted
(now the 4th birth instead of 3rd -- sequence-derived, not identity-
derived, so expected) but its dict_hash is unchanged and still
identical across arches. No compiler warnings.
Noted, not a regression: Hestia's birth log shows the same
"( Unterminated comment" HADES warning Artemis's birth has shown since
task 0.0's first baseline.
Authorized by Captain Bob ("yes").
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
|
||
|
|
ff00de9a63 |
Add Hestia to is_fleet_foundation -- FABRIC-3.6.md task 1.3
Fourth vm_name_prefix_eq_nocase(capsule_name, "Hestia") check alongside
Hera/Hermes/Artemis in src/starkernel/capsule/capsule_birth.c's
is_fleet_foundation local -- Hermes retained, per FABRIC-3.5.md
SXXXIV.4 (he stays live and fleet-foundation through Phase 1-3). The
flag's only consequence is session_set_pinned(vm_id, 1) for whichever
VM name matches.
Three-arch boot clean: amd64/aarch64/riscv64 all reach [zuse@Hera] ok>,
zero UNKNOWN WORD, byte-identical to the pre-change baseline -- expected,
since nothing births anything named "Hestia" yet (task 1.4), so the
added name never matches. Hermes confirmed still present in the check
and still born normally in all three logs.
Authorized by Captain Bob ("yes").
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
|
||
|
|
4f7cbe78fc |
Create capsules/hestia/init.4th -- FABRIC-3.6.md task 1.2
The third Tripod leg's first real file (FABRIC-3.5.md SII/SIV). Blocks
4986-4987 of the 4986-4996 allocated in task 1.1 (
|
||
|
|
b624133a28 |
capsules/MANIFEST.md: allocate Hestia's block range 4986-4996 -- FABRIC-3.6.md task 1.1
Documentation only, no capsule file created yet (that's task 1.2).
Checked against actual current occupancy rather than trusting
MANIFEST.md's own stale blanket "4853+ OPEN" line: fabric.4th already
occupies 4900-4924 and font.4th 4925-4985 (both standalone capsule
files EXEC'd by init.4th, block-numbered independently of it -- tasks
1.6/1.7 relocate which capsule EXECs them, not their own ranges), and
4997 is the console proxy's hardcoded Block 4997 string literal
(capsule_console.c:27-29). Allocated 4986-4996, the gap between the
two, avoiding 4997 as the task requires.
Split the Unassigned Ranges table's single blanket line into an
explicit claim for 4986-4996 plus a corrected "OPEN" line that excludes
the ranges actually in use. Recorded in MANIFEST.md rather than
tools/capsule-reserved.txt -- that file is for blocks owned by
non-capsule infrastructure per its own header comment; a real capsule
allocation belongs in MANIFEST.md alongside every other infrastructure
capsule's entry.
mkcapsule --lint capsules/ clean, 37 files / 0 violations, unchanged.
Authorized by Captain Bob ("YES").
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
|
||
|
|
eed2a9dfb5 |
FABRIC-3.6.md task 0.8: PLOT/FB-WIDTH/FB-HEIGHT reachability audit -- Phase 0 closed
Read-only audit, no code changed. Finding: reachable from every VM
today, not confined to one table, contrary to item 33's premise.
FORTH level matches expectation: fabric.4th/font.4th are EXEC'd only
from capsules/init.4th (Hera). C level does not: register_framebuffer_
words() (src/word_source/framebuffer_words.c:60-65) is called
unconditionally from register_forth79_words() (src/word_registry.c:
139), itself called unconditionally from vm_init()
(src/starkernel/vm/vm_bootstrap.c:263) -- the generic per-VM bootstrap
every VM goes through, no identity check.
Verified live rather than trusting the source trace alone: booted
amd64 and ran `S" FB-WIDTH ." S" Hermes" VM-EXEC` and the same against
Artemis -- both returned 1280, not UNKNOWN WORD. Neither loads
fabric.4th, so the raw C primitive itself is answering.
Not fixed here, per the task's own read-only scope. Gives task 1.8 a
concrete starting state: its own check ("a non-Hestia VM calling PLOT
gets UNKNOWN WORD") currently fails, and register_framebuffer_words()'s
call site will need to become conditional or move out of the universal
bootstrap -- not just the FORTH-level relocation tasks 1.6/1.7 already
plan for.
Phase 0 gate met across tasks 0.2-0.7 (three-arch boot, stadium_
conserved() true, zero UNKNOWN WORD, repeatedly). Phase 0 is closed;
Phase 1 (Hestia, messaging untouched) is next.
Authorized by Captain Bob ("Continue.").
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
|
||
|
|
380f0a09c9 |
Add stadium_conserved() -- FABRIC-3.6.md task 0.7, item 41
Boolean analogue of vm_physics_conserved(), for the Stadium per-VM
quota invariant rather than fleet-wide execution heat
(FABRIC-3.5.md SXXXIX.4). int stadium_conserved(VMUuid vm_id), in
src/starkernel/vm/stadium.c alongside stadium_resident_sum()/
stadium_reservoir_peek() that it's built from, declared in
include/starkernel/vm/stadium.h.
Implements the two-term form -- resident_sum(vm_id) +
reservoir_peek(vm_id) == Q48_ONE -- not the three-term form SXL.4
rules for the eventual system. That ruling's `consumed` term is a
Phase 2 kernel-Hermes ledger deliverable that doesn't exist yet:
nothing draws on any VM's Stadium quota today (task 2.2 is literally
where that wiring gets built), so consumed is honestly zero right now.
Folding it in as a placeholder would be inventing Phase 2 state ahead
of it existing -- the doc comment says so explicitly, so whoever
builds Phase 2's ledger extends this function rather than working
around it.
Wired into the existing per-VM boot diagnostic
(stadium_words_print_boot_diagnostics(), kernel_main.c:810, Hera
only -- the sole existing call site) rather than adding a new one,
printing CONSERVED/DRIFTED the same shape vm_physics_status() already
uses.
Three-arch boot clean: amd64/aarch64/riscv64 all reach [zuse@Hera] ok>,
zero UNKNOWN WORD, all print "Stadium conservation: CONSERVED" with
identical resident_sum=47641 reservoir=17895 sum=65536=Q48_ONE. No
compiler warnings on either edited file (forced recompile checked).
Authorized by Captain Bob ("Yes.").
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
|
||
|
|
9886ad5315 |
tools/capsule-reserved.txt: return freed block ranges -- FABRIC-3.6.md task 0.6
Added 4055-4059 (former common/msg.4th) and 4300-4399 (former
process.4th) as reserved, each noted as freed by this reshuffle's
strip rather than owned by non-capsule infrastructure -- the file's
usual purpose (Artemis's own block usage, etc.). Framed explicitly as
lifted, not permanent, once someone deliberately wants a range back,
per FABRIC-3.5.md SSXVIII.4/XXII.4: freed ranges should be returned
here rather than silently available for a future capsule to reclaim
without anyone noticing.
mkcapsule --lint capsules/ clean, 37 files / 0 violations.
check_reserved_conflicts() -- the hard build-gate that actually reads
this file (tools/mkcapsule.c:974) -- passes clean on a real
`make -f Makefile.starkernel ARCH=amd64 all`, confirming the new
entries don't collide with anything currently baked. Registry/
documentation only, no capsule content touched, so no 3-arch boot run
for this task.
All of Phase 0's strips and documentation corrections (0.1-0.6) are
now done. stadium_conserved() (0.7) and the PLOT/FB-WIDTH/FB-HEIGHT
registration audit (0.8) remain.
Authorized by Captain Bob ("Go for it.").
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
|
||
|
|
e233d09fa1 |
capsules/MANIFEST.md: correct blocks 4055 and 2049 -- FABRIC-3.6.md task 0.5
Block 2049's justification claimed init.4th loads compudynamics,
common:msg, fleet-k and process. Read the live file: it loads none of
these. compudynamics.4th/fleet-k.4th were already deleted (9323f776,
2026-07-05); common:msg.4th/process.4th are this reshuffle's own
Category A strips (tasks 0.2/0.3, a0e97258/aafcce43) and were never
EXEC'd from this block even before that -- the manifest entry was
already wrong prior to this pass, just not yet caught.
Removed the standalone common/msg.4th (former block 4055) and
process.4th (former blocks 4300-4301) sections, since both files no
longer exist, and folded them into "Deleted capsules (historical)"
alongside the existing compudynamics.4th/fleet-k.4th entry -- same
convention, same section. Noted that common/msg.4th's own entry had
claimed it was an immutable ABI "every messaging VM loads at birth",
a claim FABRIC-2.md:2773 had already flagged stale before this
correction landed. Updated the Unassigned Ranges table so 4055-4059
and 4300-4399 read as former-file ranges rather than "extension space"
for files that no longer exist.
Documentation only -- no capsule content touched, mkcapsule --lint
capsules/ still clean (37 files, 0 violations). Verified every
remaining ### `*.4th` section in the manifest names a file actually
present on disk.
Authorized by Captain Bob ("Keep going.").
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
|
||
|
|
bcc678e354 |
Strip SPAWN-EVENT from messaging.4th -- FABRIC-3.6.md task 0.4
Dead per task 0.1's reachability check (
|
||
|
|
aafcce4320 |
Strip capsules/process.4th and its EVENT-EMIT/-WAIT/-DRAIN -- FABRIC-3.6.md task 0.3
process.4th is dead per task 0.1's reachability check (
|
||
|
|
a0e9725883 |
Strip capsules/common/msg.4th -- FABRIC-3.6.md task 0.2
Dead per task 0.1's reachability check (
|
||
|
|
156d1642de |
FABRIC-3.6.md task 0.1: Category A reachability established
Checked the three §XXII.2 routes for capsules/common/msg.4th,
capsules/process.4th, and the SPAWN-EVENT constant, never by grep
count alone: a boot-path trace of every capsule init.4th loads at
birth, a tools/experiments/docs invocation search, and confirmation
that mere presence in the baked capsule directory doesn't make a name
reachable if nothing constructs it at runtime. All three are dead by
every route -- zero EXEC sites, zero callers of their exported words.
Two findings recorded, neither changing the punchlist's plan:
tools/hermes_smoke.sh calls EVENT-EMIT/EVENT-WAIT/EVENT-DRAIN directly,
a caller FABRIC-3.5.md SXXXIII.3 missed when it said the only other
reference was MANIFEST.md. Doesn't make them live -- the script is
already broken on its own terms (references capsules/core/init.4th and
build/amd64/standard/starforth, neither exists; calls the pre-rename
CD-INIT word). Task 0.3's plan to strip these three words alongside
process.4th stands.
PAUSE-EVENT/RESUME-EVENT/KILL-EVENT are exactly as dead-by-name as
SPAWN-EVENT -- process.4th's own calls pass bare numeric literals, never
the constant names. Task 0.4 only names SPAWN-EVENT; flagged for
whoever picks up Category A stripping next rather than expanding this
task's scope.
Investigation only, no code changed.
Authorized by Captain Bob ("begin.").
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
|
||
|
|
4544877a36 |
FABRIC-3.6.md task 0.0: three-ISA baseline smoke test, PASS
amd64/aarch64/riscv64 all boot clean to [zuse@Hera] ok>, zero UNKNOWN
WORD, and dict_hash identical across all three: Hera (PARITY:M7.1a)
0x6824fe5993239838, Hermes 0x062252c4da6858da, Artemis
0xed80117724c26f36. This is the gate task 0.0 exists for -- every later
task's acceptance in this document assumes this baseline is known-good.
Found and worked around a real confound along the way: disk/artemis.img
is deliberately shared across all three ISAs' qemu targets (FABRIC-3.md
SXXXV.2), so a same-order rerun has run 2 and 3 silently resume run 1's
already-formatted disk instead of formatting their own. The first
attempt (logs/20260919-124835 amd64, logs/20260919-124952 aarch64,
both kept for the record) shows exactly this: Artemis's dict_hash
diverges between the two runs even though Hera's and Hermes's do not,
because only Artemis's birth path branches on disk state. Restored
disk/artemis.img and disk/thumbdrives/zuse-thumb-ident.img to their
committed blank state before each of the three reruns that produced
the clean, matching baseline above, and recorded the finding in
FABRIC-3.6.md so a later task doesn't mistake the same confound for a
real architecture divergence.
capsules/BLOCK_MAP.md's timestamp header is regenerated by the build,
per .claude/CLAUDE.md's documented behavior for that generated file.
Authorized by Captain Bob ("begin", "clean new disk",
"document, commit, and push").
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
|
||
|
|
e56974e0bf |
Bleach all 9 thumbdrive identity images to blank ahead of re-mint
Prerequisite for FABRIC-3.md §XXXV.6's ratified re-mint: capsule_mint.c's capsule_mint_identity() refuses to write over an already-recognized drive (MINT_ERR_ALREADY_MINTED, mirrors WRITE(10)'s refuse-on-non-blank-media rule), so all 9 existing certs had to be destroyed before any of them could be re-minted against the new 1 GiB artemis.img. Confirmed root cause first: Zuse's genesis marker (zuse_pubkey) lives on Artemis's own disk fence (devblock_from_top=0), not on her thumbdrive -- capsule_zuse_ boot.c's capsule_zuse_boot_try_attach() silently no-ops if her drive isn't blank and the marker is missing, so the old zuse-thumb-ident.img could not re-authenticate against the fresh artemis.img regardless. All 9 images (zuse, bob, 00-06) verified all-zero, 64 MiB size intact. Re-mint itself (live QEMU boot + MINT) not yet performed. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
7c5ba7a874 |
Rebuild disk/artemis.img fresh at 1 GiB, blank; preserve old 30 MiB image
Executes the drive-resize decision ratified in FABRIC-3.md §XXXV.6: the prior 30 MiB image was too small for the block-backed per-ISA DoE persistence design (§XXXV.2). Built a new blank 1 GiB image rather than migrate the old one's top-of-device metadata fence -- re-minting Zuse and the thumbdrive identities invalidates them either way, so skip the migration entirely. - disk/artemis.img: replaced with a fresh blank 1 GiB image (was 30 MiB) - disk/artemis-30mb-pre-1gib-backup.img: the superseded 30 MiB image, preserved rather than deleted; disposal remains open (FABRIC-3.md §XXXV.7) - disk/README.md: documented both images per this repo's own convention Not yet formatted or re-minted -- that's a live-boot step, not done here. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
e4a52c01be |
FABRIC-3.md §XXXV: bump ratified artemis.img target from 256 MiB to 1 GiB
Captain Bob asked for extra safety margin over the earlier 256 MiB recommendation. Updated every reference in §XXXV.2/.4/.6 consistently. Still decision only -- no image built yet. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
c19b6febc2 |
FABRIC-3.md §XXXV.6: ratify fresh 256 MiB artemis.img + full re-mint, no fence migration
Captain Bob resolved the drive-resize fork from §XXXV.2: build a new blank 256 MiB disk/artemis.img and re-mint Zuse + all 8 thumbdrive identities fresh, rather than attempt to migrate the top-of-device fence on the existing 30 MiB image. Skips migration complexity since a re-mint invalidates prior identities either way. Decision only -- no new image built, no re-mint performed yet. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
a83cbe016f |
FABRIC-3.md §XXXV: block-backed DoE persistence design + N=2 per-ISA shape (design only, no code yet)
Scopes the self-contained (no host serial capture) persistence needed before the per-ISA campaign driver can run on real bare-metal hardware. Found disk/artemis.img is only 30 MiB (measured) against a 516 MB/8-trial raw per-tick baseline -- raw mirroring is ruled out at any realistic device size. Proposes reusing log_region.c's binary-slot/control-header pattern for a new trial-summary-only region, flags the fence-addressing risk in growing the image, and recommends N=2 reps/cell over N=3 given the added block-IO overhead. Nothing built -- ratified in conversation only. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
94215e8b47 |
Fix two catastrophically heavy workloads found running the real campaign
The full 360-trial campaign was launched, then killed after 4h39m of CPU time with zero trials completed -- still stuck on the very first touch of the very first trial. Root cause, quantified from the workload's own source, not estimated: RUN-CHAOS5 (workload-5.4th) is 1000000 0 DO CHAOS-FIELD CHAOS-RIPPLE LOOP, multiplying out to ~2.4 trillion word executions for one call -- ~495 days at the observed rate. Never designed to be called to completion as a single touch in a repeatedly-birthed worker. A second problem found computing the fix rather than discovering it mid-run again: workload-1.4th's own birth-time self-execution (20 SQUARE-WAVE + 1500 SQUARE-BURST + 100000x MICRO-BURST, all three run automatically at capsule load) sums to ~875 million words -- ~4.3 hours just to birth one worker -- and worker index 1 always maps to it in heterogeneous mode, landing in half the campaign's cells regardless of concurrency level. Fix: two new capsules, workload-1-lite.4th and workload-5-lite.4th, carrying the same word bodies verbatim but a bounded self-execution tail, comparable in scale to the campaign's other workloads. The originals are untouched; only multiuser-doe.4th's own WL-CAPSULE/ WL-ENTRY index 1 and 5 mappings were repointed. Verified live before relaunching a third time: the actual worst case in isolation (5 0 998 MU-RUN-TRIAL, concurrency=8 heterogeneous, includes both fixed workloads plus RUN-OMNI) completed in ~20 minutes wall-clock, Hera stayed healthy throughout. Clean 3-arch qemu boot. Reps reduced 60 -> 30/cell (Bob's call after seeing the real per-trial cost) -- still matches the project's "rule of 3's" DoE convention (same as ACL-RWT's own 30 reps). 6 cfgs x 30 reps = 180 trials. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K |
||
|
|
29f91f9531 |
multiuser-doe.4th: the campaign trial-loop capsule, verified live
Builds the experimental control Bob correctly identified as still missing after HB-ON/HB-OFF (§XXXIII.5): walk the shuffled matrix of concurrency-level x workload-mode cells, birth the right worker count per cell, drive each with two VM-EXEC touches, check VM-ERROR?, kill them, print a trial marker. Built entirely from existing primitives (WORKER-BIRTH, VM-EXEC, VM-HEAT, VM-ERROR?, KILL, HB-ON/HB-OFF) plus doe.4th's own RUN-MATRIX/SHUFFLE-MATRIX pattern -- no new C primitives. Real capsule-format bug found and fixed: a first draft, chunked purely by a fixed 16-line count with no regard for word boundaries, split several CASE...ENDCASE structures and one oversized colon definition across Block headers. Result was a cascading [CAPSULE][DEFER] failure from the first split forward -- every subsequent line failed to compile, and MU-RUN-TRIAL was never actually defined (confirmed: UNKNOWN WORD when called). Root cause traced to capsule_loader.c directly: a :...; word and any control structure inside it must fit entirely within one 16-line block -- the loader's per-block compile pass has no persistent record of an open CASE's (or an overlong definition's own) state across a Block boundary. Not previously documented anywhere in this project's capsule-authoring guidance. Fixed via manually curated block boundaries and factoring oversized bodies into smaller helper words. Reps ratified at 60/cell (not the 10 first drafted), matching this project's own "rule of 3's" DoE convention (ACL-RWT's 3 seeds/30 reps, std79's 3x9x3). 6 cfgs x 60 reps = 360 main-block trials. MU-MAX-REPS raised 20 -> 63 for run-matrix headroom. Verified live on amd64: one isolated trial (0 0 999 MU-RUN-TRIAL) produced two real concurrent births, two clean kills, and the exact expected marker (MU-TRIAL run=999 cfg=0 rep=0 nw=2 mode=0 fail=0). Hera stayed healthy throughout. Clean 3-arch qemu boot. Not yet run: the full 360-trial MU-EXEC-CAMPAIGN itself (a genuinely long-running action under TCG, deliberately not started without explicit confirmation) or the WIREBIND-automation fixed arm. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K |
||
|
|
41918a28a4 |
doe_log.c: add vm_name/vm_id_hex identity columns; new calibration workload
Adds the identity columns HB-ON's existing per-tick CSV logger (doe_log_tick_row(), doe_log.c) was missing. Previously every column described *a* VM's state each row, but nothing said which VM emitted it -- concurrent VMs' rows were indistinguishable by source. Two new first columns, vm_name and vm_id_hex, resolved via a reverse lookup on vm->stadium_vm_id (capsule_vm_registry_get()). Real bug found and fixed while building this: the freestanding snprintf here silently prints the literal format string instead of substituting for %016llx (width+ll+hex unsupported) -- caught by reading the actual emitted row, not assumed to work. Replaced with a hand-rolled hex nibble-table loop, the same idiom vm_uuid_format()/ MINT-SCRATCH-EMIT already use. Corrects FABRIC-3.md's own prior "still not started: CSV driver" framing: no bespoke CSV emitter is needed for the multiuser DoE at all -- this per-tick logger already exists, fires automatically inside every VM's own execution loop, and just needed HB-ON plus these identity columns to be usable for concurrent workers. Also corrects doe_log.h's own stale doc comment claiming g_doe_log_enabled defaults to 1 -- doe_log.c's own source is authoritative: 0, off by default, matching the HB-ON/HB-OFF naming. One real, honest limitation recorded rather than smoothed over: vm_name reads blank for any tick captured during a WORKER-BIRTH'd capsule's own self-execution at birth time, since capsule_birth_baby() runs that work (and its heartbeat ticks) before the registry name can be set. vm_id_hex is unaffected (reads directly from vm->stadium_vm_id) and still uniquely disambiguates every row -- confirmed live: a blank-name row's own vm_id_hex matched its later PARITY:KILL line's vm_id exactly. New workload-calib1.4th: Bob confirmed the existing 10 workload-N.4th files aren't fixed. Sized to cross HEARTBEAT_CHECK_FREQUENCY (256 word-executions/tick) many times over while staying far lighter than RUN-FIB/RUN-CHAOS5, both too slow under TCG for a bounded verification run. A real, reusable addition, not a throwaway. Verified live on amd64 via a self-contained one-shot EXEC'd test (HB-ON, WORKER-BIRTH two calibration workers, HB-OFF, VM-ERROR? on both, KILL both, completion banner) rather than interactive polling, which breaks once HB-ON makes the log grow continuously via Hera/ Hermes/Artemis's own background ticks. Clean 3-arch qemu boot on the real committed change. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K |
||
|
|
e33eb36361 |
Multiuser DoE punch list: WORKER-BIRTH + VM-ERROR? + vm_physics_init fix
Core mechanism for §XXXIII's main concurrency block, built and live-verified on amd64. Two new primitives: - WORKER-BIRTH ( capsule-c capsule-u name-c name-u -- ok? ): births a named, VM-EXEC-addressable VM from an arbitrary (p) capsule with no identity involved. Corrects the ratified design's own assumption that the main block would use UNATTENDED-BIRTH -- that requires a committed capsule per identity, impractical for dozens of trial VMs. The concurrency block never needed identity at all. - VM-ERROR? ( c-addr u -- flag ): reads a named VM's error state from Hera, mirroring VM-HEAT's silent/always-returns-a-value contract. Needed to check a VM-EXEC-driven trial VM's own fault state after the fact -- nothing existing let Hera do this. Real bug found and fixed in both WORKER-BIRTH and UNATTENDED-BIRTH: capsule_birth_baby() never calls vm_physics_init() either (same shape as the registry-name gap found building UNATTENDED-BIRTH) -- without it a born VM is never in the VM Fleet Attractor physics list, so VM-HEAT returns 0 forever regardless of work done. Fixed by adding vm_physics_init() alongside the existing registry-name call in both words. Traced (not guessed) why heat still read 0 after one VM-EXEC touch even post-fix: vm_physics_touch()'s transfer logic only fires from a VM's *second* touch onward -- the first touch just records a baseline tick. Verified live across three sequential touches: heat 0 -> 7039 -> 11333. This is a real design requirement for the DoE's heat/CV response variable (each trial must touch a worker at least twice), not a bug to route around. Also found live: all 10 existing workload-N.4th capsules self-execute their full workload at load/birth time (a bare top-level call to their own RUN-* word at file end) -- missed on an earlier, too-shallow 8-line survey of each file. WORKER-BIRTH alone already runs a worker's first pass as a side effect of birth. Verified live on amd64: two concurrent workers (fib + matrix-mul) birthed, run, measured (heat + error state), and killed cleanly; VM-HEAT/VM-ERROR? both confirmed silent-0 on an unknown name. Clean 3-arch qemu boot on the real committed change. Still open: the run-matrix/shuffle/CSV driver capsule itself, the WIREBIND-automation path for the fixed arm, and the per-VM touch-count budget's exact value. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K |
||
|
|
bc2e294c50 |
FABRIC-3.md §XXXIII: multiuser/multitasking DoE design (no code yet)
Scopes the DoE Bob asked for after §XXXII closed -- "a pretty big experiment... long running multi-factor DoE that sort of puts the OS through its paces." Design only, per this project's standing plan-before-code discipline. Three corrections found and recorded before the design could be trusted: - Console sessions are not a concurrency axis and this isn't a defect: console.c hardcodes one physical UART as static global state, not an instantiable abstraction. One physical terminal, one foreground session, by construction. Bob's reaction (eventual multi-seat/ telnet-ish/pty-multiplexer idea) recorded separately in memory, explicitly deferred behind this DoE. - The real, provable concurrency primitive is VM-EXEC (doe-campaign.4th's own THREE-VM-CAMPAIGN already proves the pattern), resolved through the full VM registry (unbounded kmalloc list), not messaging.4th's 16-slot routing table -- checked live, not assumed. - No FORTH-reachable monotonic clock exists; latency dropped as a response variable rather than building a new timing primitive under this scope. Also flagged, not chased: every acceptance boot this session produced an L8-DoE CSV despite init.4th never calling it and KERNEL_ARGS defaulting to empty -- likely a stale starforth.cfg survivng `clean`. Practical decision: the new campaign is its own driver, independent of whatever causes that. Ratified design: concurrency level (2/4/8 VM-EXEC-driven VMs) x workload assignment (uniform/heterogeneous across the existing 10 workload capsules) as the crossed factorial; identity origin (WIREBIND vs UNATTENDED-BIRTH) as a fixed arm rather than crossed, since WIREBIND's per-identity USB-attach cost doesn't scale into a full factorial. Response variables: correctness-proxy (completes without vm->error), heat/CV stability under contention, completed-trials-per- interval in place of the dropped latency variable. Punch list recorded; implementation not started per standing "plan approval is not a start signal" rule. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K |
||
|
|
af351bff92 |
Punch list item 4 (final): CONSOLE-ATTACH, full unattended-identity flow verified end-to-end
Closes FABRIC-3.md §XXXII.2's punch list. CONSOLE-ATTACH ( name-c name-u -- ok? ) pairs a fresh console VM to an already-live VM registered as "<name>~user". Deliberately a plain, unconditional primitive with no VMIdentity capability-bit check -- corrects this session's own first-pass design (§XXXII.2 amended in the same commit): identity.installed is 0 for Hera/Hermes/Artemis and for every console-proxy VM, so a capability-bit gate would be unreachable for every VM a human actually types at, Zuse included. Matches ZUSE-ELIGIBILITY-ADD's own "no bespoke gate" precedent in this file; real gating is `' CONSOLE-ATTACH ACL-PIN` in ACL.4th if ever wanted. Two real bugs found and fixed via live testing, not assumed correct: - capsule_birth_baby() never sets a VM's registry name (documented gap, same one capsule_runcap_birth()'s own history already hit) -- UNATTENDED-BIRTH gained a second `name` argument and now calls capsule_vm_registry_set_name(new_vm_id, "<name>~user") itself. - CONSOLE-ATTACH's first draft took an independent console name from the target's name. sk_repl_dispatch_line()'s pairing check (repl.c) reconstructs the target as console_get_vm_name()+"~user" -- a mismatched console name silently falls back to direct interpretation with no error. Caught live (typed `5 6 + .` at a mismatched console, got a direct `11` instead of a relay) and fixed by collapsing to one name argument, matching WIREBIND's own by-construction invariant. Full end-to-end live verification on amd64: minted a test identity, UNATTENDED-BIRTH'd it as "bob", CONSOLE-ATTACH'd a console named "bob", USE'd it, typed `5 6 + .` -- no direct output at [zuse@bob] (relay path taken), then `[zuse@bob~user] 11 ok>` appeared: the relayed command executed on the target identity VM itself and printed its own answer back through the shared console. Hera stayed healthy throughout (2 2 + . -> 4 after switching back). CONSOLE-ATTACH also verified to refuse cleanly on a nonexistent target with no orphaned VM. Test capsule reverted after capture per this project's probe convention -- never committed. FABRIC-3.md §XXXII.2 fully closed: all 4 original questions ratified, the mid-course drive_uuid and ACL-bit corrections both recorded plainly rather than silently folded in, and a doc-accuracy note left for CLAUDE.md's own stale "1024-byte block limit" framing (mkcapsule's real limits are range [2048,5120) and 16 content lines/block) -- flagged, not fixed, out of this punch list's scope. Clean 3-arch qemu boot (amd64/aarch64/riscv64) on the real committed C-only change. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K |
||
|
|
5879c8b3bc |
Punch list item 3: UNATTENDED-BIRTH call site, verified live end-to-end
Implements the unattended-birth mechanism FABRIC-3.md §XXXII.2 designed: UNATTENDED-BIRTH ( name-c name-u -- ok? ) births a VM from a named (p) capsule via capsule_birth_baby() (completely unmodified, the same generic build-time-capsule path CAPSULE-BIRTH already uses), then installs its identity the same way capsule_wirebind.c already does live for WIREBIND attaches -- vm_identity_from_cert() verification followed by a plain post-birth struct assignment -- rather than anything RUNCAP-shaped, since RUNCAP requires a real blkio_dev+ homeblocks_sig_t an unattended identity never has. The born VM's own capsule payload is expected to lay down two CREATE'd buffers (UNATTENDED-ID-UUID, UNATTENDED-ID-CERT) via MINT-SCRATCH- EMIT's own literal format; their addresses are fetched by interpreting a two-word line inside the *new* VM's own context (vm_interpret(born_vm, ...)), the same "run inside that VM's own dictionary" idiom capsule_wirebind.c already uses for VM-NAME-REG. Explicit invariant preserved: never touches g_wirebind_attached_username or any console-pairing state, births no console VM -- an unattended identity stays un-promptable (§VIII.1) until a human pairs a console to it later via the already-working VM-NAME-REG mechanism. Verified live end-to-end on amd64: minted a real test identity via MINT-SCRATCH, captured its MINT-SCRATCH-EMIT output, built a throwaway test capsule from it (discovered along the way: mkcapsule's real block constraints are range [2048,5120) and max 16 content lines per block -- neither matches this repo's own doc comment, corrected via ground truth from the tool itself, not assumed), then ran UNATTENDED-BIRTH against it: cert verified against Zuse's root pubkey, identity installed, "no console attached" reported, Hera stayed healthy afterward (5 6 + . -> 11). Test capsule reverted after capture per this project's own probe convention -- not a real identity, never committed. Clean 3-arch qemu boot (amd64/aarch64/riscv64) on the real committed C-only change. Remaining punch-list item (the ACL cap bit for console attachment) not started. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K |
||
|
|
5fc709a228 |
Punch list item 2: MINT-SCRATCH-EMIT, verified live on all 3 arches
Adds the hand-transcription mechanism FABRIC-3.md §XXXII.2's punch list item 2 calls for: MINT-SCRATCH-EMIT prints the last successful MINT-SCRATCH's drive_uuid + cert devblock as ready-to-paste FORTH source (HEX-based CREATE ... C, ... byte sequences), so packaging an unattended identity into a capsule is a mechanical copy out of the captured boot log rather than a manual hex-to-FORTH translation an operator could transpose a digit in. Refuses (no output) if no MINT-SCRATCH has ever succeeded -- printing 4112 zero bytes as if they were a real identity would be a silent, misleading success, matching this session's own error-handling audit discipline rather than adding a new silent-failure primitive right after finishing one. Verified live on amd64: MINT-SCRATCH-EMIT correctly refuses before any mint, then after MINT-SCRATCH succeeds, emits UNATTENDED-ID-UUID and UNATTENDED-ID-CERT as valid FORTH literals. Cross-checked byte-exact: the cert's own embedded ASN.1 serialNumber field matches the emitted UUID bytes exactly, confirming x509_build_user_cert()'s drive_uuid binding round-trips correctly through the scratch-device path. Clean 3-arch qemu boot (amd64/aarch64/riscv64). Remaining punch-list items (writing/committing a real capsule file for an actual named identity, the unattended-birth call site, the ACL cap bit) not started -- authoring a real committed capsule needs a name/ purpose decision that isn't mine to make. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K |
||
|
|
63864c4b01 |
Punch list item 1: scratch-device MINT-SCRATCH, verified live on all 3 arches
Implements FABRIC-3.md §XXXII.2's scratch-thumbdrive mint mechanism: capsule_mint_identity_scratch() (capsule_mint.c/.h) builds a throwaway RAM-backed blkio_dev via blkio_ram.c's backend and runs capsule_mint_identity() against it completely unmodified -- same live Zuse-signing operation, same rng_get_bytes() draw for drive_uuid a real thumbdrive gets. Reads back only drive_uuid + the cert devblock; the seed devblock is written into the scratch buffer internally but never read out (no seed is ever baked into a capsule, per the ratified no-seed decision). New FORTH word MINT-SCRATCH (mama_forth_words.c), same stack signature as MINT, mints into the scratch device instead of any attached drive and never touches sk_repl_get_attached_blk_dev() or console-pairing state. Prints the drive_uuid as hex so a live boot log itself proves each call drew fresh entropy. Build correction found along the way: blkio_ram.c was excluded from the kernel build (Makefile.starkernel VM_EXCLUDE) alongside blkio_factory.c/blkio_file.c. blkio_factory_open() unconditionally references blkio_file.c's real fopen()/fread() file I/O, which has no freestanding-kernel equivalent, so the factory function couldn't be used as-is. blkio_ram.c itself is pure memcpy over a caller buffer -- pulled it alone into the kernel build and wired it directly in capsule_mint.c, the same way blkio_factory.c's own extern declarations do internally. Verified live on amd64: two MINT-SCRATCH calls produced two genuinely different drive_uuids (b533246d.../ae11b2b2...), confirming fresh entropy per call rather than stale reuse; VM stayed healthy afterward (5 6 + . -> 11). Clean 3-arch qemu boot (amd64/aarch64/riscv64), logs and DoE CSVs committed per standing convention. Remaining punch-list items (hand-transcription into a .4th block, the unattended-birth call site, the ACL cap bit) not started. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K |
||
|
|
60bcdc09a7 |
Stage E amendment: resolve drive_uuid wrinkle via scratch-device mint (FABRIC-3.md §XXXII.2)
Bob's proposal: mint against a "sim thumbdrive" -- a scratch blkio_dev from the existing blk_subsys_add_raw_device() mechanism (same one the kernel ramdrive already uses), then hand-transcribe just drive_uuid + DER cert as FORTH literals into a normal .4th capsule block. This dissolves the wrinkle the first pass of Q2 left open: capsule_mint_identity() runs completely unmodified against the scratch device (same live Zuse-signing op, same rng_get_bytes() draw for drive_uuid), so vm_identity_from_cert() needs zero changes -- no origin-aware variant, no content-hash-derived substitute. "Sim" describes only where the bytes were written, invisible to verify. Also corrects an error in the first pass: cert production cannot be "offline" -- capsule_mint_identity() requires Zuse's live in-kernel signing key, so minting is inescapably a two-boot runtime operation (mint on boot N, package by hand, birth on boot N+1). Punch list updated to match. Still design-only, no code. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K |
||
|
|
06e2507ce4 |
Stage E ratified: unattended identity is cert+personality only, no seed (FABRIC-3.md §XXXII.2)
Settles all 4 open questions on paper, no code: - No seed baked into capsules for this pass (Bob's call); runtime-minted seed documented as a future, separately-scoped possibility - Birth via capsule_birth_baby() (unmodified, already generic), not capsule_runcap_birth() -- that path requires a real blkio_dev*+ homeblocks_sig_t* an unattended identity can't provide, and its own header says build-time-baked content is exactly the wrong case for it - Identity population is the same post-birth struct-assignment pattern capsule_wirebind.c already uses live (vm->identity = identity) - §XXIV's block-collision machinery already covers a cert-carrying capsule -- no new risk, no change needed there - ACL gate lives on the attaching human's credentials, since an unattended instance holds no secret to prove anything about itself - xHCI/live-table machinery confirmed irrelevant, independent of the seed question One real wrinkle flagged, not resolved: vm_identity_from_cert()'s serial-number check binds to a physical drive_uuid an unattended identity doesn't have -- needs Bob's call before implementation. Punch list recorded; implementation not started per standing "plan approval is not a start signal" rule. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K |
||
|
|
42d4bf3dad |
Stage D batch 7 (final): mama_forth_words.c groups 5+6 -- dictionary lookup/test words + Stadium primitives; Stage D closed (FABRIC-3.md §XXXII.6)
NAME>XT/RUNCAP-TEST/PAIR-TEST and the six STADIUM-* physics primitives (STADIUM-ADMIT/STADIUM-EVICT/STADIUM-RES-PULL/STADIUM-RES-PUSH/ STADIUM-HEAT@/STADIUM-HEAT!), 12 sites -- closes mama_forth_words.c and the entire kernel-only error-handling audit. All 105 sites from the §XXXII.3 triage now accounted for: 28 already correct, 75 silent sites fixed across repl.c/inference_words.c/ log_words.c/vm_core.c/mama_forth_words.c, 2 special cases resolved by dropping the error per their own documented contract, 1 resolved via console_println() per its own recursion constraint. Verified with awk: zero remaining vm->error=1 sites in mama_forth_words.c lack a diagnostic within the preceding three lines. Three-arch clean qemu acceptance passed. This closes Stage D and the USE/logging/audit thread opened in §XXXII; Stage E (human-vs-unattended identity model) remains open, not started this pass. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K |
||
|
|
90c86c6006 |
Stage D batch 6: mama_forth_words.c group 4 -- identity/crypto words (FABRIC-3.md §XXXII.6)
MINT + mint_pop_string()/ZUSE-ELIGIBILITY-ADD/ZUSE-ELIGIBLE?/ ELEVATE-PUBKEY-UNPACK, 10 sites. mint_pop_string() gained a field_name parameter so its diagnostics name which of MINT's four string arguments failed (phone/email/username/full_name), rather than a generic message that would leave the operator guessing. Diagnostic placement matched each function's own sibling convention where one exists (ZUSE-ELIGIBILITY-ADD), log_message() default otherwise. Three-arch clean qemu acceptance passed. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K |
||
|
|
8b5300fc4f |
Stage D batch 5: mama_forth_words.c group 3 -- cross-VM execution/dispatch, both flagged special cases resolved (FABRIC-3.md §XXXII.6)
VM-STEP/VM-EXEC/VM-CALL (5 sites) gained console_println() diagnostics
matching their own existing sibling guards.
SWITCH-MARK-WORK and VM-HEAT (3 sites) resolved per §XXXII.3's own
recommendation: dropped vm->error entirely rather than diagnosing it,
matching each function's own doc comment ("must never error or spam
the console" / "does not print/error"). SWITCH-MARK-WORK is the same
function §XXVIII.3 already fixed once for an off-by-one that fired
silently on every MSG-SEND in the system -- the guard now genuinely
cannot repeat that by contract, not just by the threshold being right.
VM-HEAT's guards now push 0 and return, matching its own
always-returns-a-value stack effect.
Three-arch clean qemu acceptance passed. Zuse's own WIREBIND attach
(every boot) drives MSG-SEND -> SWITCH-MARK-WORK, so this batch's most
safety-critical fix is exercised by standard acceptance, not just
compiled.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K
|
||
|
|
b42c3b195c |
Stage D batch 4: mama_forth_words.c groups 1+2 -- capsule/lifecycle words + BIRTH/START/KILL (FABRIC-3.md §XXXII.6)
15 of mama_forth_words.c's 48 silent sites fixed, split by functional
grouping per direct instruction: CAPSULE@/CAPSULE-HASH@/CAPSULE-FLAGS@/
CAPSULE-LEN@/CAPSULE-BIRTH/CAPSULE-RUN/EXEC (9), and BIRTH/START/KILL
(6).
Refined the diagnostic-placement rule: match whichever convention that
same function's other already-correct guards use, rather than
defaulting uniformly. BIRTH/START/KILL/EXEC each already had a
console_println() sibling guard ("name too long or empty") -- their
newly-diagnosed guards now match that, same shape as USE's own fix.
The CAPSULE*@ words have no sibling guard to match, so they keep
log_message() (defer_words.c's gold-standard default).
Three-arch clean qemu acceptance passed; BIRTH itself is exercised by
every boot (Hermes/Artemis birth).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K
|
||
|
|
66c2f3e539 |
Stage D batch 3: fix silent error sites in vm_core.c, incl. the two highest-value primitives; two real NULL-deref bugs found and fixed (FABRIC-3.md §XXXII.6)
All 13 silent vm->error=1 sites in vm_core.c now log a diagnostic
first: vm_enter_compile_mode, vm_compile_word, vm_compile_literal,
vm_compile_call, vm_exit_compile_mode, execute_colon_word (2 sites
each/combined), and the four fundamental memory primitives
vm_load_u8/vm_store_u8/vm_load_cell/vm_store_cell.
Found and fixed two real NULL-pointer-dereference risks while adding
the diagnostics: vm_compile_call() and vm_exit_compile_mode() each had
a combined `if (!vm || <cond>) { vm->error = 1; ... }` guard that
dereferenced vm->error even on the !vm branch of its own condition.
Split both, and applied the same defensive split to the four memory
primitives since vm_ptr()/vm_addr_ok() both tolerate vm==NULL
internally.
Live-verified, amd64: 999999999 @ . recovers correctly at Zuse's own
console (Stage A's mechanism holds), but the new vm_load_cell
diagnostic itself didn't print -- traced to memory_words.c's own
redundant, still-silent vm_addr_ok() pre-check in memory_word_fetch()
(and the same shape in memory_word_store()), which intercepts before
ever reaching vm_load_cell(). memory_words.c is vendored, out of this
initiative's scope, spun off to FABRIC-4.md -- recorded as a concrete
cross-reference for that future work rather than left to be
rediscovered.
Three-arch clean qemu acceptance passed, full POST suite included.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K
|
||
|
|
1a8c0e4fcf |
Stage D batch 2: fix silent error sites in log_words.c, resolve the flagged category-iii special case (FABRIC-3.md §XXXII.6)
log_word_set_level(), log_do_emit(), and log_emit_string() (7 sites total) now log a diagnostic via log_message(LOG_ERROR, ...) before setting vm->error, matching this file's own already-correct log_str_emit() sibling and defer_words.c's gold-standard pattern. log_word_append_raw() (backing (LOG-APPEND-RAW), 3 sites) resolved differently per §XXXII.3's own triage note: its doc comment forbids log_message() here (recursion into the log ring it writes to), but the word can also be invoked by hand at the console -- console_println() carries no such recursion risk and matches Stage B's policy for a manual interactive invocation. Added the console.h include this required. Three-arch clean qemu acceptance passed. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K |
||
|
|
5417ffb2bf |
Stage D batch 1: fix silent error sites in repl.c and inference_words.c (FABRIC-3.md §XXXII.6)
repl.c's sk_word_blk_attach_ack() and inference_words.c's array_ptr() helper + infer_word_run()'s allocation guard now log a diagnostic via log_message(LOG_ERROR, ...) before setting vm->error, matching defer_words.c's own gold-standard pattern (§XXXII.3) -- these are internal/background conditions (a malformed message-callback, a bad array reference or allocation failure), not interactive usage mistakes, so the fix keeps the fault and reports it rather than dropping it like USE's own fix did. Batched together (4 sites total, smaller combined than the next file) rather than two separate acceptance cycles for negligible size. Three-arch clean qemu acceptance passed. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K |
||
|
|
09959c6ca2 |
Stage C: primitive error-handling audit triage complete, 105 sites classified, no code changed (FABRIC-3.md §XXXII.3)
Read every one of the 105 vm->error=1 sites across the 8 kernel-only files in context (not sampled): 28 already correct (diagnostic before/ without erroring, e.g. defer_words.c's 15-for-15 gold-standard pattern), 75 silent (the exact USE-defect shape), 2 special cases requiring individual handling rather than a generic fix. Two findings flagged above the rest: SWITCH-MARK-WORK and VM-HEAT (mama_forth_words.c) each set vm->error on their own guard despite their own doc comments explicitly saying they must never error -- SWITCH-MARK-WORK is the same function whose off-by-one already caused a stray silent error to fire on every MSG-SEND once before (§XXVIII.3). And vm_core.c's four fundamental memory primitives (vm_load_u8/ vm_store_u8/vm_load_cell/vm_store_cell) silently fault on any out-of-bounds address -- the highest-reach fix candidates in the audit, hit by far more FORTH words than any single mama_forth_words.c site. Full per-file site list and per-category breakdown recorded for Stage D's own reference. Triage only -- no code changed this stage. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K |
||
|
|
c05b70c8d6 |
Stage B: logging policy documented, level-aware log-ring eviction built, LOG-FLUSH deferred again (FABRIC-3.md §XXXII.4)
Policy decided for the kernel-only audit scope: an interactive command's direct response stays on console_println/console_puts; everything else (state transitions, background diagnostics, audit trails) routes through log_message() at the appropriate level, matching capsule_mint.c's verify_mint() precedent. Documented, not code-swept here -- reclassifying individual sites is Stage D's job. LOG-FLUSH deferred again, explicitly: the per-VM log buffer its own doc comment presumes (vm_log_buffer.h) doesn't exist anywhere in the tree -- building it is real feature work needing its own scoped stage. Level-aware eviction built: log_region_append() now reads the oldest ring slot's own level before evicting it, protecting ERROR/WARN records from being pushed out by INFO/DEBUG churn -- drops the incoming low-priority record instead. Found and fixed an adjacent bug while making this change: the prior two-valued return contract would have made a benign "dropped by design" outcome indistinguishable from a genuine write failure to its one caller, which unconditionally set vm->error on any nonzero return. Changed to a three-valued contract (0 success, 1 dropped by design, -1 genuine failure). Three-arch clean qemu acceptance passed. Eviction path itself not live-exercised (needs 128+ LOG-APPEND calls to fill the ring) -- flagged, matching this project's own precedent for that kind of gap. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K |
||
|
|
5c5896fbc1 |
Stage A: fix USE's silent stack-underflow/bad-address guards, and the Hera fault-scoping gap they exposed (FABRIC-3.md §XXXII.1)
mama_word_use()'s two silent vm->error=1 guards (dsp<1 stack underflow, NULL from vm_ptr()) now print a diagnostic and return, matching the function's other five guards. Live verification of that fix alone surfaced a bigger problem: with the guard no longer silent, the REPL proceeds to interpret the leftover token as an unrecognized word, which independently sets vm->error, and sk_repl_step()/sk_repl_run()'s Hera-branch still hard-halted on that. Investigated kernel_main.c's boot/capsule-load paths directly: they already catch and clear mama->error entirely separately, before sk_repl_run() is ever entered -- so the "no fallthrough surface" halt in these two REPL functions was never protecting a boot-time fault, only an ordinary interactive REPL-turn one. Both functions now recover unconditionally on any VM's error, Hera included, matching how a redirected (WIREBIND/USE'd) identity's session already recovered. Removed the now-fully-unused sk_fault_handler(). Verified live, amd64: USE rajames at Zuse's own console prints the new diagnostic, then "VM fault -- session recovered, resuming", console stays interactive afterward. Three-arch clean qemu acceptance passed before and after the Hera-fault-scoping change. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K |
||
|
|
05159c9f9e |
Scope new initiative: unattended identity model, USE fail-closed-halt root cause, primitive error-handling audit, logging cleanup (FABRIC-3.md §XXXII)
USE crash fully root-caused against current source: mama_word_use() has seven guards, two of which set vm->error silently (stack underflow, bad VM address) while the other five print diagnostics; sk_repl_step() halts the whole kernel only when the faulting VM is Hera; Zuse's console runs directly on Hera's own VM (confirmed in capsule_zuse_boot.c), making a benign typo at her prompt the one deterministic path to the fail-closed halt. Three fix options named, none applied yet. Human-vs-unattended identity/console-birth model scoped: console/user VM pairing is pure name-convention + a per-console-VM VM-NAME-REG call, so after-the-fact console attach needs no new mechanism. Named the invariant an unattended birth path must not violate (never touch the global driving the "no thumbdrive, no prompt" gate). Error-handling audit scoped kernel-only (mama_forth_words.c + 5 kernel-only word_source files + vm_core.c/repl.c): 107 vm->error=1 sites across 8 files, counted directly. The vendored word_source sweep is explicitly out of scope here, spun off as future FABRIC-4.md work. Logging cleanup sequenced ahead of the audit's own fixes, with the two open §XXVII gaps (LOG-FLUSH doesn't exist; FIFO eviction is level-blind) named for an explicit in/out-of-scope call. Staged A-E plan proposed; nothing implemented this pass. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K |
||
|
|
16cc74243c |
std79 DoE campaign rerun post-Stage-4: 81/81, §XIV caught live by rerun's own harness bug (FABRIC-3.md §XXXI)
Reran the established 3x9x3 randomized full-factorial campaign (std79-doe.fth) on all 3 architectures per standing project discipline (any Stadium-adjacent change reruns the whole DoE from the top). First amd64 attempt exposed a real test-harness bug that re-triggered the already-known §XIV concurrent-attach gap: the sequential-attach wait loop checked for any recent WIREBIND-attached line instead of the specific identity requested, firing the next device_add before the kernel finished the current one. Only 2 of 8 identities attached; the campaign itself completed cleanly with graceful "VM-EXEC: VM not found" refusals rather than corrupting anything. Fixed the wait loop, discarded the invalid run's campaign result (its boot log kept for the record), reran clean. Corrected reruns: all 8 identities individually confirmed on all 3 architectures, zero VM-not-found errors, zero faults, 81/81 trials correct against established baseline values. Raw logs archived at experiments/std79-doe/results-20260915-stage4/. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K |
||
|
|
597f5a6cd8 |
Stage 4 verified: reap mechanism proven, real WIREBIND multiuser+multitasking confirmed live (FABRIC-3.md §XXX)
Reap mechanism (increments 2+3): temporary probe using Artemis as a safe stand-in parked identity, 3/3 checks PASS on all 3 architectures (refuse on SWITCHED_OUT, correct post-reap state, switch-signal slot released). Probe reverted, all 3 architectures re-verified clean. Live multiuser verification (increment 4, no code changes): a real previously-unattached identity thumbdrive attached via QMP on a running boot on all 3 architectures. VM-EXEC dispatch into her own live VM computed correctly, tagged with her own name in console output, full Tripod fleet unaffected. EJECT cleanly tore down both her VMs and released the switch-signal slot -- confirms increment 1's per-device fix under a real live attach/detach. Found and flagged, not fixed: interactive USE on a freshly-attached identity halts the kernel outright. Confirmed NOT caused by Stage 4 -- reproduced identically on the commit before any Stage 4 work. Direct VM-EXEC dispatch into the same identity works correctly; this is specific to the USE/BINDSTEP codepath, plausibly never caught before since every prior identity campaign used VM-EXEC, never interactive USE. Stage 4's original scope is complete: multitasking (Tripod) and multiuser (WIREBIND) are now genuinely composed, verified against real hardware-driven identity attach on all 3 architectures. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K |
||
|
|
d9da82b065 |
Stage 4 increments 2+3: WIREBIND VMs as switch-signal participants + mark-and-defer tombstone reap (FABRIC-3.md §XXVIII Stage 4)
Increment 2: WIREBIND user VMs (the ones that actually run FORTH work; console VMs are pure REPL proxies and never participate) register as Stage 3 switch-signal participants at attach, unregister at teardown. Slot table bumped 8 -> 16, matching messaging.4th's own VM-MAX -- a real, already-agreed ceiling, not an invented number. Added sk_vm_switch_signal_unregister() (compaction-based; Tripod VMs never needed removal, WIREBIND VMs cycle constantly and would otherwise exhaust the bounded table). Increment 3: implements the plan's own ratified option (A) for the async-detach UAF risk -- mark-and-defer via a new pending_reap flag on VMRegistryEntry, deliberately not a new VMState (capsule_vm_kill() already treats VM_STATE_DEAD as idempotent success, which would silently swallow a reap attempt; SWITCHED_OUT still accurately describes a tombstoned VM until the moment it's actually freed). unclean_detach() sets it when capsule_vm_kill() refuses a SWITCHED_OUT target; the Stage 3 checkpoint (vm_core.c) checks it before ever attempting to resume a pending switch target, and calls the new capsule_vm_force_reap() instead -- the one caller allowed to bypass capsule_vm_kill()'s own refusal, because it runs at the exact safe cooperative point the switcher itself controls. A new idle-tick sweep cleans up the WIREBIND live-table entry once the reap has actually happened. Verified clean on all 3 architectures (baseline regression -- no WIREBIND attach happens in a plain boot). The reap mechanism's own correctness under a genuinely parked context is verified separately, next, via a temporary deterministic probe. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K |
||
|
|
9f0f33dfc5 |
Stage 4 increment 1: per-device WIREBIND tracking, fixing a real multi-identity detach leak (FABRIC-3.md §XXVIII Stage 4)
capsule_wirebind_unclean_detach()/eject() tracked "the attached identity" as a single global, correct for the console-pairing UX (one physical console) but wrong for detach safety: since §XV/§XVI proved multiple identities genuinely live simultaneously via this same attach path, every attach after the first silently overwrote the singleton, so an unclean detach of any but the most-recently-attached identity was silently ignored -- that VM leaked forever, no trace in the log. Adds a per-device live-identity table, separate from the (unchanged) console-pairing singleton, so unclean-detach resolves any attached device to its own identity. Sized off messaging.4th's own VM-MAX (16) minus Tripod's 3 reserved slots, not an invented number. Corrects the stale "single-USB-device constraint" doc claim in capsule_wirebind.h, false since §XV/§XVI. Groundwork for Stage 4's real deliverable (WIREBIND VMs as switch-signal participants) -- this increment only fixes detach targeting; switch- signal registration is next. Verified clean on all 3 architectures (no WIREBIND attach happens in a plain boot, so this is a regression check on the existing Tripod-only path; live multi-identity verification comes with the switch-signal registration increment). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K |
||
|
|
1c220ad4b4 |
Close §XXVIII: defer 2 known bugs to Stage 4 (FABRIC-3.md §XXIX)
Bob's call after an honest end-to-end status check: multitasking (fixed Tripod fleet) and multiuser (Zuse/WIREBIND) each work for their own tested paths, but two known bugs remain rather than zero -- the §XIV concurrent WIREBIND attach detection gap, and the never-reconciled MSG-TICK/Stage-3-switch dual-ownership rough edge. Both deliberately deferred to be addressed during or at the close of Stage 4, since both bear directly on WIREBIND VMs joining the switch-signal population. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K |
||
|
|
05ae7aa886 |
Fix SWITCH-MARK-WORK off-by-one: was tripping vm->error on every MSG-SEND (FABRIC-3.md §XXVIII.2)
mama_word_switch_mark_work()'s stack-underflow guard checked dsp < 2, requiring 3+ items, when it only ever needs the 2 IDX>NAME leaves it (caddr u). Since dsp is index-based (2 items == dsp 1), this rejected every normal call. MSG-SEND tail-calls SWITCH-MARK-WORK unconditionally, so this fired on every message sent anywhere in the system -- visible only where a caller happened to check the target VM's error flag afterward (mama_word_vm_exec()'s "VM-EXEC: ERROR in Artemis" report). Verified clean on all 3 architectures: boot reaches Startup: Artemis live -> zuse@Hera] ok> with no VM-EXEC: ERROR line at all. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K |
||
|
|
66beae7fd4 |
Stage 3 follow-on: message-arrival eligibility hook + trampoline-blind switch-storm fix (FABRIC-3.md §XXVIII.2)
Implements the message-arrival eligibility signal FABRIC-3.md §XXVIII.1 left open (has_work per-slot flag, set via new SWITCH-MARK-WORK primitive from MSG-SEND) so an idle VM never becomes a switch target purely by waiting out the readiness threshold. Also root-causes and fixes a second, independent switch-storm: the tick's "who is current" check used vm_log_attributed_vm(), which can't see a VM parked in switch.c's own raw trampoline. Replaced with a dedicated g_switch_current_vm tracked by the switch mechanism itself, and moved target-slot eligibility reset to the switch decision point instead of relying on ISR polling to observe a window that can be only a few instructions wide. Verified live on all 3 architectures: clean boot to zuse@Hera] ok>, live cross-VM message dispatch, and (since a quiet log looks identical to a livelocked storm once the DoE probe is gone) confirmed genuine REPL liveness via QMP send-key + screendump on aarch64/riscv64, not log inspection alone. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K |
||
|
|
862d7d9c48 |
Stage 3 follow-on: fix stack-ownership corruption + DoE switch columns (FABRIC-3.md §XXVIII.1)
DoE CSV gained 6 switch-signal columns (switch_count_cumulative, switch_current_slot, switch_*_readiness, switch_ticks_since), and verifying them with a boot-time HB-ON probe surfaced a real livelock: the preemption checkpoint could fire inside a VM-EXEC-nested execute_colon_word() call and switch away from a stack it didn't own, parking a borrowed region of the caller's stack under the wrong VM's saved-context pointer. The trampoline bounce was the visible (safe) half of this; the corruption was the quiet half, live in every prior "clean" Stage 3 boot without ever showing up in the log. Fixed by gating the checkpoint on being at the outermost vm_interpret() call (g_vm_interpret_depth / sk_vm_at_outermost_interpret(), vm_core.c), per Bob's decision. Also fixed two related bugs found in the same pass: g_switch_back_to was a single global stale after first entry, now per-VM state (native_switch_back_to); note_switch_performed() fired on resume instead of switch-out, now called before the switch. Verified on all 3 architectures: steady log growth (no freeze), zero leaked QEMU processes, DoE columns internally consistent, Hermes/Artemis confirmed genuinely executing (not just trampoline-bouncing). Temporary HB-ON boot probe reverted after capture. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016UNhH1mhi52i6Qihh7ZV5S |
||
|
|
986d042aa7 |
Stage 3: timer-driven preemptive switching, live on all 3 arches (FABRIC-3.md §XXVIII)
Fourth stage of the preemptive context-switching plan, and the biggest. LithosAnanke now genuinely, continuously preempts between Hera, Hermes, and Artemis -- timer-driven, running live for the entire remainder of every boot once the Tripod fleet registers, not a bounded probe. A real design fork was resolved before writing code: the naive approach (the timer ISR calling Stage 2's sk_vm_context_switch() directly) is broken -- Stage 0's trap frame lives on whatever stack was active at interrupt time, and jumping to a different stack via Stage 2's own independent swap mid-handler would abandon that trap frame unresumed, guaranteed corruption on the first tick. Chose the safer of two named options: the ISR only ever sets a flag and returns completely normally through its own full epilogue; the actual switch happens moments later, via Stage 2's already-proven mechanism, at a safe cooperative checkpoint on the mainline (execute_colon_word()'s per-word dispatch loop, checked on literally every word, not throttled to the existing 256-word heartbeat-tuning cadence) -- confirmed with the user that word-level granularity is fine-grained enough given the eventual Zynq FPGA target where a word is a mnemonic. New capsule_vm_switch_signal.c/.h: a purpose-built run-readiness signal, deliberately separate from capsule_vm_physics.c's execution-heat engine (that one's own header documents itself as never touched from interrupt context, by design). Slot table sized with headroom (8) rather than hardcoded to today's 3 participants, so extending participation later is another register() call, not a redesign -- per direct request to leave room for swapping the participant set. Simple linear accumulate-then- threshold for this first cut; a fancier law can replace it later without touching the mechanism around it. heartbeat_tick() gains its one deliberate, documented amendment to this file's own top-half/bottom-half discipline -- the first time this codebase reaches into VM-scheduling state from real ISR context. Registration happens only after all three VMs are fully born, right before the REPL starts -- no critical-section protection yet against being switched away mid-birth-setup. Known, flagged rough edge (not reconciled this pass): MSG-TICK's own idle-pump and this new mechanism can still independently move control between the same VMs; not observed to interact badly in verification, but not fully unified either. Verified interactively at the console on all 3 architectures with continuous background preemption running throughout -- amd64 computed `1 1 + .` -> 2, aarch64 computed `1 1 + dup DUP * . CR` -> 4, both correct, REPL fully responsive, zero fault indicators over sustained runtime. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016UNhH1mhi52i6Qihh7ZV5S |
||
|
|
f790d0995e |
Stage 2: cooperative VM context switch primitive, proven on all 3 arches (FABRIC-3.md §XXVIII)
Third stage of the preemptive context-switching plan. The real save/restore switch mechanism now exists -- the first time anything has ever executed on a VM's own native stack (Stage 1 allocated them, unused). New sk_vm_switch_to() (switch.S, one per arch) is an ordinary function call, not an interrupt -- so unlike Stage 0's trap frame, the ABI already covers every caller-saved register; only the callee-saved set needs explicit save/restore (amd64: rbx/rbp/r12-r15, no FP at all since SysV has no callee-saved XMM; aarch64: x19-x28/x29/x30 + d8-d15; riscv64: s0-s11/ra + fs0-fs11, FS-gated like Stage 0 but read once and reused for both halves within one call, since FS is genuine global CPU state, not part of what's switched). A sibling sk_vm_switch_prime() in the same file builds the synthetic first-entry frame, kept in assembly so the layout can never drift out of sync with sk_vm_switch_to() itself. New switch.c/switch.h: sk_vm_context_switch(from, to) handles first-entry priming vs. resuming a parked context, and updates registry state (new VM_STATE_SWITCHED_OUT, distinct from VM_STATE_STOPPED -- STOPPED means no live frame, this means the opposite). sk_vm_switch_entry() is the minimal permanent trampoline every freshly-entered VM lands in: no production behavior defined yet, so it just yields straight back to whoever switched to it, forever. Closes the confirmed unguarded-KILL UAF found during planning: capsule_vm_kill(), mama_word_kill(), and capsule_vm_kill_all_nonmama() all now refuse (or silently leak rather than free, on the cold-restart path where arch_cold_reset() wipes everything immediately after anyway) tearing down a switched-out VM. Side effect found, not built on purpose: the existing MSG-TICK idle-pump already filters on VM_STATE_LIVE, so it automatically stopped dispatching into a switched-out VM with zero changes needed there. Verified via a temporary SWITCH-TEST probe (boot-triggered, since nothing can type interactively into a foreground-only QEMU session) that round-tripped a sentinel through 5 real Hera<->Hermes switches on all 3 architectures: 5/5 rounds, 0 failures, clean continuation to ok>. Probe fully reverted after capture; kernel_main.c shows zero diff. Also: Makefile.starkernel's LOADER_EXTRA_SRCS/LOADER_ASM needed the new files added explicitly (this project's "loader" PE binary is the full running kernel, not a thin bootstrap stage), and aarch64's switch.S needed the same #ifndef _WIN32 guard around .hidden that isr.S already carries (aarch64's loader assembles via clang targeting a PE/COFF target with no .hidden equivalent) -- caught by a build failure, fixed. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016UNhH1mhi52i6Qihh7ZV5S |
||
|
|
57ac3fc304 |
Stage 1: per-VM native stacks, allocated but not yet executed on (FABRIC-3.md §XXVIII)
Second stage of the preemptive context-switching plan. Every VM (Hera, every capsule_birth_baby()-born VM including WIREBIND identities) now gets its own dedicated 2 MiB native C stack at birth -- but nothing runs on it yet, that's Stage 2. Pure allocation-machinery proof. Design correction made before writing code: the plan called for cloning sk_vm_arena_alloc()'s guard-page pattern, but that pattern turns out to be Mama-only -- host_services.c's kernel_alloc() gives every baby VM a plain kmalloc() block for its dictionary arena, not a real guarded PMM allocation. Stacks get the real treatment instead (new sk_vm_native_stack_alloc()/_free() in arena.c): independent pmm_alloc_contiguous() + guard pages for every VM without exception, no singleton, no kmalloc fallback -- a stack overflow is exactly the failure mode guard pages exist for, and a corrupted stack could corrupt whatever saved context Stage 2 trusts. 2 MiB size matches this project's own established kernel-stack convention (g_kernel_stack/g_rpi5_native_stack), not a guess -- that one shared 2 MiB stack today already carries all VMs' combined nested VM-EXEC recursion. Three new VM struct fields, freed in vm_cleanup() alongside the existing call_stack free. Allocation failure is non-fatal to birth. All 3 architectures re-verified clean boot to ok>, no native-stack allocation failures for any Tripod-fleet VM. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016UNhH1mhi52i6Qihh7ZV5S |
||
|
|
15672ce17c |
Stage 0: trap-frame parity across all 3 arches (FABRIC-3.md §XXVIII)
First stage of the preemptive context-switching plan (see ~/.claude/plans/logical-snuggling-bear.md). Pure foundation work -- every arch's ISR now saves the full register set on interrupt entry, so a trap frame is in principle sufficient to resume execution anywhere it was taken. No FORTH-visible behavior changes. amd64: added FXSAVE/FXRSTOR, closing a genuine pre-existing correctness gap (not just future-preemption prep) -- confirmed live double-precision FP code reachable from ordinary interpreter dispatch (vm_runtime.c Loop #5/#6), and the ISR previously saved zero FP/SSE state. rbp repurposed as a fixed anchor so the 16-byte-aligned FXSAVE area can be carved out of an unpredictably-aligned rsp without disturbing existing argument reads. aarch64: extended the trap frame 672->800 bytes, adding v8-v15 (AAPCS64 callee-saved, previously excluded on call-site-only reasoning that doesn't hold for an async trap). riscv64: extended the trap frame 320->512 bytes, adding s0-s11 and fs0-fs11 (the latter still correctly gated behind sstatus.FS != Off). All 3 architectures re-verified clean boot to ok> under the new frames -- amd64 through hundreds of timer ticks with FXSAVE/FXRSTOR live on every interrupt, aarch64 through 987 ticks, riscv64 clean on the now-larger FS-conditional block. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016UNhH1mhi52i6Qihh7ZV5S |
||
|
|
2a30212bd3 |
Real per-VM log persistence: source attribution + ACL pin (FABRIC-3.md §XXVII)
Wires the previously-unused vm_log_attributed_vm() into LOG-APPEND's kernel primitive so persisted log records carry a trustworthy source (the real attributed VM's registry name, or "HADES" pseudo-source) instead of a caller-supplied, trivially forgeable string. Drops src-addr/src-u from LOG-APPEND's stack signature accordingly. Pins LOG-APPEND via bare ACL-PIN in Artemis's own init.4th, matching BIRTH/CAPSULE-BIRTH's precedent for a privileged word that can't reach the shared, host-portable ACL.4th. Also fixes two console-banner nitpicks: a mis-rendering em dash (U+2014) in the boot banner, and drops "Emergency" from the CLI banner text. Doc corrections to artemis_sig.h/zuse_eligibility_list.h reconciling the three fixed devblock ranges now in play. LOG-FLUSH (the intended normal entry point) and level-aware log eviction remain open, flagged not fixed. Re-verified clean boot to ok> on all 3 architectures after every change. riscv64 showed one new, unrelated virtio_blk write-timeout anomaly during Artemis's early physics self-test (self-recovered, boot unaffected, sector doesn't map to the log region) -- flagged, not investigated. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016UNhH1mhi52i6Qihh7ZV5S |
||
|
|
61755fde78 |
Artemis genesis stamp: fix a BAM-corrupting offset before it ever ran (FABRIC-3.md §XXVI follow-on, Step 3)
Step 3: one-time artemis_sig_t genesis stamp, written once
kernel_main.c's virtio-blk path confirms Artemis's own disk, so the disk
image is later recognizable generically (repl.c's idle-loop USB-MSC scan,
built in the prior commit) regardless of which bus found it.
Correction made before this ever touched the real disk: the signature's
first design (committed in
|
||
|
|
29b6789860 |
Artemis bus-agnostic discovery: signature format + idle-loop generalization (FABRIC-3.md §XXVI follow-on)
Step 1: new artemis_sig_t header format (magic 'ARTM', sibling to homeblocks_sig_t, distinct so a generic scan can tell Artemis's own disk apart from an identity thumbdrive by content alone) -- artemis_sig.h/.c, wired into Makefile.starkernel. Step 2: sk_repl_idle()'s existing per-USB-MSC-slot attach handling (the pattern WIREBIND already uses for identity thumbdrives) now also checks for the ARTM signature whenever a device's home-blocks check comes back BLANK. On a match, once Artemis's own storage-attach round-trip (HERA-BLK-ATTACH-REQ/BLK-ATTACH-ACK) confirms success, capsule_zuse_boot_load_root_pubkey() runs -- the same call kernel_main.c's synchronous QEMU-only virtio-blk path already makes, now reachable without a hardcoded PCI vendor/device scan. That function is already idempotent (no-op once zuse_root_pubkey_known is set), so no boot restructuring was needed despite the initial concern that deferring Artemis discovery to the idle loop would require one. Verified: clean build + QEMU boot to [zuse@Hera] ok> on all three architectures, zero regression to the existing virtio-blk/Zuse-thumbdrive attach path. Steps 3 (genesis-stamping onto disk/artemis.img) and 4 (growable production log-persistence region) not yet started. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo |
||
|
|
a8b16d41da |
Unify console prompt to [user@VM]; fix real personality-block truncation; correct §XXV's wrong lockdown conclusion (FABRIC-3.md §XXVI)
Investigating the std79 lockdown finding from FABRIC-3.md §XXV led to a real discovery: WIREBIND births TWO VMs per identity, a console proxy under the plain username and the actual restricted identity under <username>~user (capsule_wirebind.c). Every test in §XXV targeted the console proxy, which was never locked down at all. Retested against the correct target (rajames~user): the lockdown works exactly as designed. §XXV's "lockdown never engages" conclusion was wrong -- corrected here, not deleted, since the mistake and how it was caught are worth keeping (see the new feedback memory: confirm which specific VM a name resolves to before concluding anything, when a subsystem is known to birth more than one VM per identity). Two real, separate things found along the way are kept regardless of that correction: - capsule_runcap.c: the reserved personality devblock was read in full (mostly zero-padding after a short ~200-byte string) with no terminator, producing "WARN: block 4998 exceeds 1KB, truncating" on every std79-locked identity's birth, universal, since at least 2026-09-10. Fixed by trimming to the first NUL byte actually found -- real, but harmless to execution (real content sat in the truncated block's surviving head); it mattered for capsule_id/ content_hash being computed over padding instead of real content. - console.h/console.c/repl.c: unified the prompt from a separately- computed "[VMName] (user)" into a single "[user@VMName]" line prefix -- exactly the ambiguity that caused the original misdiagnosis (the prompt showed only the WIREBIND username, identical whether USE had targeted the console proxy or the real ~user identity). Implemented as a registered callback (console_set_user_prefix_provider()) rather than console.c calling into WIREBIND/session logic directly, since console.c is a clean HAL module with no prior dependency on capsule-level subsystems. Verified: clean build on all 3 architectures, zero new warnings, identical dict_hash/capsule_hash to every prior boot this session (console/prompt-only change). Full 9-identity messaging campaign re-run end to end: 202s, zero faults, all 8 identities at 99/99 tokens, zero regression. Also surfaced, not yet acted on: the full campaign's own console tags now visibly show which VM each identity's tests actually reached ([zuse@rajames], not [zuse@rajames~user]) -- messaging.4th's VM-NAMES-INIT registers identities by plain username, so std79-doe. fth's turn-attractor has been dispatching to each identity's console proxy, not the actual locked-down identity, since the messaging rewrite. Flagged for a deliberate decision, not investigated further. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo |
||
|
|
cb32e6632b |
Root-cause the std79 lockdown gap: likely never engages for any locked identity (FABRIC-3.md §XXV follow-up)
Following up on the messaging-words-not-denied finding: 'MSG-STATUS ACL-STD79-ALLOWED? .' sent into rajames came back UNKNOWN WORD, not "denied" -- acl-std79.4th's own supporting words were never compiled into the VM's dictionary at all. Root cause, confirmed via the raw boot log: "WARN: block 4998 exceeds 1KB, truncating" fires during every std79-locked identity's own birth. capsule_exec_payload() parses "Block N" content as everything from that header to the next "Block N" header or end-of-payload, truncating at 1024 bytes for both storage and execution. MINT_RESTRICTED_PERSONALITY is a short (~200 byte) string written into a zero-padded 4096-byte devblock at mint time; capsule_runcap_ birth() reads back the entire reserved multi-devblock region with no second "Block N" header anywhere in it to terminate "block 4998" early, so the parser treats the whole mostly-padding region as one oversized block. Confirmed universal, not one identity's quirk: the warning fires exactly 8 times in a full 9-identity boot -- once per std79-locked identity (rajames, 00-06; zuse runs natively on Hera, never through this path) -- and dates back to at least 2026-09-10 in this repo's own logs, well before this session. The std79 lockdown has likely never actually engaged, for any of the 8 locked identities, since it was built. Not fully closed to the byte: the real content sits at the start of the oversized block, inside the surviving truncated slice, so truncating the tail shouldn't by itself stop the head from executing -- the exact remaining mechanical step (stale leftover data, a line-boundary artifact, or something else) isn't nailed down yet. Explicitly not fixed -- documented per Bob's own "stop here, document it, and continue" call. Security-relevant (an intended lockdown restricting nothing) and deserves its own deliberate fix with real test coverage, not a bolt-on to an unrelated change. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo |
||
|
|
3c2daf50d1 |
Extend BIRTH/CAPSULE-BIRTH to all VMs symmetrically; flag a real std79 lockdown gap (FABRIC-3.md §XXV)
Scoping the workload-into-factorial design's placement-mode factor led to a real architectural improvement: rather than EXEC-ing a workload capsule into an already-running, ACL-locked identity's own persistent dictionary (filesystem-shaped, doesn't dodge the block- collision exposure just traced in §XXIV), a workload now runs as a fresh ephemeral child VM, BIRTH'd per trial and reaped after -- matching the project's own stated principle of automanagement over imposed policy. CAPSULE-BIRTH already passes vm->stadium_vm_id (who is birthing this VM) as the new child's parent, not a hardcoded Hera constant, confirmed by reading the C -- so a workload trial genuinely inherits the specific identity's own lineage when that identity does the birthing. Which surfaced a real premise: only Hera could call BIRTH/ CAPSULE-BIRTH at all (registered only in register_mama_forth_words(), confirmed directly, not part of the earlier §XX messaging-symmetry fix which deliberately kept this as one of her remaining privileges). Extended symmetrically now, agreed explicitly before touching code: - mama_forth_words.c: BIRTH and CAPSULE-BIRTH added to register_child_vm_words(), matching §XX's own pattern. - acl-std79.4th: ' BIRTH , ' CAPSULE-BIRTH , added to ACL-STD79-LIST (new block 4048) -- a deliberate, explicit, named exception to the lockdown's own "standard words only" guarantee, not a silent one. Symmetric registration alone can't weaken any lockdown on its own: ACL-LOCKDOWN-STD79 is allowlist-based, deny-by-default, so a newly registered word is auto-denied there unless explicitly added. Verified: clean build on all 3 architectures, zero new warnings. Hera's own dict_hash unchanged (expected); Hermes/Artemis show the same new dict_hash on all 3 architectures. Live-tested against a real attached std79-locked identity: CAPSULE-BIRTH executes correctly (returns vm_uuid_none() for a deliberately out-of-range capsule-id, zero fault, zero ACL denial). Found, and explicitly stopped short of fixing, a separate pre- existing gap while verifying the above: MSG-STATUS and MSG-K (messaging.4th words, not on the std79 allowlist) execute for a locked identity instead of being denied. ACL-LOCKDOWN-STD79 is confirmed to actually run; something more specific isn't reaching messaging.4th's dictionary entries. Root cause not traced -- needs its own investigation into vm_core.c's dictionary-link mechanics and whichever capsule actually loads messaging for these identities. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo |
||
|
|
cd2fda4351 |
Add mkcapsule --resolve: build-time claim registry for capsule block collisions (FABRIC-3.md §XXIV)
Traced what a "Block NNNN" collision actually means before designing a fix for it: capsule_loader.c's block-write path routes through the generic block-subsystem API, which kernel_main.c registers as two devices in a fixed order -- the volatile ramdrive first (LBN 2048-3071), then Artemis's real virtio-blk device immediately after (LBN 3072+, backed by disk/artemis.img). Every capsule this project has lands in Artemis's persistent range, not the ramdrive, and blk_update()'s dirty-marking + repl.c's idle-loop flush write that content through to the real disk file on every boot. A block-number collision is therefore a silent, persistent overwrite of real disk content surviving reboots, not a transient RAM mixup. The existing collision gate (check_block_conflicts(), already a hard non-interactive build failure) already catches capsule-vs-capsule collisions across the whole flat range. The real gap: zero visibility into blocks something other than a capsule owns (Artemis's own non-capsule persistent data), and no device-boundary/capacity awareness at all. Added, scoped step by step before writing any code: - tools/capsule-claims.txt -- derived, auto-created/regenerated, git-ignored. Lets --resolve tell "this capsule's own content changed" apart from "genuinely new collision with something else." - tools/capsule-reserved.txt -- human-authored, git-tracked, seeded with nothing yet rather than guessed at. Checked by both the plain build gate (new check_reserved_conflicts()) and --resolve. - tools/patches/ -- git-tracked, one file per accepted interactive renumber; a structured old->new block list, not a generic diff, since that's the only thing a renumber ever changes. - mkcapsule --resolve <dir> -- the only interactive mkcapsule mode, a deliberate separate invocation from the plain build path (which stays non-interactive so CI never blocks on a prompt). Suggests a renumbering that preserves a capsule's own existing block spacing, prompts y/N, rewrites the .4th source in place on acceptance. Found and fixed a real bug during verification: the registry's empty-block-list case (workload-5.4th, zero Block headers) serialized with a stray trailing space that the reader parsed back as a phantom block 0, causing spurious re-registration every run -- caught by testing idempotency directly, not assuming a clean first run meant it worked. Verified: isolated collision tests confirm both accept and reject paths, confirm a resolved collision doesn't re-prompt the other side, confirm reserved-range collisions are caught by both --resolve and the plain build gate. Full 3-architecture rebuild via the real Makefile.starkernel succeeded clean; amd64 boots with an unchanged dict_hash/capsule_hash from every prior boot this session. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo |
||
|
|
c8ba8832c4 |
Fix N-RUNS/START-REP hardcoded constants; N-REPS increase blocked on reservoir cost, not a bug (FABRIC-3.md §XXIII)
Two real bugs found and fixed while attempting to raise N-REPS from 3 (both harmless only by coincidence at N-REPS=3, since 3 happened to equal the hardcoded/literal values): - N-RUNS was `27 CONSTANT`, not derived -- now `N-ID N-REPS * CONSTANT N-RUNS`. - START-REP's run_id decode used a literal `3 *` where it meant `N-REPS *` (confirmed against EXEC-STD79-DOE, the serialized baseline, which correctly uses N-REPS for the same decode). Re-verified at N-REPS=3 (amd64): byte-identical to the already- verified baseline -- 193s, 0 faults, 99/99 tokens every identity, K conserved on all 656 rows. Raising N-REPS to 6 was then tested and found to cause real, silent data loss at full 8-identity scale (rajames 0/198, 00 66/198, 01/02 99/198, 03-06 fully complete) -- same failure class as §XXI defect 2, just past the budget again since doubling N-REPS roughly doubles total MSG-SEND volume (24->48). Measured the reservoir's replenishment directly rather than assume from source: 40 sends drained it from 21735 to 1; 300s of pure idle time (zero further sends) brought it back to 2049 -- real, but only ~6.8 units/s, meaning a full refill would take on the order of 53 minutes against a campaign's few-hundred-second runtime. Raising N-REPS further needs a cheaper per-send cost or an explicit top-up, not reliance on ambient decay. Reverted N-REPS to 3 (fixes kept, they're correctness fixes independent of the value) rather than commit a silently-lossy result. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo |
||
|
|
5c1b31669e |
Rename init-0.4th..init-9.4th to workload-0.4th..workload-9.4th; fix stale WL-HI/WL-LO doc
These are alternate boot personality capsules, not files doe.4th dispatches from -- confirmed by reading doe.4th itself (it generates its own synthetic workload internally, DOE-WORK) and mkcapsule.c (only the exact filename "init.4th" is special-cased as the active MAMA_INIT capsule). "workload-N.4th" names them for what they are without colliding with that reserved name. Renamed the 10 files (git mv, preserving history) and their own self-referential header comments (also fixed a pre-existing typo, "init-4.th" -> "workload-4.4th"), updated capsules/README.md, capsules/MANIFEST.md (21 references), .claude/CLAUDE.md, and a tools/mkcapsule.c comment. Also fixed experiments/bare_metal/README.md's "Adding a Custom Workload Capsule" section, discovered stale while doing this rename: it documented a WL-HI/WL-LO dispatch table and a wl_id CSV column that don't exist anywhere in the current capsules/ tree or doe.4th's own CSV header -- corrected to describe what's actually there (no pluggable workload dispatch; a custom workload is run by substituting it in as the boot's own init.4th). Verified: mkcapsule --lint clean (35 files, 0 violations), all 3 architectures build with zero new warnings, amd64 boots to zuse)ok> with an unchanged dict_hash/capsule_hash from every prior boot this session (0xc8f4b09e36f4fc4a / 0x1ef4939ed32ec1e6) and zero faults. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo |
||
|
|
71bcb72a59 |
Fix O(N) idle-loop messaging pump; full 3x9x3 turn-attractor campaign clean on all 3 architectures (FABRIC-3.md §XXII)
sk_repl_idle()'s messaging pump (repl.c) walked the entire live-VM registry every idle beat (~1Hz) and dispatched a full VM-EXEC "MSG-TICK" -- a dictionary lookup plus a 32-slot arena scan -- into every live VM, every tick, unconditionally, forever. Fine at Tripod's original 3-VM scale; a full 9-identity turn-attractor campaign exposed it as a genuine wall on riscv64 specifically (its TCG makes each dispatch cost more): the same campaign that completed in 194s/ 339s on amd64/aarch64 never finished on riscv64 at 9 VMs across three attempts, while 8 VMs there was fine in 159s. Ruled out capacity explanations before touching anything: bumping riscv64's QEMU RAM 1024->4096 changed nothing (reverted), and a live STADIUM-RES@/MSG-STATUS probe with all 9 VMs attached showed no depletion. Host memory pressure was also ruled out directly (one background task did get OOM-killed once during the investigation, but the identical stall reproduced again with 9.2GB free). The real mistake was three premature kills under 4 minutes with no way to tell "slow" from "stuck" from outside the guest -- fixed by having run_doe_batch.sh sample the qemu process's own /proc/<pid>/stat utime every 60s; with that signal, riscv64 at 9 VMs was unambiguously alive (climbing utime, no hang), just disproportionately slow going from 8 VMs (159s) to 9 (600s+ and climbing). This was never really a riscv64-only bug: an O(N) per-second walk over the full VM population doesn't scale to the hundreds of VMs this fleet is headed toward, on any architecture -- riscv64 just made it visible first, at N=9, because its per-dispatch cost is highest. Fixed by round-robin batching: the pump now dispatches to at most SK_MSG_PUMP_BATCH (4) live VMs per idle beat via a persistent cursor that resumes where the previous beat left off, instead of all of them every time. Bounds both the scan and dispatch cost to O(K) regardless of total VM count; any single VM's queue now drains roughly every ceil(N/K) beats instead of every beat, still bounded and still matching the pump's own existing best-effort contract. No new C primitives, no messaging/Stadium changes. Verified: clean build on all 3 architectures, then the full 3x9x3 campaign re-run on all 3 (not just riscv64) per the standing rule that a defect repair requires a clean re-run everywhere before anything counts as closed: amd64 198s 0 faults 99/99 tokens x8 656/656 K-conserved aarch64 339s 0 faults 99/99 tokens x8 659/659 K-conserved riscv64 178s 0 faults 99/99 tokens x8 659/659 K-conserved riscv64 went from "never completes" to faster than aarch64, same campaign, same seed, same identity set. No regression on amd64/ aarch64. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo |
||
|
|
19916717b2 |
Turn-attractor rebuilt on real messaging: two live defects found and fixed before the full campaign (FABRIC-3.md §XXI)
Rewrote EXEC-STD79-DOE-CD's dispatch to coordinate via MSG-SEND/ MSG-TICK instead of blocking VM-EXEC, now that Hera can genuinely message (§XX). RUN-TEST/EXEC-STD79-DOE (the serialized baseline) are untouched, kept as a byte-for-byte-reproducible historical comparison point. A small 2-identity smoke test before any multi-architecture commitment caught two real defects the design alone didn't predict: 1. Absolute VM-HEAT can't produce fine-grained interleaving -- a fresh identity starts at heat=0 against Hera's ~62000+, a gap no 1..24 divisor closes, so priority locked onto whichever identity had executed least, for its entire campaign. Fixed with BASE-HEAT: each identity's heat is snapshotted once at campaign start, and priority is computed from heat gained *this campaign*, not lifetime heat. 2. Single-test-per-message granularity silently drops most of a campaign's data: MSG-SEND no-ops on MSG-ALLOC failure, and the turn bookkeeping advanced regardless, so rows looked complete while missing most of their tests. A live reservoir probe showed Hera's own STADIUM-RES@ draining ~725/cycle with no replenishment observed -- a 648-send full campaign would exhaust it almost immediately. Fixed by dispatching a whole rep (24 tests, one concatenated command, measured 442 bytes, well under VM-EXEC's 1025-byte cap) per message instead -- 27 sends for a full campaign, not 648. Re-verified after both fixes: 99/99 expected test outputs present, zero drops, zero faults, genuine rep-level interleaving instead of either the serialized baseline's fixed order or the first cut's 72-test lock-in. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo |
||
|
|
7edccd2c35 |
Hera can't message: root-caused and fixed by symmetry (FABRIC-3.md §XX)
Hera never loaded common:messaging.4th, unlike every other VM in the fleet. The standing belief was this was deliberate, to avoid moving her dict_hash off baseline. Checked live instead of assumed: loading messaging.4th into her dictionary silently dropped every colon- definition referencing one of 8 STADIUM-* primitives that register_child_vm_words() gives every other VM but register_mama_forth_words() never gave her -- a missing-primitive gap, not a designed privilege boundary. Confirmed mama_word_birth is genuinely VM-agnostic and SPAWN-EVENT is an unwired placeholder before proposing the fix. Fix: register the same 8 STADIUM-* primitives for Hera, load messaging.4th from init.4th the same way Hermes/Artemis/console/mint already do, and give her own idle-loop context a direct MSG-TICK call (not VM-EXEC, which would hit the same reentrancy class the existing per-other-VM pump loop already guards against) so her own queued messages actually drain. Her dictionary is now a proper superset of every child VM's, plus her remaining extra privileges -- not structurally different from any other VM, just additionally privileged. Verified: dict_hash identical across amd64/aarch64/riscv64 (0xc8f4b09e36f4fc4a), Hermes/Artemis dict_hashes unchanged and still cross-arch identical, all three boot clean to zuse)ok> with zero UNKNOWN WORD faults, mkcapsule --lint clean. Unblocks rewriting the turn-attractor (FABRIC-3.md §XIX) to coordinate via real MSG-SEND/MSG-TICK instead of blocking VM-EXEC. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo |
||
|
|
5a9425b91b |
Add VM-HEAT primitive: groundwork for a compudynamic turn-attractor (FABRIC-3.md §XIX)
Prompted by reading the std79-doe K report: checked whether the campaign's
"real cross-VM dispatch load" was actually concurrent or strictly
serialized. It's serialized at two levels -- VM-EXEC's vm_interpret(target,
...) is a direct synchronous C call (Hera fully blocked until it returns),
and even the background physics tick (vm_tick(), vm_runtime.c) is driven
by each VM's own execution loop, so idle identities accrue zero ticks
between their own turns. K's perfect conservation (FABRIC-3.md §XVIII)
verifies sequential per-VM accounting correctness, not concurrent-access
safety, since there was never concurrent access to test.
Agreed direction: fix this without a scheduler, by reusing the same
least-dense-candidate judgment stadium_admit() already trusts for
eviction, applied to "whose turn is next" instead of "who gets evicted" --
a fleet-level turn-attractor giving the next turn to whichever live
identity currently has the lowest execution_heat_q48, no fixed round-robin,
no priorities, no preemption. Lives beside Stadium in
capsule_vm_physics.c (already the fleet-level consumer of Stadium
primitives, e.g. the K mechanism itself), not inside stadium.c ("the
floor" -- residency/eviction, a different concern from turn order) and
not a new subsystem.
This pass lands only the primitive the mechanism needs: VM-HEAT
( c-addr u -- heat-q48 ), pushing a named VM's current
execution_heat_q48 via vm_physics_heat_of() -- previously C-internal
only (doe_log_heat_by_name(), doe_log.c), never exposed to FORTH. Silent
0 on an unknown/dead name (no print/error), since a turn-attractor
scanning many candidates every turn shouldn't have to filter console
noise for names that simply aren't live. Registered everywhere
VM-EXEC/VM-CALL already are. Builds clean on all three architectures;
live-tested on amd64: Hera -> 65452, Hermes -> 43, unknown name -> 0, no
faults.
The turn-attractor loop itself (std79-doe.fth's trial ordering) is not
yet built -- next step, not done here.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo
|
||
|
|
238ca95b3c |
Add closing note to std79 DoE K report
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo |
||
|
|
66ba21adb4 |
Log fleet_k_q48/fleet_conserved; K holds exactly, 775/775 ticks (FABRIC-3.md §XVIII)
doe_log.c's per-heartbeat-tick CSV gains two columns: fleet_k_q48 (vm_physics_fleet_heat_sum() over ALL live VMs -- the genuine fleet-wide conservation invariant K, not reconstructable from the 3 named-Tripod- member heat columns already logged, which omit every identity VM's own heat) and fleet_conserved (vm_physics_conserved() as 0/1). Requested explicitly after the first heartbeat-telemetry analysis pass (analysis-20260912/) omitted K entirely. Kernel rebuilt on all three architectures, full 3x9x3 campaign rerun (results-20260912-with-k/). K = 1.0000000000 (Q48.16 raw 65536) on every one of 775 heartbeat-tick observations, sd(K) = 0, 100% fleet_conserved, across amd64/aarch64/riscv64, nine identities, three replicates -- zero deviation. Also a free regression check on both recent Stadium fixes (§XVI/§XVII): neither disturbed the reservoir-transfer accounting K depends on. Found and fixed a tooling wrinkle along the way: fleet_conserved, being the CSV row's very last field with nothing after it to bound a regex match, can have a resumed trial digit merge into it with zero separator on the wire -- combine.py now derives it from fleet_k_q48 directly (same epsilon vm_physics_conserved() uses) instead of trusting the raw field. fleet_k_q48 itself is unaffected either way. Full analysis, discussion, and light/dark SVG->PDF figures written up as a proper LaTeX report (report-20260912/report/std79_doe_report.pdf), following experiments/bare_metal/analysis/report/bare_metal_doe_report.tex's established style -- supersedes analysis-20260912/'s markdown-only first pass as the primary deliverable for this dataset (kept, not discarded). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo |
||
|
|
6573a6d1d5 |
Commit per-architecture correlated DoE CSVs, not just the merged one
These are data in their own right (the per-trial-labeled telemetry, one step before merging into combined.csv), not disposable scratch -- keep them alongside it rather than only documenting how to regenerate them. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo |