Commit Graph
256 Commits
Author SHA1 Message Date
Robert Allan JamesandClaude Sonnet 5 bbfd9103f4 Stage C: cut over BLK-ATTACH-EVENT alone -- FABRIC-3.6.md task 3.8
The reply leg (Artemis -> Hera ack) that used to flow through
common:messaging.4th's MSG-SEND/MSG-TICK now goes through
kernel-Hermes's sk_hermes_send_one()/sk_hermes_drain_checkpoint()
instead -- FORTH Hermes never sees a BLK-ATTACH-EVENT message again
(SXXXIV.2's partition rule). The request leg was never real FORTH
messaging traffic to begin with (a direct VM-EXEC, no type tag,
forced by Hera's own inability to load common:messaging.4th), so it
is untouched.

New KH-BLK-ATTACH-SEND (repl.c) wraps sk_hermes_send_one(), reached
from capsules/artemis/init.4th's HERA-BLK-ATTACH-REQ. Delivery reuses
task 3.4's already-wired sk_hermes_drain_checkpoint(); BLK-ATTACH-ACK
itself is unchanged, just reached by a different layer.
SK_HERMES_MSG_TYPE_BLK_ATTACH deliberately reuses BLK-ATTACH-EVENT's
own value (9) to document this as a cutover of the same message, not
a new one.

Two real bugs found on the way, both recorded in FABRIC-3.6.md's
findings log:
- A popped FORTH CREATE-buffer address was raw-cast to a host pointer
  instead of going through vm_ptr() -- silently read all-zero memory,
  no crash, no error, just a message that arrived and did nothing.
  Fixed; the rule and its exception (repl.c's own dev-addr is
  legitimately a raw pointer, formatted that way by its own pushing
  code) are written up for the next FORTH-facing C word.
- A separate, genuine hang on the very first live exercise of this
  path, never reproduced across ten subsequent boots. Reported, not
  chased -- not blocking, per the task's own check being otherwise
  fully satisfied.

Also found live: log_message() is invisible in this build's actual
serial-log capture at every level -- settled on a single
console_println in the real drain target instead, one line per real
USB attach, not a hot-path.

Final acceptance (logs/20260922-105501, -105758, -110304, disk images
reset before each): dict_hash identical across all three
architectures for every VM, zero UNKNOWN WORD, mkcapsule --lint
clean, real ledger+stadium_conserved(Artemis)=true evidence on every
boot.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-22 11:08:26 -04:00
Robert Allan JamesandClaude Sonnet 5 f37aa0fb17 Channel-open policy hook -- FABRIC-3.6.md task 3.7
Added HERMES-CHANNEL-OPEN? ( req-hi req-lo -- allow? ) at
capsules/ACL.4th block 4008 (default: approve everything) -- the one
word policy authors edit. sk_hermes_channel_open_policy(VM*, VMUuid)
(kernel_hermes.h/.c) is the C-side query that calls it via plain
word-dispatch against the target VM's own dictionary/stack, never
vm_interpret() (avoids task 3.4's input-buffer cursor hazard entirely)
and never decides the answer itself. Fails closed: no policy word,
a policy error, or stack underflow all deny, matching CLAUDE.md's
posture that absence of policy must never mean "always allow."

Two real bugs found and fixed before this was called done:
missing current_executing_entry assignment before calling the word's
func pointer (colon words silently no-op without it, vm_core.c:730 --
no crash, just a wrong answer); and a second FAIL with debug
instrumentation still in place whose precise cause isn't
reconstructable, since no intermediate commit exists for that attempt.

Self-test proves the task's check four ways against the same
unchanged C function: default approve, live redefinition to deny
(zero C change), restore, and a VM with no ACL.4th loaded at all
(fail closed). A fifth check wires the result into task 3.6's
sk_hermes_channel_respond() end to end: a denied policy produces a
NACK and no channel, ledger/stadium_conserved() holding throughout.

Scope, per Captain Bob's ruling: closes with the query built and
proven; sk_hermes_channel_respond() still takes a caller-supplied
approved bool rather than calling the policy internally. Wiring a
real channel-open call site to only this query is deferred to
whichever later task first needs a live decision.

dict_hash identical across amd64/aarch64/riscv64 for every VM, zero
UNKNOWN WORD, mkcapsule --lint clean (38 files, 0 violations).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-22 09:08:48 -04:00
Robert Allan JamesandClaude Sonnet 5 205a49ecd0 ACK/NACK and private-channel negotiation -- FABRIC-3.6.md task 3.6
Extracted sk_hermes_send_one() from sk_hermes_publish()'s own
per-subscriber body -- one code path for both point-to-point and
fan-out delivery, so the ledger can never diverge between them.
Point-to-point addressing turned out to be load-bearing, not
incidental: sk_hermes_publish()'s fan-out sets msg->to to whichever
member it is iterating, so a negotiation message "published" to the
common channel would spuriously reach every member, not just the real
target (checked with advisor() before building the naive version).
"Over the common channel" means every VM is reachable from birth (task
3.2), not that the exchange itself fans out -- messaging.4th's own
CH-REQUEST carried an explicit `to` for the same reason.

sk_hermes_channel_request/respond/close build the mechanics: request ->
grant (creates a private channel, subscribes both parties, sends
CH_GRANT + one ACK) or NACK ("a deny is a NACK", SXLV.1 -- no separate
type); close authorized by membership alone. The grant/deny decision is
a plain caller-supplied `approved` bool -- task 3.7 replaces the call
site that produces it with a real ACL.4th query, not this signature.

Self-test covers the task's own three checks plus a sibling advisor()
flagged: an approved respond() whose channel creation itself fails
(table exhausted) must still fall through to NACK, not a silent false
grant or half-open channel -- verified by exhausting the whole channel
table and confirming the fallback.

Bug found and fixed before this was called done: the first draft
dropped a message via pending_pop() alone, without releasing it first,
leaking its Stadium heat and failing the self-test's own ledger
baseline check (logs/20260922-065946/amd64/, kept as audit trail).
Fixed and re-verified PASS on all three architectures.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-22 07:25:51 -04:00
Robert Allan JamesandClaude Sonnet 5 8a2ee0fdad Payload bound and chunking -- FABRIC-3.6.md task 3.5
sk_hermes_publish() now enforces SK_HERMES_CHUNK_MAX_PAYLOAD (1024) on
every message's payload_len uniformly, chunked or not -- closing the
gap task 3.3 explicitly parked. A chunk carrier is
[SkHermesChunkHeader][content slice], slice capped at
1024 - sizeof(header) rather than 1024 itself, so every message on the
wire satisfies the same one-block bound vm_interpret()'s own drain
limit already requires -- a future chunk-aware drain never has to
special-case a carrier that can't be handed to vm_interpret() as-is.

Deliberately no chunking-sender API: building one would need
kernel-Hermes to own chunk-buffer memory with a real lifetime it has no
way to track (kept alive until every subscriber drains it). Sending is
a loop pattern a caller writes with sk_hermes_chunk_count() +
sk_hermes_publish(), demonstrated by this task's own self-test.
sk_hermes_reassemble() is pure and memory-agnostic: validates msg_id
agreement, exact seq coverage, and per-chunk slice sizes before a
single memcpy, with the total length computed once and checked against
the caller's buffer once -- never order-dependent on which chunk
happens to overflow.

Verified live on all three architectures: a 1024-byte payload as one
message, a 1025-byte send refused outright with the ledger untouched,
and a 3000-byte payload split into 3 chunks, drained, and reassembled
byte-exact against the original. dict_hash unmoved and identical
across architectures.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-22 06:51:48 -04:00
Robert Allan JamesandClaude Sonnet 5 2b1ba031a5 Drain at the outermost checkpoint -- FABRIC-3.6.md task 3.4
sk_hermes_drain_checkpoint() interprets one queued payload per checkpoint
(ruled: one message per checkpoint), reusing sk_vm_at_outermost_interpret()
and placed before the switch-signal block in vm_core.c's existing
cooperative checkpoint (sk_vm_context_switch() doesn't return until
switched back to, so drain must come first or it silently never runs on
a switching checkpoint).

Amends FABRIC-3.5.md SXLIII.5, caught by advisor() before writing the
naive version: "recursive drain is prevented for free" via
g_vm_interpret_depth is true but only for same-message re-drain -- it
doesn't cover the separate same-VM reentrancy hazard FABRIC-3.md SXX
already named for Hera specifically (VMCallState saves rsp/exit_colon/
ecw_nesting only, never input_buffer/input_length/input_pos). Draining
calls vm_interpret() on the same vm whose own vm_interpret() call is
still paused mid-word at the checkpoint; without saving and restoring
the cursor by hand, the enclosing REPL line or LOAD block would be
silently truncated. sk_hermes_drain_checkpoint() snapshots and restores
input_buffer/input_length/input_pos/mode/error/abort_requested around
the call. Not a divergence from the ruling -- cursor preservation is the
implementer's own obligation inside the ruled mechanism.

Gated behind a system-wide pending-total counter so the common
no-message-in-flight case costs one integer read per word dispatch, not
a stadium_max_vm_count()-sized queue scan (also flagged by advisor() as
a real hot-path cost, not deferred).

Verified live on all three architectures: a self-test publishes a real
payload to Hermes, proves the depth gate via VM-EXEC-ing an existing
harmless colon word into Hermes (genuine nested vm_interpret(), depth 2,
must not drain), then drains directly from genuinely-outermost context
and confirms exactly one clean drain. dict_hash unmoved and identical
across architectures.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-22 01:27:11 -04:00
Robert Allan JamesandClaude Sonnet 5 1f6343bc03 Publish path, no dispatch -- FABRIC-3.6.md task 3.3
sk_hermes_publish() allocates one SkHermesMessage per channel member
(heat-cost ruling 2026-09-21: one message per subscriber, funded by
the publisher's own reservoir) and enqueues each onto a new
per-subscriber SkHermesPendingQueue -- found-or-created lazily by
vm_id, sized from stadium_max_vm_count() like the channel/switch
tables. Best-effort across subscribers: a failed allocation or full
queue skips and rolls back just that one subscriber, not the whole
publish -- the natural reading of "ledger and stadium_conserved() hold
across N publishes to M subscribers" (the task's own check), not a
separate ruling.

Dispatches nothing -- sk_hermes_pending_count()/peek()/pop() are the
read/drain primitives task 3.4's real checkpoint-driven drain will
build on; this task's own self-test uses them directly since no
checkpoint hook exists yet.

Verified live on all three architectures: pending-queue table sized
50/202/50 slots (tracking the channel table's own per-arch sizing), a
synthetic publish self-test (2 publishes to 3 subscribers) confirms
exact per-subscriber delivery counts, and ledger/stadium_conserved()
invariants hold both mid-publish and after manually draining every
queue back to baseline.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-22 00:49:22 -04:00
Robert Allan JamesandClaude Sonnet 5 a4afdfa591 Channel table + common channel, inert (B1) -- FABRIC-3.6.md task 3.2
Adds SkHermesChannel: a channel is an index into a boot-time,
stadium_max_vm_count()-sized table (same sizing pattern task 3.1
established for the switch table -- no separate numeric rule was ruled
for this table, so task 3.1's bound is extended directly, flagged as
such rather than restated as a new ruling). No name field, mirroring
messaging.4th's own nameless CH-ARENA.

The common channel (index 0) is created at boot and permanent. Hera
subscribes explicitly in kernel_main.c (she is the one VM never born
through capsule_birth_baby()); every other VM -- Tripod fleet and
future WIREBIND identities alike -- subscribes inside
capsule_birth_baby() itself, the single choke point every other birth
already passes through.

Inert: no publish, no dispatch, no ACK/NACK, no ACL hook (tasks 3.3,
3.6, 3.7). Verified live on all three architectures: channel table
sized to 50/202/50 slots (matching switch-signal's own per-arch
sizing), common-channel fleet self-test confirms all four Tripod
members are members, and a synthetic create/subscribe/unsubscribe/
destroy round-trip against a private topic passes, including refusing
to destroy the common channel.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-22 00:25:13 -04:00
Robert Allan JamesandClaude Sonnet 5 c19ef365fe Dynamic switch table (B2) -- FABRIC-3.6.md task 3.1
Replaces the fixed SK_SWITCH_MAX_SLOTS=16 compile-time array with a
boot-time, RAM-derived allocation via a new sk_vm_switch_signal_boot_init(),
kmalloc'd to stadium_max_vm_count() entries -- the same pattern
session_boot_init() already established for Stadium-derived sizing.
Every switch-signal participant is a Stadium VM, so this reuses that
bound directly rather than deriving a separate one.

Verified live on all three architectures: switch table sized to 50
slots (amd64), 202 slots (aarch64), 50 slots (riscv64) -- all well
past the old fixed cap. All three boot to [zuse@Hera] ok> cleanly;
dict_hash for Hermes/Hestia identical across architectures, unmoved
from pre-task values.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-21 23:59:50 -04:00
Robert Allan JamesandClaude Sonnet 5 feace42397 Add diagnostic scan cross-check of the Hermes counters -- FABRIC-3.6.md task 2.8
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-21 08:46:07 -04:00
Robert Allan JamesandClaude Sonnet 5 2a2bf6eb35 Stage B proof: add per-VM consumed term to stadium_conserved -- FABRIC-3.6.md task 2.7
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-20 21:31:45 -04:00
Robert Allan JamesandClaude Sonnet 5 313ffc89e6 Add exact-equality Hermes ledger self-audit -- FABRIC-3.6.md task 2.6
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-20 21:02:27 -04:00
Robert Allan JamesandClaude Sonnet 5 d11eb2e5db Add sk_hermes_decay(), ledgering decay into consumed -- FABRIC-3.6.md task 2.5
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-20 18:48:46 -04:00
Robert Allan JamesandClaude Sonnet 5 493d410028 Add the four ledger counters (held/pulled/returned/consumed) -- FABRIC-3.6.md task 2.4
FABRIC-3.5.md SXL.4's ledger: held == pulled - returned - consumed,
epsilon zero. Added sk_hermes_ledger() (an accessor, not a mutator)
plus four static counters in kernel_hermes.c.

Each counter has exactly one increment/decrement site: held/pulled
both move at sk_hermes_alloc()'s single success path, after every
refusal branch has already returned; held/returned both move at
sk_hermes_release()'s single success path. consumed is declared and
always reads 0 -- its one increment site doesn't exist yet, and won't
until task 2.5 gives decay something to record.

sk_hermes_release() now reads the Stadium cell's live header.heat
immediately before calling stadium_evict(), rather than assuming the
original pulled amount -- stadium_evict() zeroes the header as part of
freeing the cell and its own return value is a success code, not the
credited amount, so this is the only point the true remaining heat is
available. Today this always equals the original Q.SLOT pull; once
task 2.5's decay exists, this is what keeps returned correct without
touching this function again.

Self-test (kernel_main.c) extended: snapshots the ledger before
running so it checks its own deltas, verifies held/pulled grow by
exactly got_n * Q_SLOT on allocation with returned/consumed untouched,
then verifies held returns to its starting value and returned grows by
the same amount on release, and checks the audit invariant itself as a
bonus (task 2.6 formalizes this properly).

Three-arch boot clean: amd64/aarch64/riscv64 all reach [zuse@Hera] ok>,
zero UNKNOWN WORD, dict_hash unmoved. All three print PASS with
identical final ledger: held=0 pulled=65536 returned=65536 consumed=0.
No compiler warnings.

Authorized by Captain Bob ("keep going with rhe 6.5 document").

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-20 05:59:09 -04:00
Robert Allan JamesandClaude Sonnet 5 2c1dc1753a Correct sk_hermes_alloc() to admit a real Stadium patron; add sk_hermes_release() -- FABRIC-3.6.md tasks 2.2 (amended) + 2.3
Real finding, caught before building release on a foundation that
couldn't support it: task 2.2's first cut of sk_hermes_alloc() pulled
reservoir heat but never admitted a real Stadium-floor patron -- just a
local in_use flag. FABRIC-3.5.md SXXXIII.4 item 1 says MSG-FREE-NODE
returns heat "via STADIUM-EVICT", which only means something if
allocation admitted something. SXL.4's own invariant, Sigma(resident
patron heat) + reservoir + consumed == Q48_ONE, cannot balance if held
heat is invisible to every term while held. Flagged to Captain Bob
before proceeding; authorized to correct 2.2 in the same pass as
building 2.3 on top of the fix.

sk_hermes_alloc() now calls stadium_admit() with heat = the pulled
amount, behaviour = STADIUM_BEHAVIOUR_DELIVER (matching messaging.4th's
own SB-DELIVER STADIUM-ADMIT exactly), and identity = the message's own
slot index (matching the FORTH precedent -- caught live in the first
boot of this fix that omitting this made stadium_dispatch()'s existing
DELIVER diagnostic print msg_idx=0 for every message instead of a
distinct value). The returned cell index is stored in the message's
own stadium_cell field. Stadium-floor refusal (independent of reservoir
affordability) rolls back the pull the same way the other refusal
paths already do.

sk_hermes_release() -- the function task 2.3 actually asks for -- calls
stadium_evict() on that cell, which itself returns the departing
patron's remaining heat to its owning VM's reservoir, matching
MSG-FREE-NODE's exact shape. Release does not touch the reservoir
directly.

Self-test (kernel_main.c) extended: keeps every allocated message's
pointer, allocates to exhaustion as before, releases all of them, and
checks the reservoir returns to precisely its starting value.
"Undecayed" is true by construction (no decay/TTL logic exists yet,
task 2.5) -- exact restoration, not approximate.

Three-arch boot clean: amd64/aarch64/riscv64 all reach [zuse@Hera] ok>,
zero UNKNOWN WORD, dict_hash unmoved from task 2.2. All three print
identical PASS arithmetic: reservoir0=65536, reservoir_after_alloc=0,
reservoir_final=65536. No compiler warnings.

stadium_dispatch()'s DELIVER-case console output (one line per
eviction) is pre-existing instrumentation, not new -- confirmed real
and load-bearing per stadium.c's own comment, verbose but expected.

Authorized by Captain Bob ("yes").

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-20 04:20:14 -04:00
Robert Allan JamesandClaude Sonnet 5 9d129fdb1c Add sk_hermes_alloc(), the heat-coupled allocator -- FABRIC-3.6.md task 2.2, item 28
The piece FABRIC-3.5.md SXXXIII.6 calls "what remains genuinely hard,"
built and proven first per its own recommendation. Added
src/starkernel/vm/kernel_hermes.c (wired into Makefile.starkernel's
LOADER_EXTRA_SRCS -- this repo lists vm/*.c files explicitly, no glob)
and sk_hermes_alloc()'s declaration in kernel_hermes.h.

Checks stadium_reservoir_peek(vm_id) >= SK_HERMES_Q_SLOT before
touching the reservoir at all -- refusal this way needs no rollback,
since nothing was pulled -- with an explicit rollback path
(stadium_reservoir_push) kept defensively for the pull-then-short case,
though nothing in this single-core kernel is expected to reach it.

SK_HERMES_Q_SLOT = Q48_ONE / SK_HERMES_MSG_MAX (2048), deliberately
simpler than messaging.4th's own formula, which reserves a Q.1/3 floor
for COMMON-CH's own Stadium heat -- kernel-Hermes has no such object
(SXXXIII.4/SXXXIII.5's flat membership list carries no heat of its
own), so there is nothing left for that floor to protect.

Self-test in kernel_main.c, same diagnostic-only synthetic-VM pattern
as the existing Stadium quota grant self-test (lo=3, distinct from
that test's lo=1): reads back the actual granted reservoir rather than
assuming a number, derives expected_n from it, allocates to refusal,
and checks the refusal lands at exactly expected_n, the reservoir
doesn't move on the refused attempt (rollback proven, not assumed),
and the final reservoir is exactly reservoir0 minus got_n times
Q_SLOT.

Three-arch boot clean: amd64/aarch64/riscv64 all reach [zuse@Hera] ok>,
zero UNKNOWN WORD, dict_hash unmoved from task 2.1 (pure C, no FORTH
touched). All three print identical self-test arithmetic: reservoir0=
65536 Q_SLOT=2048 expected_n=32 got_n=32 reservoir_after=0. No compiler
warnings.

Noted, not fixed: Q_SLOT's divisor and SK_HERMES_MSG_MAX are the same
32, so reservoir and arena exhaustion land at exactly the same count by
construction -- this test can't distinguish which refusal reason
fired, only that refusal is correct and rolls back correctly.

Deliberately not evidence for stadium_conserved(): allocating alone
(no release yet, task 2.3) leaves pulled heat held off the Stadium
floor, so the two-term check would correctly read false right now if
run mid-hold. That's expected, not a bug -- Stage B (task 2.7) is
defined as "before and after the alloc/free cycle," not "continuously
during." This task's self-test checks reservoir arithmetic directly
instead.

Authorized by Captain Bob ("Yes continue").

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-19 21:00:20 -04:00
Robert Allan JamesandClaude Sonnet 5 1e75bb8039 Assert Hestia's headless invariant -- FABRIC-3.6.md task 1.9, Phase 1 closed
Audited first: neither capsules/hestia/init.4th nor her birth block in
kernel_main.c references g_wirebind_attached_username, CONSOLE-ATTACH,
MINT, or any proxy-minting mechanism -- the invariant already held
structurally, by absence. Stated it explicitly anyway, per the task:
added a comment at Hestia's birth site quoting FABRIC-3.5.md SXVIII.6's
invariant verbatim, warning future edits not to add console/wirebind/
proxy code there without re-reading it first.

Verified live with the actual no-thumbdrive boot
(ARCH=<arch> qemu ZUSEDISK=), not the default. All four VMs born
successfully on all three architectures, zero UNKNOWN WORD, and zero
ok> occurrences anywhere in any of the three full logs -- genuinely
silent, matching sk_repl_headless_wait()'s own documented "no banner,
no prompt, no input surface at all." Watched each log's line count
post-birth for 8-10s to confirm it stayed flat rather than eventually
printing something late.

Confirmed no regression on the standard (with-thumbdrive) path on
amd64: dict_hash identical to task 1.8's baseline. Did not repeat that
check on aarch64/riscv64 -- the change is a comment only, cannot
diverge by compiler, and the headless invariant itself was already
proven identically on all three.

Phase 1 is now fully closed (tasks 1.1-1.9). Tripod is Hera/Artemis/
Hestia plus Hermes (retained through Phase 1-3 per SXXXIV.4); Hestia
owns the drawing fabric exclusively; headless-until-login intact with
Hestia in the fleet. Phase 2 is next.

Authorized by Captain Bob ("yes").

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-19 17:46:32 -04:00
Robert Allan JamesandClaude Sonnet 5 6c6293cd52 Move PLOT/FB-WIDTH/FB-HEIGHT registration to Hestia only -- FABRIC-3.6.md task 1.8
Real fix for task 0.8's finding. register_framebuffer_words() was
called unconditionally from register_forth79_words() (word_registry.c),
itself called unconditionally from vm_init_with_host() -- the generic
per-VM bootstrap every VM goes through, with no way to know a VM's
name at that point.

Removed the unconditional call. Added a name-gated call instead in
capsule_birth_baby() (capsule_birth.c), beside the existing
is_fleet_foundation check -- the one place in the birth path where
capsule_name and the newly-allocated VM* are both in scope together:
Hestia gets register_framebuffer_words(), nobody else does.

Positively verified live, exactly as the task's own check demands:
FB-WIDTH via VM-EXEC returns UNKNOWN WORD in Hermes and Artemis, 1280
in Hestia. Confirmed at the fundamental level too: Hera's own base
PARITY:M7.1a word count dropped from 531 to 528 -- exactly the three
words removed -- identical across all three architectures.

Three-arch boot clean: amd64/aarch64/riscv64 all reach [zuse@Hera] ok>,
zero UNKNOWN WORD during boot. mkcapsule --lint capsules/ clean, 38
files / 0 violations (pure C change, no capsule content touched). No
compiler warnings.

Side effect on the vendored hosted build, expected and not a
regression: capsule_birth.c is kernel-only, so the hosted starforth
binary has no Hestia concept and now never registers these words at
all -- confirmed live. framebuffer_words.c's own top comment already
calls this surface "kernel-only, no-op on hosted builds," so the prior
stub registration was already vestigial. Ran a plain `make` sanity
build per CLAUDE.md's own stated purpose for that target; regenerated
lfs/amd64/starforth included here rather than left stale.

Authorized by Captain Bob ("yes").

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-19 17:36:18 -04:00
Robert Allan JamesandClaude Sonnet 5 487769e18a Register Hestia for switch signals -- FABRIC-3.6.md task 1.5
Added a fourth capsule_vm_find_by_name_nocase("Hestia", ...) +
sk_vm_switch_signal_register(...) block in src/starkernel/kernel_main.c,
same shape as the existing Hermes/Artemis blocks, placed after all four
fleet members are confirmed born -- the existing comment on this block
already states why: no critical-section protection during setup, so
registering earlier risks the signal firing mid-birth.

Three-arch boot clean: amd64/aarch64/riscv64 all reach [zuse@Hera] ok>,
zero UNKNOWN WORD, hashes identical to task 1.4's baseline (this is
pure C runtime state, doesn't touch any FORTH dictionary). No compiler
warnings.

Took FABRIC-3.5.md SXXXV.0's "invisible by default" warning literally
rather than trusting a clean boot log alone: SXXVIII.2's own recorded
switch-storm signature is "QEMU pinned near 100% CPU, serial log frozen
solid," not an error message. Confirmed normal wall-clock boot time on
all three (~30s) and, since TCG itself always shows ~100% CPU
regardless of guest workload, watched each serial log's line count at
the idle prompt for 5-10s and confirmed it stopped growing rather than
flooding or silently stalling.

Authorized by Captain Bob ("yes").

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-19 17:07:19 -04:00
Robert Allan JamesandClaude Sonnet 5 10e2d654e5 Birth Hestia in kernel_main.c -- FABRIC-3.6.md task 1.4
Added a birth block immediately after Hermes's own, same shape:
S" Hestia" BIRTH followed by a registry-lookup confirmation. Fleet is
now Hera/Hermes/Artemis/Hestia, four VMs, through Phase 1-3
(FABRIC-3.5.md SXXXIV.4) until Phase 4 retires Hermes.

Three-arch boot clean: amd64/aarch64/riscv64 all reach [zuse@Hera] ok>,
zero UNKNOWN WORD. Registry shows all four (BIRTH: Hermes live, BIRTH:
Hestia live, PARITY:BIRTH for all three non-Hera VMs). Hestia's
dict_hash identical across all three architectures (0x31cab513929eea89).
Hera/Hermes hashes unchanged from task 1.3; Artemis's vm_id shifted
(now the 4th birth instead of 3rd -- sequence-derived, not identity-
derived, so expected) but its dict_hash is unchanged and still
identical across arches. No compiler warnings.

Noted, not a regression: Hestia's birth log shows the same
"( Unterminated comment" HADES warning Artemis's birth has shown since
task 0.0's first baseline.

Authorized by Captain Bob ("yes").

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-19 17:01:05 -04:00
Robert Allan JamesandClaude Sonnet 5 ff00de9a63 Add Hestia to is_fleet_foundation -- FABRIC-3.6.md task 1.3
Fourth vm_name_prefix_eq_nocase(capsule_name, "Hestia") check alongside
Hera/Hermes/Artemis in src/starkernel/capsule/capsule_birth.c's
is_fleet_foundation local -- Hermes retained, per FABRIC-3.5.md
SXXXIV.4 (he stays live and fleet-foundation through Phase 1-3). The
flag's only consequence is session_set_pinned(vm_id, 1) for whichever
VM name matches.

Three-arch boot clean: amd64/aarch64/riscv64 all reach [zuse@Hera] ok>,
zero UNKNOWN WORD, byte-identical to the pre-change baseline -- expected,
since nothing births anything named "Hestia" yet (task 1.4), so the
added name never matches. Hermes confirmed still present in the check
and still born normally in all three logs.

Authorized by Captain Bob ("yes").

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-19 16:53:40 -04:00
Robert Allan JamesandClaude Sonnet 5 380f0a09c9 Add stadium_conserved() -- FABRIC-3.6.md task 0.7, item 41
Boolean analogue of vm_physics_conserved(), for the Stadium per-VM
quota invariant rather than fleet-wide execution heat
(FABRIC-3.5.md SXXXIX.4). int stadium_conserved(VMUuid vm_id), in
src/starkernel/vm/stadium.c alongside stadium_resident_sum()/
stadium_reservoir_peek() that it's built from, declared in
include/starkernel/vm/stadium.h.

Implements the two-term form -- resident_sum(vm_id) +
reservoir_peek(vm_id) == Q48_ONE -- not the three-term form SXL.4
rules for the eventual system. That ruling's `consumed` term is a
Phase 2 kernel-Hermes ledger deliverable that doesn't exist yet:
nothing draws on any VM's Stadium quota today (task 2.2 is literally
where that wiring gets built), so consumed is honestly zero right now.
Folding it in as a placeholder would be inventing Phase 2 state ahead
of it existing -- the doc comment says so explicitly, so whoever
builds Phase 2's ledger extends this function rather than working
around it.

Wired into the existing per-VM boot diagnostic
(stadium_words_print_boot_diagnostics(), kernel_main.c:810, Hera
only -- the sole existing call site) rather than adding a new one,
printing CONSERVED/DRIFTED the same shape vm_physics_status() already
uses.

Three-arch boot clean: amd64/aarch64/riscv64 all reach [zuse@Hera] ok>,
zero UNKNOWN WORD, all print "Stadium conservation: CONSERVED" with
identical resident_sum=47641 reservoir=17895 sum=65536=Q48_ONE. No
compiler warnings on either edited file (forced recompile checked).

Authorized by Captain Bob ("Yes.").

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-19 13:56:14 -04:00
Robert Allan JamesandClaude Sonnet 5 41918a28a4 doe_log.c: add vm_name/vm_id_hex identity columns; new calibration workload
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
Adds the identity columns HB-ON's existing per-tick CSV logger
(doe_log_tick_row(), doe_log.c) was missing. Previously every column
described *a* VM's state each row, but nothing said which VM emitted
it -- concurrent VMs' rows were indistinguishable by source. Two new
first columns, vm_name and vm_id_hex, resolved via a reverse lookup on
vm->stadium_vm_id (capsule_vm_registry_get()).

Real bug found and fixed while building this: the freestanding
snprintf here silently prints the literal format string instead of
substituting for %016llx (width+ll+hex unsupported) -- caught by
reading the actual emitted row, not assumed to work. Replaced with a
hand-rolled hex nibble-table loop, the same idiom vm_uuid_format()/
MINT-SCRATCH-EMIT already use.

Corrects FABRIC-3.md's own prior "still not started: CSV driver"
framing: no bespoke CSV emitter is needed for the multiuser DoE at
all -- this per-tick logger already exists, fires automatically inside
every VM's own execution loop, and just needed HB-ON plus these
identity columns to be usable for concurrent workers. Also corrects
doe_log.h's own stale doc comment claiming g_doe_log_enabled defaults
to 1 -- doe_log.c's own source is authoritative: 0, off by default,
matching the HB-ON/HB-OFF naming.

One real, honest limitation recorded rather than smoothed over:
vm_name reads blank for any tick captured during a WORKER-BIRTH'd
capsule's own self-execution at birth time, since capsule_birth_baby()
runs that work (and its heartbeat ticks) before the registry name can
be set. vm_id_hex is unaffected (reads directly from vm->stadium_vm_id)
and still uniquely disambiguates every row -- confirmed live: a
blank-name row's own vm_id_hex matched its later PARITY:KILL line's
vm_id exactly.

New workload-calib1.4th: Bob confirmed the existing 10 workload-N.4th
files aren't fixed. Sized to cross HEARTBEAT_CHECK_FREQUENCY (256
word-executions/tick) many times over while staying far lighter than
RUN-FIB/RUN-CHAOS5, both too slow under TCG for a bounded verification
run. A real, reusable addition, not a throwaway.

Verified live on amd64 via a self-contained one-shot EXEC'd test
(HB-ON, WORKER-BIRTH two calibration workers, HB-OFF, VM-ERROR? on
both, KILL both, completion banner) rather than interactive polling,
which breaks once HB-ON makes the log grow continuously via Hera/
Hermes/Artemis's own background ticks. Clean 3-arch qemu boot on the
real committed change.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K
2026-09-16 09:53:25 -04:00
Robert Allan JamesandClaude Sonnet 5 e33eb36361 Multiuser DoE punch list: WORKER-BIRTH + VM-ERROR? + vm_physics_init fix
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
Core mechanism for §XXXIII's main concurrency block, built and
live-verified on amd64. Two new primitives:

- WORKER-BIRTH ( capsule-c capsule-u name-c name-u -- ok? ): births a
  named, VM-EXEC-addressable VM from an arbitrary (p) capsule with no
  identity involved. Corrects the ratified design's own assumption that
  the main block would use UNATTENDED-BIRTH -- that requires a
  committed capsule per identity, impractical for dozens of trial VMs.
  The concurrency block never needed identity at all.
- VM-ERROR? ( c-addr u -- flag ): reads a named VM's error state from
  Hera, mirroring VM-HEAT's silent/always-returns-a-value contract.
  Needed to check a VM-EXEC-driven trial VM's own fault state after
  the fact -- nothing existing let Hera do this.

Real bug found and fixed in both WORKER-BIRTH and UNATTENDED-BIRTH:
capsule_birth_baby() never calls vm_physics_init() either (same shape
as the registry-name gap found building UNATTENDED-BIRTH) -- without
it a born VM is never in the VM Fleet Attractor physics list, so
VM-HEAT returns 0 forever regardless of work done. Fixed by adding
vm_physics_init() alongside the existing registry-name call in both
words.

Traced (not guessed) why heat still read 0 after one VM-EXEC touch
even post-fix: vm_physics_touch()'s transfer logic only fires from a
VM's *second* touch onward -- the first touch just records a baseline
tick. Verified live across three sequential touches: heat 0 -> 7039 ->
11333. This is a real design requirement for the DoE's heat/CV
response variable (each trial must touch a worker at least twice), not
a bug to route around.

Also found live: all 10 existing workload-N.4th capsules self-execute
their full workload at load/birth time (a bare top-level call to their
own RUN-* word at file end) -- missed on an earlier, too-shallow
8-line survey of each file. WORKER-BIRTH alone already runs a worker's
first pass as a side effect of birth.

Verified live on amd64: two concurrent workers (fib + matrix-mul)
birthed, run, measured (heat + error state), and killed cleanly;
VM-HEAT/VM-ERROR? both confirmed silent-0 on an unknown name. Clean
3-arch qemu boot on the real committed change.

Still open: the run-matrix/shuffle/CSV driver capsule itself, the
WIREBIND-automation path for the fixed arm, and the per-VM touch-count
budget's exact value.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K
2026-09-16 08:51:21 -04:00
Robert Allan JamesandClaude Sonnet 5 af351bff92 Punch list item 4 (final): CONSOLE-ATTACH, full unattended-identity flow verified end-to-end
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
Closes FABRIC-3.md §XXXII.2's punch list. CONSOLE-ATTACH ( name-c
name-u -- ok? ) pairs a fresh console VM to an already-live VM
registered as "<name>~user". Deliberately a plain, unconditional
primitive with no VMIdentity capability-bit check -- corrects this
session's own first-pass design (§XXXII.2 amended in the same commit):
identity.installed is 0 for Hera/Hermes/Artemis and for every
console-proxy VM, so a capability-bit gate would be unreachable for
every VM a human actually types at, Zuse included. Matches
ZUSE-ELIGIBILITY-ADD's own "no bespoke gate" precedent in this file;
real gating is `' CONSOLE-ATTACH ACL-PIN` in ACL.4th if ever wanted.

Two real bugs found and fixed via live testing, not assumed correct:
- capsule_birth_baby() never sets a VM's registry name (documented
  gap, same one capsule_runcap_birth()'s own history already hit) --
  UNATTENDED-BIRTH gained a second `name` argument and now calls
  capsule_vm_registry_set_name(new_vm_id, "<name>~user") itself.
- CONSOLE-ATTACH's first draft took an independent console name from
  the target's name. sk_repl_dispatch_line()'s pairing check (repl.c)
  reconstructs the target as console_get_vm_name()+"~user" -- a
  mismatched console name silently falls back to direct interpretation
  with no error. Caught live (typed `5 6 + .` at a mismatched console,
  got a direct `11` instead of a relay) and fixed by collapsing to one
  name argument, matching WIREBIND's own by-construction invariant.

Full end-to-end live verification on amd64: minted a test identity,
UNATTENDED-BIRTH'd it as "bob", CONSOLE-ATTACH'd a console named "bob",
USE'd it, typed `5 6 + .` -- no direct output at [zuse@bob] (relay
path taken), then `[zuse@bob~user] 11 ok>` appeared: the relayed
command executed on the target identity VM itself and printed its own
answer back through the shared console. Hera stayed healthy throughout
(2 2 + . -> 4 after switching back). CONSOLE-ATTACH also verified to
refuse cleanly on a nonexistent target with no orphaned VM. Test
capsule reverted after capture per this project's probe convention --
never committed.

FABRIC-3.md §XXXII.2 fully closed: all 4 original questions ratified,
the mid-course drive_uuid and ACL-bit corrections both recorded
plainly rather than silently folded in, and a doc-accuracy note left
for CLAUDE.md's own stale "1024-byte block limit" framing (mkcapsule's
real limits are range [2048,5120) and 16 content lines/block) --
flagged, not fixed, out of this punch list's scope.

Clean 3-arch qemu boot (amd64/aarch64/riscv64) on the real committed
C-only change.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K
2026-09-16 07:24:27 -04:00
Robert Allan JamesandClaude Sonnet 5 5879c8b3bc Punch list item 3: UNATTENDED-BIRTH call site, verified live end-to-end
Implements the unattended-birth mechanism FABRIC-3.md §XXXII.2 designed:
UNATTENDED-BIRTH ( name-c name-u -- ok? ) births a VM from a named (p)
capsule via capsule_birth_baby() (completely unmodified, the same
generic build-time-capsule path CAPSULE-BIRTH already uses), then
installs its identity the same way capsule_wirebind.c already does
live for WIREBIND attaches -- vm_identity_from_cert() verification
followed by a plain post-birth struct assignment -- rather than
anything RUNCAP-shaped, since RUNCAP requires a real blkio_dev+
homeblocks_sig_t an unattended identity never has.

The born VM's own capsule payload is expected to lay down two CREATE'd
buffers (UNATTENDED-ID-UUID, UNATTENDED-ID-CERT) via MINT-SCRATCH-
EMIT's own literal format; their addresses are fetched by interpreting
a two-word line inside the *new* VM's own context
(vm_interpret(born_vm, ...)), the same "run inside that VM's own
dictionary" idiom capsule_wirebind.c already uses for VM-NAME-REG.

Explicit invariant preserved: never touches g_wirebind_attached_username
or any console-pairing state, births no console VM -- an unattended
identity stays un-promptable (§VIII.1) until a human pairs a console to
it later via the already-working VM-NAME-REG mechanism.

Verified live end-to-end on amd64: minted a real test identity via
MINT-SCRATCH, captured its MINT-SCRATCH-EMIT output, built a throwaway
test capsule from it (discovered along the way: mkcapsule's real block
constraints are range [2048,5120) and max 16 content lines per block --
neither matches this repo's own doc comment, corrected via ground
truth from the tool itself, not assumed), then ran UNATTENDED-BIRTH
against it: cert verified against Zuse's root pubkey, identity
installed, "no console attached" reported, Hera stayed healthy
afterward (5 6 + . -> 11). Test capsule reverted after capture per
this project's own probe convention -- not a real identity, never
committed. Clean 3-arch qemu boot (amd64/aarch64/riscv64) on the real
committed C-only change.

Remaining punch-list item (the ACL cap bit for console attachment) not
started.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K
2026-09-16 07:06:47 -04:00
Robert Allan JamesandClaude Sonnet 5 5fc709a228 Punch list item 2: MINT-SCRATCH-EMIT, verified live on all 3 arches
Adds the hand-transcription mechanism FABRIC-3.md §XXXII.2's punch list
item 2 calls for: MINT-SCRATCH-EMIT prints the last successful
MINT-SCRATCH's drive_uuid + cert devblock as ready-to-paste FORTH
source (HEX-based CREATE ... C, ... byte sequences), so packaging an
unattended identity into a capsule is a mechanical copy out of the
captured boot log rather than a manual hex-to-FORTH translation an
operator could transpose a digit in.

Refuses (no output) if no MINT-SCRATCH has ever succeeded -- printing
4112 zero bytes as if they were a real identity would be a silent,
misleading success, matching this session's own error-handling audit
discipline rather than adding a new silent-failure primitive right
after finishing one.

Verified live on amd64: MINT-SCRATCH-EMIT correctly refuses before any
mint, then after MINT-SCRATCH succeeds, emits UNATTENDED-ID-UUID and
UNATTENDED-ID-CERT as valid FORTH literals. Cross-checked byte-exact:
the cert's own embedded ASN.1 serialNumber field matches the emitted
UUID bytes exactly, confirming x509_build_user_cert()'s drive_uuid
binding round-trips correctly through the scratch-device path. Clean
3-arch qemu boot (amd64/aarch64/riscv64).

Remaining punch-list items (writing/committing a real capsule file for
an actual named identity, the unattended-birth call site, the ACL cap
bit) not started -- authoring a real committed capsule needs a name/
purpose decision that isn't mine to make.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K
2026-09-16 05:12:43 -04:00
Robert Allan JamesandClaude Sonnet 5 63864c4b01 Punch list item 1: scratch-device MINT-SCRATCH, verified live on all 3 arches
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
Implements FABRIC-3.md §XXXII.2's scratch-thumbdrive mint mechanism:
capsule_mint_identity_scratch() (capsule_mint.c/.h) builds a throwaway
RAM-backed blkio_dev via blkio_ram.c's backend and runs
capsule_mint_identity() against it completely unmodified -- same live
Zuse-signing operation, same rng_get_bytes() draw for drive_uuid a real
thumbdrive gets. Reads back only drive_uuid + the cert devblock; the
seed devblock is written into the scratch buffer internally but never
read out (no seed is ever baked into a capsule, per the ratified
no-seed decision).

New FORTH word MINT-SCRATCH (mama_forth_words.c), same stack signature
as MINT, mints into the scratch device instead of any attached drive
and never touches sk_repl_get_attached_blk_dev() or console-pairing
state. Prints the drive_uuid as hex so a live boot log itself proves
each call drew fresh entropy.

Build correction found along the way: blkio_ram.c was excluded from
the kernel build (Makefile.starkernel VM_EXCLUDE) alongside
blkio_factory.c/blkio_file.c. blkio_factory_open() unconditionally
references blkio_file.c's real fopen()/fread() file I/O, which has no
freestanding-kernel equivalent, so the factory function couldn't be
used as-is. blkio_ram.c itself is pure memcpy over a caller buffer --
pulled it alone into the kernel build and wired it directly in
capsule_mint.c, the same way blkio_factory.c's own extern declarations
do internally.

Verified live on amd64: two MINT-SCRATCH calls produced two genuinely
different drive_uuids (b533246d.../ae11b2b2...), confirming fresh
entropy per call rather than stale reuse; VM stayed healthy afterward
(5 6 + . -> 11). Clean 3-arch qemu boot (amd64/aarch64/riscv64),
logs and DoE CSVs committed per standing convention.

Remaining punch-list items (hand-transcription into a .4th block, the
unattended-birth call site, the ACL cap bit) not started.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K
2026-09-16 04:42:15 -04:00
Robert Allan JamesandClaude Sonnet 5 42d4bf3dad Stage D batch 7 (final): mama_forth_words.c groups 5+6 -- dictionary lookup/test words + Stadium primitives; Stage D closed (FABRIC-3.md §XXXII.6)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
NAME>XT/RUNCAP-TEST/PAIR-TEST and the six STADIUM-* physics primitives
(STADIUM-ADMIT/STADIUM-EVICT/STADIUM-RES-PULL/STADIUM-RES-PUSH/
STADIUM-HEAT@/STADIUM-HEAT!), 12 sites -- closes mama_forth_words.c
and the entire kernel-only error-handling audit.

All 105 sites from the §XXXII.3 triage now accounted for: 28 already
correct, 75 silent sites fixed across repl.c/inference_words.c/
log_words.c/vm_core.c/mama_forth_words.c, 2 special cases resolved by
dropping the error per their own documented contract, 1 resolved via
console_println() per its own recursion constraint. Verified with awk:
zero remaining vm->error=1 sites in mama_forth_words.c lack a
diagnostic within the preceding three lines.

Three-arch clean qemu acceptance passed. This closes Stage D and the
USE/logging/audit thread opened in §XXXII; Stage E (human-vs-unattended
identity model) remains open, not started this pass.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K
2026-09-15 20:25:30 -04:00
Robert Allan JamesandClaude Sonnet 5 90c86c6006 Stage D batch 6: mama_forth_words.c group 4 -- identity/crypto words (FABRIC-3.md §XXXII.6)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
MINT + mint_pop_string()/ZUSE-ELIGIBILITY-ADD/ZUSE-ELIGIBLE?/
ELEVATE-PUBKEY-UNPACK, 10 sites. mint_pop_string() gained a field_name
parameter so its diagnostics name which of MINT's four string
arguments failed (phone/email/username/full_name), rather than a
generic message that would leave the operator guessing. Diagnostic
placement matched each function's own sibling convention where one
exists (ZUSE-ELIGIBILITY-ADD), log_message() default otherwise.

Three-arch clean qemu acceptance passed.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K
2026-09-15 19:53:47 -04:00
Robert Allan JamesandClaude Sonnet 5 8b5300fc4f Stage D batch 5: mama_forth_words.c group 3 -- cross-VM execution/dispatch, both flagged special cases resolved (FABRIC-3.md §XXXII.6)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
VM-STEP/VM-EXEC/VM-CALL (5 sites) gained console_println() diagnostics
matching their own existing sibling guards.

SWITCH-MARK-WORK and VM-HEAT (3 sites) resolved per §XXXII.3's own
recommendation: dropped vm->error entirely rather than diagnosing it,
matching each function's own doc comment ("must never error or spam
the console" / "does not print/error"). SWITCH-MARK-WORK is the same
function §XXVIII.3 already fixed once for an off-by-one that fired
silently on every MSG-SEND in the system -- the guard now genuinely
cannot repeat that by contract, not just by the threshold being right.
VM-HEAT's guards now push 0 and return, matching its own
always-returns-a-value stack effect.

Three-arch clean qemu acceptance passed. Zuse's own WIREBIND attach
(every boot) drives MSG-SEND -> SWITCH-MARK-WORK, so this batch's most
safety-critical fix is exercised by standard acceptance, not just
compiled.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K
2026-09-15 19:23:12 -04:00
Robert Allan JamesandClaude Sonnet 5 b42c3b195c Stage D batch 4: mama_forth_words.c groups 1+2 -- capsule/lifecycle words + BIRTH/START/KILL (FABRIC-3.md §XXXII.6)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
15 of mama_forth_words.c's 48 silent sites fixed, split by functional
grouping per direct instruction: CAPSULE@/CAPSULE-HASH@/CAPSULE-FLAGS@/
CAPSULE-LEN@/CAPSULE-BIRTH/CAPSULE-RUN/EXEC (9), and BIRTH/START/KILL
(6).

Refined the diagnostic-placement rule: match whichever convention that
same function's other already-correct guards use, rather than
defaulting uniformly. BIRTH/START/KILL/EXEC each already had a
console_println() sibling guard ("name too long or empty") -- their
newly-diagnosed guards now match that, same shape as USE's own fix.
The CAPSULE*@ words have no sibling guard to match, so they keep
log_message() (defer_words.c's gold-standard default).

Three-arch clean qemu acceptance passed; BIRTH itself is exercised by
every boot (Hermes/Artemis birth).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K
2026-09-15 18:51:54 -04:00
Robert Allan JamesandClaude Sonnet 5 66c2f3e539 Stage D batch 3: fix silent error sites in vm_core.c, incl. the two highest-value primitives; two real NULL-deref bugs found and fixed (FABRIC-3.md §XXXII.6)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
All 13 silent vm->error=1 sites in vm_core.c now log a diagnostic
first: vm_enter_compile_mode, vm_compile_word, vm_compile_literal,
vm_compile_call, vm_exit_compile_mode, execute_colon_word (2 sites
each/combined), and the four fundamental memory primitives
vm_load_u8/vm_store_u8/vm_load_cell/vm_store_cell.

Found and fixed two real NULL-pointer-dereference risks while adding
the diagnostics: vm_compile_call() and vm_exit_compile_mode() each had
a combined `if (!vm || <cond>) { vm->error = 1; ... }` guard that
dereferenced vm->error even on the !vm branch of its own condition.
Split both, and applied the same defensive split to the four memory
primitives since vm_ptr()/vm_addr_ok() both tolerate vm==NULL
internally.

Live-verified, amd64: 999999999 @ . recovers correctly at Zuse's own
console (Stage A's mechanism holds), but the new vm_load_cell
diagnostic itself didn't print -- traced to memory_words.c's own
redundant, still-silent vm_addr_ok() pre-check in memory_word_fetch()
(and the same shape in memory_word_store()), which intercepts before
ever reaching vm_load_cell(). memory_words.c is vendored, out of this
initiative's scope, spun off to FABRIC-4.md -- recorded as a concrete
cross-reference for that future work rather than left to be
rediscovered.

Three-arch clean qemu acceptance passed, full POST suite included.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K
2026-09-15 17:29:37 -04:00
Robert Allan JamesandClaude Sonnet 5 1a8c0e4fcf Stage D batch 2: fix silent error sites in log_words.c, resolve the flagged category-iii special case (FABRIC-3.md §XXXII.6)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
log_word_set_level(), log_do_emit(), and log_emit_string() (7 sites
total) now log a diagnostic via log_message(LOG_ERROR, ...) before
setting vm->error, matching this file's own already-correct
log_str_emit() sibling and defer_words.c's gold-standard pattern.

log_word_append_raw() (backing (LOG-APPEND-RAW), 3 sites) resolved
differently per §XXXII.3's own triage note: its doc comment forbids
log_message() here (recursion into the log ring it writes to), but the
word can also be invoked by hand at the console -- console_println()
carries no such recursion risk and matches Stage B's policy for a
manual interactive invocation. Added the console.h include this
required.

Three-arch clean qemu acceptance passed.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K
2026-09-15 16:53:18 -04:00
Robert Allan JamesandClaude Sonnet 5 5417ffb2bf Stage D batch 1: fix silent error sites in repl.c and inference_words.c (FABRIC-3.md §XXXII.6)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
repl.c's sk_word_blk_attach_ack() and inference_words.c's array_ptr()
helper + infer_word_run()'s allocation guard now log a diagnostic via
log_message(LOG_ERROR, ...) before setting vm->error, matching
defer_words.c's own gold-standard pattern (§XXXII.3) -- these are
internal/background conditions (a malformed message-callback, a bad
array reference or allocation failure), not interactive usage mistakes,
so the fix keeps the fault and reports it rather than dropping it like
USE's own fix did.

Batched together (4 sites total, smaller combined than the next file)
rather than two separate acceptance cycles for negligible size.

Three-arch clean qemu acceptance passed.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K
2026-09-15 16:23:14 -04:00
Robert Allan JamesandClaude Sonnet 5 c05b70c8d6 Stage B: logging policy documented, level-aware log-ring eviction built, LOG-FLUSH deferred again (FABRIC-3.md §XXXII.4)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
Policy decided for the kernel-only audit scope: an interactive
command's direct response stays on console_println/console_puts;
everything else (state transitions, background diagnostics, audit
trails) routes through log_message() at the appropriate level,
matching capsule_mint.c's verify_mint() precedent. Documented, not
code-swept here -- reclassifying individual sites is Stage D's job.

LOG-FLUSH deferred again, explicitly: the per-VM log buffer its own
doc comment presumes (vm_log_buffer.h) doesn't exist anywhere in the
tree -- building it is real feature work needing its own scoped stage.

Level-aware eviction built: log_region_append() now reads the oldest
ring slot's own level before evicting it, protecting ERROR/WARN
records from being pushed out by INFO/DEBUG churn -- drops the
incoming low-priority record instead. Found and fixed an adjacent bug
while making this change: the prior two-valued return contract would
have made a benign "dropped by design" outcome indistinguishable from
a genuine write failure to its one caller, which unconditionally set
vm->error on any nonzero return. Changed to a three-valued contract (0
success, 1 dropped by design, -1 genuine failure).

Three-arch clean qemu acceptance passed. Eviction path itself not
live-exercised (needs 128+ LOG-APPEND calls to fill the ring) --
flagged, matching this project's own precedent for that kind of gap.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K
2026-09-15 14:01:36 -04:00
Robert Allan JamesandClaude Sonnet 5 5c5896fbc1 Stage A: fix USE's silent stack-underflow/bad-address guards, and the Hera fault-scoping gap they exposed (FABRIC-3.md §XXXII.1)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
mama_word_use()'s two silent vm->error=1 guards (dsp<1 stack underflow,
NULL from vm_ptr()) now print a diagnostic and return, matching the
function's other five guards.

Live verification of that fix alone surfaced a bigger problem: with the
guard no longer silent, the REPL proceeds to interpret the leftover
token as an unrecognized word, which independently sets vm->error, and
sk_repl_step()/sk_repl_run()'s Hera-branch still hard-halted on that.
Investigated kernel_main.c's boot/capsule-load paths directly: they
already catch and clear mama->error entirely separately, before
sk_repl_run() is ever entered -- so the "no fallthrough surface" halt
in these two REPL functions was never protecting a boot-time fault, only
an ordinary interactive REPL-turn one. Both functions now recover
unconditionally on any VM's error, Hera included, matching how a
redirected (WIREBIND/USE'd) identity's session already recovered.
Removed the now-fully-unused sk_fault_handler().

Verified live, amd64: USE rajames at Zuse's own console prints the new
diagnostic, then "VM fault -- session recovered, resuming", console
stays interactive afterward. Three-arch clean qemu acceptance passed
before and after the Hera-fault-scoping change.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K
2026-09-15 09:06:33 -04:00
Robert Allan JamesandClaude Sonnet 5 d9da82b065 Stage 4 increments 2+3: WIREBIND VMs as switch-signal participants + mark-and-defer tombstone reap (FABRIC-3.md §XXVIII Stage 4)
Increment 2: WIREBIND user VMs (the ones that actually run FORTH work;
console VMs are pure REPL proxies and never participate) register as
Stage 3 switch-signal participants at attach, unregister at teardown.
Slot table bumped 8 -> 16, matching messaging.4th's own VM-MAX -- a real,
already-agreed ceiling, not an invented number. Added
sk_vm_switch_signal_unregister() (compaction-based; Tripod VMs never
needed removal, WIREBIND VMs cycle constantly and would otherwise
exhaust the bounded table).

Increment 3: implements the plan's own ratified option (A) for the
async-detach UAF risk -- mark-and-defer via a new pending_reap flag on
VMRegistryEntry, deliberately not a new VMState (capsule_vm_kill()
already treats VM_STATE_DEAD as idempotent success, which would silently
swallow a reap attempt; SWITCHED_OUT still accurately describes a
tombstoned VM until the moment it's actually freed). unclean_detach()
sets it when capsule_vm_kill() refuses a SWITCHED_OUT target; the Stage 3
checkpoint (vm_core.c) checks it before ever attempting to resume a
pending switch target, and calls the new capsule_vm_force_reap() instead
-- the one caller allowed to bypass capsule_vm_kill()'s own refusal,
because it runs at the exact safe cooperative point the switcher itself
controls. A new idle-tick sweep cleans up the WIREBIND live-table entry
once the reap has actually happened.

Verified clean on all 3 architectures (baseline regression -- no
WIREBIND attach happens in a plain boot). The reap mechanism's own
correctness under a genuinely parked context is verified separately,
next, via a temporary deterministic probe.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K
2026-09-15 00:58:34 -04:00
Robert Allan JamesandClaude Sonnet 5 9f0f33dfc5 Stage 4 increment 1: per-device WIREBIND tracking, fixing a real multi-identity detach leak (FABRIC-3.md §XXVIII Stage 4)
capsule_wirebind_unclean_detach()/eject() tracked "the attached identity"
as a single global, correct for the console-pairing UX (one physical
console) but wrong for detach safety: since §XV/§XVI proved multiple
identities genuinely live simultaneously via this same attach path, every
attach after the first silently overwrote the singleton, so an unclean
detach of any but the most-recently-attached identity was silently
ignored -- that VM leaked forever, no trace in the log.

Adds a per-device live-identity table, separate from the (unchanged)
console-pairing singleton, so unclean-detach resolves any attached
device to its own identity. Sized off messaging.4th's own VM-MAX (16)
minus Tripod's 3 reserved slots, not an invented number. Corrects the
stale "single-USB-device constraint" doc claim in capsule_wirebind.h,
false since §XV/§XVI.

Groundwork for Stage 4's real deliverable (WIREBIND VMs as switch-signal
participants) -- this increment only fixes detach targeting; switch-
signal registration is next.

Verified clean on all 3 architectures (no WIREBIND attach happens in a
plain boot, so this is a regression check on the existing Tripod-only
path; live multi-identity verification comes with the switch-signal
registration increment).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K
2026-09-15 00:02:20 -04:00
Robert Allan JamesandClaude Sonnet 5 05ae7aa886 Fix SWITCH-MARK-WORK off-by-one: was tripping vm->error on every MSG-SEND (FABRIC-3.md §XXVIII.2)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
mama_word_switch_mark_work()'s stack-underflow guard checked dsp < 2,
requiring 3+ items, when it only ever needs the 2 IDX>NAME leaves it
(caddr u). Since dsp is index-based (2 items == dsp 1), this rejected
every normal call. MSG-SEND tail-calls SWITCH-MARK-WORK unconditionally,
so this fired on every message sent anywhere in the system -- visible
only where a caller happened to check the target VM's error flag
afterward (mama_word_vm_exec()'s "VM-EXEC: ERROR in Artemis" report).

Verified clean on all 3 architectures: boot reaches Startup: Artemis
live -> zuse@Hera] ok> with no VM-EXEC: ERROR line at all.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K
2026-09-14 20:17:01 -04:00
Robert Allan JamesandClaude Sonnet 5 66beae7fd4 Stage 3 follow-on: message-arrival eligibility hook + trampoline-blind switch-storm fix (FABRIC-3.md §XXVIII.2)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
Implements the message-arrival eligibility signal FABRIC-3.md §XXVIII.1 left
open (has_work per-slot flag, set via new SWITCH-MARK-WORK primitive from
MSG-SEND) so an idle VM never becomes a switch target purely by waiting out
the readiness threshold.

Also root-causes and fixes a second, independent switch-storm: the tick's
"who is current" check used vm_log_attributed_vm(), which can't see a VM
parked in switch.c's own raw trampoline. Replaced with a dedicated
g_switch_current_vm tracked by the switch mechanism itself, and moved
target-slot eligibility reset to the switch decision point instead of
relying on ISR polling to observe a window that can be only a few
instructions wide.

Verified live on all 3 architectures: clean boot to zuse@Hera] ok>, live
cross-VM message dispatch, and (since a quiet log looks identical to a
livelocked storm once the DoE probe is gone) confirmed genuine REPL
liveness via QMP send-key + screendump on aarch64/riscv64, not log
inspection alone.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K
2026-09-14 17:39:20 -04:00
Robert Allan JamesandClaude Sonnet 5 862d7d9c48 Stage 3 follow-on: fix stack-ownership corruption + DoE switch columns (FABRIC-3.md §XXVIII.1)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
DoE CSV gained 6 switch-signal columns (switch_count_cumulative,
switch_current_slot, switch_*_readiness, switch_ticks_since), and verifying
them with a boot-time HB-ON probe surfaced a real livelock: the preemption
checkpoint could fire inside a VM-EXEC-nested execute_colon_word() call and
switch away from a stack it didn't own, parking a borrowed region of the
caller's stack under the wrong VM's saved-context pointer. The trampoline
bounce was the visible (safe) half of this; the corruption was the quiet
half, live in every prior "clean" Stage 3 boot without ever showing up in
the log.

Fixed by gating the checkpoint on being at the outermost vm_interpret()
call (g_vm_interpret_depth / sk_vm_at_outermost_interpret(), vm_core.c),
per Bob's decision. Also fixed two related bugs found in the same pass:
g_switch_back_to was a single global stale after first entry, now per-VM
state (native_switch_back_to); note_switch_performed() fired on resume
instead of switch-out, now called before the switch.

Verified on all 3 architectures: steady log growth (no freeze), zero
leaked QEMU processes, DoE columns internally consistent, Hermes/Artemis
confirmed genuinely executing (not just trampoline-bouncing). Temporary
HB-ON boot probe reverted after capture.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016UNhH1mhi52i6Qihh7ZV5S
2026-09-13 23:58:59 -04:00
Robert Allan JamesandClaude Sonnet 5 986d042aa7 Stage 3: timer-driven preemptive switching, live on all 3 arches (FABRIC-3.md §XXVIII)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
Fourth stage of the preemptive context-switching plan, and the biggest.
LithosAnanke now genuinely, continuously preempts between Hera, Hermes,
and Artemis -- timer-driven, running live for the entire remainder of
every boot once the Tripod fleet registers, not a bounded probe.

A real design fork was resolved before writing code: the naive approach
(the timer ISR calling Stage 2's sk_vm_context_switch() directly) is
broken -- Stage 0's trap frame lives on whatever stack was active at
interrupt time, and jumping to a different stack via Stage 2's own
independent swap mid-handler would abandon that trap frame unresumed,
guaranteed corruption on the first tick. Chose the safer of two named
options: the ISR only ever sets a flag and returns completely normally
through its own full epilogue; the actual switch happens moments later,
via Stage 2's already-proven mechanism, at a safe cooperative checkpoint
on the mainline (execute_colon_word()'s per-word dispatch loop, checked
on literally every word, not throttled to the existing 256-word
heartbeat-tuning cadence) -- confirmed with the user that word-level
granularity is fine-grained enough given the eventual Zynq FPGA target
where a word is a mnemonic.

New capsule_vm_switch_signal.c/.h: a purpose-built run-readiness signal,
deliberately separate from capsule_vm_physics.c's execution-heat engine
(that one's own header documents itself as never touched from interrupt
context, by design). Slot table sized with headroom (8) rather than
hardcoded to today's 3 participants, so extending participation later is
another register() call, not a redesign -- per direct request to leave
room for swapping the participant set. Simple linear accumulate-then-
threshold for this first cut; a fancier law can replace it later without
touching the mechanism around it. heartbeat_tick() gains its one
deliberate, documented amendment to this file's own top-half/bottom-half
discipline -- the first time this codebase reaches into VM-scheduling
state from real ISR context.

Registration happens only after all three VMs are fully born, right
before the REPL starts -- no critical-section protection yet against
being switched away mid-birth-setup.

Known, flagged rough edge (not reconciled this pass): MSG-TICK's own
idle-pump and this new mechanism can still independently move control
between the same VMs; not observed to interact badly in verification,
but not fully unified either.

Verified interactively at the console on all 3 architectures with
continuous background preemption running throughout -- amd64 computed
`1 1 + .` -> 2, aarch64 computed `1 1 + dup DUP * . CR` -> 4, both
correct, REPL fully responsive, zero fault indicators over sustained
runtime.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016UNhH1mhi52i6Qihh7ZV5S
2026-09-13 19:32:36 -04:00
Robert Allan JamesandClaude Sonnet 5 f790d0995e Stage 2: cooperative VM context switch primitive, proven on all 3 arches (FABRIC-3.md §XXVIII)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
Third stage of the preemptive context-switching plan. The real
save/restore switch mechanism now exists -- the first time anything has
ever executed on a VM's own native stack (Stage 1 allocated them,
unused).

New sk_vm_switch_to() (switch.S, one per arch) is an ordinary function
call, not an interrupt -- so unlike Stage 0's trap frame, the ABI already
covers every caller-saved register; only the callee-saved set needs
explicit save/restore (amd64: rbx/rbp/r12-r15, no FP at all since SysV
has no callee-saved XMM; aarch64: x19-x28/x29/x30 + d8-d15; riscv64:
s0-s11/ra + fs0-fs11, FS-gated like Stage 0 but read once and reused for
both halves within one call, since FS is genuine global CPU state, not
part of what's switched). A sibling sk_vm_switch_prime() in the same file
builds the synthetic first-entry frame, kept in assembly so the layout
can never drift out of sync with sk_vm_switch_to() itself.

New switch.c/switch.h: sk_vm_context_switch(from, to) handles first-entry
priming vs. resuming a parked context, and updates registry state (new
VM_STATE_SWITCHED_OUT, distinct from VM_STATE_STOPPED -- STOPPED means no
live frame, this means the opposite). sk_vm_switch_entry() is the minimal
permanent trampoline every freshly-entered VM lands in: no production
behavior defined yet, so it just yields straight back to whoever switched
to it, forever.

Closes the confirmed unguarded-KILL UAF found during planning:
capsule_vm_kill(), mama_word_kill(), and capsule_vm_kill_all_nonmama()
all now refuse (or silently leak rather than free, on the cold-restart
path where arch_cold_reset() wipes everything immediately after anyway)
tearing down a switched-out VM. Side effect found, not built on purpose:
the existing MSG-TICK idle-pump already filters on VM_STATE_LIVE, so it
automatically stopped dispatching into a switched-out VM with zero
changes needed there.

Verified via a temporary SWITCH-TEST probe (boot-triggered, since nothing
can type interactively into a foreground-only QEMU session) that
round-tripped a sentinel through 5 real Hera<->Hermes switches on all 3
architectures: 5/5 rounds, 0 failures, clean continuation to ok>. Probe
fully reverted after capture; kernel_main.c shows zero diff.

Also: Makefile.starkernel's LOADER_EXTRA_SRCS/LOADER_ASM needed the new
files added explicitly (this project's "loader" PE binary is the full
running kernel, not a thin bootstrap stage), and aarch64's switch.S
needed the same #ifndef _WIN32 guard around .hidden that isr.S already
carries (aarch64's loader assembles via clang targeting a PE/COFF
target with no .hidden equivalent) -- caught by a build failure, fixed.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016UNhH1mhi52i6Qihh7ZV5S
2026-09-13 15:25:39 -04:00
Robert Allan JamesandClaude Sonnet 5 57ac3fc304 Stage 1: per-VM native stacks, allocated but not yet executed on (FABRIC-3.md §XXVIII)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
Second stage of the preemptive context-switching plan. Every VM (Hera,
every capsule_birth_baby()-born VM including WIREBIND identities) now
gets its own dedicated 2 MiB native C stack at birth -- but nothing runs
on it yet, that's Stage 2. Pure allocation-machinery proof.

Design correction made before writing code: the plan called for cloning
sk_vm_arena_alloc()'s guard-page pattern, but that pattern turns out to
be Mama-only -- host_services.c's kernel_alloc() gives every baby VM a
plain kmalloc() block for its dictionary arena, not a real guarded PMM
allocation. Stacks get the real treatment instead (new
sk_vm_native_stack_alloc()/_free() in arena.c): independent
pmm_alloc_contiguous() + guard pages for every VM without exception, no
singleton, no kmalloc fallback -- a stack overflow is exactly the
failure mode guard pages exist for, and a corrupted stack could corrupt
whatever saved context Stage 2 trusts.

2 MiB size matches this project's own established kernel-stack
convention (g_kernel_stack/g_rpi5_native_stack), not a guess -- that one
shared 2 MiB stack today already carries all VMs' combined nested
VM-EXEC recursion.

Three new VM struct fields, freed in vm_cleanup() alongside the existing
call_stack free. Allocation failure is non-fatal to birth.

All 3 architectures re-verified clean boot to ok>, no native-stack
allocation failures for any Tripod-fleet VM.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016UNhH1mhi52i6Qihh7ZV5S
2026-09-13 14:30:46 -04:00
Robert Allan JamesandClaude Sonnet 5 15672ce17c Stage 0: trap-frame parity across all 3 arches (FABRIC-3.md §XXVIII)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
First stage of the preemptive context-switching plan (see
~/.claude/plans/logical-snuggling-bear.md). Pure foundation work -- every
arch's ISR now saves the full register set on interrupt entry, so a trap
frame is in principle sufficient to resume execution anywhere it was
taken. No FORTH-visible behavior changes.

amd64: added FXSAVE/FXRSTOR, closing a genuine pre-existing correctness
gap (not just future-preemption prep) -- confirmed live double-precision
FP code reachable from ordinary interpreter dispatch (vm_runtime.c Loop
#5/#6), and the ISR previously saved zero FP/SSE state. rbp repurposed as
a fixed anchor so the 16-byte-aligned FXSAVE area can be carved out of an
unpredictably-aligned rsp without disturbing existing argument reads.

aarch64: extended the trap frame 672->800 bytes, adding v8-v15 (AAPCS64
callee-saved, previously excluded on call-site-only reasoning that
doesn't hold for an async trap).

riscv64: extended the trap frame 320->512 bytes, adding s0-s11 and
fs0-fs11 (the latter still correctly gated behind sstatus.FS != Off).

All 3 architectures re-verified clean boot to ok> under the new frames --
amd64 through hundreds of timer ticks with FXSAVE/FXRSTOR live on every
interrupt, aarch64 through 987 ticks, riscv64 clean on the now-larger
FS-conditional block.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016UNhH1mhi52i6Qihh7ZV5S
2026-09-13 09:23:28 -04:00
Robert Allan JamesandClaude Sonnet 5 2a30212bd3 Real per-VM log persistence: source attribution + ACL pin (FABRIC-3.md §XXVII)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
Wires the previously-unused vm_log_attributed_vm() into LOG-APPEND's kernel
primitive so persisted log records carry a trustworthy source (the real
attributed VM's registry name, or "HADES" pseudo-source) instead of a
caller-supplied, trivially forgeable string. Drops src-addr/src-u from
LOG-APPEND's stack signature accordingly. Pins LOG-APPEND via bare ACL-PIN
in Artemis's own init.4th, matching BIRTH/CAPSULE-BIRTH's precedent for a
privileged word that can't reach the shared, host-portable ACL.4th.

Also fixes two console-banner nitpicks: a mis-rendering em dash (U+2014)
in the boot banner, and drops "Emergency" from the CLI banner text.

Doc corrections to artemis_sig.h/zuse_eligibility_list.h reconciling the
three fixed devblock ranges now in play. LOG-FLUSH (the intended normal
entry point) and level-aware log eviction remain open, flagged not fixed.

Re-verified clean boot to ok> on all 3 architectures after every change.
riscv64 showed one new, unrelated virtio_blk write-timeout anomaly during
Artemis's early physics self-test (self-recovered, boot unaffected,
sector doesn't map to the log region) -- flagged, not investigated.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016UNhH1mhi52i6Qihh7ZV5S
2026-09-13 08:56:33 -04:00
Robert Allan JamesandClaude Sonnet 5 61755fde78 Artemis genesis stamp: fix a BAM-corrupting offset before it ever ran (FABRIC-3.md §XXVI follow-on, Step 3)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
Step 3: one-time artemis_sig_t genesis stamp, written once
kernel_main.c's virtio-blk path confirms Artemis's own disk, so the disk
image is later recognizable generically (repl.c's idle-loop USB-MSC scan,
built in the prior commit) regardless of which bus found it.

Correction made before this ever touched the real disk: the signature's
first design (committed in 29b6789) placed it at a fixed bottom-of-device
forth-block (4, devblock 1) -- copying homeblocks_sig_t's own convention,
which is safe for a raw identity thumbdrive but not for Artemis's own
disk. Artemis's disk is block_subsystem.c's own STFR/v2-formatted volume:
devblock 0 holds that format's header and devblock 1 is the FIRST
DEVBLOCK OF THE LIVE BAM (blk_compute_fresh_geometry(): bam_start = 1).
The original design would have overwritten Artemis's live allocation map
on the very first real boot. Caught via direct cross-reference against
block_subsystem.c before the genesis-stamp call site was ever run against
the real image -- no corruption occurred.

Fixed by moving the header to a fixed offset from the END of the device
instead (ARTEMIS_SIG_DEVBLOCK_FROM_TOP=64), the same top-of-device region
block_subsystem.c's own meta_fence_blocks reservation (128 devblocks)
already carves out for system metadata, and where Zuse's genesis marker/
eligibility list already live -- but computed independently via
blkio_info() rather than through blk_meta_zone_*(), since that accessor
needs an already-attached, format-detected slot, which is exactly the
state pre-attach generic discovery doesn't have yet. Picked well clear of
Zuse's two tenants (devblock_from_top 0 and 1+, open-ended) so the two
subsystems' independent math can never collide.

Also reordered kernel_main.c: rng_init() now runs before the Artemis
virtio-blk block (was after) -- the genesis stamp needs rng_get_bytes()
for disk_uuid, and the original order would have failed the stamp on
every boot.

Verified live: booted amd64 against the real disk/artemis.img twice --
first boot logs "Artemis: genesis signature stamped" (confirmed blank at
the target offset beforehand via a host-side read), second boot on the
now-stamped image logs no re-stamp (idempotent, CRC/read-back verified)
-- both boots and aarch64/riscv64 (against the same now-stamped image)
all still report "Artemis: 22998 data blocks" / "PASS: persist-read"
unchanged, confirming the BAM and data pool were never touched.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo
2026-09-13 07:31:21 -04:00
Robert Allan JamesandClaude Sonnet 5 29b6789860 Artemis bus-agnostic discovery: signature format + idle-loop generalization (FABRIC-3.md §XXVI follow-on)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
Step 1: new artemis_sig_t header format (magic 'ARTM', sibling to
homeblocks_sig_t, distinct so a generic scan can tell Artemis's own disk
apart from an identity thumbdrive by content alone) -- artemis_sig.h/.c,
wired into Makefile.starkernel.

Step 2: sk_repl_idle()'s existing per-USB-MSC-slot attach handling (the
pattern WIREBIND already uses for identity thumbdrives) now also checks
for the ARTM signature whenever a device's home-blocks check comes back
BLANK. On a match, once Artemis's own storage-attach round-trip
(HERA-BLK-ATTACH-REQ/BLK-ATTACH-ACK) confirms success,
capsule_zuse_boot_load_root_pubkey() runs -- the same call kernel_main.c's
synchronous QEMU-only virtio-blk path already makes, now reachable
without a hardcoded PCI vendor/device scan. That function is already
idempotent (no-op once zuse_root_pubkey_known is set), so no boot
restructuring was needed despite the initial concern that deferring
Artemis discovery to the idle loop would require one.

Verified: clean build + QEMU boot to [zuse@Hera] ok> on all three
architectures, zero regression to the existing virtio-blk/Zuse-thumbdrive
attach path.

Steps 3 (genesis-stamping onto disk/artemis.img) and 4 (growable
production log-persistence region) not yet started.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo
2026-09-13 07:16:48 -04:00
Robert Allan JamesandClaude Sonnet 5 a8b16d41da Unify console prompt to [user@VM]; fix real personality-block truncation; correct §XXV's wrong lockdown conclusion (FABRIC-3.md §XXVI)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
Investigating the std79 lockdown finding from FABRIC-3.md §XXV led
to a real discovery: WIREBIND births TWO VMs per identity, a console
proxy under the plain username and the actual restricted identity
under <username>~user (capsule_wirebind.c). Every test in §XXV
targeted the console proxy, which was never locked down at all.
Retested against the correct target (rajames~user): the lockdown
works exactly as designed. §XXV's "lockdown never engages" conclusion
was wrong -- corrected here, not deleted, since the mistake and how
it was caught are worth keeping (see the new feedback memory:
confirm which specific VM a name resolves to before concluding
anything, when a subsystem is known to birth more than one VM per
identity).

Two real, separate things found along the way are kept regardless
of that correction:

- capsule_runcap.c: the reserved personality devblock was read in
  full (mostly zero-padding after a short ~200-byte string) with no
  terminator, producing "WARN: block 4998 exceeds 1KB, truncating"
  on every std79-locked identity's birth, universal, since at least
  2026-09-10. Fixed by trimming to the first NUL byte actually found
  -- real, but harmless to execution (real content sat in the
  truncated block's surviving head); it mattered for capsule_id/
  content_hash being computed over padding instead of real content.

- console.h/console.c/repl.c: unified the prompt from a separately-
  computed "[VMName] (user)" into a single "[user@VMName]" line
  prefix -- exactly the ambiguity that caused the original
  misdiagnosis (the prompt showed only the WIREBIND username,
  identical whether USE had targeted the console proxy or the real
  ~user identity). Implemented as a registered callback
  (console_set_user_prefix_provider()) rather than console.c calling
  into WIREBIND/session logic directly, since console.c is a clean
  HAL module with no prior dependency on capsule-level subsystems.

Verified: clean build on all 3 architectures, zero new warnings,
identical dict_hash/capsule_hash to every prior boot this session
(console/prompt-only change). Full 9-identity messaging campaign
re-run end to end: 202s, zero faults, all 8 identities at 99/99
tokens, zero regression.

Also surfaced, not yet acted on: the full campaign's own console
tags now visibly show which VM each identity's tests actually
reached ([zuse@rajames], not [zuse@rajames~user]) -- messaging.4th's
VM-NAMES-INIT registers identities by plain username, so std79-doe.
fth's turn-attractor has been dispatching to each identity's console
proxy, not the actual locked-down identity, since the messaging
rewrite. Flagged for a deliberate decision, not investigated further.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo
2026-09-13 06:44:58 -04:00
Robert Allan JamesandClaude Sonnet 5 3c2daf50d1 Extend BIRTH/CAPSULE-BIRTH to all VMs symmetrically; flag a real std79 lockdown gap (FABRIC-3.md §XXV)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
Scoping the workload-into-factorial design's placement-mode factor
led to a real architectural improvement: rather than EXEC-ing a
workload capsule into an already-running, ACL-locked identity's own
persistent dictionary (filesystem-shaped, doesn't dodge the block-
collision exposure just traced in §XXIV), a workload now runs as a
fresh ephemeral child VM, BIRTH'd per trial and reaped after --
matching the project's own stated principle of automanagement over
imposed policy. CAPSULE-BIRTH already passes vm->stadium_vm_id (who
is birthing this VM) as the new child's parent, not a hardcoded Hera
constant, confirmed by reading the C -- so a workload trial genuinely
inherits the specific identity's own lineage when that identity does
the birthing.

Which surfaced a real premise: only Hera could call BIRTH/
CAPSULE-BIRTH at all (registered only in register_mama_forth_words(),
confirmed directly, not part of the earlier §XX messaging-symmetry
fix which deliberately kept this as one of her remaining privileges).
Extended symmetrically now, agreed explicitly before touching code:

- mama_forth_words.c: BIRTH and CAPSULE-BIRTH added to
  register_child_vm_words(), matching §XX's own pattern.
- acl-std79.4th: ' BIRTH , ' CAPSULE-BIRTH , added to ACL-STD79-LIST
  (new block 4048) -- a deliberate, explicit, named exception to the
  lockdown's own "standard words only" guarantee, not a silent one.
  Symmetric registration alone can't weaken any lockdown on its own:
  ACL-LOCKDOWN-STD79 is allowlist-based, deny-by-default, so a newly
  registered word is auto-denied there unless explicitly added.

Verified: clean build on all 3 architectures, zero new warnings.
Hera's own dict_hash unchanged (expected); Hermes/Artemis show the
same new dict_hash on all 3 architectures. Live-tested against a
real attached std79-locked identity: CAPSULE-BIRTH executes
correctly (returns vm_uuid_none() for a deliberately out-of-range
capsule-id, zero fault, zero ACL denial).

Found, and explicitly stopped short of fixing, a separate pre-
existing gap while verifying the above: MSG-STATUS and MSG-K
(messaging.4th words, not on the std79 allowlist) execute for a
locked identity instead of being denied. ACL-LOCKDOWN-STD79 is
confirmed to actually run; something more specific isn't reaching
messaging.4th's dictionary entries. Root cause not traced -- needs
its own investigation into vm_core.c's dictionary-link mechanics and
whichever capsule actually loads messaging for these identities.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo
2026-09-12 22:31:52 -04:00