41918a28a4e3ddaa1ba05efb9727888aca3111fa
235
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
41918a28a4 |
doe_log.c: add vm_name/vm_id_hex identity columns; new calibration workload
Adds the identity columns HB-ON's existing per-tick CSV logger (doe_log_tick_row(), doe_log.c) was missing. Previously every column described *a* VM's state each row, but nothing said which VM emitted it -- concurrent VMs' rows were indistinguishable by source. Two new first columns, vm_name and vm_id_hex, resolved via a reverse lookup on vm->stadium_vm_id (capsule_vm_registry_get()). Real bug found and fixed while building this: the freestanding snprintf here silently prints the literal format string instead of substituting for %016llx (width+ll+hex unsupported) -- caught by reading the actual emitted row, not assumed to work. Replaced with a hand-rolled hex nibble-table loop, the same idiom vm_uuid_format()/ MINT-SCRATCH-EMIT already use. Corrects FABRIC-3.md's own prior "still not started: CSV driver" framing: no bespoke CSV emitter is needed for the multiuser DoE at all -- this per-tick logger already exists, fires automatically inside every VM's own execution loop, and just needed HB-ON plus these identity columns to be usable for concurrent workers. Also corrects doe_log.h's own stale doc comment claiming g_doe_log_enabled defaults to 1 -- doe_log.c's own source is authoritative: 0, off by default, matching the HB-ON/HB-OFF naming. One real, honest limitation recorded rather than smoothed over: vm_name reads blank for any tick captured during a WORKER-BIRTH'd capsule's own self-execution at birth time, since capsule_birth_baby() runs that work (and its heartbeat ticks) before the registry name can be set. vm_id_hex is unaffected (reads directly from vm->stadium_vm_id) and still uniquely disambiguates every row -- confirmed live: a blank-name row's own vm_id_hex matched its later PARITY:KILL line's vm_id exactly. New workload-calib1.4th: Bob confirmed the existing 10 workload-N.4th files aren't fixed. Sized to cross HEARTBEAT_CHECK_FREQUENCY (256 word-executions/tick) many times over while staying far lighter than RUN-FIB/RUN-CHAOS5, both too slow under TCG for a bounded verification run. A real, reusable addition, not a throwaway. Verified live on amd64 via a self-contained one-shot EXEC'd test (HB-ON, WORKER-BIRTH two calibration workers, HB-OFF, VM-ERROR? on both, KILL both, completion banner) rather than interactive polling, which breaks once HB-ON makes the log grow continuously via Hera/ Hermes/Artemis's own background ticks. Clean 3-arch qemu boot on the real committed change. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K |
||
|
|
e33eb36361 |
Multiuser DoE punch list: WORKER-BIRTH + VM-ERROR? + vm_physics_init fix
Core mechanism for §XXXIII's main concurrency block, built and live-verified on amd64. Two new primitives: - WORKER-BIRTH ( capsule-c capsule-u name-c name-u -- ok? ): births a named, VM-EXEC-addressable VM from an arbitrary (p) capsule with no identity involved. Corrects the ratified design's own assumption that the main block would use UNATTENDED-BIRTH -- that requires a committed capsule per identity, impractical for dozens of trial VMs. The concurrency block never needed identity at all. - VM-ERROR? ( c-addr u -- flag ): reads a named VM's error state from Hera, mirroring VM-HEAT's silent/always-returns-a-value contract. Needed to check a VM-EXEC-driven trial VM's own fault state after the fact -- nothing existing let Hera do this. Real bug found and fixed in both WORKER-BIRTH and UNATTENDED-BIRTH: capsule_birth_baby() never calls vm_physics_init() either (same shape as the registry-name gap found building UNATTENDED-BIRTH) -- without it a born VM is never in the VM Fleet Attractor physics list, so VM-HEAT returns 0 forever regardless of work done. Fixed by adding vm_physics_init() alongside the existing registry-name call in both words. Traced (not guessed) why heat still read 0 after one VM-EXEC touch even post-fix: vm_physics_touch()'s transfer logic only fires from a VM's *second* touch onward -- the first touch just records a baseline tick. Verified live across three sequential touches: heat 0 -> 7039 -> 11333. This is a real design requirement for the DoE's heat/CV response variable (each trial must touch a worker at least twice), not a bug to route around. Also found live: all 10 existing workload-N.4th capsules self-execute their full workload at load/birth time (a bare top-level call to their own RUN-* word at file end) -- missed on an earlier, too-shallow 8-line survey of each file. WORKER-BIRTH alone already runs a worker's first pass as a side effect of birth. Verified live on amd64: two concurrent workers (fib + matrix-mul) birthed, run, measured (heat + error state), and killed cleanly; VM-HEAT/VM-ERROR? both confirmed silent-0 on an unknown name. Clean 3-arch qemu boot on the real committed change. Still open: the run-matrix/shuffle/CSV driver capsule itself, the WIREBIND-automation path for the fixed arm, and the per-VM touch-count budget's exact value. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K |
||
|
|
af351bff92 |
Punch list item 4 (final): CONSOLE-ATTACH, full unattended-identity flow verified end-to-end
Closes FABRIC-3.md §XXXII.2's punch list. CONSOLE-ATTACH ( name-c name-u -- ok? ) pairs a fresh console VM to an already-live VM registered as "<name>~user". Deliberately a plain, unconditional primitive with no VMIdentity capability-bit check -- corrects this session's own first-pass design (§XXXII.2 amended in the same commit): identity.installed is 0 for Hera/Hermes/Artemis and for every console-proxy VM, so a capability-bit gate would be unreachable for every VM a human actually types at, Zuse included. Matches ZUSE-ELIGIBILITY-ADD's own "no bespoke gate" precedent in this file; real gating is `' CONSOLE-ATTACH ACL-PIN` in ACL.4th if ever wanted. Two real bugs found and fixed via live testing, not assumed correct: - capsule_birth_baby() never sets a VM's registry name (documented gap, same one capsule_runcap_birth()'s own history already hit) -- UNATTENDED-BIRTH gained a second `name` argument and now calls capsule_vm_registry_set_name(new_vm_id, "<name>~user") itself. - CONSOLE-ATTACH's first draft took an independent console name from the target's name. sk_repl_dispatch_line()'s pairing check (repl.c) reconstructs the target as console_get_vm_name()+"~user" -- a mismatched console name silently falls back to direct interpretation with no error. Caught live (typed `5 6 + .` at a mismatched console, got a direct `11` instead of a relay) and fixed by collapsing to one name argument, matching WIREBIND's own by-construction invariant. Full end-to-end live verification on amd64: minted a test identity, UNATTENDED-BIRTH'd it as "bob", CONSOLE-ATTACH'd a console named "bob", USE'd it, typed `5 6 + .` -- no direct output at [zuse@bob] (relay path taken), then `[zuse@bob~user] 11 ok>` appeared: the relayed command executed on the target identity VM itself and printed its own answer back through the shared console. Hera stayed healthy throughout (2 2 + . -> 4 after switching back). CONSOLE-ATTACH also verified to refuse cleanly on a nonexistent target with no orphaned VM. Test capsule reverted after capture per this project's probe convention -- never committed. FABRIC-3.md §XXXII.2 fully closed: all 4 original questions ratified, the mid-course drive_uuid and ACL-bit corrections both recorded plainly rather than silently folded in, and a doc-accuracy note left for CLAUDE.md's own stale "1024-byte block limit" framing (mkcapsule's real limits are range [2048,5120) and 16 content lines/block) -- flagged, not fixed, out of this punch list's scope. Clean 3-arch qemu boot (amd64/aarch64/riscv64) on the real committed C-only change. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K |
||
|
|
5879c8b3bc |
Punch list item 3: UNATTENDED-BIRTH call site, verified live end-to-end
Implements the unattended-birth mechanism FABRIC-3.md §XXXII.2 designed: UNATTENDED-BIRTH ( name-c name-u -- ok? ) births a VM from a named (p) capsule via capsule_birth_baby() (completely unmodified, the same generic build-time-capsule path CAPSULE-BIRTH already uses), then installs its identity the same way capsule_wirebind.c already does live for WIREBIND attaches -- vm_identity_from_cert() verification followed by a plain post-birth struct assignment -- rather than anything RUNCAP-shaped, since RUNCAP requires a real blkio_dev+ homeblocks_sig_t an unattended identity never has. The born VM's own capsule payload is expected to lay down two CREATE'd buffers (UNATTENDED-ID-UUID, UNATTENDED-ID-CERT) via MINT-SCRATCH- EMIT's own literal format; their addresses are fetched by interpreting a two-word line inside the *new* VM's own context (vm_interpret(born_vm, ...)), the same "run inside that VM's own dictionary" idiom capsule_wirebind.c already uses for VM-NAME-REG. Explicit invariant preserved: never touches g_wirebind_attached_username or any console-pairing state, births no console VM -- an unattended identity stays un-promptable (§VIII.1) until a human pairs a console to it later via the already-working VM-NAME-REG mechanism. Verified live end-to-end on amd64: minted a real test identity via MINT-SCRATCH, captured its MINT-SCRATCH-EMIT output, built a throwaway test capsule from it (discovered along the way: mkcapsule's real block constraints are range [2048,5120) and max 16 content lines per block -- neither matches this repo's own doc comment, corrected via ground truth from the tool itself, not assumed), then ran UNATTENDED-BIRTH against it: cert verified against Zuse's root pubkey, identity installed, "no console attached" reported, Hera stayed healthy afterward (5 6 + . -> 11). Test capsule reverted after capture per this project's own probe convention -- not a real identity, never committed. Clean 3-arch qemu boot (amd64/aarch64/riscv64) on the real committed C-only change. Remaining punch-list item (the ACL cap bit for console attachment) not started. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K |
||
|
|
5fc709a228 |
Punch list item 2: MINT-SCRATCH-EMIT, verified live on all 3 arches
Adds the hand-transcription mechanism FABRIC-3.md §XXXII.2's punch list item 2 calls for: MINT-SCRATCH-EMIT prints the last successful MINT-SCRATCH's drive_uuid + cert devblock as ready-to-paste FORTH source (HEX-based CREATE ... C, ... byte sequences), so packaging an unattended identity into a capsule is a mechanical copy out of the captured boot log rather than a manual hex-to-FORTH translation an operator could transpose a digit in. Refuses (no output) if no MINT-SCRATCH has ever succeeded -- printing 4112 zero bytes as if they were a real identity would be a silent, misleading success, matching this session's own error-handling audit discipline rather than adding a new silent-failure primitive right after finishing one. Verified live on amd64: MINT-SCRATCH-EMIT correctly refuses before any mint, then after MINT-SCRATCH succeeds, emits UNATTENDED-ID-UUID and UNATTENDED-ID-CERT as valid FORTH literals. Cross-checked byte-exact: the cert's own embedded ASN.1 serialNumber field matches the emitted UUID bytes exactly, confirming x509_build_user_cert()'s drive_uuid binding round-trips correctly through the scratch-device path. Clean 3-arch qemu boot (amd64/aarch64/riscv64). Remaining punch-list items (writing/committing a real capsule file for an actual named identity, the unattended-birth call site, the ACL cap bit) not started -- authoring a real committed capsule needs a name/ purpose decision that isn't mine to make. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K |
||
|
|
63864c4b01 |
Punch list item 1: scratch-device MINT-SCRATCH, verified live on all 3 arches
Implements FABRIC-3.md §XXXII.2's scratch-thumbdrive mint mechanism: capsule_mint_identity_scratch() (capsule_mint.c/.h) builds a throwaway RAM-backed blkio_dev via blkio_ram.c's backend and runs capsule_mint_identity() against it completely unmodified -- same live Zuse-signing operation, same rng_get_bytes() draw for drive_uuid a real thumbdrive gets. Reads back only drive_uuid + the cert devblock; the seed devblock is written into the scratch buffer internally but never read out (no seed is ever baked into a capsule, per the ratified no-seed decision). New FORTH word MINT-SCRATCH (mama_forth_words.c), same stack signature as MINT, mints into the scratch device instead of any attached drive and never touches sk_repl_get_attached_blk_dev() or console-pairing state. Prints the drive_uuid as hex so a live boot log itself proves each call drew fresh entropy. Build correction found along the way: blkio_ram.c was excluded from the kernel build (Makefile.starkernel VM_EXCLUDE) alongside blkio_factory.c/blkio_file.c. blkio_factory_open() unconditionally references blkio_file.c's real fopen()/fread() file I/O, which has no freestanding-kernel equivalent, so the factory function couldn't be used as-is. blkio_ram.c itself is pure memcpy over a caller buffer -- pulled it alone into the kernel build and wired it directly in capsule_mint.c, the same way blkio_factory.c's own extern declarations do internally. Verified live on amd64: two MINT-SCRATCH calls produced two genuinely different drive_uuids (b533246d.../ae11b2b2...), confirming fresh entropy per call rather than stale reuse; VM stayed healthy afterward (5 6 + . -> 11). Clean 3-arch qemu boot (amd64/aarch64/riscv64), logs and DoE CSVs committed per standing convention. Remaining punch-list items (hand-transcription into a .4th block, the unattended-birth call site, the ACL cap bit) not started. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K |
||
|
|
42d4bf3dad |
Stage D batch 7 (final): mama_forth_words.c groups 5+6 -- dictionary lookup/test words + Stadium primitives; Stage D closed (FABRIC-3.md §XXXII.6)
NAME>XT/RUNCAP-TEST/PAIR-TEST and the six STADIUM-* physics primitives (STADIUM-ADMIT/STADIUM-EVICT/STADIUM-RES-PULL/STADIUM-RES-PUSH/ STADIUM-HEAT@/STADIUM-HEAT!), 12 sites -- closes mama_forth_words.c and the entire kernel-only error-handling audit. All 105 sites from the §XXXII.3 triage now accounted for: 28 already correct, 75 silent sites fixed across repl.c/inference_words.c/ log_words.c/vm_core.c/mama_forth_words.c, 2 special cases resolved by dropping the error per their own documented contract, 1 resolved via console_println() per its own recursion constraint. Verified with awk: zero remaining vm->error=1 sites in mama_forth_words.c lack a diagnostic within the preceding three lines. Three-arch clean qemu acceptance passed. This closes Stage D and the USE/logging/audit thread opened in §XXXII; Stage E (human-vs-unattended identity model) remains open, not started this pass. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K |
||
|
|
90c86c6006 |
Stage D batch 6: mama_forth_words.c group 4 -- identity/crypto words (FABRIC-3.md §XXXII.6)
MINT + mint_pop_string()/ZUSE-ELIGIBILITY-ADD/ZUSE-ELIGIBLE?/ ELEVATE-PUBKEY-UNPACK, 10 sites. mint_pop_string() gained a field_name parameter so its diagnostics name which of MINT's four string arguments failed (phone/email/username/full_name), rather than a generic message that would leave the operator guessing. Diagnostic placement matched each function's own sibling convention where one exists (ZUSE-ELIGIBILITY-ADD), log_message() default otherwise. Three-arch clean qemu acceptance passed. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K |
||
|
|
8b5300fc4f |
Stage D batch 5: mama_forth_words.c group 3 -- cross-VM execution/dispatch, both flagged special cases resolved (FABRIC-3.md §XXXII.6)
VM-STEP/VM-EXEC/VM-CALL (5 sites) gained console_println() diagnostics
matching their own existing sibling guards.
SWITCH-MARK-WORK and VM-HEAT (3 sites) resolved per §XXXII.3's own
recommendation: dropped vm->error entirely rather than diagnosing it,
matching each function's own doc comment ("must never error or spam
the console" / "does not print/error"). SWITCH-MARK-WORK is the same
function §XXVIII.3 already fixed once for an off-by-one that fired
silently on every MSG-SEND in the system -- the guard now genuinely
cannot repeat that by contract, not just by the threshold being right.
VM-HEAT's guards now push 0 and return, matching its own
always-returns-a-value stack effect.
Three-arch clean qemu acceptance passed. Zuse's own WIREBIND attach
(every boot) drives MSG-SEND -> SWITCH-MARK-WORK, so this batch's most
safety-critical fix is exercised by standard acceptance, not just
compiled.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K
|
||
|
|
b42c3b195c |
Stage D batch 4: mama_forth_words.c groups 1+2 -- capsule/lifecycle words + BIRTH/START/KILL (FABRIC-3.md §XXXII.6)
15 of mama_forth_words.c's 48 silent sites fixed, split by functional
grouping per direct instruction: CAPSULE@/CAPSULE-HASH@/CAPSULE-FLAGS@/
CAPSULE-LEN@/CAPSULE-BIRTH/CAPSULE-RUN/EXEC (9), and BIRTH/START/KILL
(6).
Refined the diagnostic-placement rule: match whichever convention that
same function's other already-correct guards use, rather than
defaulting uniformly. BIRTH/START/KILL/EXEC each already had a
console_println() sibling guard ("name too long or empty") -- their
newly-diagnosed guards now match that, same shape as USE's own fix.
The CAPSULE*@ words have no sibling guard to match, so they keep
log_message() (defer_words.c's gold-standard default).
Three-arch clean qemu acceptance passed; BIRTH itself is exercised by
every boot (Hermes/Artemis birth).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K
|
||
|
|
66c2f3e539 |
Stage D batch 3: fix silent error sites in vm_core.c, incl. the two highest-value primitives; two real NULL-deref bugs found and fixed (FABRIC-3.md §XXXII.6)
All 13 silent vm->error=1 sites in vm_core.c now log a diagnostic
first: vm_enter_compile_mode, vm_compile_word, vm_compile_literal,
vm_compile_call, vm_exit_compile_mode, execute_colon_word (2 sites
each/combined), and the four fundamental memory primitives
vm_load_u8/vm_store_u8/vm_load_cell/vm_store_cell.
Found and fixed two real NULL-pointer-dereference risks while adding
the diagnostics: vm_compile_call() and vm_exit_compile_mode() each had
a combined `if (!vm || <cond>) { vm->error = 1; ... }` guard that
dereferenced vm->error even on the !vm branch of its own condition.
Split both, and applied the same defensive split to the four memory
primitives since vm_ptr()/vm_addr_ok() both tolerate vm==NULL
internally.
Live-verified, amd64: 999999999 @ . recovers correctly at Zuse's own
console (Stage A's mechanism holds), but the new vm_load_cell
diagnostic itself didn't print -- traced to memory_words.c's own
redundant, still-silent vm_addr_ok() pre-check in memory_word_fetch()
(and the same shape in memory_word_store()), which intercepts before
ever reaching vm_load_cell(). memory_words.c is vendored, out of this
initiative's scope, spun off to FABRIC-4.md -- recorded as a concrete
cross-reference for that future work rather than left to be
rediscovered.
Three-arch clean qemu acceptance passed, full POST suite included.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K
|
||
|
|
1a8c0e4fcf |
Stage D batch 2: fix silent error sites in log_words.c, resolve the flagged category-iii special case (FABRIC-3.md §XXXII.6)
log_word_set_level(), log_do_emit(), and log_emit_string() (7 sites total) now log a diagnostic via log_message(LOG_ERROR, ...) before setting vm->error, matching this file's own already-correct log_str_emit() sibling and defer_words.c's gold-standard pattern. log_word_append_raw() (backing (LOG-APPEND-RAW), 3 sites) resolved differently per §XXXII.3's own triage note: its doc comment forbids log_message() here (recursion into the log ring it writes to), but the word can also be invoked by hand at the console -- console_println() carries no such recursion risk and matches Stage B's policy for a manual interactive invocation. Added the console.h include this required. Three-arch clean qemu acceptance passed. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K |
||
|
|
5417ffb2bf |
Stage D batch 1: fix silent error sites in repl.c and inference_words.c (FABRIC-3.md §XXXII.6)
repl.c's sk_word_blk_attach_ack() and inference_words.c's array_ptr() helper + infer_word_run()'s allocation guard now log a diagnostic via log_message(LOG_ERROR, ...) before setting vm->error, matching defer_words.c's own gold-standard pattern (§XXXII.3) -- these are internal/background conditions (a malformed message-callback, a bad array reference or allocation failure), not interactive usage mistakes, so the fix keeps the fault and reports it rather than dropping it like USE's own fix did. Batched together (4 sites total, smaller combined than the next file) rather than two separate acceptance cycles for negligible size. Three-arch clean qemu acceptance passed. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K |
||
|
|
c05b70c8d6 |
Stage B: logging policy documented, level-aware log-ring eviction built, LOG-FLUSH deferred again (FABRIC-3.md §XXXII.4)
Policy decided for the kernel-only audit scope: an interactive command's direct response stays on console_println/console_puts; everything else (state transitions, background diagnostics, audit trails) routes through log_message() at the appropriate level, matching capsule_mint.c's verify_mint() precedent. Documented, not code-swept here -- reclassifying individual sites is Stage D's job. LOG-FLUSH deferred again, explicitly: the per-VM log buffer its own doc comment presumes (vm_log_buffer.h) doesn't exist anywhere in the tree -- building it is real feature work needing its own scoped stage. Level-aware eviction built: log_region_append() now reads the oldest ring slot's own level before evicting it, protecting ERROR/WARN records from being pushed out by INFO/DEBUG churn -- drops the incoming low-priority record instead. Found and fixed an adjacent bug while making this change: the prior two-valued return contract would have made a benign "dropped by design" outcome indistinguishable from a genuine write failure to its one caller, which unconditionally set vm->error on any nonzero return. Changed to a three-valued contract (0 success, 1 dropped by design, -1 genuine failure). Three-arch clean qemu acceptance passed. Eviction path itself not live-exercised (needs 128+ LOG-APPEND calls to fill the ring) -- flagged, matching this project's own precedent for that kind of gap. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K |
||
|
|
5c5896fbc1 |
Stage A: fix USE's silent stack-underflow/bad-address guards, and the Hera fault-scoping gap they exposed (FABRIC-3.md §XXXII.1)
mama_word_use()'s two silent vm->error=1 guards (dsp<1 stack underflow, NULL from vm_ptr()) now print a diagnostic and return, matching the function's other five guards. Live verification of that fix alone surfaced a bigger problem: with the guard no longer silent, the REPL proceeds to interpret the leftover token as an unrecognized word, which independently sets vm->error, and sk_repl_step()/sk_repl_run()'s Hera-branch still hard-halted on that. Investigated kernel_main.c's boot/capsule-load paths directly: they already catch and clear mama->error entirely separately, before sk_repl_run() is ever entered -- so the "no fallthrough surface" halt in these two REPL functions was never protecting a boot-time fault, only an ordinary interactive REPL-turn one. Both functions now recover unconditionally on any VM's error, Hera included, matching how a redirected (WIREBIND/USE'd) identity's session already recovered. Removed the now-fully-unused sk_fault_handler(). Verified live, amd64: USE rajames at Zuse's own console prints the new diagnostic, then "VM fault -- session recovered, resuming", console stays interactive afterward. Three-arch clean qemu acceptance passed before and after the Hera-fault-scoping change. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K |
||
|
|
d9da82b065 |
Stage 4 increments 2+3: WIREBIND VMs as switch-signal participants + mark-and-defer tombstone reap (FABRIC-3.md §XXVIII Stage 4)
Increment 2: WIREBIND user VMs (the ones that actually run FORTH work; console VMs are pure REPL proxies and never participate) register as Stage 3 switch-signal participants at attach, unregister at teardown. Slot table bumped 8 -> 16, matching messaging.4th's own VM-MAX -- a real, already-agreed ceiling, not an invented number. Added sk_vm_switch_signal_unregister() (compaction-based; Tripod VMs never needed removal, WIREBIND VMs cycle constantly and would otherwise exhaust the bounded table). Increment 3: implements the plan's own ratified option (A) for the async-detach UAF risk -- mark-and-defer via a new pending_reap flag on VMRegistryEntry, deliberately not a new VMState (capsule_vm_kill() already treats VM_STATE_DEAD as idempotent success, which would silently swallow a reap attempt; SWITCHED_OUT still accurately describes a tombstoned VM until the moment it's actually freed). unclean_detach() sets it when capsule_vm_kill() refuses a SWITCHED_OUT target; the Stage 3 checkpoint (vm_core.c) checks it before ever attempting to resume a pending switch target, and calls the new capsule_vm_force_reap() instead -- the one caller allowed to bypass capsule_vm_kill()'s own refusal, because it runs at the exact safe cooperative point the switcher itself controls. A new idle-tick sweep cleans up the WIREBIND live-table entry once the reap has actually happened. Verified clean on all 3 architectures (baseline regression -- no WIREBIND attach happens in a plain boot). The reap mechanism's own correctness under a genuinely parked context is verified separately, next, via a temporary deterministic probe. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K |
||
|
|
9f0f33dfc5 |
Stage 4 increment 1: per-device WIREBIND tracking, fixing a real multi-identity detach leak (FABRIC-3.md §XXVIII Stage 4)
capsule_wirebind_unclean_detach()/eject() tracked "the attached identity" as a single global, correct for the console-pairing UX (one physical console) but wrong for detach safety: since §XV/§XVI proved multiple identities genuinely live simultaneously via this same attach path, every attach after the first silently overwrote the singleton, so an unclean detach of any but the most-recently-attached identity was silently ignored -- that VM leaked forever, no trace in the log. Adds a per-device live-identity table, separate from the (unchanged) console-pairing singleton, so unclean-detach resolves any attached device to its own identity. Sized off messaging.4th's own VM-MAX (16) minus Tripod's 3 reserved slots, not an invented number. Corrects the stale "single-USB-device constraint" doc claim in capsule_wirebind.h, false since §XV/§XVI. Groundwork for Stage 4's real deliverable (WIREBIND VMs as switch-signal participants) -- this increment only fixes detach targeting; switch- signal registration is next. Verified clean on all 3 architectures (no WIREBIND attach happens in a plain boot, so this is a regression check on the existing Tripod-only path; live multi-identity verification comes with the switch-signal registration increment). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K |
||
|
|
05ae7aa886 |
Fix SWITCH-MARK-WORK off-by-one: was tripping vm->error on every MSG-SEND (FABRIC-3.md §XXVIII.2)
mama_word_switch_mark_work()'s stack-underflow guard checked dsp < 2, requiring 3+ items, when it only ever needs the 2 IDX>NAME leaves it (caddr u). Since dsp is index-based (2 items == dsp 1), this rejected every normal call. MSG-SEND tail-calls SWITCH-MARK-WORK unconditionally, so this fired on every message sent anywhere in the system -- visible only where a caller happened to check the target VM's error flag afterward (mama_word_vm_exec()'s "VM-EXEC: ERROR in Artemis" report). Verified clean on all 3 architectures: boot reaches Startup: Artemis live -> zuse@Hera] ok> with no VM-EXEC: ERROR line at all. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K |
||
|
|
66beae7fd4 |
Stage 3 follow-on: message-arrival eligibility hook + trampoline-blind switch-storm fix (FABRIC-3.md §XXVIII.2)
Implements the message-arrival eligibility signal FABRIC-3.md §XXVIII.1 left open (has_work per-slot flag, set via new SWITCH-MARK-WORK primitive from MSG-SEND) so an idle VM never becomes a switch target purely by waiting out the readiness threshold. Also root-causes and fixes a second, independent switch-storm: the tick's "who is current" check used vm_log_attributed_vm(), which can't see a VM parked in switch.c's own raw trampoline. Replaced with a dedicated g_switch_current_vm tracked by the switch mechanism itself, and moved target-slot eligibility reset to the switch decision point instead of relying on ISR polling to observe a window that can be only a few instructions wide. Verified live on all 3 architectures: clean boot to zuse@Hera] ok>, live cross-VM message dispatch, and (since a quiet log looks identical to a livelocked storm once the DoE probe is gone) confirmed genuine REPL liveness via QMP send-key + screendump on aarch64/riscv64, not log inspection alone. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K |
||
|
|
862d7d9c48 |
Stage 3 follow-on: fix stack-ownership corruption + DoE switch columns (FABRIC-3.md §XXVIII.1)
DoE CSV gained 6 switch-signal columns (switch_count_cumulative, switch_current_slot, switch_*_readiness, switch_ticks_since), and verifying them with a boot-time HB-ON probe surfaced a real livelock: the preemption checkpoint could fire inside a VM-EXEC-nested execute_colon_word() call and switch away from a stack it didn't own, parking a borrowed region of the caller's stack under the wrong VM's saved-context pointer. The trampoline bounce was the visible (safe) half of this; the corruption was the quiet half, live in every prior "clean" Stage 3 boot without ever showing up in the log. Fixed by gating the checkpoint on being at the outermost vm_interpret() call (g_vm_interpret_depth / sk_vm_at_outermost_interpret(), vm_core.c), per Bob's decision. Also fixed two related bugs found in the same pass: g_switch_back_to was a single global stale after first entry, now per-VM state (native_switch_back_to); note_switch_performed() fired on resume instead of switch-out, now called before the switch. Verified on all 3 architectures: steady log growth (no freeze), zero leaked QEMU processes, DoE columns internally consistent, Hermes/Artemis confirmed genuinely executing (not just trampoline-bouncing). Temporary HB-ON boot probe reverted after capture. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016UNhH1mhi52i6Qihh7ZV5S |
||
|
|
986d042aa7 |
Stage 3: timer-driven preemptive switching, live on all 3 arches (FABRIC-3.md §XXVIII)
Fourth stage of the preemptive context-switching plan, and the biggest. LithosAnanke now genuinely, continuously preempts between Hera, Hermes, and Artemis -- timer-driven, running live for the entire remainder of every boot once the Tripod fleet registers, not a bounded probe. A real design fork was resolved before writing code: the naive approach (the timer ISR calling Stage 2's sk_vm_context_switch() directly) is broken -- Stage 0's trap frame lives on whatever stack was active at interrupt time, and jumping to a different stack via Stage 2's own independent swap mid-handler would abandon that trap frame unresumed, guaranteed corruption on the first tick. Chose the safer of two named options: the ISR only ever sets a flag and returns completely normally through its own full epilogue; the actual switch happens moments later, via Stage 2's already-proven mechanism, at a safe cooperative checkpoint on the mainline (execute_colon_word()'s per-word dispatch loop, checked on literally every word, not throttled to the existing 256-word heartbeat-tuning cadence) -- confirmed with the user that word-level granularity is fine-grained enough given the eventual Zynq FPGA target where a word is a mnemonic. New capsule_vm_switch_signal.c/.h: a purpose-built run-readiness signal, deliberately separate from capsule_vm_physics.c's execution-heat engine (that one's own header documents itself as never touched from interrupt context, by design). Slot table sized with headroom (8) rather than hardcoded to today's 3 participants, so extending participation later is another register() call, not a redesign -- per direct request to leave room for swapping the participant set. Simple linear accumulate-then- threshold for this first cut; a fancier law can replace it later without touching the mechanism around it. heartbeat_tick() gains its one deliberate, documented amendment to this file's own top-half/bottom-half discipline -- the first time this codebase reaches into VM-scheduling state from real ISR context. Registration happens only after all three VMs are fully born, right before the REPL starts -- no critical-section protection yet against being switched away mid-birth-setup. Known, flagged rough edge (not reconciled this pass): MSG-TICK's own idle-pump and this new mechanism can still independently move control between the same VMs; not observed to interact badly in verification, but not fully unified either. Verified interactively at the console on all 3 architectures with continuous background preemption running throughout -- amd64 computed `1 1 + .` -> 2, aarch64 computed `1 1 + dup DUP * . CR` -> 4, both correct, REPL fully responsive, zero fault indicators over sustained runtime. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016UNhH1mhi52i6Qihh7ZV5S |
||
|
|
f790d0995e |
Stage 2: cooperative VM context switch primitive, proven on all 3 arches (FABRIC-3.md §XXVIII)
Third stage of the preemptive context-switching plan. The real save/restore switch mechanism now exists -- the first time anything has ever executed on a VM's own native stack (Stage 1 allocated them, unused). New sk_vm_switch_to() (switch.S, one per arch) is an ordinary function call, not an interrupt -- so unlike Stage 0's trap frame, the ABI already covers every caller-saved register; only the callee-saved set needs explicit save/restore (amd64: rbx/rbp/r12-r15, no FP at all since SysV has no callee-saved XMM; aarch64: x19-x28/x29/x30 + d8-d15; riscv64: s0-s11/ra + fs0-fs11, FS-gated like Stage 0 but read once and reused for both halves within one call, since FS is genuine global CPU state, not part of what's switched). A sibling sk_vm_switch_prime() in the same file builds the synthetic first-entry frame, kept in assembly so the layout can never drift out of sync with sk_vm_switch_to() itself. New switch.c/switch.h: sk_vm_context_switch(from, to) handles first-entry priming vs. resuming a parked context, and updates registry state (new VM_STATE_SWITCHED_OUT, distinct from VM_STATE_STOPPED -- STOPPED means no live frame, this means the opposite). sk_vm_switch_entry() is the minimal permanent trampoline every freshly-entered VM lands in: no production behavior defined yet, so it just yields straight back to whoever switched to it, forever. Closes the confirmed unguarded-KILL UAF found during planning: capsule_vm_kill(), mama_word_kill(), and capsule_vm_kill_all_nonmama() all now refuse (or silently leak rather than free, on the cold-restart path where arch_cold_reset() wipes everything immediately after anyway) tearing down a switched-out VM. Side effect found, not built on purpose: the existing MSG-TICK idle-pump already filters on VM_STATE_LIVE, so it automatically stopped dispatching into a switched-out VM with zero changes needed there. Verified via a temporary SWITCH-TEST probe (boot-triggered, since nothing can type interactively into a foreground-only QEMU session) that round-tripped a sentinel through 5 real Hera<->Hermes switches on all 3 architectures: 5/5 rounds, 0 failures, clean continuation to ok>. Probe fully reverted after capture; kernel_main.c shows zero diff. Also: Makefile.starkernel's LOADER_EXTRA_SRCS/LOADER_ASM needed the new files added explicitly (this project's "loader" PE binary is the full running kernel, not a thin bootstrap stage), and aarch64's switch.S needed the same #ifndef _WIN32 guard around .hidden that isr.S already carries (aarch64's loader assembles via clang targeting a PE/COFF target with no .hidden equivalent) -- caught by a build failure, fixed. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016UNhH1mhi52i6Qihh7ZV5S |
||
|
|
57ac3fc304 |
Stage 1: per-VM native stacks, allocated but not yet executed on (FABRIC-3.md §XXVIII)
Second stage of the preemptive context-switching plan. Every VM (Hera, every capsule_birth_baby()-born VM including WIREBIND identities) now gets its own dedicated 2 MiB native C stack at birth -- but nothing runs on it yet, that's Stage 2. Pure allocation-machinery proof. Design correction made before writing code: the plan called for cloning sk_vm_arena_alloc()'s guard-page pattern, but that pattern turns out to be Mama-only -- host_services.c's kernel_alloc() gives every baby VM a plain kmalloc() block for its dictionary arena, not a real guarded PMM allocation. Stacks get the real treatment instead (new sk_vm_native_stack_alloc()/_free() in arena.c): independent pmm_alloc_contiguous() + guard pages for every VM without exception, no singleton, no kmalloc fallback -- a stack overflow is exactly the failure mode guard pages exist for, and a corrupted stack could corrupt whatever saved context Stage 2 trusts. 2 MiB size matches this project's own established kernel-stack convention (g_kernel_stack/g_rpi5_native_stack), not a guess -- that one shared 2 MiB stack today already carries all VMs' combined nested VM-EXEC recursion. Three new VM struct fields, freed in vm_cleanup() alongside the existing call_stack free. Allocation failure is non-fatal to birth. All 3 architectures re-verified clean boot to ok>, no native-stack allocation failures for any Tripod-fleet VM. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016UNhH1mhi52i6Qihh7ZV5S |
||
|
|
15672ce17c |
Stage 0: trap-frame parity across all 3 arches (FABRIC-3.md §XXVIII)
First stage of the preemptive context-switching plan (see ~/.claude/plans/logical-snuggling-bear.md). Pure foundation work -- every arch's ISR now saves the full register set on interrupt entry, so a trap frame is in principle sufficient to resume execution anywhere it was taken. No FORTH-visible behavior changes. amd64: added FXSAVE/FXRSTOR, closing a genuine pre-existing correctness gap (not just future-preemption prep) -- confirmed live double-precision FP code reachable from ordinary interpreter dispatch (vm_runtime.c Loop #5/#6), and the ISR previously saved zero FP/SSE state. rbp repurposed as a fixed anchor so the 16-byte-aligned FXSAVE area can be carved out of an unpredictably-aligned rsp without disturbing existing argument reads. aarch64: extended the trap frame 672->800 bytes, adding v8-v15 (AAPCS64 callee-saved, previously excluded on call-site-only reasoning that doesn't hold for an async trap). riscv64: extended the trap frame 320->512 bytes, adding s0-s11 and fs0-fs11 (the latter still correctly gated behind sstatus.FS != Off). All 3 architectures re-verified clean boot to ok> under the new frames -- amd64 through hundreds of timer ticks with FXSAVE/FXRSTOR live on every interrupt, aarch64 through 987 ticks, riscv64 clean on the now-larger FS-conditional block. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016UNhH1mhi52i6Qihh7ZV5S |
||
|
|
2a30212bd3 |
Real per-VM log persistence: source attribution + ACL pin (FABRIC-3.md §XXVII)
Wires the previously-unused vm_log_attributed_vm() into LOG-APPEND's kernel primitive so persisted log records carry a trustworthy source (the real attributed VM's registry name, or "HADES" pseudo-source) instead of a caller-supplied, trivially forgeable string. Drops src-addr/src-u from LOG-APPEND's stack signature accordingly. Pins LOG-APPEND via bare ACL-PIN in Artemis's own init.4th, matching BIRTH/CAPSULE-BIRTH's precedent for a privileged word that can't reach the shared, host-portable ACL.4th. Also fixes two console-banner nitpicks: a mis-rendering em dash (U+2014) in the boot banner, and drops "Emergency" from the CLI banner text. Doc corrections to artemis_sig.h/zuse_eligibility_list.h reconciling the three fixed devblock ranges now in play. LOG-FLUSH (the intended normal entry point) and level-aware log eviction remain open, flagged not fixed. Re-verified clean boot to ok> on all 3 architectures after every change. riscv64 showed one new, unrelated virtio_blk write-timeout anomaly during Artemis's early physics self-test (self-recovered, boot unaffected, sector doesn't map to the log region) -- flagged, not investigated. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016UNhH1mhi52i6Qihh7ZV5S |
||
|
|
61755fde78 |
Artemis genesis stamp: fix a BAM-corrupting offset before it ever ran (FABRIC-3.md §XXVI follow-on, Step 3)
Step 3: one-time artemis_sig_t genesis stamp, written once
kernel_main.c's virtio-blk path confirms Artemis's own disk, so the disk
image is later recognizable generically (repl.c's idle-loop USB-MSC scan,
built in the prior commit) regardless of which bus found it.
Correction made before this ever touched the real disk: the signature's
first design (committed in
|
||
|
|
29b6789860 |
Artemis bus-agnostic discovery: signature format + idle-loop generalization (FABRIC-3.md §XXVI follow-on)
Step 1: new artemis_sig_t header format (magic 'ARTM', sibling to homeblocks_sig_t, distinct so a generic scan can tell Artemis's own disk apart from an identity thumbdrive by content alone) -- artemis_sig.h/.c, wired into Makefile.starkernel. Step 2: sk_repl_idle()'s existing per-USB-MSC-slot attach handling (the pattern WIREBIND already uses for identity thumbdrives) now also checks for the ARTM signature whenever a device's home-blocks check comes back BLANK. On a match, once Artemis's own storage-attach round-trip (HERA-BLK-ATTACH-REQ/BLK-ATTACH-ACK) confirms success, capsule_zuse_boot_load_root_pubkey() runs -- the same call kernel_main.c's synchronous QEMU-only virtio-blk path already makes, now reachable without a hardcoded PCI vendor/device scan. That function is already idempotent (no-op once zuse_root_pubkey_known is set), so no boot restructuring was needed despite the initial concern that deferring Artemis discovery to the idle loop would require one. Verified: clean build + QEMU boot to [zuse@Hera] ok> on all three architectures, zero regression to the existing virtio-blk/Zuse-thumbdrive attach path. Steps 3 (genesis-stamping onto disk/artemis.img) and 4 (growable production log-persistence region) not yet started. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo |
||
|
|
a8b16d41da |
Unify console prompt to [user@VM]; fix real personality-block truncation; correct §XXV's wrong lockdown conclusion (FABRIC-3.md §XXVI)
Investigating the std79 lockdown finding from FABRIC-3.md §XXV led to a real discovery: WIREBIND births TWO VMs per identity, a console proxy under the plain username and the actual restricted identity under <username>~user (capsule_wirebind.c). Every test in §XXV targeted the console proxy, which was never locked down at all. Retested against the correct target (rajames~user): the lockdown works exactly as designed. §XXV's "lockdown never engages" conclusion was wrong -- corrected here, not deleted, since the mistake and how it was caught are worth keeping (see the new feedback memory: confirm which specific VM a name resolves to before concluding anything, when a subsystem is known to birth more than one VM per identity). Two real, separate things found along the way are kept regardless of that correction: - capsule_runcap.c: the reserved personality devblock was read in full (mostly zero-padding after a short ~200-byte string) with no terminator, producing "WARN: block 4998 exceeds 1KB, truncating" on every std79-locked identity's birth, universal, since at least 2026-09-10. Fixed by trimming to the first NUL byte actually found -- real, but harmless to execution (real content sat in the truncated block's surviving head); it mattered for capsule_id/ content_hash being computed over padding instead of real content. - console.h/console.c/repl.c: unified the prompt from a separately- computed "[VMName] (user)" into a single "[user@VMName]" line prefix -- exactly the ambiguity that caused the original misdiagnosis (the prompt showed only the WIREBIND username, identical whether USE had targeted the console proxy or the real ~user identity). Implemented as a registered callback (console_set_user_prefix_provider()) rather than console.c calling into WIREBIND/session logic directly, since console.c is a clean HAL module with no prior dependency on capsule-level subsystems. Verified: clean build on all 3 architectures, zero new warnings, identical dict_hash/capsule_hash to every prior boot this session (console/prompt-only change). Full 9-identity messaging campaign re-run end to end: 202s, zero faults, all 8 identities at 99/99 tokens, zero regression. Also surfaced, not yet acted on: the full campaign's own console tags now visibly show which VM each identity's tests actually reached ([zuse@rajames], not [zuse@rajames~user]) -- messaging.4th's VM-NAMES-INIT registers identities by plain username, so std79-doe. fth's turn-attractor has been dispatching to each identity's console proxy, not the actual locked-down identity, since the messaging rewrite. Flagged for a deliberate decision, not investigated further. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo |
||
|
|
3c2daf50d1 |
Extend BIRTH/CAPSULE-BIRTH to all VMs symmetrically; flag a real std79 lockdown gap (FABRIC-3.md §XXV)
Scoping the workload-into-factorial design's placement-mode factor led to a real architectural improvement: rather than EXEC-ing a workload capsule into an already-running, ACL-locked identity's own persistent dictionary (filesystem-shaped, doesn't dodge the block- collision exposure just traced in §XXIV), a workload now runs as a fresh ephemeral child VM, BIRTH'd per trial and reaped after -- matching the project's own stated principle of automanagement over imposed policy. CAPSULE-BIRTH already passes vm->stadium_vm_id (who is birthing this VM) as the new child's parent, not a hardcoded Hera constant, confirmed by reading the C -- so a workload trial genuinely inherits the specific identity's own lineage when that identity does the birthing. Which surfaced a real premise: only Hera could call BIRTH/ CAPSULE-BIRTH at all (registered only in register_mama_forth_words(), confirmed directly, not part of the earlier §XX messaging-symmetry fix which deliberately kept this as one of her remaining privileges). Extended symmetrically now, agreed explicitly before touching code: - mama_forth_words.c: BIRTH and CAPSULE-BIRTH added to register_child_vm_words(), matching §XX's own pattern. - acl-std79.4th: ' BIRTH , ' CAPSULE-BIRTH , added to ACL-STD79-LIST (new block 4048) -- a deliberate, explicit, named exception to the lockdown's own "standard words only" guarantee, not a silent one. Symmetric registration alone can't weaken any lockdown on its own: ACL-LOCKDOWN-STD79 is allowlist-based, deny-by-default, so a newly registered word is auto-denied there unless explicitly added. Verified: clean build on all 3 architectures, zero new warnings. Hera's own dict_hash unchanged (expected); Hermes/Artemis show the same new dict_hash on all 3 architectures. Live-tested against a real attached std79-locked identity: CAPSULE-BIRTH executes correctly (returns vm_uuid_none() for a deliberately out-of-range capsule-id, zero fault, zero ACL denial). Found, and explicitly stopped short of fixing, a separate pre- existing gap while verifying the above: MSG-STATUS and MSG-K (messaging.4th words, not on the std79 allowlist) execute for a locked identity instead of being denied. ACL-LOCKDOWN-STD79 is confirmed to actually run; something more specific isn't reaching messaging.4th's dictionary entries. Root cause not traced -- needs its own investigation into vm_core.c's dictionary-link mechanics and whichever capsule actually loads messaging for these identities. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo |
||
|
|
71bcb72a59 |
Fix O(N) idle-loop messaging pump; full 3x9x3 turn-attractor campaign clean on all 3 architectures (FABRIC-3.md §XXII)
sk_repl_idle()'s messaging pump (repl.c) walked the entire live-VM registry every idle beat (~1Hz) and dispatched a full VM-EXEC "MSG-TICK" -- a dictionary lookup plus a 32-slot arena scan -- into every live VM, every tick, unconditionally, forever. Fine at Tripod's original 3-VM scale; a full 9-identity turn-attractor campaign exposed it as a genuine wall on riscv64 specifically (its TCG makes each dispatch cost more): the same campaign that completed in 194s/ 339s on amd64/aarch64 never finished on riscv64 at 9 VMs across three attempts, while 8 VMs there was fine in 159s. Ruled out capacity explanations before touching anything: bumping riscv64's QEMU RAM 1024->4096 changed nothing (reverted), and a live STADIUM-RES@/MSG-STATUS probe with all 9 VMs attached showed no depletion. Host memory pressure was also ruled out directly (one background task did get OOM-killed once during the investigation, but the identical stall reproduced again with 9.2GB free). The real mistake was three premature kills under 4 minutes with no way to tell "slow" from "stuck" from outside the guest -- fixed by having run_doe_batch.sh sample the qemu process's own /proc/<pid>/stat utime every 60s; with that signal, riscv64 at 9 VMs was unambiguously alive (climbing utime, no hang), just disproportionately slow going from 8 VMs (159s) to 9 (600s+ and climbing). This was never really a riscv64-only bug: an O(N) per-second walk over the full VM population doesn't scale to the hundreds of VMs this fleet is headed toward, on any architecture -- riscv64 just made it visible first, at N=9, because its per-dispatch cost is highest. Fixed by round-robin batching: the pump now dispatches to at most SK_MSG_PUMP_BATCH (4) live VMs per idle beat via a persistent cursor that resumes where the previous beat left off, instead of all of them every time. Bounds both the scan and dispatch cost to O(K) regardless of total VM count; any single VM's queue now drains roughly every ceil(N/K) beats instead of every beat, still bounded and still matching the pump's own existing best-effort contract. No new C primitives, no messaging/Stadium changes. Verified: clean build on all 3 architectures, then the full 3x9x3 campaign re-run on all 3 (not just riscv64) per the standing rule that a defect repair requires a clean re-run everywhere before anything counts as closed: amd64 198s 0 faults 99/99 tokens x8 656/656 K-conserved aarch64 339s 0 faults 99/99 tokens x8 659/659 K-conserved riscv64 178s 0 faults 99/99 tokens x8 659/659 K-conserved riscv64 went from "never completes" to faster than aarch64, same campaign, same seed, same identity set. No regression on amd64/ aarch64. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo |
||
|
|
7edccd2c35 |
Hera can't message: root-caused and fixed by symmetry (FABRIC-3.md §XX)
Hera never loaded common:messaging.4th, unlike every other VM in the fleet. The standing belief was this was deliberate, to avoid moving her dict_hash off baseline. Checked live instead of assumed: loading messaging.4th into her dictionary silently dropped every colon- definition referencing one of 8 STADIUM-* primitives that register_child_vm_words() gives every other VM but register_mama_forth_words() never gave her -- a missing-primitive gap, not a designed privilege boundary. Confirmed mama_word_birth is genuinely VM-agnostic and SPAWN-EVENT is an unwired placeholder before proposing the fix. Fix: register the same 8 STADIUM-* primitives for Hera, load messaging.4th from init.4th the same way Hermes/Artemis/console/mint already do, and give her own idle-loop context a direct MSG-TICK call (not VM-EXEC, which would hit the same reentrancy class the existing per-other-VM pump loop already guards against) so her own queued messages actually drain. Her dictionary is now a proper superset of every child VM's, plus her remaining extra privileges -- not structurally different from any other VM, just additionally privileged. Verified: dict_hash identical across amd64/aarch64/riscv64 (0xc8f4b09e36f4fc4a), Hermes/Artemis dict_hashes unchanged and still cross-arch identical, all three boot clean to zuse)ok> with zero UNKNOWN WORD faults, mkcapsule --lint clean. Unblocks rewriting the turn-attractor (FABRIC-3.md §XIX) to coordinate via real MSG-SEND/MSG-TICK instead of blocking VM-EXEC. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo |
||
|
|
5a9425b91b |
Add VM-HEAT primitive: groundwork for a compudynamic turn-attractor (FABRIC-3.md §XIX)
Prompted by reading the std79-doe K report: checked whether the campaign's
"real cross-VM dispatch load" was actually concurrent or strictly
serialized. It's serialized at two levels -- VM-EXEC's vm_interpret(target,
...) is a direct synchronous C call (Hera fully blocked until it returns),
and even the background physics tick (vm_tick(), vm_runtime.c) is driven
by each VM's own execution loop, so idle identities accrue zero ticks
between their own turns. K's perfect conservation (FABRIC-3.md §XVIII)
verifies sequential per-VM accounting correctness, not concurrent-access
safety, since there was never concurrent access to test.
Agreed direction: fix this without a scheduler, by reusing the same
least-dense-candidate judgment stadium_admit() already trusts for
eviction, applied to "whose turn is next" instead of "who gets evicted" --
a fleet-level turn-attractor giving the next turn to whichever live
identity currently has the lowest execution_heat_q48, no fixed round-robin,
no priorities, no preemption. Lives beside Stadium in
capsule_vm_physics.c (already the fleet-level consumer of Stadium
primitives, e.g. the K mechanism itself), not inside stadium.c ("the
floor" -- residency/eviction, a different concern from turn order) and
not a new subsystem.
This pass lands only the primitive the mechanism needs: VM-HEAT
( c-addr u -- heat-q48 ), pushing a named VM's current
execution_heat_q48 via vm_physics_heat_of() -- previously C-internal
only (doe_log_heat_by_name(), doe_log.c), never exposed to FORTH. Silent
0 on an unknown/dead name (no print/error), since a turn-attractor
scanning many candidates every turn shouldn't have to filter console
noise for names that simply aren't live. Registered everywhere
VM-EXEC/VM-CALL already are. Builds clean on all three architectures;
live-tested on amd64: Hera -> 65452, Hermes -> 43, unknown name -> 0, no
faults.
The turn-attractor loop itself (std79-doe.fth's trial ordering) is not
yet built -- next step, not done here.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo
|
||
|
|
66ba21adb4 |
Log fleet_k_q48/fleet_conserved; K holds exactly, 775/775 ticks (FABRIC-3.md §XVIII)
doe_log.c's per-heartbeat-tick CSV gains two columns: fleet_k_q48 (vm_physics_fleet_heat_sum() over ALL live VMs -- the genuine fleet-wide conservation invariant K, not reconstructable from the 3 named-Tripod- member heat columns already logged, which omit every identity VM's own heat) and fleet_conserved (vm_physics_conserved() as 0/1). Requested explicitly after the first heartbeat-telemetry analysis pass (analysis-20260912/) omitted K entirely. Kernel rebuilt on all three architectures, full 3x9x3 campaign rerun (results-20260912-with-k/). K = 1.0000000000 (Q48.16 raw 65536) on every one of 775 heartbeat-tick observations, sd(K) = 0, 100% fleet_conserved, across amd64/aarch64/riscv64, nine identities, three replicates -- zero deviation. Also a free regression check on both recent Stadium fixes (§XVI/§XVII): neither disturbed the reservoir-transfer accounting K depends on. Found and fixed a tooling wrinkle along the way: fleet_conserved, being the CSV row's very last field with nothing after it to bound a regex match, can have a resumed trial digit merge into it with zero separator on the wire -- combine.py now derives it from fleet_k_q48 directly (same epsilon vm_physics_conserved() uses) instead of trusting the raw field. fleet_k_q48 itself is unaffected either way. Full analysis, discussion, and light/dark SVG->PDF figures written up as a proper LaTeX report (report-20260912/report/std79_doe_report.pdf), following experiments/bare_metal/analysis/report/bare_metal_doe_report.tex's established style -- supersedes analysis-20260912/'s markdown-only first pass as the primary deliverable for this dataset (kept, not discarded). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo |
||
|
|
e51a8d229e |
Fix stadium_grant_quota() donor floor; rerun std79 DoE clean, 81/81 (FABRIC-3.md §XVII)
capsule_birth.c hardcoded every new VM's initial Stadium quota grant to split from Hera specifically. Since a grant always halves whatever the donor currently has, Hera's own free list converges toward empty after a bounded number of grants — independent of whether the Stadium as a whole still had spare capacity, since VMs she'd granted to earlier typically still held nearly all of their own share untouched. Past that point every subsequent VM birth's Stadium grant would be silently refused (soft-failed, non-fatal by existing design), even with plenty of capacity sitting idle elsewhere. Fixed by adding an O(1)-maintained free_count to StadiumVMQuota (incremented in stadium_evict(), decremented at both of stadium_admit()'s free-list-pop sites, set/adjusted in stadium_grant_quota()'s own split — this also let grant_quota drop its old O(free-list length) counting walk in favor of an O(1) read) and stadium_best_donor(), an O(live VM count) scan over quota slots returning whichever in-use VM currently has the most free cells. capsule_birth.c's birth path now splits from that VM instead of unconditionally vm_uuid_hera(). Verified with another full rerun of the 3x9x3 std79 DoE campaign from scratch — same discipline as the prior Stadium fix (any defect repair reruns the whole DoE from the top) — one continuous boot per architecture, all 9 identities simultaneously live throughout. 81/81 trials correct, 0 mismatches, DOE-RUN header sequence md5-identical to every prior run. aarch64 ~280s total (vs ~290s for the O(ncells)-scan fix alone — confirms no regression). Both known Stadium defects are now closed together on one clean campaign rerun. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo |
||
|
|
9eff122090 |
Fix aarch64 Stadium/COOL O(ncells) scan; rerun std79 DoE clean, 81/81 (FABRIC-3.md §XVI)
Root cause of the 90+ minute aarch64 VM-birth stall found in §XV: stadium_admit()'s eviction-fallback scan iterated the entire stadium_ncells array filtered by owner, not the calling VM's own resident cells as its own doc comment claimed. Combined with stadium_grant_quota() always splitting from Hera's shrinking free list and stadium_word_dispatch() calling stadium_admit() per distinct word a VM's capsule executes, this compounded into a real O(n) blowup — catastrophic specifically on aarch64 because its -m 4096 (vs 1024 on amd64/riscv64) inflates the kmalloc heap kmalloc_init() bisects down to, which inflates stadium_ncells 4x (335,544 vs 83,886 cells, measured from boot logs). Fixed by threading a real per-VM doubly-linked resident-cell list (StadiumVMQuota.resident_head + stadium_resident_next[]/stadium_resident_prev[]) so the fallback scan is bounded by that VM's own resident count, not the global cell array size. Verified with a full rerun of the 3x9x3 std79 DoE campaign from scratch: one continuous boot per architecture, all 9 identities simultaneously live throughout (the 3-boot aarch64/riscv64 batching workaround is no longer needed). 81/81 trials correct, 0 mismatches, DOE-RUN header sequence md5-identical across all three raw logs. Identity 04's attach on aarch64, which stalled 90+ minutes before, now completes in ~34s; full boot-to-DoE-complete in ~290s. Corrects an earlier misreading (carried into §XV, std79-doe.fth's comments, and the project memory note) that described the symptom as a runaway "335,000+ cycles" dispatch counter — those were cell array indices, not an event count. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo |
||
|
|
70db955ac9 |
Fix M/MOD: hand-rolled 128/64 bit-serial division (FABRIC-3.md §XII.4)
mixed_math_word_m_slash_mod() had the identical bug class already fixed
in M* (commit
|
||
|
|
9a09949c69 |
Fix M*: use __int128 for a genuine 128-bit double-cell product (FABRIC-3.md §XII.4)
mixed_math_word_m_star() confused "double" (two full cell_t-width cells,
128 bits total on this 64-bit build -- what D+/D-/D./etc. all actually
expect) with "the low/high 32-bit halves of a single 64-bit product" --
code clearly written assuming cell_t is 32-bit. It computed an ordinary
64-bit `long long` product (already wrong for any true product exceeding
64 bits, since long long is the same width as one cell here) and split
that into 32-bit halves via `result & 0xFFFFFFFF` / `result >> 32`. For a
small negative product like -56088, this produced a positive, zero-
extended low cell paired with a correctly-looking dhigh=-1 -- D.'s
overflow check (correctly) rejected the resulting malformed double, on
every architecture, every time (this bug was never architecture-specific,
unlike the D+/D-/DNEGATE/d_compare family already fixed in bea8d74/
|
||
|
|
1a716c8048 |
Fix d_compare: use ucell_t instead of unsigned long (FABRIC-3.md §XII.4)
d_compare() (double_words.c), the static helper backing DMAX/DMIN/D</D=,
had the same bare-unsigned-long pattern already fixed in D+/D-/DNEGATE
(commit
|
||
|
|
bea8d7436a |
Fix D+/D-/DNEGATE: use ucell_t instead of unsigned long (FABRIC-3.md §XII.4)
double_word_d_plus(), double_word_d_minus(), and double_word_dnegate() (double_words.c) all cast through plain `unsigned long` for their carry/ borrow-detection arithmetic. On this aarch64 bare-metal cross-compile target, unsigned long is 32-bit (confirmed: sizeof(unsigned long)==4) -- amd64 and riscv64 both happen to have a 64-bit long, so the identical code only broke on aarch64. The low-cell arithmetic silently truncated to 32 bits, then widened back to cell_t via ordinary (non-sign-extending) conversion, producing a wrong result whenever the true 64-bit result was negative -- D. then correctly, faithfully reported DOUBLE-OVERFLOW on the resulting malformed double. vm.h already defines ucell_t for exactly this: same conditional as cell_t, guaranteed width-matched on every target. print_number_formatted() (format_words.c) already used it correctly; these three words didn't. Switched all three to ucell_t -- a one-word-class fix, no logic change. Verified: rebuilt and booted all three architectures clean. T19 (D+) on aarch64 now correctly prints -2, matching amd64/riscv64; T20 (DNEGATE) unaffected everywhere. Additional manual cases beyond the original exerciser, run live on aarch64 to specifically exercise the >32-bit-magnitude path the old bug depended on: D- (-5-3=-8), DNEGATE on 2^33 (8589934592 -> -8589934592), D+ crossing the same boundary (3+8589934592=8589934595) -- all correct. d_compare() (backing DMAX/DMIN/D</D=) has the identical latent pattern but is out of scope for this fix (not named in the request, never exercised by the campaign) -- left open, flagged in FABRIC-3.md. M*'s separate, universal-across-all-three-architectures DOUBLE-OVERFLOW bug is also untouched -- unrelated defect, not part of this fix. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo |
||
|
|
9142dda2d6 |
Fix use-after-free in sk_repl_idle()'s idle-tick VM resolution (FABRIC-3.md §XIII)
Root-caused a heap-corruption bug that reliably failed WIREBIND identity attach on the third attach/detach cycle in one boot. sk_console_readline() and sk_console_getkey() captured `active_vm` once from their caller and kept passing that same (possibly long-stale) pointer to sk_repl_idle() on every idle tick serviced while blocked waiting for input. If the VM it pointed at was killed (WIREBIND detach) mid-block, the existing bailout only checked a generic "is anyone attached" boolean -- masked as soon as a different identity attached next -- so blk_vm_flush_all() kept writing into a freed VM struct sitting on kmalloc's own free list, corrupting the free list's linked-list metadata itself. Both idle branches now re-resolve the live active VM fresh from g_repl_active_vm on every tick, matching the dispatch-side fix already made for the sibling bug in §XII.3. Verified: rebuilt amd64, reran the exact three-cycle repro that reliably corrupted the heap before the fix -- free-list census stayed stable through the same idle window that previously collapsed to zero. All three architectures (amd64/aarch64/riscv64) boot clean to the zuse)ok> prompt. kmalloc_debug_census()/kmalloc_debug_census_bytes() kept as permanent diagnostic infrastructure; every other temporary probe added during the investigation was reverted. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo |
||
|
|
70421bdd43 |
Fix real WIREBIND crash: stale active-VM pointer dispatched after blocking read (FABRIC-3.md §XII.3)
The interpreter_enabled guard added in the previous commit (
|
||
|
|
662ef44e59 |
Fix EXEC/LOAD block-persistence gap and WIREBIND/USE interpreter-race panic (FABRIC-3.md §XII)
Found live while building a cross-ISA FORTH-79 dictionary exerciser: - capsule_exec_init() zeroed a capsule's block content immediately after running it, so LOAD (a genuine FORTH-79 standard word, ACL-allowed even for locked identities) could never actually read back what EXEC had just written. Removed the clear from capsule_exec_init(); block content now persists like any other Standard BLOCK/BUFFER/UPDATE write. kernel_main.c's own explicit post-birth clear of Mama's init.4th range is untouched. - USE could redirect the console to a WIREBIND identity's VM before that VM's vm_enable_interpreter() step of its own birth sequence had run, causing the next typed line to hit vm_assert_interpreter_enabled() and panic the entire machine -- not the per-session-recoverable ACL-fault path a redirected VM otherwise gets. USE now checks interpreter_enabled first and refuses with a retry message instead. Also includes the amd64/aarch64/riscv64 acceptance boot logs and DoE CSVs from this session. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo |
||
|
|
b301317902 |
xHCI: drive Port Reset on port reuse; WIREBIND: kill the console VM too (FABRIC-3.md §XI.5)
Two independent bugs that together caused a reliable hotplug wedge: reusing an xHCI port for a second identity right after an unclean detach of a first would leave no further hotplug events reaching the guest at all. Bug 1 (xhci.c/xhci_driver.h): the xHCI driver never drove PORTSC.PR -- a known, named gap since Milestone 2e (the code's own comment flagged it, PORTSC_PR/PRC were defined but never referenced). A port's first connect each boot reads PED already set, so skipping the reset happened to work; a second device on the same port after a prior disconnect reads PED clear, and Address Device reliably failed without an explicit reset cycle. New XHCI_CONN_AWAIT_PORT_RESET state drives PR and waits for PED to read set before proceeding to Enable Slot. Bug 2 (capsule_wirebind.c): capsule_wirebind_eject()/unclean_detach() compared g_repl_active_vm against the *user* VM's pointer (g_wirebind_attached_vm_id tracks that one, not the console VM) -- never equal, since USE/g_repl_active_vm always points at the console VM. The guard never fired and the console VM was never killed at all, only orphaned -- paired to a dead user VM but still the REPL's active session. New wirebind_teardown_console() helper resolves and tears down the console VM by its own tracked bare username. Verified live on amd64: the exact repro (identity 01 on port 2, unclean detach, identity 02 on the same port immediately after) -- previously wedged with "xhci: address device failed" and no further hotplug activity; now attaches cleanly and fast, both VMs' KILL messages appear, console is immediately interactive on the new identity. Three-arch clean qemu acceptance passed, all clean on the first attempt. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014Ec88YKxxhZGG1RNnune78 |
||
|
|
d6661b5eed |
Scope VM fault halt to the faulting identity's own session (FABRIC-3.md §XI.4)
A standalone WIREBIND identity (no Zuse, USE'd in directly) hitting an ACL-denied word halted the entire machine -- Hera, Hermes, Artemis, all of it -- instead of just that identity's own session. sk_fault_handler() was being called unconditionally on whichever VM's ->error was set, with no distinction between Hera's own root session (where "no fallthrough surface" is the correct, deliberate fail-closed behavior) and a USE'd-in guest identity (which should recover and resume at its own prompt instead of taking the fleet down with it). Both call sites (sk_repl_step, sk_repl_run) now compare the faulting VM against Hera before deciding: Hera's own session still halts by design; any other VM prints a recovery message, clears its fault state, and continues. Also: mint identities 01-06 with the same FORTH-79/83 restricted personality identity 00 already had, verified via the fixed fault scoping above (which this verification pass surfaced). Verified live on amd64 (both the Hera-halts and identity-recovers branches); three-arch clean qemu acceptance passed (riscv64's first attempt hit an unrelated virtio_blk I/O timeout hang, a known QEMU/TCG flake -- a clean retry booted normally). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014Ec88YKxxhZGG1RNnune78 |
||
|
|
1839a2b0c3 |
WIREBIND cert verification: load Zuse's root pubkey independently of her live session
Root cause of the remaining "identity attach doesn't complete when Zuse never attaches this boot" issue: capsule_wirebind_verify_cert() gated on mama_vm->zuse_cert_installed, which is only ever set when Zuse's own drive attaches and authenticates this specific boot (capsule_zuse_boot_try_attach() -> install_and_activate() -> vm_zuse_cert_install()). Without her, any other identity's WIREBIND cert verification silently refused -- correctly, by the old design, but that design conflated two genuinely different things: "can mint new identities" (needs Zuse's live private seed, a real privileged operation) and "can verify an existing identity's cert" (needs nothing but her already-public key). That public key was already being persisted independently of her live session: zuse_genesis_marker_t (zuse_genesis_marker.h) stores it in the kernel's own top-of-device metadata fence (Artemis's resident storage), written once at genesis, specifically *not* alongside her private seed (which stays only on her own removable thumbdrive) -- the type's own doc comment says as much. It just wasn't being loaded for anything but confirming which drive is genuinely hers. Fix: a new capsule_zuse_boot_load_root_pubkey() (capsule_zuse_boot.c) reads that marker and populates two new VM fields, zuse_root_pubkey_known / zuse_root_pubkey (vm.h) -- deliberately separate from zuse_cert_installed/zuse_cert_seed/zuse_cert_pubkey, which stay untouched and still gate MINT exactly as before. Called once from kernel_main.c as soon as Artemis's own storage attaches, unconditionally, independent of whether Zuse's own drive is ever attached this boot. capsule_wirebind_verify_cert()/capsule_wirebind_try_attach() now check zuse_root_pubkey_known instead of zuse_cert_installed. One identity's attach must not depend on another identity's live presence -- each identity stands on its own once the fleet's root of trust has been established once, ever. Verified live, amd64: identity 00 (disk/thumbdrives/00-thumb-ident.img) now attaches and completes WIREBIND in 19 seconds with Zuse's own drive never attached this boot at all (previously: unbounded, many real minutes or effectively never, before today's other fixes; still slow/ stuck after those, stuck specifically on this silent refusal). Zuse's own attach flow re-verified unaffected (regression check, amd64). Three-arch clean qemu acceptance (amd64/aarch64/riscv64) passed with this change included. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014Ec88YKxxhZGG1RNnune78 |
||
|
|
1a263555e2 |
Fix pathological migration scan that stalled WIREBIND identity attach
Root cause (found by a fresh subagent after an extended live-debugging investigation into "identity 00 attaches slowly/stalls when Zuse never attached first this boot"): blk_migration_idle_check() was generalized earlier today to walk every attached device slot uniformly instead of hardcoding first_disk_slot() (Artemis's own disk). But its per-slot scan can only early-exit once it finds a devblock that is BOTH "hot" (claimed and worn) AND "free" -- and a just-attached, never-claimed USB identity drive can never satisfy the "hot" half by design (claiming only ever happens via blk_firsttouch_claim(), which only ever targets first_disk_slot()). So the scan ran to completion -- the drive's entire ~16,000 devblocks, mostly cache misses over slow emulated USB/BOT -- every single idle tick, forever, blocking sk_repl_idle() (and therefore the console and the storage-attach message round-trip) each time. Fix, in src/block_subsystem.c: a new has_ever_claimed flag on blk_dev_slot_t (set in blk_set_meta(), the single choke point every BLK_FLAG_CLAIMED transition passes through) skips the scan entirely, O(1), for any slot nothing has ever claimed -- the common case for a freshly-attached drive. A new migration_scan_lbn resume cursor bounds *any* slot's per-tick cost to MIGRATION_SCAN_BUDGET (256) devblocks examined, picking up where the previous tick left off instead of restarting from start_lbn every time -- restores this function's own documented "coarse cadence, cheap early-exit" design intent for every device, not just the one it used to hardcode. Also along the way (kept, all real improvements, verified live): - src/starkernel/usb/xhci.c: xhci_wait_bit()/xhci_bot_wait_for_idle() had zero yield hints in their MMIO-polling loops; added arch_relax() to both (matches virtio_blk.c below) -- a tight loop of nothing but MMIO reads can starve TCG's own host-side timer injection under QEMU. - src/starkernel/virtio/virtio_blk.c: vblk_io()'s spin bound was 33M iterations with zero logging on timeout; a single real (still not fully root-caused) timeout cost 31+ minutes of CPU before this was caught. Reduced to 1M and added a log line naming the failing sector, turning a silent, effectively-unbounded stall into a fast, loud failure -- callers already tolerate BLKIO_EIO. - src/starkernel/repl.c: blk_migration_idle_check() deferred for any idle tick where a storage-attach message round-trip is still pending, to keep the two block-subsystem-touching paths from interleaving; the existing MSG-TICK pump now checks the target VM's own dictionary for MSG-TICK (vm_find_word + acl_allow) before VM-EXECing into it, instead of spamming "VM-EXEC: ERROR" every tick for a VM that doesn't have -- or isn't allowed -- the word (e.g. a FORTH-79/83-locked-down identity); fixed a real -Wunused-variable build break in the EMERGENCY_CONSOLE_ ENABLED=1 path (unused today, but needed live to reproduce this bug with no identity attached at all). - src/starkernel/capsule/capsule_mint.c: dropped the dead S" common:messaging.4th" EXEC / MSG-CD-INIT lines from MINT_RESTRICTED_PERSONALITY -- a FORTH-79/83-locked-down identity has no legitimate use for a messaging vocabulary it can never call. Status: the pathological CPU-climbing scan is confirmed fixed (verified live: CPU stays flat across an extended run instead of climbing without bound). The WIREBIND storage-attach message round-trip still does not complete promptly in the "Zuse never attached, other identity attaches first" scenario -- a separate, still-open issue in the message-delivery path itself, not the scan. Tracked as follow-on work. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014Ec88YKxxhZGG1RNnune78 |
||
|
|
63b8b3bc29 |
Migrate Hera->Artemis storage-attach to a real message round-trip
Hera still polls xHCI and sig-checks attached drives, but the storage
registration step (blk_subsys_attach_device(), now wrapped as the
BLK-ATTACH primitive) moves to Artemis's own dictionary, reached via
HERA-BLK-ATTACH-REQ/BLK-ATTACH-ACK (VM-EXEC, since Hera can't load her
own messaging.4th -- see the doc comment in repl.c). Identity birth
(Zuse genesis / WIREBIND) is deferred until the ack confirms storage
actually succeeded, instead of running synchronously underneath a
storage call that might fail ("wait for ack, safer for identity data").
Caught and fixed a real bug live during acceptance testing: Artemis's
ACK-APPEND-NUM fed a single-cell value into <# #S #> (which expects a
double-cell pair), causing a stack underflow the first time
HERA-BLK-ATTACH-REQ ran. Fixed with the same `0 SWAP` convention every
other numeric-append helper in this codebase already uses.
Verified booting clean to (zuse) ok> with no VM-EXEC errors on all
three architectures (amd64/aarch64/riscv64).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014Ec88YKxxhZGG1RNnune78
|
||
|
|
f4ded3e1a8 |
MINT: add a FORTH-79/83-standard-words-only lockdown personality
Captain Bob, 2026-09-07: "starting with that 00 user we created, we're
going to give access only to FORTH 79 and 83 standard words. everything
else is locked down."
New capsules/acl-std79.4th (blocks 4023-4047): walks a VM's own
dictionary (>LINK/LINK> traversal, same as ACL-INIT-PRIMITIVES/WORDS
already use) and permanently denies+pins every word not on an explicit
FORTH-79/83 allowlist, extracted from the real registered word set
(stack_words.c through control_words.c), not recited from memory.
Deliberately excludes, beyond plain non-standard words: BYE (100% ACL
bypass to the emergency console -- "needs more discussion, exclude for
now"), COLD/WARM/REBOOT/SAVE-SYSTEM (system lifecycle), the block/screen
editor L/S/SHOW/EDIT/UPDATE/SAVE-BUFFERS (lets a session rewrite
persistent block/capsule content, defeating the lockdown even though
nominally standard), BLK-ACL-*/BLK-OWNER@ (StarForth-specific), and
FORGET/FENCE (flagged as an unrestricted superpower word, 2026-09-03
audit). Keeps WORDS/VLIST/SEE (introspection only -- ACL is enforced
per-target-word at execution time regardless of how an XT was
obtained) and the parenthesized control-flow runtime primitives
((BRANCH) etc. -- IF/DO/LOOP compile calls to these; denying them
breaks ordinary control flow, not security).
MintPersonality enum (capsule_mint.h) lets capsule_mint_identity()
select which personality-source template gets written to a new
identity's devblock -- MINT_PERSONALITY_DEFAULT (unchanged) or
MINT_PERSONALITY_STD79_LOCKDOWN (EXECs acl-std79.4th then
ACL-LOCKDOWN-STD79 as the VM's own last bootstrap step). The actual
restriction logic stays entirely in FORTH per .claude/CLAUDE.md's
Word-Level ACL System rules ("ACL policy belongs in ACL.4th, never in
C") -- capsule_mint.c only picks which few-line bootstrap stub to
write. MINT's own stack signature gains a trailing restrict? flag;
capsule_zuse_boot.c's genesis mint (Zuse herself) explicitly passes
MINT_PERSONALITY_DEFAULT -- the superuser is never restricted.
Two real bugs found and fixed live during testing, both the same class
of self-referential fault: ACL-LOCKDOWN-STD79's own walk loop calls
ACL-STD79-ALLOWED?/ACL-STD79-LIST/ACL-ALLOW!/ACL-PIN on every single
iteration to do its job -- none of those are FORTH-79/83 standard
words, so the walk was denying its own load-bearing infrastructure
partway through and then faulting the next time it tried to call it
("VM fault -- emergency console disabled; halting", reproduced twice
live). Fixed by explicitly protecting all four in the allowlist
(block 4047) -- they must stay allowed for the walk to finish, not
because they belong on a "standard words" list.
Verified live end-to-end: minted a throwaway test identity with the
restrict? flag, confirmed her WIREBIND birth completes cleanly (no
faults, no shadow conflicts) on a single real attach, then USE'd into
her VM and confirmed standard arithmetic and user-defined words work
(1 2 + . -> 3; : X 5 5 * . ; X -> 25) while KILL is entirely unknown to
her dictionary and VM-EXEC is denied. One real, non-fatal side effect
found and left as-is (not asked to fix): the fleet's inter-VM messaging
pump (MSG-ARENA) is also denied by the lockdown, logging a harmless
per-idle-tick warning -- a fully locked-down VM doesn't participate in
message routing.
Not yet applied to the real identity 00 -- this commit is the
mechanism, verified against a disposable test identity only.
Three-arch clean qemu acceptance (single Zuse device, standard
regression case) passed on amd64, aarch64, and riscv64.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014Ec88YKxxhZGG1RNnune78
|
||
|
|
56e19a00f9 |
Route VM-dictionary sf_malloc/sf_free through the kernel's kmalloc heap
Captain Bob's call after seeing the identity-heap-capacity findings (FABRIC-3.md §X.4): the number of concurrently-running VMs is not known in advance, and once this is a complete operating system the heap should be able to use whatever memory is actually available -- not a hardcoded compile-time ceiling. This was already half-built and just not wired up. src/starkernel/vm/alloc_kernel.c previously implemented sf_malloc()/ sf_free() (platform_alloc.h's allocator abstraction -- what vm_create_word() calls for every VM's word dictionary) as its own isolated static 4MB arena: first-fit free list, no splitting or coalescing. That's exactly the allocator that topped out around 6 concurrent WIREBIND-born identities, failing from fragmentation before true capacity exhaustion (§X.4's own measurements). Sitting right next to it, unused for this purpose: src/starkernel/memory/ kmalloc.c, the kernel's general heap. Already initialized at boot (M6, kernel_main.c, well before any VM is ever born), reserved from real PMM-tracked physical memory rather than a fixed array, defaults to a 2 GiB floor explicitly sized "for 256+ baby VMs" per its own comment, overridable via the --heap= boot flag, and its free list actually coalesces neighboring blocks on every free. Change: alloc_kernel.c's sf_malloc()/sf_free() now delegate to kmalloc_aligned()/kfree() instead of managing a separate arena. sf_alloc_init() becomes a no-op (kmalloc is already initialized by the time any VM allocation can happen, and "resetting" a heap now shared by every kernel subsystem would be actively wrong -- confirmed no external caller depended on its old reset semantics). sf_alloc_get_stats() reads kmalloc_get_stats() fresh rather than shadowing byte counts locally; alloc_count/free_count (which kmalloc.c doesn't track) stay as simple local counters. sf_calloc()/sf_realloc() are otherwise unchanged. Kernel- only: the hosted (non-kernel) StarForth build keeps its own separate alloc_host.c implementation, untouched. Verified live: replaying the exact hotplug sequence that previously topped out at 6 identities (Zuse + 8 identities, one at a time via QMP device_add) now succeeds for all 9, where identity 05 specifically used to fail. Three-arch clean qemu acceptance (single Zuse device, the standard regression case) passed on amd64, aarch64, and riscv64 -- one aarch64 attempt hit an unrelated, already-documented one-off QEMU hiccup (empty log, boot never progressed past firmware) and passed cleanly on retry with no rebuild. Not addressed here: the underlying free-list itself is still first-fit without splitting (only coalescing changed, inherited from kmalloc.c); per-VM dictionary sizing (shrinking what each WIREBIND VM's word set actually needs) is a separate, still-open lever from FABRIC-3.md §X.4's open architecture question. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014Ec88YKxxhZGG1RNnune78 |
||
|
|
d8a195b8d8 |
WIREBIND: kill the orphaned console VM when the user VM birth fails
Found live during identity-heap-capacity testing (2026-09-07, hotplugging Zuse + 8 identities one at a time and measuring the kernel heap arena via a temporary allocator-stats probe, since reverted): every WIREBIND identity attach births two VMs in sequence -- a "console" VM, then the real "user" VM. When the second birth failed (arena fragmentation under concurrent VM load, a separate, not-yet-fixed capacity issue), capsule_wirebind_try_attach() logged the failure and returned, but the console VM that had *already succeeded* was never torn down. It stays live and registered under the identity's username, consuming its own ~228KB of the fixed 4MB kernel heap arena forever -- nothing ever points a real user at it, since WIREBIND only ever hands the caller the user VM's id. This turns every failed identity attach into a permanent net loss of heap rather than a neutral retry: confirmed live that a failed attach left the arena 228,576 bytes worse off than before the attempt, and every subsequent attempt starts from that worse baseline, compounding. Fix: call capsule_vm_kill(username) on the now-orphaned console VM before returning from the failure path -- the same teardown capsule_wirebind_eject()/capsule_wirebind_unclean_detach() already use elsewhere in this file (vm_cleanup() + sf_free(), confirmed live to actually reclaim per-word dictionary allocations, FABRIC-3.md §IX.2/§IX.3). Verified live with the same allocator-stats probe (written, captured, reverted -- not part of this commit): after the fix, a forced user-VM birth failure now returns the arena to exactly its pre-attempt byte count (3,526,256, matching the baseline precisely) instead of leaking 228,576 bytes. Three-arch clean qemu acceptance (single Zuse device, the standard regression case) passed on amd64, aarch64, and riscv64. The underlying capacity/fragmentation question (why the 7th concurrent identity's arena allocation fails at all despite technically-sufficient free bytes) is a separate, open architecture question -- not addressed here. See project memory for the full measured numbers. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014Ec88YKxxhZGG1RNnune78 |