5417ffb2bf381e1e328f0b7df76e6402dc53ef75
223
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
5417ffb2bf |
Stage D batch 1: fix silent error sites in repl.c and inference_words.c (FABRIC-3.md §XXXII.6)
repl.c's sk_word_blk_attach_ack() and inference_words.c's array_ptr() helper + infer_word_run()'s allocation guard now log a diagnostic via log_message(LOG_ERROR, ...) before setting vm->error, matching defer_words.c's own gold-standard pattern (§XXXII.3) -- these are internal/background conditions (a malformed message-callback, a bad array reference or allocation failure), not interactive usage mistakes, so the fix keeps the fault and reports it rather than dropping it like USE's own fix did. Batched together (4 sites total, smaller combined than the next file) rather than two separate acceptance cycles for negligible size. Three-arch clean qemu acceptance passed. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K |
||
|
|
c05b70c8d6 |
Stage B: logging policy documented, level-aware log-ring eviction built, LOG-FLUSH deferred again (FABRIC-3.md §XXXII.4)
Policy decided for the kernel-only audit scope: an interactive command's direct response stays on console_println/console_puts; everything else (state transitions, background diagnostics, audit trails) routes through log_message() at the appropriate level, matching capsule_mint.c's verify_mint() precedent. Documented, not code-swept here -- reclassifying individual sites is Stage D's job. LOG-FLUSH deferred again, explicitly: the per-VM log buffer its own doc comment presumes (vm_log_buffer.h) doesn't exist anywhere in the tree -- building it is real feature work needing its own scoped stage. Level-aware eviction built: log_region_append() now reads the oldest ring slot's own level before evicting it, protecting ERROR/WARN records from being pushed out by INFO/DEBUG churn -- drops the incoming low-priority record instead. Found and fixed an adjacent bug while making this change: the prior two-valued return contract would have made a benign "dropped by design" outcome indistinguishable from a genuine write failure to its one caller, which unconditionally set vm->error on any nonzero return. Changed to a three-valued contract (0 success, 1 dropped by design, -1 genuine failure). Three-arch clean qemu acceptance passed. Eviction path itself not live-exercised (needs 128+ LOG-APPEND calls to fill the ring) -- flagged, matching this project's own precedent for that kind of gap. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K |
||
|
|
5c5896fbc1 |
Stage A: fix USE's silent stack-underflow/bad-address guards, and the Hera fault-scoping gap they exposed (FABRIC-3.md §XXXII.1)
mama_word_use()'s two silent vm->error=1 guards (dsp<1 stack underflow, NULL from vm_ptr()) now print a diagnostic and return, matching the function's other five guards. Live verification of that fix alone surfaced a bigger problem: with the guard no longer silent, the REPL proceeds to interpret the leftover token as an unrecognized word, which independently sets vm->error, and sk_repl_step()/sk_repl_run()'s Hera-branch still hard-halted on that. Investigated kernel_main.c's boot/capsule-load paths directly: they already catch and clear mama->error entirely separately, before sk_repl_run() is ever entered -- so the "no fallthrough surface" halt in these two REPL functions was never protecting a boot-time fault, only an ordinary interactive REPL-turn one. Both functions now recover unconditionally on any VM's error, Hera included, matching how a redirected (WIREBIND/USE'd) identity's session already recovered. Removed the now-fully-unused sk_fault_handler(). Verified live, amd64: USE rajames at Zuse's own console prints the new diagnostic, then "VM fault -- session recovered, resuming", console stays interactive afterward. Three-arch clean qemu acceptance passed before and after the Hera-fault-scoping change. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K |
||
|
|
d9da82b065 |
Stage 4 increments 2+3: WIREBIND VMs as switch-signal participants + mark-and-defer tombstone reap (FABRIC-3.md §XXVIII Stage 4)
Increment 2: WIREBIND user VMs (the ones that actually run FORTH work; console VMs are pure REPL proxies and never participate) register as Stage 3 switch-signal participants at attach, unregister at teardown. Slot table bumped 8 -> 16, matching messaging.4th's own VM-MAX -- a real, already-agreed ceiling, not an invented number. Added sk_vm_switch_signal_unregister() (compaction-based; Tripod VMs never needed removal, WIREBIND VMs cycle constantly and would otherwise exhaust the bounded table). Increment 3: implements the plan's own ratified option (A) for the async-detach UAF risk -- mark-and-defer via a new pending_reap flag on VMRegistryEntry, deliberately not a new VMState (capsule_vm_kill() already treats VM_STATE_DEAD as idempotent success, which would silently swallow a reap attempt; SWITCHED_OUT still accurately describes a tombstoned VM until the moment it's actually freed). unclean_detach() sets it when capsule_vm_kill() refuses a SWITCHED_OUT target; the Stage 3 checkpoint (vm_core.c) checks it before ever attempting to resume a pending switch target, and calls the new capsule_vm_force_reap() instead -- the one caller allowed to bypass capsule_vm_kill()'s own refusal, because it runs at the exact safe cooperative point the switcher itself controls. A new idle-tick sweep cleans up the WIREBIND live-table entry once the reap has actually happened. Verified clean on all 3 architectures (baseline regression -- no WIREBIND attach happens in a plain boot). The reap mechanism's own correctness under a genuinely parked context is verified separately, next, via a temporary deterministic probe. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K |
||
|
|
9f0f33dfc5 |
Stage 4 increment 1: per-device WIREBIND tracking, fixing a real multi-identity detach leak (FABRIC-3.md §XXVIII Stage 4)
capsule_wirebind_unclean_detach()/eject() tracked "the attached identity" as a single global, correct for the console-pairing UX (one physical console) but wrong for detach safety: since §XV/§XVI proved multiple identities genuinely live simultaneously via this same attach path, every attach after the first silently overwrote the singleton, so an unclean detach of any but the most-recently-attached identity was silently ignored -- that VM leaked forever, no trace in the log. Adds a per-device live-identity table, separate from the (unchanged) console-pairing singleton, so unclean-detach resolves any attached device to its own identity. Sized off messaging.4th's own VM-MAX (16) minus Tripod's 3 reserved slots, not an invented number. Corrects the stale "single-USB-device constraint" doc claim in capsule_wirebind.h, false since §XV/§XVI. Groundwork for Stage 4's real deliverable (WIREBIND VMs as switch-signal participants) -- this increment only fixes detach targeting; switch- signal registration is next. Verified clean on all 3 architectures (no WIREBIND attach happens in a plain boot, so this is a regression check on the existing Tripod-only path; live multi-identity verification comes with the switch-signal registration increment). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K |
||
|
|
05ae7aa886 |
Fix SWITCH-MARK-WORK off-by-one: was tripping vm->error on every MSG-SEND (FABRIC-3.md §XXVIII.2)
mama_word_switch_mark_work()'s stack-underflow guard checked dsp < 2, requiring 3+ items, when it only ever needs the 2 IDX>NAME leaves it (caddr u). Since dsp is index-based (2 items == dsp 1), this rejected every normal call. MSG-SEND tail-calls SWITCH-MARK-WORK unconditionally, so this fired on every message sent anywhere in the system -- visible only where a caller happened to check the target VM's error flag afterward (mama_word_vm_exec()'s "VM-EXEC: ERROR in Artemis" report). Verified clean on all 3 architectures: boot reaches Startup: Artemis live -> zuse@Hera] ok> with no VM-EXEC: ERROR line at all. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K |
||
|
|
66beae7fd4 |
Stage 3 follow-on: message-arrival eligibility hook + trampoline-blind switch-storm fix (FABRIC-3.md §XXVIII.2)
Implements the message-arrival eligibility signal FABRIC-3.md §XXVIII.1 left open (has_work per-slot flag, set via new SWITCH-MARK-WORK primitive from MSG-SEND) so an idle VM never becomes a switch target purely by waiting out the readiness threshold. Also root-causes and fixes a second, independent switch-storm: the tick's "who is current" check used vm_log_attributed_vm(), which can't see a VM parked in switch.c's own raw trampoline. Replaced with a dedicated g_switch_current_vm tracked by the switch mechanism itself, and moved target-slot eligibility reset to the switch decision point instead of relying on ISR polling to observe a window that can be only a few instructions wide. Verified live on all 3 architectures: clean boot to zuse@Hera] ok>, live cross-VM message dispatch, and (since a quiet log looks identical to a livelocked storm once the DoE probe is gone) confirmed genuine REPL liveness via QMP send-key + screendump on aarch64/riscv64, not log inspection alone. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K |
||
|
|
862d7d9c48 |
Stage 3 follow-on: fix stack-ownership corruption + DoE switch columns (FABRIC-3.md §XXVIII.1)
DoE CSV gained 6 switch-signal columns (switch_count_cumulative, switch_current_slot, switch_*_readiness, switch_ticks_since), and verifying them with a boot-time HB-ON probe surfaced a real livelock: the preemption checkpoint could fire inside a VM-EXEC-nested execute_colon_word() call and switch away from a stack it didn't own, parking a borrowed region of the caller's stack under the wrong VM's saved-context pointer. The trampoline bounce was the visible (safe) half of this; the corruption was the quiet half, live in every prior "clean" Stage 3 boot without ever showing up in the log. Fixed by gating the checkpoint on being at the outermost vm_interpret() call (g_vm_interpret_depth / sk_vm_at_outermost_interpret(), vm_core.c), per Bob's decision. Also fixed two related bugs found in the same pass: g_switch_back_to was a single global stale after first entry, now per-VM state (native_switch_back_to); note_switch_performed() fired on resume instead of switch-out, now called before the switch. Verified on all 3 architectures: steady log growth (no freeze), zero leaked QEMU processes, DoE columns internally consistent, Hermes/Artemis confirmed genuinely executing (not just trampoline-bouncing). Temporary HB-ON boot probe reverted after capture. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016UNhH1mhi52i6Qihh7ZV5S |
||
|
|
986d042aa7 |
Stage 3: timer-driven preemptive switching, live on all 3 arches (FABRIC-3.md §XXVIII)
Fourth stage of the preemptive context-switching plan, and the biggest. LithosAnanke now genuinely, continuously preempts between Hera, Hermes, and Artemis -- timer-driven, running live for the entire remainder of every boot once the Tripod fleet registers, not a bounded probe. A real design fork was resolved before writing code: the naive approach (the timer ISR calling Stage 2's sk_vm_context_switch() directly) is broken -- Stage 0's trap frame lives on whatever stack was active at interrupt time, and jumping to a different stack via Stage 2's own independent swap mid-handler would abandon that trap frame unresumed, guaranteed corruption on the first tick. Chose the safer of two named options: the ISR only ever sets a flag and returns completely normally through its own full epilogue; the actual switch happens moments later, via Stage 2's already-proven mechanism, at a safe cooperative checkpoint on the mainline (execute_colon_word()'s per-word dispatch loop, checked on literally every word, not throttled to the existing 256-word heartbeat-tuning cadence) -- confirmed with the user that word-level granularity is fine-grained enough given the eventual Zynq FPGA target where a word is a mnemonic. New capsule_vm_switch_signal.c/.h: a purpose-built run-readiness signal, deliberately separate from capsule_vm_physics.c's execution-heat engine (that one's own header documents itself as never touched from interrupt context, by design). Slot table sized with headroom (8) rather than hardcoded to today's 3 participants, so extending participation later is another register() call, not a redesign -- per direct request to leave room for swapping the participant set. Simple linear accumulate-then- threshold for this first cut; a fancier law can replace it later without touching the mechanism around it. heartbeat_tick() gains its one deliberate, documented amendment to this file's own top-half/bottom-half discipline -- the first time this codebase reaches into VM-scheduling state from real ISR context. Registration happens only after all three VMs are fully born, right before the REPL starts -- no critical-section protection yet against being switched away mid-birth-setup. Known, flagged rough edge (not reconciled this pass): MSG-TICK's own idle-pump and this new mechanism can still independently move control between the same VMs; not observed to interact badly in verification, but not fully unified either. Verified interactively at the console on all 3 architectures with continuous background preemption running throughout -- amd64 computed `1 1 + .` -> 2, aarch64 computed `1 1 + dup DUP * . CR` -> 4, both correct, REPL fully responsive, zero fault indicators over sustained runtime. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016UNhH1mhi52i6Qihh7ZV5S |
||
|
|
f790d0995e |
Stage 2: cooperative VM context switch primitive, proven on all 3 arches (FABRIC-3.md §XXVIII)
Third stage of the preemptive context-switching plan. The real save/restore switch mechanism now exists -- the first time anything has ever executed on a VM's own native stack (Stage 1 allocated them, unused). New sk_vm_switch_to() (switch.S, one per arch) is an ordinary function call, not an interrupt -- so unlike Stage 0's trap frame, the ABI already covers every caller-saved register; only the callee-saved set needs explicit save/restore (amd64: rbx/rbp/r12-r15, no FP at all since SysV has no callee-saved XMM; aarch64: x19-x28/x29/x30 + d8-d15; riscv64: s0-s11/ra + fs0-fs11, FS-gated like Stage 0 but read once and reused for both halves within one call, since FS is genuine global CPU state, not part of what's switched). A sibling sk_vm_switch_prime() in the same file builds the synthetic first-entry frame, kept in assembly so the layout can never drift out of sync with sk_vm_switch_to() itself. New switch.c/switch.h: sk_vm_context_switch(from, to) handles first-entry priming vs. resuming a parked context, and updates registry state (new VM_STATE_SWITCHED_OUT, distinct from VM_STATE_STOPPED -- STOPPED means no live frame, this means the opposite). sk_vm_switch_entry() is the minimal permanent trampoline every freshly-entered VM lands in: no production behavior defined yet, so it just yields straight back to whoever switched to it, forever. Closes the confirmed unguarded-KILL UAF found during planning: capsule_vm_kill(), mama_word_kill(), and capsule_vm_kill_all_nonmama() all now refuse (or silently leak rather than free, on the cold-restart path where arch_cold_reset() wipes everything immediately after anyway) tearing down a switched-out VM. Side effect found, not built on purpose: the existing MSG-TICK idle-pump already filters on VM_STATE_LIVE, so it automatically stopped dispatching into a switched-out VM with zero changes needed there. Verified via a temporary SWITCH-TEST probe (boot-triggered, since nothing can type interactively into a foreground-only QEMU session) that round-tripped a sentinel through 5 real Hera<->Hermes switches on all 3 architectures: 5/5 rounds, 0 failures, clean continuation to ok>. Probe fully reverted after capture; kernel_main.c shows zero diff. Also: Makefile.starkernel's LOADER_EXTRA_SRCS/LOADER_ASM needed the new files added explicitly (this project's "loader" PE binary is the full running kernel, not a thin bootstrap stage), and aarch64's switch.S needed the same #ifndef _WIN32 guard around .hidden that isr.S already carries (aarch64's loader assembles via clang targeting a PE/COFF target with no .hidden equivalent) -- caught by a build failure, fixed. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016UNhH1mhi52i6Qihh7ZV5S |
||
|
|
57ac3fc304 |
Stage 1: per-VM native stacks, allocated but not yet executed on (FABRIC-3.md §XXVIII)
Second stage of the preemptive context-switching plan. Every VM (Hera, every capsule_birth_baby()-born VM including WIREBIND identities) now gets its own dedicated 2 MiB native C stack at birth -- but nothing runs on it yet, that's Stage 2. Pure allocation-machinery proof. Design correction made before writing code: the plan called for cloning sk_vm_arena_alloc()'s guard-page pattern, but that pattern turns out to be Mama-only -- host_services.c's kernel_alloc() gives every baby VM a plain kmalloc() block for its dictionary arena, not a real guarded PMM allocation. Stacks get the real treatment instead (new sk_vm_native_stack_alloc()/_free() in arena.c): independent pmm_alloc_contiguous() + guard pages for every VM without exception, no singleton, no kmalloc fallback -- a stack overflow is exactly the failure mode guard pages exist for, and a corrupted stack could corrupt whatever saved context Stage 2 trusts. 2 MiB size matches this project's own established kernel-stack convention (g_kernel_stack/g_rpi5_native_stack), not a guess -- that one shared 2 MiB stack today already carries all VMs' combined nested VM-EXEC recursion. Three new VM struct fields, freed in vm_cleanup() alongside the existing call_stack free. Allocation failure is non-fatal to birth. All 3 architectures re-verified clean boot to ok>, no native-stack allocation failures for any Tripod-fleet VM. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016UNhH1mhi52i6Qihh7ZV5S |
||
|
|
15672ce17c |
Stage 0: trap-frame parity across all 3 arches (FABRIC-3.md §XXVIII)
First stage of the preemptive context-switching plan (see ~/.claude/plans/logical-snuggling-bear.md). Pure foundation work -- every arch's ISR now saves the full register set on interrupt entry, so a trap frame is in principle sufficient to resume execution anywhere it was taken. No FORTH-visible behavior changes. amd64: added FXSAVE/FXRSTOR, closing a genuine pre-existing correctness gap (not just future-preemption prep) -- confirmed live double-precision FP code reachable from ordinary interpreter dispatch (vm_runtime.c Loop #5/#6), and the ISR previously saved zero FP/SSE state. rbp repurposed as a fixed anchor so the 16-byte-aligned FXSAVE area can be carved out of an unpredictably-aligned rsp without disturbing existing argument reads. aarch64: extended the trap frame 672->800 bytes, adding v8-v15 (AAPCS64 callee-saved, previously excluded on call-site-only reasoning that doesn't hold for an async trap). riscv64: extended the trap frame 320->512 bytes, adding s0-s11 and fs0-fs11 (the latter still correctly gated behind sstatus.FS != Off). All 3 architectures re-verified clean boot to ok> under the new frames -- amd64 through hundreds of timer ticks with FXSAVE/FXRSTOR live on every interrupt, aarch64 through 987 ticks, riscv64 clean on the now-larger FS-conditional block. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016UNhH1mhi52i6Qihh7ZV5S |
||
|
|
2a30212bd3 |
Real per-VM log persistence: source attribution + ACL pin (FABRIC-3.md §XXVII)
Wires the previously-unused vm_log_attributed_vm() into LOG-APPEND's kernel primitive so persisted log records carry a trustworthy source (the real attributed VM's registry name, or "HADES" pseudo-source) instead of a caller-supplied, trivially forgeable string. Drops src-addr/src-u from LOG-APPEND's stack signature accordingly. Pins LOG-APPEND via bare ACL-PIN in Artemis's own init.4th, matching BIRTH/CAPSULE-BIRTH's precedent for a privileged word that can't reach the shared, host-portable ACL.4th. Also fixes two console-banner nitpicks: a mis-rendering em dash (U+2014) in the boot banner, and drops "Emergency" from the CLI banner text. Doc corrections to artemis_sig.h/zuse_eligibility_list.h reconciling the three fixed devblock ranges now in play. LOG-FLUSH (the intended normal entry point) and level-aware log eviction remain open, flagged not fixed. Re-verified clean boot to ok> on all 3 architectures after every change. riscv64 showed one new, unrelated virtio_blk write-timeout anomaly during Artemis's early physics self-test (self-recovered, boot unaffected, sector doesn't map to the log region) -- flagged, not investigated. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016UNhH1mhi52i6Qihh7ZV5S |
||
|
|
61755fde78 |
Artemis genesis stamp: fix a BAM-corrupting offset before it ever ran (FABRIC-3.md §XXVI follow-on, Step 3)
Step 3: one-time artemis_sig_t genesis stamp, written once
kernel_main.c's virtio-blk path confirms Artemis's own disk, so the disk
image is later recognizable generically (repl.c's idle-loop USB-MSC scan,
built in the prior commit) regardless of which bus found it.
Correction made before this ever touched the real disk: the signature's
first design (committed in
|
||
|
|
29b6789860 |
Artemis bus-agnostic discovery: signature format + idle-loop generalization (FABRIC-3.md §XXVI follow-on)
Step 1: new artemis_sig_t header format (magic 'ARTM', sibling to homeblocks_sig_t, distinct so a generic scan can tell Artemis's own disk apart from an identity thumbdrive by content alone) -- artemis_sig.h/.c, wired into Makefile.starkernel. Step 2: sk_repl_idle()'s existing per-USB-MSC-slot attach handling (the pattern WIREBIND already uses for identity thumbdrives) now also checks for the ARTM signature whenever a device's home-blocks check comes back BLANK. On a match, once Artemis's own storage-attach round-trip (HERA-BLK-ATTACH-REQ/BLK-ATTACH-ACK) confirms success, capsule_zuse_boot_load_root_pubkey() runs -- the same call kernel_main.c's synchronous QEMU-only virtio-blk path already makes, now reachable without a hardcoded PCI vendor/device scan. That function is already idempotent (no-op once zuse_root_pubkey_known is set), so no boot restructuring was needed despite the initial concern that deferring Artemis discovery to the idle loop would require one. Verified: clean build + QEMU boot to [zuse@Hera] ok> on all three architectures, zero regression to the existing virtio-blk/Zuse-thumbdrive attach path. Steps 3 (genesis-stamping onto disk/artemis.img) and 4 (growable production log-persistence region) not yet started. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo |
||
|
|
a8b16d41da |
Unify console prompt to [user@VM]; fix real personality-block truncation; correct §XXV's wrong lockdown conclusion (FABRIC-3.md §XXVI)
Investigating the std79 lockdown finding from FABRIC-3.md §XXV led to a real discovery: WIREBIND births TWO VMs per identity, a console proxy under the plain username and the actual restricted identity under <username>~user (capsule_wirebind.c). Every test in §XXV targeted the console proxy, which was never locked down at all. Retested against the correct target (rajames~user): the lockdown works exactly as designed. §XXV's "lockdown never engages" conclusion was wrong -- corrected here, not deleted, since the mistake and how it was caught are worth keeping (see the new feedback memory: confirm which specific VM a name resolves to before concluding anything, when a subsystem is known to birth more than one VM per identity). Two real, separate things found along the way are kept regardless of that correction: - capsule_runcap.c: the reserved personality devblock was read in full (mostly zero-padding after a short ~200-byte string) with no terminator, producing "WARN: block 4998 exceeds 1KB, truncating" on every std79-locked identity's birth, universal, since at least 2026-09-10. Fixed by trimming to the first NUL byte actually found -- real, but harmless to execution (real content sat in the truncated block's surviving head); it mattered for capsule_id/ content_hash being computed over padding instead of real content. - console.h/console.c/repl.c: unified the prompt from a separately- computed "[VMName] (user)" into a single "[user@VMName]" line prefix -- exactly the ambiguity that caused the original misdiagnosis (the prompt showed only the WIREBIND username, identical whether USE had targeted the console proxy or the real ~user identity). Implemented as a registered callback (console_set_user_prefix_provider()) rather than console.c calling into WIREBIND/session logic directly, since console.c is a clean HAL module with no prior dependency on capsule-level subsystems. Verified: clean build on all 3 architectures, zero new warnings, identical dict_hash/capsule_hash to every prior boot this session (console/prompt-only change). Full 9-identity messaging campaign re-run end to end: 202s, zero faults, all 8 identities at 99/99 tokens, zero regression. Also surfaced, not yet acted on: the full campaign's own console tags now visibly show which VM each identity's tests actually reached ([zuse@rajames], not [zuse@rajames~user]) -- messaging.4th's VM-NAMES-INIT registers identities by plain username, so std79-doe. fth's turn-attractor has been dispatching to each identity's console proxy, not the actual locked-down identity, since the messaging rewrite. Flagged for a deliberate decision, not investigated further. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo |
||
|
|
3c2daf50d1 |
Extend BIRTH/CAPSULE-BIRTH to all VMs symmetrically; flag a real std79 lockdown gap (FABRIC-3.md §XXV)
Scoping the workload-into-factorial design's placement-mode factor led to a real architectural improvement: rather than EXEC-ing a workload capsule into an already-running, ACL-locked identity's own persistent dictionary (filesystem-shaped, doesn't dodge the block- collision exposure just traced in §XXIV), a workload now runs as a fresh ephemeral child VM, BIRTH'd per trial and reaped after -- matching the project's own stated principle of automanagement over imposed policy. CAPSULE-BIRTH already passes vm->stadium_vm_id (who is birthing this VM) as the new child's parent, not a hardcoded Hera constant, confirmed by reading the C -- so a workload trial genuinely inherits the specific identity's own lineage when that identity does the birthing. Which surfaced a real premise: only Hera could call BIRTH/ CAPSULE-BIRTH at all (registered only in register_mama_forth_words(), confirmed directly, not part of the earlier §XX messaging-symmetry fix which deliberately kept this as one of her remaining privileges). Extended symmetrically now, agreed explicitly before touching code: - mama_forth_words.c: BIRTH and CAPSULE-BIRTH added to register_child_vm_words(), matching §XX's own pattern. - acl-std79.4th: ' BIRTH , ' CAPSULE-BIRTH , added to ACL-STD79-LIST (new block 4048) -- a deliberate, explicit, named exception to the lockdown's own "standard words only" guarantee, not a silent one. Symmetric registration alone can't weaken any lockdown on its own: ACL-LOCKDOWN-STD79 is allowlist-based, deny-by-default, so a newly registered word is auto-denied there unless explicitly added. Verified: clean build on all 3 architectures, zero new warnings. Hera's own dict_hash unchanged (expected); Hermes/Artemis show the same new dict_hash on all 3 architectures. Live-tested against a real attached std79-locked identity: CAPSULE-BIRTH executes correctly (returns vm_uuid_none() for a deliberately out-of-range capsule-id, zero fault, zero ACL denial). Found, and explicitly stopped short of fixing, a separate pre- existing gap while verifying the above: MSG-STATUS and MSG-K (messaging.4th words, not on the std79 allowlist) execute for a locked identity instead of being denied. ACL-LOCKDOWN-STD79 is confirmed to actually run; something more specific isn't reaching messaging.4th's dictionary entries. Root cause not traced -- needs its own investigation into vm_core.c's dictionary-link mechanics and whichever capsule actually loads messaging for these identities. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo |
||
|
|
71bcb72a59 |
Fix O(N) idle-loop messaging pump; full 3x9x3 turn-attractor campaign clean on all 3 architectures (FABRIC-3.md §XXII)
sk_repl_idle()'s messaging pump (repl.c) walked the entire live-VM registry every idle beat (~1Hz) and dispatched a full VM-EXEC "MSG-TICK" -- a dictionary lookup plus a 32-slot arena scan -- into every live VM, every tick, unconditionally, forever. Fine at Tripod's original 3-VM scale; a full 9-identity turn-attractor campaign exposed it as a genuine wall on riscv64 specifically (its TCG makes each dispatch cost more): the same campaign that completed in 194s/ 339s on amd64/aarch64 never finished on riscv64 at 9 VMs across three attempts, while 8 VMs there was fine in 159s. Ruled out capacity explanations before touching anything: bumping riscv64's QEMU RAM 1024->4096 changed nothing (reverted), and a live STADIUM-RES@/MSG-STATUS probe with all 9 VMs attached showed no depletion. Host memory pressure was also ruled out directly (one background task did get OOM-killed once during the investigation, but the identical stall reproduced again with 9.2GB free). The real mistake was three premature kills under 4 minutes with no way to tell "slow" from "stuck" from outside the guest -- fixed by having run_doe_batch.sh sample the qemu process's own /proc/<pid>/stat utime every 60s; with that signal, riscv64 at 9 VMs was unambiguously alive (climbing utime, no hang), just disproportionately slow going from 8 VMs (159s) to 9 (600s+ and climbing). This was never really a riscv64-only bug: an O(N) per-second walk over the full VM population doesn't scale to the hundreds of VMs this fleet is headed toward, on any architecture -- riscv64 just made it visible first, at N=9, because its per-dispatch cost is highest. Fixed by round-robin batching: the pump now dispatches to at most SK_MSG_PUMP_BATCH (4) live VMs per idle beat via a persistent cursor that resumes where the previous beat left off, instead of all of them every time. Bounds both the scan and dispatch cost to O(K) regardless of total VM count; any single VM's queue now drains roughly every ceil(N/K) beats instead of every beat, still bounded and still matching the pump's own existing best-effort contract. No new C primitives, no messaging/Stadium changes. Verified: clean build on all 3 architectures, then the full 3x9x3 campaign re-run on all 3 (not just riscv64) per the standing rule that a defect repair requires a clean re-run everywhere before anything counts as closed: amd64 198s 0 faults 99/99 tokens x8 656/656 K-conserved aarch64 339s 0 faults 99/99 tokens x8 659/659 K-conserved riscv64 178s 0 faults 99/99 tokens x8 659/659 K-conserved riscv64 went from "never completes" to faster than aarch64, same campaign, same seed, same identity set. No regression on amd64/ aarch64. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo |
||
|
|
7edccd2c35 |
Hera can't message: root-caused and fixed by symmetry (FABRIC-3.md §XX)
Hera never loaded common:messaging.4th, unlike every other VM in the fleet. The standing belief was this was deliberate, to avoid moving her dict_hash off baseline. Checked live instead of assumed: loading messaging.4th into her dictionary silently dropped every colon- definition referencing one of 8 STADIUM-* primitives that register_child_vm_words() gives every other VM but register_mama_forth_words() never gave her -- a missing-primitive gap, not a designed privilege boundary. Confirmed mama_word_birth is genuinely VM-agnostic and SPAWN-EVENT is an unwired placeholder before proposing the fix. Fix: register the same 8 STADIUM-* primitives for Hera, load messaging.4th from init.4th the same way Hermes/Artemis/console/mint already do, and give her own idle-loop context a direct MSG-TICK call (not VM-EXEC, which would hit the same reentrancy class the existing per-other-VM pump loop already guards against) so her own queued messages actually drain. Her dictionary is now a proper superset of every child VM's, plus her remaining extra privileges -- not structurally different from any other VM, just additionally privileged. Verified: dict_hash identical across amd64/aarch64/riscv64 (0xc8f4b09e36f4fc4a), Hermes/Artemis dict_hashes unchanged and still cross-arch identical, all three boot clean to zuse)ok> with zero UNKNOWN WORD faults, mkcapsule --lint clean. Unblocks rewriting the turn-attractor (FABRIC-3.md §XIX) to coordinate via real MSG-SEND/MSG-TICK instead of blocking VM-EXEC. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo |
||
|
|
5a9425b91b |
Add VM-HEAT primitive: groundwork for a compudynamic turn-attractor (FABRIC-3.md §XIX)
Prompted by reading the std79-doe K report: checked whether the campaign's
"real cross-VM dispatch load" was actually concurrent or strictly
serialized. It's serialized at two levels -- VM-EXEC's vm_interpret(target,
...) is a direct synchronous C call (Hera fully blocked until it returns),
and even the background physics tick (vm_tick(), vm_runtime.c) is driven
by each VM's own execution loop, so idle identities accrue zero ticks
between their own turns. K's perfect conservation (FABRIC-3.md §XVIII)
verifies sequential per-VM accounting correctness, not concurrent-access
safety, since there was never concurrent access to test.
Agreed direction: fix this without a scheduler, by reusing the same
least-dense-candidate judgment stadium_admit() already trusts for
eviction, applied to "whose turn is next" instead of "who gets evicted" --
a fleet-level turn-attractor giving the next turn to whichever live
identity currently has the lowest execution_heat_q48, no fixed round-robin,
no priorities, no preemption. Lives beside Stadium in
capsule_vm_physics.c (already the fleet-level consumer of Stadium
primitives, e.g. the K mechanism itself), not inside stadium.c ("the
floor" -- residency/eviction, a different concern from turn order) and
not a new subsystem.
This pass lands only the primitive the mechanism needs: VM-HEAT
( c-addr u -- heat-q48 ), pushing a named VM's current
execution_heat_q48 via vm_physics_heat_of() -- previously C-internal
only (doe_log_heat_by_name(), doe_log.c), never exposed to FORTH. Silent
0 on an unknown/dead name (no print/error), since a turn-attractor
scanning many candidates every turn shouldn't have to filter console
noise for names that simply aren't live. Registered everywhere
VM-EXEC/VM-CALL already are. Builds clean on all three architectures;
live-tested on amd64: Hera -> 65452, Hermes -> 43, unknown name -> 0, no
faults.
The turn-attractor loop itself (std79-doe.fth's trial ordering) is not
yet built -- next step, not done here.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo
|
||
|
|
66ba21adb4 |
Log fleet_k_q48/fleet_conserved; K holds exactly, 775/775 ticks (FABRIC-3.md §XVIII)
doe_log.c's per-heartbeat-tick CSV gains two columns: fleet_k_q48 (vm_physics_fleet_heat_sum() over ALL live VMs -- the genuine fleet-wide conservation invariant K, not reconstructable from the 3 named-Tripod- member heat columns already logged, which omit every identity VM's own heat) and fleet_conserved (vm_physics_conserved() as 0/1). Requested explicitly after the first heartbeat-telemetry analysis pass (analysis-20260912/) omitted K entirely. Kernel rebuilt on all three architectures, full 3x9x3 campaign rerun (results-20260912-with-k/). K = 1.0000000000 (Q48.16 raw 65536) on every one of 775 heartbeat-tick observations, sd(K) = 0, 100% fleet_conserved, across amd64/aarch64/riscv64, nine identities, three replicates -- zero deviation. Also a free regression check on both recent Stadium fixes (§XVI/§XVII): neither disturbed the reservoir-transfer accounting K depends on. Found and fixed a tooling wrinkle along the way: fleet_conserved, being the CSV row's very last field with nothing after it to bound a regex match, can have a resumed trial digit merge into it with zero separator on the wire -- combine.py now derives it from fleet_k_q48 directly (same epsilon vm_physics_conserved() uses) instead of trusting the raw field. fleet_k_q48 itself is unaffected either way. Full analysis, discussion, and light/dark SVG->PDF figures written up as a proper LaTeX report (report-20260912/report/std79_doe_report.pdf), following experiments/bare_metal/analysis/report/bare_metal_doe_report.tex's established style -- supersedes analysis-20260912/'s markdown-only first pass as the primary deliverable for this dataset (kept, not discarded). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo |
||
|
|
e51a8d229e |
Fix stadium_grant_quota() donor floor; rerun std79 DoE clean, 81/81 (FABRIC-3.md §XVII)
capsule_birth.c hardcoded every new VM's initial Stadium quota grant to split from Hera specifically. Since a grant always halves whatever the donor currently has, Hera's own free list converges toward empty after a bounded number of grants — independent of whether the Stadium as a whole still had spare capacity, since VMs she'd granted to earlier typically still held nearly all of their own share untouched. Past that point every subsequent VM birth's Stadium grant would be silently refused (soft-failed, non-fatal by existing design), even with plenty of capacity sitting idle elsewhere. Fixed by adding an O(1)-maintained free_count to StadiumVMQuota (incremented in stadium_evict(), decremented at both of stadium_admit()'s free-list-pop sites, set/adjusted in stadium_grant_quota()'s own split — this also let grant_quota drop its old O(free-list length) counting walk in favor of an O(1) read) and stadium_best_donor(), an O(live VM count) scan over quota slots returning whichever in-use VM currently has the most free cells. capsule_birth.c's birth path now splits from that VM instead of unconditionally vm_uuid_hera(). Verified with another full rerun of the 3x9x3 std79 DoE campaign from scratch — same discipline as the prior Stadium fix (any defect repair reruns the whole DoE from the top) — one continuous boot per architecture, all 9 identities simultaneously live throughout. 81/81 trials correct, 0 mismatches, DOE-RUN header sequence md5-identical to every prior run. aarch64 ~280s total (vs ~290s for the O(ncells)-scan fix alone — confirms no regression). Both known Stadium defects are now closed together on one clean campaign rerun. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo |
||
|
|
9eff122090 |
Fix aarch64 Stadium/COOL O(ncells) scan; rerun std79 DoE clean, 81/81 (FABRIC-3.md §XVI)
Root cause of the 90+ minute aarch64 VM-birth stall found in §XV: stadium_admit()'s eviction-fallback scan iterated the entire stadium_ncells array filtered by owner, not the calling VM's own resident cells as its own doc comment claimed. Combined with stadium_grant_quota() always splitting from Hera's shrinking free list and stadium_word_dispatch() calling stadium_admit() per distinct word a VM's capsule executes, this compounded into a real O(n) blowup — catastrophic specifically on aarch64 because its -m 4096 (vs 1024 on amd64/riscv64) inflates the kmalloc heap kmalloc_init() bisects down to, which inflates stadium_ncells 4x (335,544 vs 83,886 cells, measured from boot logs). Fixed by threading a real per-VM doubly-linked resident-cell list (StadiumVMQuota.resident_head + stadium_resident_next[]/stadium_resident_prev[]) so the fallback scan is bounded by that VM's own resident count, not the global cell array size. Verified with a full rerun of the 3x9x3 std79 DoE campaign from scratch: one continuous boot per architecture, all 9 identities simultaneously live throughout (the 3-boot aarch64/riscv64 batching workaround is no longer needed). 81/81 trials correct, 0 mismatches, DOE-RUN header sequence md5-identical across all three raw logs. Identity 04's attach on aarch64, which stalled 90+ minutes before, now completes in ~34s; full boot-to-DoE-complete in ~290s. Corrects an earlier misreading (carried into §XV, std79-doe.fth's comments, and the project memory note) that described the symptom as a runaway "335,000+ cycles" dispatch counter — those were cell array indices, not an event count. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo |
||
|
|
70db955ac9 |
Fix M/MOD: hand-rolled 128/64 bit-serial division (FABRIC-3.md §XII.4)
mixed_math_word_m_slash_mod() had the identical bug class already fixed
in M* (commit
|
||
|
|
9a09949c69 |
Fix M*: use __int128 for a genuine 128-bit double-cell product (FABRIC-3.md §XII.4)
mixed_math_word_m_star() confused "double" (two full cell_t-width cells,
128 bits total on this 64-bit build -- what D+/D-/D./etc. all actually
expect) with "the low/high 32-bit halves of a single 64-bit product" --
code clearly written assuming cell_t is 32-bit. It computed an ordinary
64-bit `long long` product (already wrong for any true product exceeding
64 bits, since long long is the same width as one cell here) and split
that into 32-bit halves via `result & 0xFFFFFFFF` / `result >> 32`. For a
small negative product like -56088, this produced a positive, zero-
extended low cell paired with a correctly-looking dhigh=-1 -- D.'s
overflow check (correctly) rejected the resulting malformed double, on
every architecture, every time (this bug was never architecture-specific,
unlike the D+/D-/DNEGATE/d_compare family already fixed in bea8d74/
|
||
|
|
1a716c8048 |
Fix d_compare: use ucell_t instead of unsigned long (FABRIC-3.md §XII.4)
d_compare() (double_words.c), the static helper backing DMAX/DMIN/D</D=,
had the same bare-unsigned-long pattern already fixed in D+/D-/DNEGATE
(commit
|
||
|
|
bea8d7436a |
Fix D+/D-/DNEGATE: use ucell_t instead of unsigned long (FABRIC-3.md §XII.4)
double_word_d_plus(), double_word_d_minus(), and double_word_dnegate() (double_words.c) all cast through plain `unsigned long` for their carry/ borrow-detection arithmetic. On this aarch64 bare-metal cross-compile target, unsigned long is 32-bit (confirmed: sizeof(unsigned long)==4) -- amd64 and riscv64 both happen to have a 64-bit long, so the identical code only broke on aarch64. The low-cell arithmetic silently truncated to 32 bits, then widened back to cell_t via ordinary (non-sign-extending) conversion, producing a wrong result whenever the true 64-bit result was negative -- D. then correctly, faithfully reported DOUBLE-OVERFLOW on the resulting malformed double. vm.h already defines ucell_t for exactly this: same conditional as cell_t, guaranteed width-matched on every target. print_number_formatted() (format_words.c) already used it correctly; these three words didn't. Switched all three to ucell_t -- a one-word-class fix, no logic change. Verified: rebuilt and booted all three architectures clean. T19 (D+) on aarch64 now correctly prints -2, matching amd64/riscv64; T20 (DNEGATE) unaffected everywhere. Additional manual cases beyond the original exerciser, run live on aarch64 to specifically exercise the >32-bit-magnitude path the old bug depended on: D- (-5-3=-8), DNEGATE on 2^33 (8589934592 -> -8589934592), D+ crossing the same boundary (3+8589934592=8589934595) -- all correct. d_compare() (backing DMAX/DMIN/D</D=) has the identical latent pattern but is out of scope for this fix (not named in the request, never exercised by the campaign) -- left open, flagged in FABRIC-3.md. M*'s separate, universal-across-all-three-architectures DOUBLE-OVERFLOW bug is also untouched -- unrelated defect, not part of this fix. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo |
||
|
|
9142dda2d6 |
Fix use-after-free in sk_repl_idle()'s idle-tick VM resolution (FABRIC-3.md §XIII)
Root-caused a heap-corruption bug that reliably failed WIREBIND identity attach on the third attach/detach cycle in one boot. sk_console_readline() and sk_console_getkey() captured `active_vm` once from their caller and kept passing that same (possibly long-stale) pointer to sk_repl_idle() on every idle tick serviced while blocked waiting for input. If the VM it pointed at was killed (WIREBIND detach) mid-block, the existing bailout only checked a generic "is anyone attached" boolean -- masked as soon as a different identity attached next -- so blk_vm_flush_all() kept writing into a freed VM struct sitting on kmalloc's own free list, corrupting the free list's linked-list metadata itself. Both idle branches now re-resolve the live active VM fresh from g_repl_active_vm on every tick, matching the dispatch-side fix already made for the sibling bug in §XII.3. Verified: rebuilt amd64, reran the exact three-cycle repro that reliably corrupted the heap before the fix -- free-list census stayed stable through the same idle window that previously collapsed to zero. All three architectures (amd64/aarch64/riscv64) boot clean to the zuse)ok> prompt. kmalloc_debug_census()/kmalloc_debug_census_bytes() kept as permanent diagnostic infrastructure; every other temporary probe added during the investigation was reverted. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo |
||
|
|
70421bdd43 |
Fix real WIREBIND crash: stale active-VM pointer dispatched after blocking read (FABRIC-3.md §XII.3)
The interpreter_enabled guard added in the previous commit (
|
||
|
|
662ef44e59 |
Fix EXEC/LOAD block-persistence gap and WIREBIND/USE interpreter-race panic (FABRIC-3.md §XII)
Found live while building a cross-ISA FORTH-79 dictionary exerciser: - capsule_exec_init() zeroed a capsule's block content immediately after running it, so LOAD (a genuine FORTH-79 standard word, ACL-allowed even for locked identities) could never actually read back what EXEC had just written. Removed the clear from capsule_exec_init(); block content now persists like any other Standard BLOCK/BUFFER/UPDATE write. kernel_main.c's own explicit post-birth clear of Mama's init.4th range is untouched. - USE could redirect the console to a WIREBIND identity's VM before that VM's vm_enable_interpreter() step of its own birth sequence had run, causing the next typed line to hit vm_assert_interpreter_enabled() and panic the entire machine -- not the per-session-recoverable ACL-fault path a redirected VM otherwise gets. USE now checks interpreter_enabled first and refuses with a retry message instead. Also includes the amd64/aarch64/riscv64 acceptance boot logs and DoE CSVs from this session. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo |
||
|
|
b301317902 |
xHCI: drive Port Reset on port reuse; WIREBIND: kill the console VM too (FABRIC-3.md §XI.5)
Two independent bugs that together caused a reliable hotplug wedge: reusing an xHCI port for a second identity right after an unclean detach of a first would leave no further hotplug events reaching the guest at all. Bug 1 (xhci.c/xhci_driver.h): the xHCI driver never drove PORTSC.PR -- a known, named gap since Milestone 2e (the code's own comment flagged it, PORTSC_PR/PRC were defined but never referenced). A port's first connect each boot reads PED already set, so skipping the reset happened to work; a second device on the same port after a prior disconnect reads PED clear, and Address Device reliably failed without an explicit reset cycle. New XHCI_CONN_AWAIT_PORT_RESET state drives PR and waits for PED to read set before proceeding to Enable Slot. Bug 2 (capsule_wirebind.c): capsule_wirebind_eject()/unclean_detach() compared g_repl_active_vm against the *user* VM's pointer (g_wirebind_attached_vm_id tracks that one, not the console VM) -- never equal, since USE/g_repl_active_vm always points at the console VM. The guard never fired and the console VM was never killed at all, only orphaned -- paired to a dead user VM but still the REPL's active session. New wirebind_teardown_console() helper resolves and tears down the console VM by its own tracked bare username. Verified live on amd64: the exact repro (identity 01 on port 2, unclean detach, identity 02 on the same port immediately after) -- previously wedged with "xhci: address device failed" and no further hotplug activity; now attaches cleanly and fast, both VMs' KILL messages appear, console is immediately interactive on the new identity. Three-arch clean qemu acceptance passed, all clean on the first attempt. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014Ec88YKxxhZGG1RNnune78 |
||
|
|
d6661b5eed |
Scope VM fault halt to the faulting identity's own session (FABRIC-3.md §XI.4)
A standalone WIREBIND identity (no Zuse, USE'd in directly) hitting an ACL-denied word halted the entire machine -- Hera, Hermes, Artemis, all of it -- instead of just that identity's own session. sk_fault_handler() was being called unconditionally on whichever VM's ->error was set, with no distinction between Hera's own root session (where "no fallthrough surface" is the correct, deliberate fail-closed behavior) and a USE'd-in guest identity (which should recover and resume at its own prompt instead of taking the fleet down with it). Both call sites (sk_repl_step, sk_repl_run) now compare the faulting VM against Hera before deciding: Hera's own session still halts by design; any other VM prints a recovery message, clears its fault state, and continues. Also: mint identities 01-06 with the same FORTH-79/83 restricted personality identity 00 already had, verified via the fixed fault scoping above (which this verification pass surfaced). Verified live on amd64 (both the Hera-halts and identity-recovers branches); three-arch clean qemu acceptance passed (riscv64's first attempt hit an unrelated virtio_blk I/O timeout hang, a known QEMU/TCG flake -- a clean retry booted normally). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014Ec88YKxxhZGG1RNnune78 |
||
|
|
1839a2b0c3 |
WIREBIND cert verification: load Zuse's root pubkey independently of her live session
Root cause of the remaining "identity attach doesn't complete when Zuse never attaches this boot" issue: capsule_wirebind_verify_cert() gated on mama_vm->zuse_cert_installed, which is only ever set when Zuse's own drive attaches and authenticates this specific boot (capsule_zuse_boot_try_attach() -> install_and_activate() -> vm_zuse_cert_install()). Without her, any other identity's WIREBIND cert verification silently refused -- correctly, by the old design, but that design conflated two genuinely different things: "can mint new identities" (needs Zuse's live private seed, a real privileged operation) and "can verify an existing identity's cert" (needs nothing but her already-public key). That public key was already being persisted independently of her live session: zuse_genesis_marker_t (zuse_genesis_marker.h) stores it in the kernel's own top-of-device metadata fence (Artemis's resident storage), written once at genesis, specifically *not* alongside her private seed (which stays only on her own removable thumbdrive) -- the type's own doc comment says as much. It just wasn't being loaded for anything but confirming which drive is genuinely hers. Fix: a new capsule_zuse_boot_load_root_pubkey() (capsule_zuse_boot.c) reads that marker and populates two new VM fields, zuse_root_pubkey_known / zuse_root_pubkey (vm.h) -- deliberately separate from zuse_cert_installed/zuse_cert_seed/zuse_cert_pubkey, which stay untouched and still gate MINT exactly as before. Called once from kernel_main.c as soon as Artemis's own storage attaches, unconditionally, independent of whether Zuse's own drive is ever attached this boot. capsule_wirebind_verify_cert()/capsule_wirebind_try_attach() now check zuse_root_pubkey_known instead of zuse_cert_installed. One identity's attach must not depend on another identity's live presence -- each identity stands on its own once the fleet's root of trust has been established once, ever. Verified live, amd64: identity 00 (disk/thumbdrives/00-thumb-ident.img) now attaches and completes WIREBIND in 19 seconds with Zuse's own drive never attached this boot at all (previously: unbounded, many real minutes or effectively never, before today's other fixes; still slow/ stuck after those, stuck specifically on this silent refusal). Zuse's own attach flow re-verified unaffected (regression check, amd64). Three-arch clean qemu acceptance (amd64/aarch64/riscv64) passed with this change included. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014Ec88YKxxhZGG1RNnune78 |
||
|
|
1a263555e2 |
Fix pathological migration scan that stalled WIREBIND identity attach
Root cause (found by a fresh subagent after an extended live-debugging investigation into "identity 00 attaches slowly/stalls when Zuse never attached first this boot"): blk_migration_idle_check() was generalized earlier today to walk every attached device slot uniformly instead of hardcoding first_disk_slot() (Artemis's own disk). But its per-slot scan can only early-exit once it finds a devblock that is BOTH "hot" (claimed and worn) AND "free" -- and a just-attached, never-claimed USB identity drive can never satisfy the "hot" half by design (claiming only ever happens via blk_firsttouch_claim(), which only ever targets first_disk_slot()). So the scan ran to completion -- the drive's entire ~16,000 devblocks, mostly cache misses over slow emulated USB/BOT -- every single idle tick, forever, blocking sk_repl_idle() (and therefore the console and the storage-attach message round-trip) each time. Fix, in src/block_subsystem.c: a new has_ever_claimed flag on blk_dev_slot_t (set in blk_set_meta(), the single choke point every BLK_FLAG_CLAIMED transition passes through) skips the scan entirely, O(1), for any slot nothing has ever claimed -- the common case for a freshly-attached drive. A new migration_scan_lbn resume cursor bounds *any* slot's per-tick cost to MIGRATION_SCAN_BUDGET (256) devblocks examined, picking up where the previous tick left off instead of restarting from start_lbn every time -- restores this function's own documented "coarse cadence, cheap early-exit" design intent for every device, not just the one it used to hardcode. Also along the way (kept, all real improvements, verified live): - src/starkernel/usb/xhci.c: xhci_wait_bit()/xhci_bot_wait_for_idle() had zero yield hints in their MMIO-polling loops; added arch_relax() to both (matches virtio_blk.c below) -- a tight loop of nothing but MMIO reads can starve TCG's own host-side timer injection under QEMU. - src/starkernel/virtio/virtio_blk.c: vblk_io()'s spin bound was 33M iterations with zero logging on timeout; a single real (still not fully root-caused) timeout cost 31+ minutes of CPU before this was caught. Reduced to 1M and added a log line naming the failing sector, turning a silent, effectively-unbounded stall into a fast, loud failure -- callers already tolerate BLKIO_EIO. - src/starkernel/repl.c: blk_migration_idle_check() deferred for any idle tick where a storage-attach message round-trip is still pending, to keep the two block-subsystem-touching paths from interleaving; the existing MSG-TICK pump now checks the target VM's own dictionary for MSG-TICK (vm_find_word + acl_allow) before VM-EXECing into it, instead of spamming "VM-EXEC: ERROR" every tick for a VM that doesn't have -- or isn't allowed -- the word (e.g. a FORTH-79/83-locked-down identity); fixed a real -Wunused-variable build break in the EMERGENCY_CONSOLE_ ENABLED=1 path (unused today, but needed live to reproduce this bug with no identity attached at all). - src/starkernel/capsule/capsule_mint.c: dropped the dead S" common:messaging.4th" EXEC / MSG-CD-INIT lines from MINT_RESTRICTED_PERSONALITY -- a FORTH-79/83-locked-down identity has no legitimate use for a messaging vocabulary it can never call. Status: the pathological CPU-climbing scan is confirmed fixed (verified live: CPU stays flat across an extended run instead of climbing without bound). The WIREBIND storage-attach message round-trip still does not complete promptly in the "Zuse never attached, other identity attaches first" scenario -- a separate, still-open issue in the message-delivery path itself, not the scan. Tracked as follow-on work. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014Ec88YKxxhZGG1RNnune78 |
||
|
|
63b8b3bc29 |
Migrate Hera->Artemis storage-attach to a real message round-trip
Hera still polls xHCI and sig-checks attached drives, but the storage
registration step (blk_subsys_attach_device(), now wrapped as the
BLK-ATTACH primitive) moves to Artemis's own dictionary, reached via
HERA-BLK-ATTACH-REQ/BLK-ATTACH-ACK (VM-EXEC, since Hera can't load her
own messaging.4th -- see the doc comment in repl.c). Identity birth
(Zuse genesis / WIREBIND) is deferred until the ack confirms storage
actually succeeded, instead of running synchronously underneath a
storage call that might fail ("wait for ack, safer for identity data").
Caught and fixed a real bug live during acceptance testing: Artemis's
ACK-APPEND-NUM fed a single-cell value into <# #S #> (which expects a
double-cell pair), causing a stack underflow the first time
HERA-BLK-ATTACH-REQ ran. Fixed with the same `0 SWAP` convention every
other numeric-append helper in this codebase already uses.
Verified booting clean to (zuse) ok> with no VM-EXEC errors on all
three architectures (amd64/aarch64/riscv64).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014Ec88YKxxhZGG1RNnune78
|
||
|
|
f4ded3e1a8 |
MINT: add a FORTH-79/83-standard-words-only lockdown personality
Captain Bob, 2026-09-07: "starting with that 00 user we created, we're
going to give access only to FORTH 79 and 83 standard words. everything
else is locked down."
New capsules/acl-std79.4th (blocks 4023-4047): walks a VM's own
dictionary (>LINK/LINK> traversal, same as ACL-INIT-PRIMITIVES/WORDS
already use) and permanently denies+pins every word not on an explicit
FORTH-79/83 allowlist, extracted from the real registered word set
(stack_words.c through control_words.c), not recited from memory.
Deliberately excludes, beyond plain non-standard words: BYE (100% ACL
bypass to the emergency console -- "needs more discussion, exclude for
now"), COLD/WARM/REBOOT/SAVE-SYSTEM (system lifecycle), the block/screen
editor L/S/SHOW/EDIT/UPDATE/SAVE-BUFFERS (lets a session rewrite
persistent block/capsule content, defeating the lockdown even though
nominally standard), BLK-ACL-*/BLK-OWNER@ (StarForth-specific), and
FORGET/FENCE (flagged as an unrestricted superpower word, 2026-09-03
audit). Keeps WORDS/VLIST/SEE (introspection only -- ACL is enforced
per-target-word at execution time regardless of how an XT was
obtained) and the parenthesized control-flow runtime primitives
((BRANCH) etc. -- IF/DO/LOOP compile calls to these; denying them
breaks ordinary control flow, not security).
MintPersonality enum (capsule_mint.h) lets capsule_mint_identity()
select which personality-source template gets written to a new
identity's devblock -- MINT_PERSONALITY_DEFAULT (unchanged) or
MINT_PERSONALITY_STD79_LOCKDOWN (EXECs acl-std79.4th then
ACL-LOCKDOWN-STD79 as the VM's own last bootstrap step). The actual
restriction logic stays entirely in FORTH per .claude/CLAUDE.md's
Word-Level ACL System rules ("ACL policy belongs in ACL.4th, never in
C") -- capsule_mint.c only picks which few-line bootstrap stub to
write. MINT's own stack signature gains a trailing restrict? flag;
capsule_zuse_boot.c's genesis mint (Zuse herself) explicitly passes
MINT_PERSONALITY_DEFAULT -- the superuser is never restricted.
Two real bugs found and fixed live during testing, both the same class
of self-referential fault: ACL-LOCKDOWN-STD79's own walk loop calls
ACL-STD79-ALLOWED?/ACL-STD79-LIST/ACL-ALLOW!/ACL-PIN on every single
iteration to do its job -- none of those are FORTH-79/83 standard
words, so the walk was denying its own load-bearing infrastructure
partway through and then faulting the next time it tried to call it
("VM fault -- emergency console disabled; halting", reproduced twice
live). Fixed by explicitly protecting all four in the allowlist
(block 4047) -- they must stay allowed for the walk to finish, not
because they belong on a "standard words" list.
Verified live end-to-end: minted a throwaway test identity with the
restrict? flag, confirmed her WIREBIND birth completes cleanly (no
faults, no shadow conflicts) on a single real attach, then USE'd into
her VM and confirmed standard arithmetic and user-defined words work
(1 2 + . -> 3; : X 5 5 * . ; X -> 25) while KILL is entirely unknown to
her dictionary and VM-EXEC is denied. One real, non-fatal side effect
found and left as-is (not asked to fix): the fleet's inter-VM messaging
pump (MSG-ARENA) is also denied by the lockdown, logging a harmless
per-idle-tick warning -- a fully locked-down VM doesn't participate in
message routing.
Not yet applied to the real identity 00 -- this commit is the
mechanism, verified against a disposable test identity only.
Three-arch clean qemu acceptance (single Zuse device, standard
regression case) passed on amd64, aarch64, and riscv64.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014Ec88YKxxhZGG1RNnune78
|
||
|
|
56e19a00f9 |
Route VM-dictionary sf_malloc/sf_free through the kernel's kmalloc heap
Captain Bob's call after seeing the identity-heap-capacity findings (FABRIC-3.md §X.4): the number of concurrently-running VMs is not known in advance, and once this is a complete operating system the heap should be able to use whatever memory is actually available -- not a hardcoded compile-time ceiling. This was already half-built and just not wired up. src/starkernel/vm/alloc_kernel.c previously implemented sf_malloc()/ sf_free() (platform_alloc.h's allocator abstraction -- what vm_create_word() calls for every VM's word dictionary) as its own isolated static 4MB arena: first-fit free list, no splitting or coalescing. That's exactly the allocator that topped out around 6 concurrent WIREBIND-born identities, failing from fragmentation before true capacity exhaustion (§X.4's own measurements). Sitting right next to it, unused for this purpose: src/starkernel/memory/ kmalloc.c, the kernel's general heap. Already initialized at boot (M6, kernel_main.c, well before any VM is ever born), reserved from real PMM-tracked physical memory rather than a fixed array, defaults to a 2 GiB floor explicitly sized "for 256+ baby VMs" per its own comment, overridable via the --heap= boot flag, and its free list actually coalesces neighboring blocks on every free. Change: alloc_kernel.c's sf_malloc()/sf_free() now delegate to kmalloc_aligned()/kfree() instead of managing a separate arena. sf_alloc_init() becomes a no-op (kmalloc is already initialized by the time any VM allocation can happen, and "resetting" a heap now shared by every kernel subsystem would be actively wrong -- confirmed no external caller depended on its old reset semantics). sf_alloc_get_stats() reads kmalloc_get_stats() fresh rather than shadowing byte counts locally; alloc_count/free_count (which kmalloc.c doesn't track) stay as simple local counters. sf_calloc()/sf_realloc() are otherwise unchanged. Kernel- only: the hosted (non-kernel) StarForth build keeps its own separate alloc_host.c implementation, untouched. Verified live: replaying the exact hotplug sequence that previously topped out at 6 identities (Zuse + 8 identities, one at a time via QMP device_add) now succeeds for all 9, where identity 05 specifically used to fail. Three-arch clean qemu acceptance (single Zuse device, the standard regression case) passed on amd64, aarch64, and riscv64 -- one aarch64 attempt hit an unrelated, already-documented one-off QEMU hiccup (empty log, boot never progressed past firmware) and passed cleanly on retry with no rebuild. Not addressed here: the underlying free-list itself is still first-fit without splitting (only coalescing changed, inherited from kmalloc.c); per-VM dictionary sizing (shrinking what each WIREBIND VM's word set actually needs) is a separate, still-open lever from FABRIC-3.md §X.4's open architecture question. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014Ec88YKxxhZGG1RNnune78 |
||
|
|
d8a195b8d8 |
WIREBIND: kill the orphaned console VM when the user VM birth fails
Found live during identity-heap-capacity testing (2026-09-07, hotplugging Zuse + 8 identities one at a time and measuring the kernel heap arena via a temporary allocator-stats probe, since reverted): every WIREBIND identity attach births two VMs in sequence -- a "console" VM, then the real "user" VM. When the second birth failed (arena fragmentation under concurrent VM load, a separate, not-yet-fixed capacity issue), capsule_wirebind_try_attach() logged the failure and returned, but the console VM that had *already succeeded* was never torn down. It stays live and registered under the identity's username, consuming its own ~228KB of the fixed 4MB kernel heap arena forever -- nothing ever points a real user at it, since WIREBIND only ever hands the caller the user VM's id. This turns every failed identity attach into a permanent net loss of heap rather than a neutral retry: confirmed live that a failed attach left the arena 228,576 bytes worse off than before the attempt, and every subsequent attempt starts from that worse baseline, compounding. Fix: call capsule_vm_kill(username) on the now-orphaned console VM before returning from the failure path -- the same teardown capsule_wirebind_eject()/capsule_wirebind_unclean_detach() already use elsewhere in this file (vm_cleanup() + sf_free(), confirmed live to actually reclaim per-word dictionary allocations, FABRIC-3.md §IX.2/§IX.3). Verified live with the same allocator-stats probe (written, captured, reverted -- not part of this commit): after the fix, a forced user-VM birth failure now returns the arena to exactly its pre-attempt byte count (3,526,256, matching the baseline precisely) instead of leaking 228,576 bytes. Three-arch clean qemu acceptance (single Zuse device, the standard regression case) passed on amd64, aarch64, and riscv64. The underlying capacity/fragmentation question (why the 7th concurrent identity's arena allocation fails at all despite technically-sufficient free bytes) is a separate, open architecture question -- not addressed here. See project memory for the full measured numbers. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014Ec88YKxxhZGG1RNnune78 |
||
|
|
e10fb76fb2 |
xhci: fix Configure Endpoint completion drop under concurrent multi-device enumeration
Root cause of the FABRIC-3.md §IX.5 follow-on: with 9 devices attached concurrently at boot (Zuse + 8 identities), only 1 of 9 ever completed enumeration and reached blkio_usb: MSC device ready -- the other 8 produced no error and no success, just silence. xhci_poll_events()'s deferred per-slot dispatch loop submitted a Configure Endpoint command (a Command Ring op) unconditionally for every slot with that action pending in a single pass -- unlike every other Command Ring op in this driver (Enable Slot, Address Device, Disable Slot), which is correctly gated behind dev->connect_state == XHCI_CONN_IDLE before ever submitting. With 2+ devices enumerating concurrently, this let multiple Configure Endpoint commands sit outstanding on the Command Ring at once. Their completion is correlated purely via the single shared dev->connect_state field (== XHCI_CONN_AWAIT_CONFIGURE_ENDPOINT), not the completion event's own Slot ID -- so whichever slot's completion happened to land while connect_state still read AWAIT_CONFIGURE_ENDPOINT got correctly chained into SET_CONFIG, and every other slot's completion arrived after connect_state had already moved on, silently swallowed by the handler's generic "unrelated command completion" catch-all. No error path exists for this, which is why it produced total silence rather than a diagnosable failure. Root-caused live via temporary WARN-level diagnostic probes (written, captured, and fully reverted per the project's own probe convention -- this commit contains only the functional fix and its explanatory comment, no probe code) added at four points: the initial port scan, the connect handler, the Command Completion Event handler, and the deferred dispatch loop itself. The probes showed all 9 devices correctly completing Enable Slot + Address Device (ruling out the connect-state queue as the cause, the original hypothesis), then all 9 correctly submitting Configure Endpoint and all 9 commands completing successfully in hardware (code= SUCCESS, no errors logged) -- but only 1 of 9 ever got its next_action chained to SET_CONFIG. Fix: apply the same single-in-flight discipline this driver already uses for every other Command Ring op. If the Ring isn't free when a slot's Configure Endpoint action is due, put the action back on that slot instead of submitting a second command onto a busy Ring -- the next tick's dispatch pass retries it once the Ring frees up. Verified live: booting Zuse + all 8 identity drives concurrently (9 devices, one per real xHCI port via the XHCI_PORTS fix from the previous commit) now produces 9 "MSC device ready" lines and zero xHCI errors, where it previously produced exactly 1. Three-arch clean qemu acceptance (single Zuse device, the standard regression case) passed on amd64, aarch64, and riscv64 -- no change in that baseline behavior. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014Ec88YKxxhZGG1RNnune78 |
||
|
|
2c1b3cd695 |
Four bugs found live verifying the 8 identity thumbdrives (FABRIC-3.md §IX)
All found by actually running the identity workflow §VII/§VIII made possible, not by code review: 1. Zuse/WIREBIND cross-contamination on detach: capsule_zuse_boot_logout() and capsule_wirebind_unclean_detach() both had no device parameter, so an unrelated device detaching (while the real owner's own stayed attached) incorrectly tore down the wrong session. Both now compare the departing device against their own tracked one, mirroring capsule_wirebind.c's pre-existing g_wirebind_attached_dev precedent. 2. Dictionary-entry memory leak: vm_create_word()'s sf_malloc()'d DictEntry (plus a second per-entry allocation for transition_metrics) was never freed by vm_cleanup(), in both the hosted and kernel implementations. Caused a real kernel PANIC after 8-9 repeated VM birth/kill cycles in one boot. Fixed by walking vm->latest in both. 3. sf_malloc/sf_free (alloc_kernel.c) was a 4MB bump arena with a deliberate no-op free, sized on "VM born once, never killed" -- fix #2 alone didn't stop the panic because free() itself discarded the pointer regardless. Given a real free list (first-fit reuse). 4. Headless-console gate didn't re-engage after a mid-boot logout: the original fix (sk_console_mark_login(), one-way sticky) only gated the first login of the boot. Replaced with a live check (sk_console_identity_present()) re-evaluated continuously, including inside sk_console_readline()'s own blocking idle loop -- the console is normally sitting blocked there when a hot-unplug logout happens, so checking only at the top of the REPL loop wasn't enough. Also: MINT now verifies its own write (verify_mint(), capsule_mint.c) by reading back through the same check a real attach performs, rather than trusting blkio_write()'s BLK_OK alone -- logged via log_message(), not console_println(), per direct instruction. Verified live, amd64: the full 8-identity repeated attach/detach cycle that previously panicked at the same point every time now completes clean, and a full serial-log sweep found zero bare unauthenticated prompts anywhere in the run. Three-arch clean-qemu acceptance passed. Still open, not fixed here: a 3+-simultaneous-device USB enumeration failure found in a separate live test, not yet root-caused. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018EjXFo7mPXjUMjfJeuUUz4 |
||
|
|
9e81de3f43 |
xHCI/BOT driver: genuine multi-device support (FABRIC-3.md §VII)
Per-slot registry (xhci_msc_slot_t/dev->msc_slots, sized off the controller's own reported max_slots) replaces the single-device scalar fields the driver carried since Milestones 2e-2h. Boot-time port scan no longer stops at the first connected device; a connect/disconnect that arrives while the Command Ring is busy is now queued and drained instead of dropped. blkio_usb.c and repl.c's own single-device state (device descriptor buffers, blkio_dev_t, attach bookkeeping) became per-slot registries the same way. Live multi-device testing (not just compiling) surfaced a second, more severe bug outside the original plan: transfer_purpose and next_action were also single scalars shared across the whole controller. Two devices enumerating concurrently could have one's completion silently overwrite the other's still-outstanding one, permanently stalling it with no error. Fixed by moving both per-slot and, critically, reading the Transfer Event TRB's own real Slot ID field instead of trusting external bookkeeping. Verified live, all three architectures, mandatory clean-qemu acceptance: existing single-device path unchanged, and two devices attached simultaneously (amd64) both progress independently through enumeration without corrupting or stalling each other. Also in this pass (implemented and verified in earlier turns this session, committed together per direct instruction): - Headless-until-login console policy: no prompt/banner until a real identity logs in via an attached thumbdrive (WIREBIND or Zuse, neither special), reusing EMERGENCY_CONSOLE_ENABLED as the debug/recovery escape hatch (now default-off). - KILL/g_repl_active_vm dangling-pointer fix: killing the VM the console is currently USE'd onto now detaches back to Hera first, matching the existing EJECT/UNCLEAN precedent. FABRIC-3.md §VII/§VIII carry full closure notes for all three. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018EjXFo7mPXjUMjfJeuUUz4 |
||
|
|
4d4ab59189 |
Build the FIRSTTOUCH overflow trigger flagged in FABRIC-2.md §I.2
The overflow-triggered migration path (a WIREBIND-attached identity's own drive running low on space) was scoped but never built -- only the trigger-detection call site was missing, per this section's own text. - capsule_wirebind.c now tracks the attached blkio_dev* alongside the already-tracked VM id, set in try_attach() and cleared in both EJECT/UNCLEAN paths. - New capsule_wirebind_overflow_idle_check(), called once per idle tick in repl.c right alongside blk_migration_idle_check() (same cadence): reads the attached drive's free/total via blk_get_device_free_blocks(), and if free space is below a fixed 10% threshold, extends the identity's pool with a one-time blk_firsttouch_claim() of 8 additional devblocks on Artemis's system-resident device. - New blk_owner_has_claim(owner_fp) in block_subsystem.c answers the debounce question blk_firsttouch_claim()'s own doc comment had left open: a disk scan, not a RAM flag, so the already-extended answer survives reboot/reattach, matching BMAPFMT's "ownership travels with the block" model. Premise checked before building (does a WIREBIND-attached drive actually give a real free/total signal, or does it stay PROVISIONAL/raw): traced repl.c's attach sequence and confirmed blk_subsys_attach_device() runs on the same dev pointer right after WIREBIND, and a WIREBIND-eligible drive is always already STFR/v2-formatted, so the signal is real. Premise held, unlike the BAM item's overstated one. Verified with the mandatory 3-arch QEMU acceptance (identical dictionary hashes, no regression) plus a live logic test of blk_owner_has_claim(): a temporary TEST-OWNER-CLAIM word, run once via SK_CMD and reverted, confirmed it correctly detects the claiming owner and rejects an unrelated one. The low-disk-space-triggers-a-claim path itself is not verified end-to-end -- that needs a real minted WIREBIND-user thumbdrive with deliberately tiny capacity, out of scope for this pass; noted as such in the FABRIC-2.md §I.2 closure note rather than overclaimed. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018EjXFo7mPXjUMjfJeuUUz4 |
||
|
|
aa33aedca8 |
Fix blk_meta_t/BAM accounting reconciliation flagged in FABRIC-2.md §I.2
FIRSTTOUCH ownership (blk_meta_t's BLK_FLAG_CLAIMED/owner_fp) and the generic block-allocation bitmap (BAM, blk_bam_entry_t) were two parallel, unreconciled accounting systems: blk_firsttouch_claim() never touched the BAM, and blk_allocate()/devblock_is_free() never checked BLK_FLAG_CLAIMED. A FIRSTTOUCH claim could be silently overwritten by a later blk_allocate() call, or could itself steal a devblock already in ordinary use via BLOCK/UPDATE. - devblock_is_free() (shared by blk_firsttouch_claim() and blk_migration_idle_check()) now also checks the BAM entries of all BLK_PACK_RATIO member LBNs, not just blk_meta_t. - New devblock_claimed_by_lbn() helper wired into blk_allocate()'s free-scan, so it skips any LBN whose devblock is BLK_FLAG_CLAIMED. - blk_firsttouch_claim() now marks the BAM allocated for all 3 member LBNs of each devblock it claims, which also fixes vol_meta.free_blocks never decrementing for FIRSTTOUCH claims. - blk_meta_relocate_devblock() traced and confirmed NOT part of the bug — it already keeps BAM in sync via blk_subsys_relocate_block()'s own blk_mark_free()/blk_update() calls. Both boundary cases (the reserved/user LBN split at a slot's start_lbn, and BAM-array bounds) are guarded explicitly. Verified with the mandatory 3-arch QEMU acceptance (identical dictionary hashes, clean BYE) plus a live logic test: a temporary TEST-BAM-RECON word, run once via SK_CMD and fully reverted, confirmed on running code that ordinary allocation and a FIRSTTOUCH claim land on disjoint LBN ranges in both directions. FABRIC-2.md §I.2 updated in place with the closure note, per this project's documentation discipline. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018EjXFo7mPXjUMjfJeuUUz4 |
||
|
|
d6895bdd15 |
Move vocabulary and control-flow state off file-scope statics onto VM
proof/FINDINGS.md's Isabelle/HOL word-source sweep (§1) found the two defects severe enough to actively corrupt the live Tripod multi-VM fleet: file-scope C statics standing in for state that belongs on struct VM. - vocabulary_words.c (highest severity in the sweep): forth_vocab/ context_vocab/current_vocab, context_var_addr/current_var_addr, the ctx_fc/forth_fc first-char search index, and the `initialized` guard were all process-wide statics. Only the first VM to touch any vocabulary word ever ran setup; every VM after that silently shared VM #1's dictionary-chain pointers and reused VM #1's byte-offset addresses as if valid in its own vm->memory. One VM's VOCABULARY/ DEFINITIONS/FORTH silently changed where every other VM looked up and defined words. - control_words.c: cf_stack/cf_sp/cf_last_mode (IF/THEN/BEGIN/DO/CASE compile-time nesting) and the LEAVE/ENDOF patch-site bookkeeping (leave_addrs/leave_sp/leave_mark_*, endof_addrs/endof_sp/endof_mark_*) were also process-wide statics. Two VMs compiling colon definitions at overlapping times would corrupt each other's nesting state. Both moved onto struct VM, following the existing hold_addr/hold_pos precedent in include/vm.h ("lives in each VM's own memory... so child VMs never alias Hera's buffer"): - New VocabularyState struct (vm->vocab): chain heads, VM-cell addresses, first-char index, initialized flag. - New ControlFlowState struct (vm->cf): cf_stack/cf_sp/cf_last_mode plus the LEAVE/ENDOF patch-site stacks. cf_tag_t/cf_item_t/CF_STACK_MAX moved from control_words.c into include/vm.h since they're now part of the struct VM field's type. - Sentinel fields (-1/-999, meaning "empty") explicitly initialized in both vm_init_with_host() implementations (hosted src/vm_bootstrap.c and kernel src/starkernel/vm/vm_bootstrap.c) alongside the existing dsp/rsp = -1 initialization, since the preceding zero-init leaves them at 0 rather than their empty sentinel. Every word function in both files already took VM *vm, so no call sites outside these two files needed to change; cf_push_item/cf_pop_item/ cf_peek_item gained a VM* parameter to reach vm->cf. Verified: hosted (amd64) and kernel (amd64, __STARKERNEL__) both build clean with -Wall -Werror after a full clean rebuild (struct VM's layout changed size, and this Makefile has no header-dependency tracking, so a stale incremental build would have linked mismatched object layouts). Hosted POST suite 1012/1012 passing (0 regressions). Manually exercised VOCABULARY/DEFINITIONS/FORTH/ORDER, and IF/ELSE, DO/LOOP/LEAVE, BEGIN/WHILE/REPEAT, and CASE/OF/ENDOF/ENDCASE (including nested DO with I/J) in the REPL -- all correct and unchanged from pre-refactor behavior. Note: a pre-existing CASE/ENDCASE default-clause bug (the code after the last OF...ENDOF pair does not correctly become the "default" value once DROP runs) was found while testing this refactor and confirmed present on unmodified master too -- not touched here, out of scope for this pass. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014Qf6YcnHgaEtEygq3knx19 |
||
|
|
c36bd99e1e |
Fix EXECUTE/?/DUMP/TYPE/DECIMAL-HEX-OCTAL/ALIGN defects from proof sweep
proof/FINDINGS.md's Isabelle/HOL word-source sweep (§4) flagged five real defects; this fixes all five and records resolution in that doc: - EXECUTE (system_words.c): cast a popped cell straight to a DictEntry* and called through it with only a null check. Now validates via a new shared vm_dict_entry_ok(), promoted out of starforth_words.c's ENTROPY@/ENTROPY! guard (dictionary_management.c) so EXECUTE gets the same live-entry check. - ? and DUMP (format_words.c): dereferenced the popped cell as a raw host pointer, bypassing vm_addr_ok entirely (out-of-VM-bounds read). Both now go through VM_ADDR/vm_addr_ok/vm_load_cell/vm_ptr like every other memory word (@, `,`, editor_words.c). - TYPE (io_words.c): bounds check computed addr+count in signed 64-bit arithmetic, which can overflow and bypass the check on large operands. Replaced with vm_addr_ok(), which is written to avoid that overflow. - DECIMAL/HEX/OCTAL (format_words.c): wrote only the BASE memory cell, never vm->base, the host-mirror field number-output words actually read via current_base() -- so these words silently affected number parsing but never printing. Now call the existing vm_set_base() (previously only used at boot init), which updates both. vm_get_base/vm_set_base promoted to public declarations in include/vm.h. - ALIGN vs ALLOT/,/C,/2, (dictionary_words.c): disagreed on dictionary growth ceiling (2MB vs 5MB). Investigated which was correct rather than blindly widening: vm_get_block_addr() maps block N to vm->memory + N*BLOCK_SIZE across the full 5MB arena, and USER_BLOCKS_START (block 2048) lines up exactly with DICTIONARY_MEMORY_SIZE -- so ALLOT/,/C,/2, letting `here` grow past 2MB could silently corrupt live block/user data sharing that memory. Tightened ALLOT/,/C,/2, to DICTIONARY_MEMORY_SIZE to match ALIGN. Verified: hosted (amd64) and kernel (amd64, __STARKERNEL__) both build clean with -Wall -Werror; hosted POST suite 1012/1012 passing (0 regressions); manually exercised EXECUTE, ?/DUMP, TYPE, HEX/DECIMAL/OCTAL, and large-ALLOT rejection in the REPL. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014Qf6YcnHgaEtEygq3knx19 |
||
|
|
70dc8beba4 |
FABRIC-2.md §I.9 follow-on: blinking | cursor instead of static block
Captain Bob asked for the framebuffer cursor to render as a blinking vertical bar rather than the previous static solid-block glyph. vt100_draw_cursor() (hal/vt100.c) now fills a thin bar (cell_w()/8, min 1px, full cell height) at the cursor's left edge instead of the whole cell -- an I-beam shape. vt100_erase_cursor() is unchanged (clearing the whole cell already safely covers the narrower bar). Blinking is new in repl.c: sk_console_readline()'s idle branch toggles the cursor on/off every SK_CURSOR_BLINK_INTERVAL (50 ticks, 500ms at 100Hz) via alternating console_fb_draw_cursor()/console_fb_erase_cursor() calls, independent of the heartbeat/idle-beat mechanism the §I.9 fix just touched (deliberately not reused, to avoid recoupling to that path). Runs regardless of n, so it blinks whether sitting at a bare prompt or paused mid-edit. Every deterministic draw site (initial prompt, prompt reanchor, backspace, character echo) now goes through a new helper, sk_cursor_show(), which resets the blink cycle to "on" and redraws -- typing always shows a solid cursor, never mid-blink. Verified via the mandatory foreground 3-arch QEMU acceptance boot: amd64 (logs/20260905-021054, extensive live interactive typing including multi-line : / ; word definitions and error cases, prompts stayed correctly attached throughout), aarch64 (logs/20260905-021551), riscv64 (logs/20260905-022324) -- all three reached (zuse) ok> and shut down cleanly via BYE. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019YcT3H2PQeyujrzjqS3Var |
||
|
|
8edb95b65d |
FABRIC-2.md §I.9: fix the terminal phantom-linebreak defect
console_ensure_line_start() (hal/console.c) used to emit its newline
immediately via a path that deliberately skipped the tx-byte counter, so
that sk_repl_idle()'s unconditional per-beat call to it (repl.c, ~1s idle
heartbeat) could force a real newline the REPL's own prompt-reanchor logic
never noticed -- the prompt was never reprinted, and the next real
keystroke echoed onto the now-blank line, indistinguishable from Enter
having already been pressed at a bare prompt. Root-caused in the previous
commit (
|
||
|
|
dbaead0af2 |
xhci: route driver chatter through log_message(), silence at default log level
xhci.c and repl.c's USB attach/detach path printed every xHCI command submission, completion, and BOT transfer step unconditionally via console_println()/console_puts() -- floods the serial log on every boot, regardless of whether anyone is debugging the USB stack. Converted every "xhci:"-prefixed line to log_message() with a level chosen by what it reports, not blanket debug: - LOG_ERROR: allocation/mapping failures, timeouts, command failures, CSW signature/tag mismatches, CSW FAILED/PHASE ERROR, "not implemented" refusals, every "deferred ... setup failed" path - LOG_WARN: dropped/skipped conditions (command ring busy, tracked-port range exceeded), unrecognized media (bad version/CRC), TUR retry - LOG_DEBUG: routine progress (command submitted, succeeded, transfer completed, port connected) and expected outcomes (recognized/blank media) Default log level is LOG_INFO, so the LOG_DEBUG chatter that was the actual complaint is now silent by default and re-enabled with --log-level=debug; LOG_ERROR/LOG_WARN stay visible so real faults aren't buried. xhci_log_hex32() now routes through log_message(LOG_DEBUG, ...) instead of console_puts()/console_println() directly -- kept its own zero-padded 8-digit hex formatting rather than switching to log_message()'s %x (which has no width control), since register values lining up in the log is the reason this helper exists. console.h dropped from xhci.c, no longer used directly. Verified functionally unchanged, not just "still boots": all three architectures reach zuse)ok>, log_message()-instrumented lines are gone from the default-level log (grep -c xhci == 0 on all three, versus dozens before), and the USB thumbdrive path still works end to end -- "Zuse: identity confirmed from attached thumbdrive" appears on all three boots exactly as before. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019YcT3H2PQeyujrzjqS3Var |
||
|
|
9b6de5d6c7 |
riscv64: PLIC base address DTB-discovered, QEMU-virt constant as fallback (§V.3 item 3)
plic_init() now takes boot_info->dtb (threaded through apic_init()) and tries fdt_find_node_by_compatible(dtb, "sifive,plic-1.0.0") -> fdt_find_prop_in_node(..., "reg", ...) before falling back to the QEMU-virt-specific constant it previously hardcoded unconditionally. Reuses the node-scoped DTB lookup primitive built for the aarch64 GIC base fix unchanged. s_plic_base is now a runtime uintptr_t, same shape as apic.c's s_gicd_base/s_gicc_base. This system's QEMU/UEFI riscv64 firmware does not forward a DTB to the guest (timer.c's own timebase-frequency read falls back too, confirmed in this boot's own log), so only the no-DTB fallback branch is exercised here -- the success branch (a real DTB with a matching PLIC node) stays unverified until real Milk-V Mars hardware. FABRIC-3.md's first-drafted claim that the success branch would run (based on a stale comment in plic.c's own pre-fix header) was checked against the actual log and corrected before this commit. 3-arch acceptance: amd64/aarch64 don't compile these files, so their runs are non-regression on untouched files only. riscv64's own boot log confirms the fallback path prints exactly as designed and boot reaches zuse)ok> unchanged. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019YcT3H2PQeyujrzjqS3Var |
||
|
|
90ee8deb6d |
riscv64: skip satp Bare-mode switch when already Bare (§V.3 item 7 audit fix)
arch_early_init() unconditionally cleared satp on every boot, justified only by behavior observed under QEMU/EDK2 firmware (satp.MODE=10/Sv57, kernel identity-mapped within it). That reasoning never applied to the native U-Boot+OpenSBI boot path, where satp is conventionally already 0 at S-mode handoff -- the unconditional clear was likely a harmless no-op there, but on an unverified assumption. Fix: read satp.MODE first and return early when it's already 0 (nothing to switch away from, no safety argument needed). The unconditional csrw/sfence pair still runs unchanged for the confirmed QEMU/EDK2 case. No Sv39/Sv48/Sv57 page-table walker built -- out of proportion to this finding's severity. 3-arch acceptance: amd64/aarch64 don't compile this file, so their runs are non-regression on untouched files only. riscv64's own boot log confirms satp.MODE = 0xa at entry, so the mode != 0 branch ran and "satp cleared -- Bare mode, explicit" printed exactly as before the fix. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019YcT3H2PQeyujrzjqS3Var |