Root cause of the remaining "identity attach doesn't complete when Zuse
never attaches this boot" issue: capsule_wirebind_verify_cert() gated on
mama_vm->zuse_cert_installed, which is only ever set when Zuse's own
drive attaches and authenticates this specific boot
(capsule_zuse_boot_try_attach() -> install_and_activate() ->
vm_zuse_cert_install()). Without her, any other identity's WIREBIND cert
verification silently refused -- correctly, by the old design, but that
design conflated two genuinely different things: "can mint new
identities" (needs Zuse's live private seed, a real privileged
operation) and "can verify an existing identity's cert" (needs nothing
but her already-public key).
That public key was already being persisted independently of her live
session: zuse_genesis_marker_t (zuse_genesis_marker.h) stores it in the
kernel's own top-of-device metadata fence (Artemis's resident storage),
written once at genesis, specifically *not* alongside her private seed
(which stays only on her own removable thumbdrive) -- the type's own doc
comment says as much. It just wasn't being loaded for anything but
confirming which drive is genuinely hers.
Fix: a new capsule_zuse_boot_load_root_pubkey() (capsule_zuse_boot.c)
reads that marker and populates two new VM fields, zuse_root_pubkey_known
/ zuse_root_pubkey (vm.h) -- deliberately separate from
zuse_cert_installed/zuse_cert_seed/zuse_cert_pubkey, which stay
untouched and still gate MINT exactly as before. Called once from
kernel_main.c as soon as Artemis's own storage attaches, unconditionally,
independent of whether Zuse's own drive is ever attached this boot.
capsule_wirebind_verify_cert()/capsule_wirebind_try_attach() now check
zuse_root_pubkey_known instead of zuse_cert_installed.
One identity's attach must not depend on another identity's live
presence -- each identity stands on its own once the fleet's root of
trust has been established once, ever.
Verified live, amd64: identity 00 (disk/thumbdrives/00-thumb-ident.img)
now attaches and completes WIREBIND in 19 seconds with Zuse's own drive
never attached this boot at all (previously: unbounded, many real
minutes or effectively never, before today's other fixes; still slow/
stuck after those, stuck specifically on this silent refusal). Zuse's
own attach flow re-verified unaffected (regression check, amd64).
Three-arch clean qemu acceptance (amd64/aarch64/riscv64) passed with
this change included.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014Ec88YKxxhZGG1RNnune78
All found by actually running the identity workflow §VII/§VIII made
possible, not by code review:
1. Zuse/WIREBIND cross-contamination on detach: capsule_zuse_boot_logout()
and capsule_wirebind_unclean_detach() both had no device parameter, so
an unrelated device detaching (while the real owner's own stayed
attached) incorrectly tore down the wrong session. Both now compare
the departing device against their own tracked one, mirroring
capsule_wirebind.c's pre-existing g_wirebind_attached_dev precedent.
2. Dictionary-entry memory leak: vm_create_word()'s sf_malloc()'d
DictEntry (plus a second per-entry allocation for transition_metrics)
was never freed by vm_cleanup(), in both the hosted and kernel
implementations. Caused a real kernel PANIC after 8-9 repeated VM
birth/kill cycles in one boot. Fixed by walking vm->latest in both.
3. sf_malloc/sf_free (alloc_kernel.c) was a 4MB bump arena with a
deliberate no-op free, sized on "VM born once, never killed" -- fix#2
alone didn't stop the panic because free() itself discarded the
pointer regardless. Given a real free list (first-fit reuse).
4. Headless-console gate didn't re-engage after a mid-boot logout: the
original fix (sk_console_mark_login(), one-way sticky) only gated the
first login of the boot. Replaced with a live check
(sk_console_identity_present()) re-evaluated continuously, including
inside sk_console_readline()'s own blocking idle loop -- the console
is normally sitting blocked there when a hot-unplug logout happens, so
checking only at the top of the REPL loop wasn't enough.
Also: MINT now verifies its own write (verify_mint(), capsule_mint.c) by
reading back through the same check a real attach performs, rather than
trusting blkio_write()'s BLK_OK alone -- logged via log_message(), not
console_println(), per direct instruction.
Verified live, amd64: the full 8-identity repeated attach/detach cycle
that previously panicked at the same point every time now completes
clean, and a full serial-log sweep found zero bare unauthenticated
prompts anywhere in the run. Three-arch clean-qemu acceptance passed.
Still open, not fixed here: a 3+-simultaneous-device USB enumeration
failure found in a separate live test, not yet root-caused.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018EjXFo7mPXjUMjfJeuUUz4
Per-slot registry (xhci_msc_slot_t/dev->msc_slots, sized off the
controller's own reported max_slots) replaces the single-device scalar
fields the driver carried since Milestones 2e-2h. Boot-time port scan no
longer stops at the first connected device; a connect/disconnect that
arrives while the Command Ring is busy is now queued and drained instead
of dropped. blkio_usb.c and repl.c's own single-device state (device
descriptor buffers, blkio_dev_t, attach bookkeeping) became per-slot
registries the same way.
Live multi-device testing (not just compiling) surfaced a second, more
severe bug outside the original plan: transfer_purpose and next_action
were also single scalars shared across the whole controller. Two devices
enumerating concurrently could have one's completion silently overwrite
the other's still-outstanding one, permanently stalling it with no error.
Fixed by moving both per-slot and, critically, reading the Transfer Event
TRB's own real Slot ID field instead of trusting external bookkeeping.
Verified live, all three architectures, mandatory clean-qemu acceptance:
existing single-device path unchanged, and two devices attached
simultaneously (amd64) both progress independently through enumeration
without corrupting or stalling each other.
Also in this pass (implemented and verified in earlier turns this
session, committed together per direct instruction):
- Headless-until-login console policy: no prompt/banner until a real
identity logs in via an attached thumbdrive (WIREBIND or Zuse, neither
special), reusing EMERGENCY_CONSOLE_ENABLED as the debug/recovery
escape hatch (now default-off).
- KILL/g_repl_active_vm dangling-pointer fix: killing the VM the console
is currently USE'd onto now detaches back to Hera first, matching the
existing EJECT/UNCLEAN precedent.
FABRIC-3.md §VII/§VIII carry full closure notes for all three.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018EjXFo7mPXjUMjfJeuUUz4
FABRIC.md -> FABRIC-0.md
FABRIC-2.md -> FABRIC-1.md
FABRIC-3.md -> FABRIC-2.md (the current/living document)
FABRIC-4.md unchanged (new #3 to follow separately)
Every cross-reference repo-wide updated to match, including doc-comment
citations inside kernel source (.c/.h) files -- done via an ordered
placeholder substitution (FABRIC-3.md->placeholder2, FABRIC-2.md->
placeholder1, FABRIC.md->placeholder0, then placeholders resolved to
final names) in a single pass per file to avoid double-shifting
already-renamed references.
One line in capsules/font.4th grew past the 64-char block-format limit
as a side effect of the longer filename; shortened it and reverified
with mkcapsule --lint (34/34 pass) before rebuilding.
Verified 3-arch boot to ok> (amd64/aarch64/riscv64, each in the
foreground) after the fix; logs and DoE CSVs from this session's
verification runs included per this repo's own audit-artifact
convention.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019YcT3H2PQeyujrzjqS3Var
Rewired stadium_birth_hera() to admit unpinned then register through
session_register()/session_set_pinned() instead of setting
STADIUM_FLAG_PIN directly on the candidate header. Self-referential
parent (vm_uuid_hera(), vm_uuid_hera()), matching capsule_run.h's
parent_vm_id == vm_id root convention. Soft-fail, non-fatal, if
session_register() fails -- Hera's actual Stadium admission is what the
patron-zero invariant is about. Wired session_boot_init() into
kernel_main.c right after stadium_boot_init(), before stadium_birth_hera().
Also converted §H.12's punch list from bold "DONE" markers to this
document's established - [ ]/- [x] checkbox convention (already used
throughout §A), for consistency.
Verified 3-arch boot to ok> (amd64/aarch64/riscv64), no soft-fail message
on any arch, Hermes/Artemis births unaffected.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The QEMU-verifiable slice of the real-hardware RNG driver (per FABRIC-3.md
§G.2). New include/starkernel/rng.h + src/starkernel/rng/rng.c provide the
single entropy entry point: rng_init() probes the backend set (v2.0.0:
virtio-rng only) and, on no backend, prints a loud boot-time warning while
rng_get_bytes() returns RNG_ERR_NO_BACKEND - never silently degrading to a
deterministic seed. The backend-selection switch in rng.c is the exact seam
v2.5.0's per-arch drivers (amd64 RDRAND, riscv64 Zkr, aarch64 peripheral) plug
into without touching the call path.
Consumers route through the unified layer instead of virtio-rng directly:
capsule_mint.c (identity seed + drive_uuid) and kernel_main.c phase 8
(rng_init()). virtio_rng.c stays as the sole backend. Built clean on
amd64/aarch64/riscv64. QEMU amd64 boot: POST 1012/0/0 + ok>, "rng: backend =
virtio-rng" + "entropy: ready", Zuse identity confirmed from thumbdrive -
mint/cert behavior unchanged.
FABRIC-3.md §G.2 v2.0.0 slice marked BUILT+VERIFIED.
Code review fixes, all compile clean (hosted gcc + aarch64/riscv64 kernel flags):
- repl.c (H1): reentrancy guards on the MSG-TICK idle pump. sk_repl_idle()
now defers when Hera is mid-interpret (g_mama_interpreting) or when its
own vm_interpret is on the stack (g_idle_pump_active), so a blocking
KEY/EXPECT/QUERY inside a dispatched line can no longer re-enter the
interpreter and clobber the in-flight input buffer.
- virtio_rng.c: clamp device-returned used_len to VRNG_BUF_SIZE before the
caller's data_buf copy, closing a device-controlled OOB read.
- block_subsystem.c: first-write path now keys off created_time==0 instead
of dead magic==0 so fresh blocks get a real created_time stamp; first_free/
last_allocated fixed to absolute Forth LBNs (set in blk_compute_fresh_geometry
from slot->start_lbn, no longer the wrong physical-BAM-index values from
compute_totals_from_B); physical-bounds guard on blk_meta_zone_read/write
prevents unsigned underflow on a corrupt fence >= device size.
- capsule_zuse_boot.c / capsule_wirebind.c: identity seed validated magic ->
version -> CRC-64 (compute_crc64 over offsetof(crc)) before trusting it,
so a corrupt/format-mismatched record is refused, never loaded.
- log.h / starkernel/log.h: unused LOG_LINE_MAX 256 renamed LOG_MSG_LINE_MAX
to lift the include-order collision with vm.h's LOG_LINE_MAX 64; stale
include-order comments dropped (kernel_main.c, shim.c, capsule_birth.c).
- FABRIC-3.md: three stale-doc carry-forward items closed [x] with cbe7b49
notes.
Real KEY/?TERMINAL/QUERY/EXPECT bodies (console WIP):
- repl.h/repl.c: sk_console_getkey()/sk_console_key_available()/
sk_console_readline() public bodies; non-destructive peek buffers the
found byte so a following KEY returns it.
- shim.c: getchar()/fgetc()/fgets()/sf_terminal_ready() routed through the
real console paths instead of stubs; sf_terminal_ready() in platform_io.h
with sf_terminal_ready() implemented for the hosted build (linux/io.c,
POSIX select on fd 0) wired into Makefile.
- io_words.c: ?TERMINAL now returns actual terminal-readiness, not constant 0.
Artifacts: minted disk/artemis.img + rebuilt lfs kernel; BLOCK_MAP.md,
doe csv + qemu log regenerated.
Three tightly-coupled changes, verified together per Captain Bob's own
"getting rid of the emergency cli" direction:
1. Zuse's identity is thumbdrive-resident, never system-resident. New
zuse_genesis_marker_t (magic/version/zuse_pubkey[32]/crc) replaces
zuse_cert_devblock_t's slot in the top-of-device fence -- the system
now remembers only that a root identity exists and its pubkey, never
a seed. zuse_cert_devblock_t is kept in the repo, marked superseded,
no longer written by any code path.
capsule_mint_identity() grows a genesis mode (issuer_vm=NULL): no
cert is built or written (Zuse isn't verified against a separate
signer -- she's recognized by pubkey match against the marker) and
two new optional out-params (out_pubkey/out_seed) let the caller
install the cert immediately after a genesis mint.
New capsule_zuse_boot_try_attach() (capsule_zuse_boot.c), called
from sk_repl_idle() on every fresh USB attach (the only point in the
boot lifecycle a thumbdrive can actually be detected -- attach
polling doesn't exist yet at kernel_main.c's old one-shot mint point,
which is why that whole block is gone): no marker + blank drive ->
genesis-mint; marker present + matching drive -> read its own
user_identity_seed_t, install the cert. Either way, re-runs
ACL-ZUSE-BOOT (zuse.4th) so zuse_session activates exactly like it
always has for a same-boot cert install -- ACL-PIN only blocks
redefinition, not re-execution, so no new C-side auth logic needed.
2. ACL.4th activated (capsules/init.4th) -- inactive all session until
now. Found and fixed a real bug this immediately surfaced: zuse.4th's
ACL-ZUSE-BOOT tried `['] ACL-ZUSE-BOOT ACL-PIN` from inside its own
still-compiling definition -- the word isn't findable yet at that
point, so the whole definition silently failed to compile every
previous boot this session (dormant, since ACL.4th never loaded).
Fixed: pin after the definition closes, not from within it -- it
only needs to happen once anyway, and pinning doesn't block the
re-invocation genesis/attach needs.
3. The unauthenticated emergency-CLI ACL bypass is retired
(repl.c): `emergency_console = is_hera ? (zuse_session ? 0 : 1) : 0`
deleted from both sk_repl_step and sk_repl_run. Every word run from
Hera's own bare prompt now goes through ordinary ACL enforcement;
emergency_console is driven only by the genuine C-level fault
handler again.
Added ZUSE-SESSION? (starforth_words.c), a read-only diagnostic
matching ZUSE-PUBKEY@'s own precedent, to verify the whole chain
directly rather than by inference.
Verified end-to-end live in QEMU: fresh boot, no thumbdrive ->
ZUSE-SESSION? reads 0. Attach a genuinely blank drive via QMP -> genesis
mint fires automatically (no typing) -> ZUSE-SESSION? reads -1 (true).
Hermes/Artemis both birth clean on all three architectures with ACL
now actually enforced for the first time all session -- no denials, no
UNKNOWN WORD beyond the deliberate POST self-test cases.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019ZGkimpfyh63EZyRkNbkPD
Extract the messaging vocabulary (arenas, MSG-*/CH-*/MBR-* words) out of
capsules/hermes/init.4th into a new shared capsules/common/messaging.4th
that Hermes and Artemis each load at birth, giving every VM its own
private MSG-ARENA/CH-ARENA instead of only Hermes having one. Hermes
stays the owner of the one real, canonical COMMON-CH; Artemis subscribes
into it via VM-EXEC at her own birth, and Hermes proactively subscribes
Hera (idx 0) since Hera always exists first.
Hera does NOT get her own copy: register_child_vm_words()'s own doc
comment explains why the STADIUM-* primitives common:messaging.4th
depends on are deliberately never registered in her dictionary (keeps
her dict_hash off item 4.1's baseline). Confirmed live by loading it
into her dictionary anyway first -- every colon-definition referencing
an unregistered Stadium primitive was silently dropped (MSG-HEAT@/!,
CH-HEAT@/!, MSG-COOL-ALL, MSG-TICK all missing after boot). Reverted
that path; she orchestrates via BIRTH/VM-EXEC/VM-CALL instead.
Added capsule_vm_registry_get_by_index() (capsule_birth.c/.h) for
registry enumeration by birth-order position, and a pump in repl.c's
existing idle hook that walks every live VM once per idle beat and
VM-EXECs MSG-TICK into each one except Hera's own entry.
Verified clean (no UNKNOWN WORD / VM-EXEC errors after birth) on all
three architectures.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019ZGkimpfyh63EZyRkNbkPD
Per FABRIC-3.md D.7 (birth-by-message-only): the Tripod legs must be alive
session-less before any thumbdrive-attach flow has a running Hermes/Artemis
to message. kernel_main.c's item 4.2/4.6 self-tests previously birthed
both, exercised diagnostics, then explicitly KILLed them before ok> every
boot -- production boot never actually kept either alive. Stripped both
blocks down to birth-only, diagnostics and KILL removed.
Found and fixed a real regression this surfaced, not left broken: Hermes's
CD-INIT (message/channel arena init, and the thing that loads lib.4th into
her own dictionary) was only ever invoked by the self-test just removed --
nothing in hermes/init.4th itself called it. Added an unconditional CD-INIT
call at the end of her own init.4th so a real birth actually initializes
her arena, matching how Artemis's own ART-BOOT-ENTRY already runs
unconditionally at her own capsule load.
Added STARTUP-BANNER (lib.4th): a shared word reading a common
IDENTITY-BANNER buffer, printing "(no identity)" until CERTVERIFY/MINT
exist to populate it from a thumbdrive's PKI fields -- scaffolding only,
per direct instruction. Wired into each Tripod leg's own init.4th
(Hera/Hermes/Artemis); Artemis didn't load lib.4th before, added that too.
Verified live on all three architectures: BIRTH for both, no KILL anywhere
in any log, and an interactive USE round-trip confirmed Hermes is
genuinely reachable post-boot, not just logged as born.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019ZGkimpfyh63EZyRkNbkPD
Replaces the crashed NVRAM approach entirely. New
include/starkernel/zuse_cert_devblock.h: a standalone on-disk record
(magic + version + 32-byte seed + 32-byte pubkey + a real CRC-64/ISO
from day one, same discipline homeblocks_sig_t established) occupying
devblock_from_top=0 of the fence. Its own header, not inlined at the
boot call site, since the still-open MINT word will be a second
consumer of this exact format.
kernel_main.c's mint-or-load logic now reads the fence, installs an
existing valid cert, or mints fresh via virtio_rng+ed25519_keygen and
writes it. Runs right after virtio_rng_init(), before
capsule_birth_mama() -- unlike the crashed NVRAM attempt, raw block I/O
against Artemis's already-proven device has no boot-timing risk, so the
earlier "re-invoke ACL-ZUSE-BOOT after Mama birth" workaround is gone;
ACL.4th's self-activating ACL-ZUSE-BOOT sees a correct cert on its one
ordinary pass.
Verified independently across every real scenario, never trusting the
kernel's own report: fresh mint decodes correctly on disk with a CRC
confirmed by a from-scratch Python re-implementation of the algorithm;
a reboot without reformatting loads back byte-for-byte identical
seed/pubkey (genuinely "mint once, ever"); a pre-fence volume refuses
cleanly (no crash, no silent data loss, honest "not persistent"
reporting); the real, untouched disk/artemis.img exercises the same
graceful-refusal path identically on all three architectures.
Phase 8's core arc is now functionally complete: real entropy -> real
signing -> real anti-file block-native persistence -> a first-boot mint
that survives reboots. Still open: the ongoing MINT word for minting
additional regular users. Documented in FABRIC-3.md.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01U14ET9CWAtbQMbYqomKgXd
Cert storage expanded from the old 16-byte placeholder to a real
32-byte seed + 32-byte pubkey. vm_zuse_cert_install() now has a
kernel-side duplicate in src/starkernel/vm/vm_core.c -- the kernel
build's VM_EXCLUDE list drops src/vm.c entirely (same reason
vm_set_base() already has two independent copies), so the hosted-only
version added earlier this session was never actually linked into the
kernel. FORTH-side ZUSE-CERT-LO@/HI@ replaced with ZUSE-PUBKEY@ (i -- u)
over the public half only; ACL-ZUSE-BOOT now checks
ZUSE-CERT-INSTALLED? before authenticating instead of unconditionally.
Attempted NVRAM-based persistence (GetVariable/SetVariable) for the
first-boot mint flow: page-faulted inside OVMF's variable service
(CR2 in the flash MMIO window). Moving the call site to match the one
proven-safe existing SetVariable call site in this codebase produced
the identical crash -- not a timing issue. Localized with debug
markers (one boot): GetVariable works; SetVariable with real data
never returns. The existing "working" precedent call is actually a
delete-of-nonexistent-variable (size=0, data=NULL), a cheaper path
that never touches flash, so it proved nothing about real writes.
Root cause: this kernel's VMM never maps the region OVMF's variable
service needs for real flash writes -- a genuine gap in UEFI runtime-
services support, not Zuse-specific, and not obviously fixable in a
3-arch-uniform way (flash window location is firmware/arch-specific).
Independently, storing the raw seed in RUNTIME_ACCESS NVRAM would have
been a real security defect regardless of the crash -- readable by any
later-loaded UEFI app or the booted OS.
Reverted to a known-safe state: all NVRAM/mint code removed from
kernel_main.c, init.4th's ACL.4th line back to its documented
commented-out default. Verified clean compile and clean boot on all
three architectures. Cert storage expansion (the part that works)
stays. A dedicated system-identity disk (virtio-blk, already proven
for writes via Artemis) is the recommended next substrate -- not yet
decided or built. Full investigation documented in FABRIC-3.md.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01U14ET9CWAtbQMbYqomKgXd
The kernel's ed25519_verify() is deliberately verify-only -- no signing,
no keygen, no entropy source. That conflicts with the on-device MINT
word vision (Zuse signing new user certs live at runtime), so this
reopens that constraint on request rather than reshaping MINT around
verify-only.
vm_uuid.h already found the real gap: amd64 has RDRAND, riscv64 has Zkr,
but QEMU's aarch64 CPU models have neither -- confirmed against QEMU
10.2.1. A deterministic PRNG (fine for VM UUIDs) is not safe for key
generation, so this adds a virtio-rng device instead of a per-arch split:
real host entropy, identical guest-side protocol on all three arches.
New src/starkernel/virtio/virtio_rng.c + include/starkernel/virtio_rng.h,
transport plumbing mirroring the existing virtio_blk.c driver exactly.
Wired into kernel_main.c boot, -device virtio-rng-pci added to all three
QEMU targets.
Verified live (temp probe, written/run/captured/reverted): 16 real bytes
pulled through the full request/notify/poll round trip on all three
arches, three different values confirming real entropy. Final boot
against the reverted, permanent code: clean compile, clean boot to ok>
on amd64/aarch64/riscv64, Stadium conservation intact, no panics or
guest errors.
Ed25519 keygen/signing itself (Phase B) and the MINT word design
(Phase C) remain open, documented in FABRIC-3.md.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01U14ET9CWAtbQMbYqomKgXd
stadium_admit()'s mass==1 refusal looked like a hard blocker for 1024-byte
blocks, but stadium_word_dispatch()'s real candidate construction proves
Stadium cells carry pure identity/heat/bookkeeping, never the resident's
actual content -- a block patron follows the same shape (identity=LBN,
payload unused), so this was real, scoped work, not a case for stubbing.
New stadium_blocks.h/.c mirror stadium_words.c's admission/cooling shape,
keyed by (quota_slot, lbn) in a fixed-capacity open-addressing hash table
(tombstone deletion) instead of a dense array, since LBN space isn't
densely bounded like word_id. Wired into block_word_block()/buffer()/
update() (block_words.c), __STARKERNEL__-guarded. stadium_dispatch()'s
MIGRATE case now calls blk_flush(lbn) for real instead of printing
"(stub)". Three new Kconfig constants (STADIUM_BLOCK_HEAT_QUANTUM/
STADIUM_BLOCK_COOL_RATE_Q48/STADIUM_BLOCK_TRACK_CAP_MULT) mirror the
word-patron ones, same three-layer wiring.
VM-COOL/DELIVER/EXPIRE stay explicit punch-list items -- VM-COOL
deferred pending the still-iterating Tripod/Zuse/messaging vision,
DELIVER/EXPIRE are their own future subsystem integrations per
FABRIC.md's own "open, not resolved" notes.
Verified clean compile (zero warnings) and clean boot to REPL with
conservation intact (resident_sum + reservoir == Q48_ONE) on all three
architectures (amd64/aarch64/riscv64); BLOCK/BUFFER touches exercised
live from the REPL with no crash; a 22,000-distinct-block flood loop
against an artificially shrunk Stadium ran clean under heavy admission
load. A live MIGRATE console fire was not directly observed this
session (root-caused to a pre-existing reservoir-floor/density-eviction
interaction unrelated to this change, documented in FABRIC-3.md) --
flagged as an honest follow-up, not silently claimed.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CXjAPTEKrgY2Mrk25KoLDn
Implements Event Ring TRB parsing and ERDP dequeue-pointer update
(xhci_poll_events(), src/starkernel/usb/xhci.c), called from
sk_repl_idle()'s existing ~1s idle cadence rather than a per-arch
interrupt handler.
A first attempt wired real interrupt delivery (PCI->IOAPIC GSI routing,
a dedicated isr_stub34/vector 0x22, GIC/PLIC routing mirroring
virtio_input.c). Checked live via QMP query-pci before trusting it: the
amd64 PIRQ swizzle formula predicted GSI 16 for the xHCI controller at
PCI slot 4; the real QEMU-assigned IRQ was 10, and embedded ICH9
functions contradicted the same formula too. Reverted all of it back to
the exact committed baseline rather than chasing chipset PIRQ routing
further, and reframed around Section U item 6's own design intent
("interrupt-driven, coarse cadence, cheap early-exit... quick check
blocks... done") via sk_repl_idle() instead -- USB insertion is a
human-timescale event, not a hot path.
Added -device qemu-xhci to all three QEMU launch targets (required for
any of this to be testable). Verified end to end via genuine post-boot
hotplug (QMP device_add/device_del usb-storage): all three architectures
detect a live attach within seconds. A false-alarm heartbeat "freeze"
found mid-verification traced to querying the wrong counter
(vm->heartbeat.tick_count, which only advances during word execution,
not the kernel's real ISR-driven heartbeat_ticks()) -- confirmed via a
temporary diagnostic word, captured and reverted.
Full writeup, including the discarded interrupt-routing attempt and the
false-alarm investigation, in FABRIC-2.md's Milestone 2c/2d entries.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HZ8kNoTuP63pbQtro4qvrm
Continued investigating the aarch64 BYE cold-restart exception (FABRIC-2.md
Section I). Added permanent boot diagnostics: kmalloc_heap_base_addr()/
kmalloc_heap_end_addr() now print in print_heap_stats(), confirming the
fault address is provably inside the kmalloc heap (not kernel code, not
firmware). Bumped aarch64 QEMU RAM to 4096MB to test heap-placement
sensitivity (no effect -- heap size is a fixed 2GiB default, independent
of total RAM once "enough" exists).
Three separate live gdb debugging attempts (software breakpoint, hardware
breakpoint on arch_cold_reset, hardware breakpoint on mama_word_bye's
entry) all silently failed to fire despite disassembly-confirmed-correct
addresses and confirmed execution reaching those points. A sanity check
(hbreak on console_println, called thousands of times per boot) also never
fired even 8802 lines into a serial log -- conclusively a gdbstub/QEMU
tooling limitation for this aarch64 target, not a kernel-side finding.
Live single-stepping is not currently viable here; documented so it isn't
re-attempted the same way.
Root cause still open. Full trail in FABRIC-2.md Section I.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Artemis's 30-rep surface stress campaign was failing 100% of trials on all
three architectures: stadium_grant_quota() ran after IDENTITY exec in
capsule_birth.c, but Artemis's init.4th auto-runs the stress campaign as
part of that same IDENTITY exec, so every STADIUM-ADMIT call during it hit
a nonexistent quota slot and refused unconditionally. Moved the grant call
before IDENTITY exec. Verified 30/30 reps PASS on amd64, aarch64, and
riscv64 post-fix (was 30/30 FAIL on all three pre-fix).
Also fixed an independent, real bug found during the same acceptance pass:
aarch64's arch_cold_reset() issued PSCI SYSTEM_RESET using the SMC64
calling convention (0xC4000009), which is not a valid PSCI function ID --
SYSTEM_RESET has no SMC64 variant. Corrected to the SMC32 encoding
(0x84000009). This did not resolve the separate aarch64 BYE cold-restart
exception also found in this pass (root cause not yet found, tested and
refuted an interrupt-race hypothesis, documented in FABRIC-2.md Section I
for follow-up) but is a genuine spec fix worth keeping regardless.
Full writeup, evidence, and the still-open aarch64 crash investigation in
FABRIC-2.md Sections H and I.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Default log level dropped from info to warn so the per-word ECW dispatch
trace doesn't flood REPL output after POST (--log-level=info/debug still
re-enables it). qemu target gains QEMU_DISPLAY (default gtk) so the
framebuffer window shows by default; serial log is tee'd live via
`tail -f` instead of dumped with `cat` at the end. Includes regenerated
BLOCK_MAP.md/amd64.csv/artemis.img and this morning's boot logs/DoE runs
from the sessions that produced this WIP.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Moves the console_fb_init() call in kernel_main.c from after
capsule_birth_mama() to before it, so the fleet-birth/self-test
transcript (Hermes x2, Artemis births, Stadium self-tests -- currently
serial-only) is also framebuffer-visible, not just the small post-birth
tail.
4.5f's -O2 experiment already showed this doesn't hang under
optimization, just costs roughly 12x more boot-time heartbeat ticks
(one-shot, paid only during fleet birth, never repeated at runtime).
Captain Bob's call: worth it, since the serial log was never the
problem -- this is about the same transcript also reaching a real
screen.
Three-arch verified: amd64/aarch64/riscv64 all reach ok>, POST
Failed: 0, identical dict-hashes across all three. amd64 screendump
confirms the framebuffer now carries the full transcript.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
fb_draw_orientation_test() was a one-time diagnostic for 4.3.1 (raw GOP
framebuffer wiring), already [x] done. With 4.4c's vt100 console now live,
the corner blocks just obscure real console output in every screendump.
Function definition left in framebuffer.c/framebuffer.h for future reuse;
only the boot-time call site is removed.
Also commits routine log/manifest housekeeping: capsules/BLOCK_MAP.md
(regenerated by mkcapsule on each build), the amd64 probe logs from before
today's outage, and the screendump logs + evidence screenshot from this
session's 4.4c/4.4d/4.4e verification.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
console_fb_init() (which calls vt100_init()) was never called anywhere in the
boot sequence -- only raw fb_init() ran, so vt100_putc() no-op'd on every
call and no console output (REPL, POST, boot logs) ever reached the
framebuffer, only the corner orientation-test blocks. Replace the raw
fb_init() call in kernel_main_deep() with console_fb_init() so the
framebuffer console actually comes up (4.4c).
emit_prefix() also wrote the "[VMName] " bracket text via raw_putc() only
(serial), never vt100_putc(), so the bracketed VM name never reached the
screen even once the framebuffer console was live. Mirror console_putc()'s
existing serial/framebuffer split there too (4.4d).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Punch list §25 item 4.3.5c complete.
Amended from a nonexistent MMIO transport to PCI (matching the board's
actual virtio-blk-pci precedent). New virtio-input driver: eventq with
pre-posted buffers, PLIC source computed at runtime from PCI slot/pin
(derived live from this host's QEMU riscv64 DTB), mandatory ISR-status
read, PCI interrupt-disable-bit check. New VKBD-EVENT/VKBD-DEBUG FORTH
words. Verified with a real QEMU sendkey keypress: exact KEY_A/press
match, two real interrupts serviced, zero exceptions. Three-arch
acceptance boot clean.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Punch list §25 item 4.3.5 complete.
New ioapic.c/i8042.c drivers (MADT-derived I/O APIC base, no hardcoded
constants) plus a KBD-SCAN/KBD-DEBUG diagnostic word pair. Three real
bugs found and fixed en route, all blocking this item's own acceptance:
a fatal LAPIC spurious-vector crash (nothing had driven a real external
interrupt through the I/O APIC before), OVMF leaving the keyboard device
itself scanning-disabled (0xF4 fix), and isr.S's stub table only having
individually-numbered stubs through vector 32 -- everything above that,
including our IRQ1 vector 33, silently reported as vector 255 regardless
of which IDT slot actually fired. Verified live via QEMU sendkey against
KBD-SCAN: correct XT Set-1 make/break codes for two different keys.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Adds fb_draw_orientation_test() (framebuffer.c/.h): fills the four raster
corners RED/GREEN/BLUE/YELLOW via fb_fill_rect. Wired into kernel_main.c
calling fb_init() directly -- console_fb_init()/vt100_init() removed from
the boot path, since vt100.c/console.c are superseded by the Console
drawing-fabric redesign (FABRIC.md ss27) and should not be exercised even
incidentally.
The diagnostic caught a real, pre-existing bug on its first run: framebuffer.c's
pack_pixel() had its FB_PIXEL_RGBX32/FB_PIXEL_BGRX32 branches swapped relative
to UEFI GOP's own byte-order naming convention, producing a clean R<->B channel
swap (G unaffected). Spatial placement was already correct -- no flip/rotation.
Fixed by swapping pack_pixel's two return bodies to match framebuffer.h's
already-correct doc comments; kernel_main.c's GOP-format switch needed no change.
Also item 4.3.2 -- QEMU screenshot capability. scripts/qemu_screenshot.sh
already existed (monitor socket + socat + HMP screendump), just unwired and
unused this session. Redirected its PNG output to a new top-level fb/
directory (tracked in git, not logs/, not a gitignored temp dir) and added a
python3+PIL fallback for PPM->PNG conversion since imagemagick isn't
installed here. Left as a standalone script for now, not wired into a
Makefile target.
FABRIC.md items 4.3.1 and 4.3.2 marked done with acceptance evidence.
Migrates Hermes's message/channel lifecycle onto the Stadium's unified
heat/capacity economy: MSG-ALLOC/FREE-NODE and CH-ALLOC/FREE-NODE now
route entirely through stadium_admit()/stadium_evict(), replacing the
old local free-list + independent heat-field mechanism. Eight
kernel-only STADIUM-* FORTH primitives (ADMIT, EVICT, RES@, RES-PULL,
RES-PUSH, HEAT@, HEAT!, WORD-HEAT), VM.stadium_vm_id threaded through
all three vm_core.c dispatch sites (replacing item 4.1's hardcoded
vm_uuid_hera()), and the stadium_owner[idx] fix so evict-credit lands
in the VM that actually admitted a patron, not whoever owned cell 0.
This session's own contribution, on top of that pre-existing
implementation: found and fixed two bugs blocking the item's own K≡1.0
conservation self-check (HERMES-K was reading 0, not 65536):
- Q.SLOT admission-heat fix (capsules/hermes/init.4th): MSG-SEND/
CH-ACCEPT admitted with Q.1 (the entire fleet-wide "1.0" unit) per
item, a leftover from before the Stadium migration when each
message/channel had its own unconstrained heat field. Instantly
drained the shared, finite reservoir.
- Reservoir floor for word-execution admission (stadium_words.c):
stadium_word_dispatch() (item 4.1) pulls STADIUM_WORD_HEAT_QUANTUM on
every word dispatch, not just first admission -- exhausts a VM's
entire reservoir in ~32 dispatches, starving any application-level
economy sharing that VM's reservoir before it gets a chance to pull
anything. word_dispatch_pull() now clamps word-execution's own pulls
to leave a Q48_ONE/3 floor (same fair-share figure COMMON-CH's own
floor already uses); application-level pulls are unaffected.
- STADIUM-WORD-HEAT primitive + stadium_words_resident_heat(): the
floor deliberately leaves word-execution residents holding real
heat, invisible to HERMES-K's original formula (MSG+CH+reservoir,
no term for word patrons). Adding this term closes K to exactly
65536 on all three architectures.
Also rules on two open scope questions in FABRIC.md: MBR-ALLOC/
MBR-FREE-NODE stay off the Stadium (membership records have no heat
field, never did -- the acceptance bullet's inclusion of them was a
completeness gesture predating a check of the actual layout), and
records the effort number (12 implementation files, +759/-120 lines).
Verified: all three architectures boot clean, full self-test passes,
Stadium conservation closes exactly (resident_sum + reservoir =
Q48_ONE) at both the C/Stadium level and the FORTH-level HERMES-K
check.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Punch list §25 item 4.1a complete.
New prerequisite item, found while scoping 4.2: no quota-granting mechanism
existed at all. Adds stadium_grant_quota(new_vm_id, from_vm_id) -- a
one-time initial grant at birth, distinct from item 1.3's still-unbuilt
recurring capacity-transfer arbitration. Splits the donor's free list evenly
by cell count, reassigns stadium_owner[] for every moved cell, and grants
the new VM a fresh Q48_ONE reservoir (not a split of the donor's -- per-VM
conservation, same pattern as Hera's own boot grant). Wired into every baby
VM's birth in capsule_birth.c.
Verified via a boot-time self-test in kernel_main.c using a synthetic
identity (not the real UUID pool, not a real capsule birth -- item 0.1's
Hera-alone pruning stays intact). All three architectures booted to ok> with
identical output: grant OK, Hera reservoir=0 (already fully committed to
resident words, correctly unchanged), test-vm reservoir=65536 (fresh
Q48_ONE). dict_hash identical across all three and unchanged from item 4.1's
baseline (0x3d4e1daf289da94f) -- confirms no dictionary word was added.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Punch list §25 item 4.1 complete.
Replaces the round-robin hotwords cache with Stadium density-ranked
admission/eviction on the kernel side, via the §17.7 reservoir mechanism and a
kernel-side word_id -> cell_index map (no DictEntry change, dict_hash
untouched). Adds stadium_birth_hera() to close the cell-0 panic hazard,
STADIUM_WORD_HEAT_QUANTUM/STADIUM_WORD_COOL_RATE_Q48 Kconfig knobs (flagged
untuned), and a stadium_word_forget() FORGET coherence hook to close a
recycled-word_id aliasing gap.
Verified: all five hotwords_cache_* call sites in dictionary_management.c
bypassed under __STARKERNEL__; word dispatch feeds the Stadium at all three
vm_core.c physics_execution_heat_increment() sites; hosted make unaffected;
all three architectures booted to ok> with matching dict_hash
(0x3d4e1daf289da94f) and matching conservation stats (promotions=354
evictions=0, resident_sum=65536 reservoir=0 sum=65536).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Punch list §25 item 3.2 complete.
stadium_boot_init() (src/starkernel/vm/stadium.c) sizes the global
cell array at boot from a real memory-budget query rather than a
hardcoded count: pmm_get_stats().free_bytes at the point of
allocation, times the new STADIUM_MEMORY_PERCENT Kconfig symbol
(default 1%), rounded down to whole 64-byte cells. Matches §17.6's
position (b) literally. Also allocates the header/continuation
discriminator bitmap item 3.1 declared but did not allocate. Both are
kmalloc'd and explicitly zero-filled (kmalloc does not zero).
Called from kernel_main.c immediately before sk_vm_bootstrap_parity(),
i.e. before any VM exists (§6). Failure is soft -- logs and continues,
does not halt boot -- matching the existing precedent one line below
it (VM bootstrap parity failure does the same).
Added a "Stadium: N cells (M KB)" boot console line at the allocation
site so the acceptance logs are evidence the array was actually
allocated, not just that the kernel still boots -- the same blind spot
item 3.1's uncompiled-header gap exposed.
Verified: three-architecture boot (amd64, aarch64, riscv64), all
reaching ok> with identical dict_hash=0x3d4e1daf289da94f matching the
item-3.1 baseline, and the Stadium boot line confirmed present in all
three serial logs (amd64: 74234 cells/4639 KB, aarch64: 161329
cells/10083 KB, riscv64: 76122 cells/4757 KB).
Not built here, reported per §25.0 rule 3: per-VM free lists (§22.3)
-- granted when Hera assigns quota, not this item's scope.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Adds a one-time boot diagnostic in kernel_main.c, right before sk_repl()
is entered: bounded wait for 3 real heartbeat ticks, then prints tick
count, TIME-TRUST, and variance. Needed because printing immediately
after apic_timer_start() (as first tried) measured 1 tick on amd64 and
0 on riscv64 -- not evidence the heartbeat doesn't work, just that
almost no wall time elapses between arming the timer and that point in
boot; report it honestly rather than let it stand as a false negative.
Verified this session (logs/20260804-001727, -001805, -001850,
-001948, -002021):
- All three architectures boot to ok>.
- Tick count non-zero: amd64 4, riscv64 3, aarch64 3.
- riscv64: trust=Q48_ONE exactly, variance=0 -- architecturally
invariant counter, as designed.
- amd64: dict_hash=0x3d4e1daf289da94f, identical to the pre-item-0.8
baseline (logs/20260803-231322) -- unchanged output, satisfying the
GAP-A1 control.
- Two consecutive amd64 boots produced the identical dict hash --
reproducible, no wall-clock leakage into patron state.
Phase 0 (Substrate) is complete.
Punch list §25 item 0.10 complete.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
kernel_main on riscv64 ran directly on EDK2's UEFI boot-time stack, with
no dedicated stack switch — amd64 has always had a kernel_entry.S
trampoline for exactly this reason (its own comment: "the FORTH
interpreter + DOE experiment loop can easily exceed that depth").
aarch64 happens to get away without one because its firmware's default
stack is apparently larger, but that was never a guarantee.
On riscv64 the VM bootstrap's call depth (27 word-registration modules
-> physics/SSM init -> Tripod capsule birth) overflowed that small
stack, corrupting a return address and producing a wild jump / page
fault right after vm_init_with_host() returned — reproduced consistently
across the 2026-08-01 DoE campaign logs.
- src/starkernel/arch/riscv64/kernel_entry.S (new): RISC-V stack-switch
trampoline mirroring amd64's, giving the kernel a dedicated 2 MiB BSS
stack before anything deep runs.
- kernel_main.c: riscv64 now builds kernel_main_impl (invoked via the
trampoline) instead of kernel_main directly, same pattern as amd64.
- Makefile.starkernel: wires the new file into the riscv64 build.
- uefi_loader.c: RAW_LOG() was silently a no-op on every non-amd64 arch;
added a real raw-UART writer for riscv64 (QEMU virt's uart8250 at MMIO
0x10000000) so existing loader diagnostics actually produce output.
Verified: all three architectures boot clean to [Hera] ok> in the
required order (amd64, aarch64, riscv64); logs and DoE CSVs from these
runs included.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>