d0c4f60194c4e57bdc2197fb6e93705bb592f7f7
661
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
d0c4f60194 |
FABRIC-3.5.md §XXIII: the sinking latch settled -- design phase complete
The last open design question. The signal is a plain scalar in kernel-Hermes's own translation unit, written only by Hermes, read by every VM through a registered primitive. That invents nothing: it is the same registration path §XX already established when register_child_vm_ words() hands the eight STADIUM-* primitives to every VM and Hera's two sites mirror them. Reading a scalar needs no queue, channel, routing table or working arbiter, which is precisely the requirement -- the transport is the thing that failed. Names it accurately. "Semaphore" was the original framing but the mechanism is a one-way latch, and the distinction carries correctness. Per §VII.1 a VM attempts its own recovery before announcing anything, so raising the latch means recovery already failed and there is no path back to not-sinking. Single writer, many readers, one irreversible transition, both observable values valid -- which is the discipline heartbeat.c:33-34 already states verbatim and §XXVIII Stage 3 already reused once rather than inventing new locking. This is the third such variable and takes the same route for the same reason. Calling it a semaphore in code would imply counting and blocking semantics it must not acquire. Settles who reads it, which is the part §XXVIII's shared C stack constrains: a VM has no thread and cannot poll. So the check happens at the dispatch point, and the split is the design -- the kernel delivers the fact, the VM runs its own flush-and-BYE in its own context. The kernel publishes, never decides, so §XV.3 and §XIV.5 hold, and each VM running its own shutdown is what §IX.1 already ruled. The window needs no deadline because Hera's halt is physically terminal. Completes §VIII.2's two-mechanism rule with no exceptions list: floor VMs send SOS through routing, the arbiter beneath raises the latch, and which one applies is decided entirely by which side of the routing layer the sender sits on. Every design question this document opened is now ruled. What remains is build sequencing plus two small local decisions, items 17 and 18. No code is authorized. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VkM1zHGvBerLF6aqkHPweP |
||
|
|
3f65a86993 |
FABRIC-3.5.md §XXII: the surgical strip is punch item 1, and it is iterative
Captain Bob, 2026-09-19: strip the anticipated-dead FORTH before coding; the real gaps reveal themselves as we code/test iterate, and further dead FORTH falls out in the same loop. Re-orders the punch list to put the strip first. Draws the ordering distinction that protects it. Category A is dead today independent of this reshuffle and can be stripped immediately. Category B -- hermes/init.4th, most of messaging.4th, the routing table, the slot-3 pairing -- is dead only on arrival of kernel-Hermes and is load-bearing until the replacement boots. "Strip before we code" must not be read as deleting the working messaging layer with nothing behind it. Records the trap this document nearly walked into while building the inventory. A reference count over capsules/ and src/ put the six init-l8-* personality capsules and sdk.4th at zero references. They are not dead -- including experiments/, tools/ and docs/ finds 18-19 hits each. Reachability cannot be established by grep in either direction: it under-reports because capsules are birthed by name from runtime strings, and over-reports because a capsule named for a common word returns prose noise. So no strip list is authoritative, including any this document produces; reachability is per capsule via boot path, tooling invocation, or the baked capsule directory. Names the asymmetry that makes "surgical" the right word. The three-arch boot is the only acceptance test and it exercises the boot path, so a bad Category A strip fails loudly. It does not exercise DoE capsules, so deleting one passes all three architectures cleanly and surfaces months later during a campaign -- and that infrastructure has an evidentiary role, since CLAUDE.md records the logs as committed audit artifacts and the campaign report as patent support material. Proposes treating every experiments-reachable capsule as live by default. Sets the loop: strip Category A in isolated commits, boot three arches before writing anything new, then code, then strip each Category B item in its own commit as it is proven dead. Notes dict_hash moves every time and cross-arch identity is the property that must hold. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VkM1zHGvBerLF6aqkHPweP |
||
|
|
8bdbfd1aad |
FABRIC-3.5.md §XXI: item 9 settled -- SOS travels up the birth graph
Actionability is a property of the relationship, not of the message. To the sender's birther an SOS is actionable: the parent may re-birth, exercising its own birth authority rather than inheriting the child's, so §IX.2 is untouched. To every other VM it is advisory -- a peer may shed its own dependence on the sender but may not act on its behalf. One rule, no exceptions list: an SOS is actionable exactly where the birth graph already grants authority. Pressure-tested against the supervisor risk §XIV.5 warned about, and it survives: message-driven rather than polling (the doctrine's own sanctioned form), optional rather than obligatory, no declarer since the sender reports its own trouble, and no authority transfer. What this actually closes is larger than a message type. §VII.1's ladder had no actor for rung 2, and now each rung has exactly one: the VM itself warm-restarts while alive, the birther re-births on SOS, and nobody acts on compudynamic death. That also shows §IX.1's fleet-shutdown rule is not a special case bolted on for Hera -- she has no birther, so rung 2 has no actor and the ladder falls through to rung 3. The ladder and §IX are the same rule from two ends: rung 2 exists iff you have a birther. Flags a hazard specific to this system. Payloads here are FORTH text executed by the receiver, per messaging.4th block 5041's own comment. An SOS carrying an executable payload would let a failing VM drive its parent, inverting the authority direction at the moment the sender is least trustworthy. SOS payloads are descriptive only -- a deliberate departure from the convention, and the reason the message exists. Number: 10, the next in sequence. Verified 5 and 6 are genuinely unallocated, but monotonic allocation means a number in any log means one thing forever, which matters while legacy constants and the new kernel enum coexist during cleanup. Names the consumer explicitly so this does not become another SPAWN-EVENT, which §I.4 recorded as having none. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VkM1zHGvBerLF6aqkHPweP |
||
|
|
d984b16102 |
FABRIC-3.5.md §XX: item 3 settled -- delete the suffix from the mechanism
§XVIII.5 posed agent binding as widen-the-convention vs. add-a-word. Both accept a premise the code does not support: the pairing is already recorded explicitly at bind time, and the ~user reconstruction is a redundant second derivation of information the mechanism already holds. sk_repl_dispatch_line() does two separable things. The guard builds console_get_vm_name() + "~user" and checks it is live. The actual relay then sends to index 3 -- and repl.c's own comment says that slot is "set once at pairing time," which PAIR-TEST's header confirms from the other side. Routing never uses the derived string; only the liveness guard does, and the stored pairing can answer that question without deriving anything. This is also the documented cause of a known silent failure. CONSOLE-ATTACH's own comment records that a console named anything other than the identity's base name reconstructs the wrong target and quietly falls back to direct interpretation with no error. Its "takes exactly one name" constraint is a workaround for that derivation, not a requirement of the problem. Ruling: the guard reads the recorded pairing, CONSOLE-ATTACH resolves the target as given, and ~user survives as a naming convention rather than a mechanism. Agent binding then stops being a special case -- a GPIO VM binds by the same unchanged path, because nothing appends anything to its name. Punch item 19 is retired rather than fixed: with no string to mismatch there is nothing to make loud. The six memcpy sites split cleanly -- three name a VM at birth and stay, two re-derive a partner for lookup and go. Only the derivation sites change, which is why the convention can survive without the mechanism. Also separates what §V.1 conflated: reachability is kernel-Hermes's, fabric access is vocabulary and ACL, and binding proper is the console path. An agent VM does not routinely bind at all. Scoped per §XIV.1: this states the principle kernel-Hermes must carry forward, not a proposal to patch the legacy FORTH path in place. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VkM1zHGvBerLF6aqkHPweP |
||
|
|
7a9a8ed698 |
FABRIC-3.5.md §XIX: the third Tripod leg is named Hestia
The Tripod is Hera, Artemis, Hestia. 99 occurrences renamed; three verbatim quotations deliberately left saying Console. Records the naming convention, which had never been written down and was being re-litigated for want of it: a Greek deity name, evocative rather than literal. Captain Bob's reasoning, and it generalizes -- Hera and Hermes are close fits, Artemis for block storage and freemap arbitration is a loose one that has never caused anyone a problem. Fidelity was never the standard. antiprosopos was considered first and rejected for being the wrong kind of word: a common noun straining for literal accuracy, which is the opposite of how the other three work. Hestia is a deity, is evocative -- the hearth is the fixed centre where everyone gathers, which is what a bind point is -- and carries an incidental that was not the reason for choosing it: Hestia is the one who never leaves, which is the pinned session model FABRIC-2 §H.1 already describes. It also retires both costs the longer name carried, restoring the pattern fully and cutting twelve characters of unusual spelling to six in a path where a name mismatch fails silently. The rename itself is structural rather than cosmetic. §XVII.1 found three distinct things all called console, and that collision is what led §XIII.5 to the wrong conclusion about whether a singular Console existed. "Hestia owns console.c" is unambiguous in a way "Console owns console.c" cannot be. Deliberately narrow: the HAL stays console.c, the proxies stay console proxies, and CONSOLE-ATTACH keeps its name since CLAUDE.md's rule against modifying a registered tested word applies with more force to renaming one. Keeps the ~user suffix question separate and still open as item 3, with the buffer analysis recorded so the two are not conflated later. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VkM1zHGvBerLF6aqkHPweP |
||
|
|
38735e52c2 |
FABRIC-3.5.md §XVIII: the Console component designed in full
Console is three layers and only the middle one is new: the HAL fabric (~98 KB of C) is unchanged, the per-attach proxies are unchanged, and Console the VM is a capsule rather than a mechanism. The finding that shapes the design is that 1,051 lines of Console vocabulary already exist in the wrong dictionary. fabric.4th's own header calls it "Console drawing-fabric coordinate machinery" and says outright that it is the FORTH-side policy over the C pixel primitives -- and it loads into Hera, along with font.4th, from init.4th block 2049. Moving both to capsules/console/init.4th is a pure relocation and the clearest evidence that the fabric wants an owner: it already has the vocabulary, just attached to the wrong VM. PLOT/FB-WIDTH/FB-HEIGHT registration moves with it; the implementations do not change. Records the invariant that constrains everything else. console.h:220-224 states that console.c must have no dependency on capsule or WIREBIND logic, calling it a real layering violation rather than a style preference. So Console the VM cannot be built by making console.c VM-aware: ownership flows one way, and where the HAL genuinely needs something VM-shaped it takes the shape already established by console_set_user_prefix_provider() -- a registered callback, never an include. Console's ownership is therefore by vocabulary and registration, not enforcement, which is weaker than it sounds and exactly right. Carries the headless-until-login invariant onto this second path: Console is born at boot, a console session is not, and Console's birth must not touch g_wirebind_attached_username or mint a proxy. Same rule §XXXII.2 imposed on unattended birth, cheap to hold and easy to violate with a well-meaning startup line. Notes what Console does not own, since a component owning "the bind point" attracts responsibilities belonging elsewhere, and records that the three Tripod legs have three different failure characters -- Console degrades gracefully because the fabric is HAL C and survives its owner, which is why only Hermes needed the semaphore. One genuinely open design question remains: whether agent-VM binding widens the ~user convention or takes a sibling word. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VkM1zHGvBerLF6aqkHPweP |
||
|
|
887b34d0a7 |
FABRIC-3.5.md §XVII: item 12 CLOSED -- Console already exists as a singleton, it just has no VM face
Closes the largest remaining design item, the one gating §IV. §XIII.5 got this wrong by comparing the wrong object. It measured the Tripod legs against capsule_console_birth()'s per-attach proxies, found Console "plural, ephemeral and capsule-less," and concluded a singular Console did not exist -- so §IV.1 would be inventing one. Traced: there are three distinct things called console here, not two. The HAL drawing fabric (hal/console.c, framebuffer.c, vt100.c, font_8x16.c, ~98 KB of kernel C) is already singular, already permanent, already kernel-resident. It is not a VM and has no owner. So the Console leg is that fabric given a VM face, not a promoted proxy. The proxies stay plural and ephemeral, which is correct -- they are what binding produces, not what Console is. §V.1's bind point then falls out rather than being asserted: the VM owning the fabric is necessarily what you bind to. §XIII.5's proposed shape was right; its justification was wrong, and is corrected in place. Records the evidence that the fabric actually wants an owner, which is an argument for the reshuffle this document had not previously identified. The ownerless console_set_vm_name() global has already caused a live bug, documented in console.h:193-203 -- get_vm_name() is unsafe for save-then-restore because an intervening set overwrites the saved pointer's own bytes before the restore runs; found live 2026-08-28, silently no-opping, with console_save_vm_name() existing solely to work around it. That is the signature of shared mutable state with no owner. Notes what the ruling commits to: a new capsules/console/init.4th (the one genuinely new artifact, and it is a capsule rather than a mechanism), the is_fleet_foundation triple changing membership, kernel_main.c's Hermes birth becoming Console's, the DoE CSV schema change, and Console needing BIRTH registered. §IV is unblocked. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VkM1zHGvBerLF6aqkHPweP |
||
|
|
bd0c23286c |
FABRIC-3.5.md §XVI: item 10 settled -- clean shutdown is forced flush + BYE, Hera's is suicide
Captain Bob, 2026-09-18. A clean shutdown is a VM termination where flush is forced and the VM says BYE; Hera is the special case that halts the processor instead. Traced, and the ruling is almost entirely existing mechanism. blk_flush(0) is already the documented flush-all sentinel (block_subsystem.c:988-990) and already the Stadium's own write-back action. A child's BYE already terminates it -- system_word_bye sets vm->halted, which mama_forth_words.c's own doc comment describes as the child path. arch_halt() exists on all three architectures (amd64:207, aarch64:126, riscv64:170), so three-arch parity holds without new per-arch work. The one real finding is a behavioural change rather than a gap: Hera's BYE does not halt today, it cold-restarts. mama_word_bye() reaps children then calls arch_cold_reset(). The reaping half already matches the ruling; the terminal action is the opposite of it. Flagged so the change is made knowingly -- CLAUDE.md forbids modifying a registered tested word to "fix" it, and this is a ruled change rather than a fix, but it should not surface first in a diff. Whether suicide replaces BYE or becomes a separate word is left open as item 17, noting cold-restart-on-BYE is plausibly wanted for an operator typing BYE at Hera's REPL. Records a constraint on where the flush goes: system_word_bye lives in shared vendored source that must keep working in the hosted build, so the forced flush belongs in the kernel-side shutdown path, not inside BYE. §VIII.3's bounded-window question answers itself -- the window ends when Hera halts, a bound by construction with no timeout to tune, which is consistent with §XV.3's no-owner invariant. Item 9 is now half-answered. Raises one residual edge as analysis only: if Hera is already dead, nobody performs the halt, leaving a live kernel over an empty floor. An empty floor is arguably a kernel-observable condition rather than a VM's decision, which would terminate without any VM inheriting her authority. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VkM1zHGvBerLF6aqkHPweP |
||
|
|
a842789ae6 |
FABRIC-3.5.md §XV: the authority questions answered, three by rejecting the premise
Captain Bob, 2026-09-18. Any VM may birth a near-clone of itself carrying its own ACL DNA; nobody declares a VM dead, it dies a compudynamic death; nobody owns turn order, turn order is compudynamic. These close six punch items. Birth-by-near-clone settles what "ACL/DNA" refers to -- inheritance by copy, using the existing ACL-INHERIT semantics, needing no fifth DictEntry field -- and settles "standing" as decorative, confirming the minimal reading. The authority bound holds literally rather than aspirationally: a child cannot exceed its parent because it is built from it, so there is no gate to bypass. Compudynamic death retires the largest scheduler-shaped risk in the design. §VII's ladder appeared to need a supervisor watching liveness, and a supervisor with a timer is a scheduler's twin; with death as a physical outcome rather than a judgment, that component is not needed and does not exist. §VII.3 and §IX.3 both asked the wrong question and are marked withdrawn and premise-rejected in place. Turn order gets a stronger invariant than the firewall previously proposed: that one named an owner, and naming an owner is what a conventional scheduler is. Nothing owns turn order. Flags one honest seam for later, predating this work and belonging to §XXVIII: SK_SWITCH_READINESS_THRESHOLD is a hardcoded 50, the same category of fixed policy §XIX rejected. Reframes the message-durability question that was posed badly enough to be unintelligible. It is not persistence but conservation: a message holds Q.SLOT of heat (MSG-ALLOC pulls, MSG-FREE-NODE evicts), so a dying arbiter holds N x Q.SLOT of fleet heat. Payloads may vanish; the heat may not, or K breaks. Nothing survives except K. This dissolves the Artemis-at-death-time hazard entirely rather than trading against it -- there is nothing to persist, and since death is already a Stadium eviction, returning heat is what dying consists of. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VkM1zHGvBerLF6aqkHPweP |
||
|
|
9b69c1f447 |
FABRIC-3.5.md §XIV: two rulings collapse the root dependency
Captain Bob, 2026-09-18. The existing FORTH messaging layer is legacy -- kernel Hermes is greenfield C, not a migration of messaging.4th, and the dead FORTH gets retired by a separate cleanup pass. Persistence goes through Artemis wherever possible; that part is ratified. Together these dissolve rather than answer four open questions. §III.4, named the root dependency throughout §X, is resolved in the "full centralization" direction but without the migration cost that made it frightening: with the FORTH legacy there is no live state to port. §III.3 loses its subject, and §XIII.1's 14-site inventory changes purpose rather than content -- it now feeds the cleanup. §IV.2 dissolves via its own option 3, exactly as that section predicted. The MANIFEST corrections are absorbed into the cleanup pass. All four marked superseded in place with pointers rather than rewritten. Traced the layering question the persistence ruling raises, and it answers cleanly: storage is already two layers, so no inversion is needed. The block subsystem is kernel C below the floor (block_subsystem.c, blkio_*.c, virtio_blk.c; capsule_loader.c:98-100 calls it with no Artemis involved), while Artemis is the storage policy arbiter on the floor. Fleet persistence goes through Artemis; anything kernel Hermes needs uses the same block layer Artemis sits on. Flags the one hazard that ruling carries: Hermes must not depend on Artemis for persistence at its own death, since reaching Artemis requires the mechanism that is broken -- the same trap §VIII solved for signalling with a semaphore. Recommends a stateless arbiter holding only registry-reconstructible state, offered as analysis, not a ruling. Also records the scheduler firewall as raised-but-unruled, and why the legacy ruling makes it more urgent: eligibility marking moves inside the arbiter by default once kernel Hermes replaces FORTH MSG-SEND as the caller of sk_vm_switch_signal_mark_work(). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VkM1zHGvBerLF6aqkHPweP |
||
|
|
bcf41e5618 |
FABRIC-3.5.md §XIII: work punch items 1 and 2, CLOSED with evidence
Items 1 and 2 were the two pure-investigation entries, carrying no design
commitment, so they were executable without a ruling. Traced against the
tree at
|
||
|
|
0e4e6673a9 |
FABRIC-3.5.md: open the Tripod/kernel reshuffle as its own document
Opened by direct instruction as a standalone document rather than a FABRIC-4 scratchpad entry: the reshuffle is a ratified architectural direction needing full FABRIC discipline, but it is not bare-metal boot, so it does not belong in FABRIC-3 and does not close it. FABRIC-3 stays open, living and authoritative for its own topic. Records what was decided 2026-09-18 (Hermes into the kernel as arbiter rather than client; Tripod reconstituted as Hera/Artemis/Console; Console as the fleet bind point; BIRTH general with authority bounded by inheritance; one gentle -> from-scratch -> brutal recovery ladder; SOS as a standard message type with Hermes excepted via a sinking semaphore; no provisional handoff of a failed birther's authority), and separates it from what is still open. Grounded against the tree rather than recalled, with provenance marked per claim. Notes the concrete consequences the move runs into: Artemis's two by-name S" Hermes" VM-EXEC call sites, the routing table's written "never renumbered" commitment on slots 0/1/2, and the per-VM CREATE/ALLOT message arenas that decide whether this stays a reshuffle or becomes a rewrite. Design only. No code authorized, nothing else in the tree touched. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VkM1zHGvBerLF6aqkHPweP |
||
|
|
e56974e0bf |
Bleach all 9 thumbdrive identity images to blank ahead of re-mint
Prerequisite for FABRIC-3.md §XXXV.6's ratified re-mint: capsule_mint.c's capsule_mint_identity() refuses to write over an already-recognized drive (MINT_ERR_ALREADY_MINTED, mirrors WRITE(10)'s refuse-on-non-blank-media rule), so all 9 existing certs had to be destroyed before any of them could be re-minted against the new 1 GiB artemis.img. Confirmed root cause first: Zuse's genesis marker (zuse_pubkey) lives on Artemis's own disk fence (devblock_from_top=0), not on her thumbdrive -- capsule_zuse_ boot.c's capsule_zuse_boot_try_attach() silently no-ops if her drive isn't blank and the marker is missing, so the old zuse-thumb-ident.img could not re-authenticate against the fresh artemis.img regardless. All 9 images (zuse, bob, 00-06) verified all-zero, 64 MiB size intact. Re-mint itself (live QEMU boot + MINT) not yet performed. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
7c5ba7a874 |
Rebuild disk/artemis.img fresh at 1 GiB, blank; preserve old 30 MiB image
Executes the drive-resize decision ratified in FABRIC-3.md §XXXV.6: the prior 30 MiB image was too small for the block-backed per-ISA DoE persistence design (§XXXV.2). Built a new blank 1 GiB image rather than migrate the old one's top-of-device metadata fence -- re-minting Zuse and the thumbdrive identities invalidates them either way, so skip the migration entirely. - disk/artemis.img: replaced with a fresh blank 1 GiB image (was 30 MiB) - disk/artemis-30mb-pre-1gib-backup.img: the superseded 30 MiB image, preserved rather than deleted; disposal remains open (FABRIC-3.md §XXXV.7) - disk/README.md: documented both images per this repo's own convention Not yet formatted or re-minted -- that's a live-boot step, not done here. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
e4a52c01be |
FABRIC-3.md §XXXV: bump ratified artemis.img target from 256 MiB to 1 GiB
Captain Bob asked for extra safety margin over the earlier 256 MiB recommendation. Updated every reference in §XXXV.2/.4/.6 consistently. Still decision only -- no image built yet. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
c19b6febc2 |
FABRIC-3.md §XXXV.6: ratify fresh 256 MiB artemis.img + full re-mint, no fence migration
Captain Bob resolved the drive-resize fork from §XXXV.2: build a new blank 256 MiB disk/artemis.img and re-mint Zuse + all 8 thumbdrive identities fresh, rather than attempt to migrate the top-of-device fence on the existing 30 MiB image. Skips migration complexity since a re-mint invalidates prior identities either way. Decision only -- no new image built, no re-mint performed yet. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
a83cbe016f |
FABRIC-3.md §XXXV: block-backed DoE persistence design + N=2 per-ISA shape (design only, no code yet)
Scopes the self-contained (no host serial capture) persistence needed before the per-ISA campaign driver can run on real bare-metal hardware. Found disk/artemis.img is only 30 MiB (measured) against a 516 MB/8-trial raw per-tick baseline -- raw mirroring is ruled out at any realistic device size. Proposes reusing log_region.c's binary-slot/control-header pattern for a new trial-summary-only region, flags the fence-addressing risk in growing the image, and recommends N=2 reps/cell over N=3 given the added block-IO overhead. Nothing built -- ratified in conversation only. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
94215e8b47 |
Fix two catastrophically heavy workloads found running the real campaign
The full 360-trial campaign was launched, then killed after 4h39m of CPU time with zero trials completed -- still stuck on the very first touch of the very first trial. Root cause, quantified from the workload's own source, not estimated: RUN-CHAOS5 (workload-5.4th) is 1000000 0 DO CHAOS-FIELD CHAOS-RIPPLE LOOP, multiplying out to ~2.4 trillion word executions for one call -- ~495 days at the observed rate. Never designed to be called to completion as a single touch in a repeatedly-birthed worker. A second problem found computing the fix rather than discovering it mid-run again: workload-1.4th's own birth-time self-execution (20 SQUARE-WAVE + 1500 SQUARE-BURST + 100000x MICRO-BURST, all three run automatically at capsule load) sums to ~875 million words -- ~4.3 hours just to birth one worker -- and worker index 1 always maps to it in heterogeneous mode, landing in half the campaign's cells regardless of concurrency level. Fix: two new capsules, workload-1-lite.4th and workload-5-lite.4th, carrying the same word bodies verbatim but a bounded self-execution tail, comparable in scale to the campaign's other workloads. The originals are untouched; only multiuser-doe.4th's own WL-CAPSULE/ WL-ENTRY index 1 and 5 mappings were repointed. Verified live before relaunching a third time: the actual worst case in isolation (5 0 998 MU-RUN-TRIAL, concurrency=8 heterogeneous, includes both fixed workloads plus RUN-OMNI) completed in ~20 minutes wall-clock, Hera stayed healthy throughout. Clean 3-arch qemu boot. Reps reduced 60 -> 30/cell (Bob's call after seeing the real per-trial cost) -- still matches the project's "rule of 3's" DoE convention (same as ACL-RWT's own 30 reps). 6 cfgs x 30 reps = 180 trials. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K |
||
|
|
29f91f9531 |
multiuser-doe.4th: the campaign trial-loop capsule, verified live
Builds the experimental control Bob correctly identified as still missing after HB-ON/HB-OFF (§XXXIII.5): walk the shuffled matrix of concurrency-level x workload-mode cells, birth the right worker count per cell, drive each with two VM-EXEC touches, check VM-ERROR?, kill them, print a trial marker. Built entirely from existing primitives (WORKER-BIRTH, VM-EXEC, VM-HEAT, VM-ERROR?, KILL, HB-ON/HB-OFF) plus doe.4th's own RUN-MATRIX/SHUFFLE-MATRIX pattern -- no new C primitives. Real capsule-format bug found and fixed: a first draft, chunked purely by a fixed 16-line count with no regard for word boundaries, split several CASE...ENDCASE structures and one oversized colon definition across Block headers. Result was a cascading [CAPSULE][DEFER] failure from the first split forward -- every subsequent line failed to compile, and MU-RUN-TRIAL was never actually defined (confirmed: UNKNOWN WORD when called). Root cause traced to capsule_loader.c directly: a :...; word and any control structure inside it must fit entirely within one 16-line block -- the loader's per-block compile pass has no persistent record of an open CASE's (or an overlong definition's own) state across a Block boundary. Not previously documented anywhere in this project's capsule-authoring guidance. Fixed via manually curated block boundaries and factoring oversized bodies into smaller helper words. Reps ratified at 60/cell (not the 10 first drafted), matching this project's own "rule of 3's" DoE convention (ACL-RWT's 3 seeds/30 reps, std79's 3x9x3). 6 cfgs x 60 reps = 360 main-block trials. MU-MAX-REPS raised 20 -> 63 for run-matrix headroom. Verified live on amd64: one isolated trial (0 0 999 MU-RUN-TRIAL) produced two real concurrent births, two clean kills, and the exact expected marker (MU-TRIAL run=999 cfg=0 rep=0 nw=2 mode=0 fail=0). Hera stayed healthy throughout. Clean 3-arch qemu boot. Not yet run: the full 360-trial MU-EXEC-CAMPAIGN itself (a genuinely long-running action under TCG, deliberately not started without explicit confirmation) or the WIREBIND-automation fixed arm. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K |
||
|
|
41918a28a4 |
doe_log.c: add vm_name/vm_id_hex identity columns; new calibration workload
Adds the identity columns HB-ON's existing per-tick CSV logger (doe_log_tick_row(), doe_log.c) was missing. Previously every column described *a* VM's state each row, but nothing said which VM emitted it -- concurrent VMs' rows were indistinguishable by source. Two new first columns, vm_name and vm_id_hex, resolved via a reverse lookup on vm->stadium_vm_id (capsule_vm_registry_get()). Real bug found and fixed while building this: the freestanding snprintf here silently prints the literal format string instead of substituting for %016llx (width+ll+hex unsupported) -- caught by reading the actual emitted row, not assumed to work. Replaced with a hand-rolled hex nibble-table loop, the same idiom vm_uuid_format()/ MINT-SCRATCH-EMIT already use. Corrects FABRIC-3.md's own prior "still not started: CSV driver" framing: no bespoke CSV emitter is needed for the multiuser DoE at all -- this per-tick logger already exists, fires automatically inside every VM's own execution loop, and just needed HB-ON plus these identity columns to be usable for concurrent workers. Also corrects doe_log.h's own stale doc comment claiming g_doe_log_enabled defaults to 1 -- doe_log.c's own source is authoritative: 0, off by default, matching the HB-ON/HB-OFF naming. One real, honest limitation recorded rather than smoothed over: vm_name reads blank for any tick captured during a WORKER-BIRTH'd capsule's own self-execution at birth time, since capsule_birth_baby() runs that work (and its heartbeat ticks) before the registry name can be set. vm_id_hex is unaffected (reads directly from vm->stadium_vm_id) and still uniquely disambiguates every row -- confirmed live: a blank-name row's own vm_id_hex matched its later PARITY:KILL line's vm_id exactly. New workload-calib1.4th: Bob confirmed the existing 10 workload-N.4th files aren't fixed. Sized to cross HEARTBEAT_CHECK_FREQUENCY (256 word-executions/tick) many times over while staying far lighter than RUN-FIB/RUN-CHAOS5, both too slow under TCG for a bounded verification run. A real, reusable addition, not a throwaway. Verified live on amd64 via a self-contained one-shot EXEC'd test (HB-ON, WORKER-BIRTH two calibration workers, HB-OFF, VM-ERROR? on both, KILL both, completion banner) rather than interactive polling, which breaks once HB-ON makes the log grow continuously via Hera/ Hermes/Artemis's own background ticks. Clean 3-arch qemu boot on the real committed change. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K |
||
|
|
e33eb36361 |
Multiuser DoE punch list: WORKER-BIRTH + VM-ERROR? + vm_physics_init fix
Core mechanism for §XXXIII's main concurrency block, built and live-verified on amd64. Two new primitives: - WORKER-BIRTH ( capsule-c capsule-u name-c name-u -- ok? ): births a named, VM-EXEC-addressable VM from an arbitrary (p) capsule with no identity involved. Corrects the ratified design's own assumption that the main block would use UNATTENDED-BIRTH -- that requires a committed capsule per identity, impractical for dozens of trial VMs. The concurrency block never needed identity at all. - VM-ERROR? ( c-addr u -- flag ): reads a named VM's error state from Hera, mirroring VM-HEAT's silent/always-returns-a-value contract. Needed to check a VM-EXEC-driven trial VM's own fault state after the fact -- nothing existing let Hera do this. Real bug found and fixed in both WORKER-BIRTH and UNATTENDED-BIRTH: capsule_birth_baby() never calls vm_physics_init() either (same shape as the registry-name gap found building UNATTENDED-BIRTH) -- without it a born VM is never in the VM Fleet Attractor physics list, so VM-HEAT returns 0 forever regardless of work done. Fixed by adding vm_physics_init() alongside the existing registry-name call in both words. Traced (not guessed) why heat still read 0 after one VM-EXEC touch even post-fix: vm_physics_touch()'s transfer logic only fires from a VM's *second* touch onward -- the first touch just records a baseline tick. Verified live across three sequential touches: heat 0 -> 7039 -> 11333. This is a real design requirement for the DoE's heat/CV response variable (each trial must touch a worker at least twice), not a bug to route around. Also found live: all 10 existing workload-N.4th capsules self-execute their full workload at load/birth time (a bare top-level call to their own RUN-* word at file end) -- missed on an earlier, too-shallow 8-line survey of each file. WORKER-BIRTH alone already runs a worker's first pass as a side effect of birth. Verified live on amd64: two concurrent workers (fib + matrix-mul) birthed, run, measured (heat + error state), and killed cleanly; VM-HEAT/VM-ERROR? both confirmed silent-0 on an unknown name. Clean 3-arch qemu boot on the real committed change. Still open: the run-matrix/shuffle/CSV driver capsule itself, the WIREBIND-automation path for the fixed arm, and the per-VM touch-count budget's exact value. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K |
||
|
|
bc2e294c50 |
FABRIC-3.md §XXXIII: multiuser/multitasking DoE design (no code yet)
Scopes the DoE Bob asked for after §XXXII closed -- "a pretty big experiment... long running multi-factor DoE that sort of puts the OS through its paces." Design only, per this project's standing plan-before-code discipline. Three corrections found and recorded before the design could be trusted: - Console sessions are not a concurrency axis and this isn't a defect: console.c hardcodes one physical UART as static global state, not an instantiable abstraction. One physical terminal, one foreground session, by construction. Bob's reaction (eventual multi-seat/ telnet-ish/pty-multiplexer idea) recorded separately in memory, explicitly deferred behind this DoE. - The real, provable concurrency primitive is VM-EXEC (doe-campaign.4th's own THREE-VM-CAMPAIGN already proves the pattern), resolved through the full VM registry (unbounded kmalloc list), not messaging.4th's 16-slot routing table -- checked live, not assumed. - No FORTH-reachable monotonic clock exists; latency dropped as a response variable rather than building a new timing primitive under this scope. Also flagged, not chased: every acceptance boot this session produced an L8-DoE CSV despite init.4th never calling it and KERNEL_ARGS defaulting to empty -- likely a stale starforth.cfg survivng `clean`. Practical decision: the new campaign is its own driver, independent of whatever causes that. Ratified design: concurrency level (2/4/8 VM-EXEC-driven VMs) x workload assignment (uniform/heterogeneous across the existing 10 workload capsules) as the crossed factorial; identity origin (WIREBIND vs UNATTENDED-BIRTH) as a fixed arm rather than crossed, since WIREBIND's per-identity USB-attach cost doesn't scale into a full factorial. Response variables: correctness-proxy (completes without vm->error), heat/CV stability under contention, completed-trials-per- interval in place of the dropped latency variable. Punch list recorded; implementation not started per standing "plan approval is not a start signal" rule. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K |
||
|
|
af351bff92 |
Punch list item 4 (final): CONSOLE-ATTACH, full unattended-identity flow verified end-to-end
Closes FABRIC-3.md §XXXII.2's punch list. CONSOLE-ATTACH ( name-c name-u -- ok? ) pairs a fresh console VM to an already-live VM registered as "<name>~user". Deliberately a plain, unconditional primitive with no VMIdentity capability-bit check -- corrects this session's own first-pass design (§XXXII.2 amended in the same commit): identity.installed is 0 for Hera/Hermes/Artemis and for every console-proxy VM, so a capability-bit gate would be unreachable for every VM a human actually types at, Zuse included. Matches ZUSE-ELIGIBILITY-ADD's own "no bespoke gate" precedent in this file; real gating is `' CONSOLE-ATTACH ACL-PIN` in ACL.4th if ever wanted. Two real bugs found and fixed via live testing, not assumed correct: - capsule_birth_baby() never sets a VM's registry name (documented gap, same one capsule_runcap_birth()'s own history already hit) -- UNATTENDED-BIRTH gained a second `name` argument and now calls capsule_vm_registry_set_name(new_vm_id, "<name>~user") itself. - CONSOLE-ATTACH's first draft took an independent console name from the target's name. sk_repl_dispatch_line()'s pairing check (repl.c) reconstructs the target as console_get_vm_name()+"~user" -- a mismatched console name silently falls back to direct interpretation with no error. Caught live (typed `5 6 + .` at a mismatched console, got a direct `11` instead of a relay) and fixed by collapsing to one name argument, matching WIREBIND's own by-construction invariant. Full end-to-end live verification on amd64: minted a test identity, UNATTENDED-BIRTH'd it as "bob", CONSOLE-ATTACH'd a console named "bob", USE'd it, typed `5 6 + .` -- no direct output at [zuse@bob] (relay path taken), then `[zuse@bob~user] 11 ok>` appeared: the relayed command executed on the target identity VM itself and printed its own answer back through the shared console. Hera stayed healthy throughout (2 2 + . -> 4 after switching back). CONSOLE-ATTACH also verified to refuse cleanly on a nonexistent target with no orphaned VM. Test capsule reverted after capture per this project's probe convention -- never committed. FABRIC-3.md §XXXII.2 fully closed: all 4 original questions ratified, the mid-course drive_uuid and ACL-bit corrections both recorded plainly rather than silently folded in, and a doc-accuracy note left for CLAUDE.md's own stale "1024-byte block limit" framing (mkcapsule's real limits are range [2048,5120) and 16 content lines/block) -- flagged, not fixed, out of this punch list's scope. Clean 3-arch qemu boot (amd64/aarch64/riscv64) on the real committed C-only change. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K |
||
|
|
5879c8b3bc |
Punch list item 3: UNATTENDED-BIRTH call site, verified live end-to-end
Implements the unattended-birth mechanism FABRIC-3.md §XXXII.2 designed: UNATTENDED-BIRTH ( name-c name-u -- ok? ) births a VM from a named (p) capsule via capsule_birth_baby() (completely unmodified, the same generic build-time-capsule path CAPSULE-BIRTH already uses), then installs its identity the same way capsule_wirebind.c already does live for WIREBIND attaches -- vm_identity_from_cert() verification followed by a plain post-birth struct assignment -- rather than anything RUNCAP-shaped, since RUNCAP requires a real blkio_dev+ homeblocks_sig_t an unattended identity never has. The born VM's own capsule payload is expected to lay down two CREATE'd buffers (UNATTENDED-ID-UUID, UNATTENDED-ID-CERT) via MINT-SCRATCH- EMIT's own literal format; their addresses are fetched by interpreting a two-word line inside the *new* VM's own context (vm_interpret(born_vm, ...)), the same "run inside that VM's own dictionary" idiom capsule_wirebind.c already uses for VM-NAME-REG. Explicit invariant preserved: never touches g_wirebind_attached_username or any console-pairing state, births no console VM -- an unattended identity stays un-promptable (§VIII.1) until a human pairs a console to it later via the already-working VM-NAME-REG mechanism. Verified live end-to-end on amd64: minted a real test identity via MINT-SCRATCH, captured its MINT-SCRATCH-EMIT output, built a throwaway test capsule from it (discovered along the way: mkcapsule's real block constraints are range [2048,5120) and max 16 content lines per block -- neither matches this repo's own doc comment, corrected via ground truth from the tool itself, not assumed), then ran UNATTENDED-BIRTH against it: cert verified against Zuse's root pubkey, identity installed, "no console attached" reported, Hera stayed healthy afterward (5 6 + . -> 11). Test capsule reverted after capture per this project's own probe convention -- not a real identity, never committed. Clean 3-arch qemu boot (amd64/aarch64/riscv64) on the real committed C-only change. Remaining punch-list item (the ACL cap bit for console attachment) not started. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K |
||
|
|
5fc709a228 |
Punch list item 2: MINT-SCRATCH-EMIT, verified live on all 3 arches
Adds the hand-transcription mechanism FABRIC-3.md §XXXII.2's punch list item 2 calls for: MINT-SCRATCH-EMIT prints the last successful MINT-SCRATCH's drive_uuid + cert devblock as ready-to-paste FORTH source (HEX-based CREATE ... C, ... byte sequences), so packaging an unattended identity into a capsule is a mechanical copy out of the captured boot log rather than a manual hex-to-FORTH translation an operator could transpose a digit in. Refuses (no output) if no MINT-SCRATCH has ever succeeded -- printing 4112 zero bytes as if they were a real identity would be a silent, misleading success, matching this session's own error-handling audit discipline rather than adding a new silent-failure primitive right after finishing one. Verified live on amd64: MINT-SCRATCH-EMIT correctly refuses before any mint, then after MINT-SCRATCH succeeds, emits UNATTENDED-ID-UUID and UNATTENDED-ID-CERT as valid FORTH literals. Cross-checked byte-exact: the cert's own embedded ASN.1 serialNumber field matches the emitted UUID bytes exactly, confirming x509_build_user_cert()'s drive_uuid binding round-trips correctly through the scratch-device path. Clean 3-arch qemu boot (amd64/aarch64/riscv64). Remaining punch-list items (writing/committing a real capsule file for an actual named identity, the unattended-birth call site, the ACL cap bit) not started -- authoring a real committed capsule needs a name/ purpose decision that isn't mine to make. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K |
||
|
|
63864c4b01 |
Punch list item 1: scratch-device MINT-SCRATCH, verified live on all 3 arches
Implements FABRIC-3.md §XXXII.2's scratch-thumbdrive mint mechanism: capsule_mint_identity_scratch() (capsule_mint.c/.h) builds a throwaway RAM-backed blkio_dev via blkio_ram.c's backend and runs capsule_mint_identity() against it completely unmodified -- same live Zuse-signing operation, same rng_get_bytes() draw for drive_uuid a real thumbdrive gets. Reads back only drive_uuid + the cert devblock; the seed devblock is written into the scratch buffer internally but never read out (no seed is ever baked into a capsule, per the ratified no-seed decision). New FORTH word MINT-SCRATCH (mama_forth_words.c), same stack signature as MINT, mints into the scratch device instead of any attached drive and never touches sk_repl_get_attached_blk_dev() or console-pairing state. Prints the drive_uuid as hex so a live boot log itself proves each call drew fresh entropy. Build correction found along the way: blkio_ram.c was excluded from the kernel build (Makefile.starkernel VM_EXCLUDE) alongside blkio_factory.c/blkio_file.c. blkio_factory_open() unconditionally references blkio_file.c's real fopen()/fread() file I/O, which has no freestanding-kernel equivalent, so the factory function couldn't be used as-is. blkio_ram.c itself is pure memcpy over a caller buffer -- pulled it alone into the kernel build and wired it directly in capsule_mint.c, the same way blkio_factory.c's own extern declarations do internally. Verified live on amd64: two MINT-SCRATCH calls produced two genuinely different drive_uuids (b533246d.../ae11b2b2...), confirming fresh entropy per call rather than stale reuse; VM stayed healthy afterward (5 6 + . -> 11). Clean 3-arch qemu boot (amd64/aarch64/riscv64), logs and DoE CSVs committed per standing convention. Remaining punch-list items (hand-transcription into a .4th block, the unattended-birth call site, the ACL cap bit) not started. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K |
||
|
|
60bcdc09a7 |
Stage E amendment: resolve drive_uuid wrinkle via scratch-device mint (FABRIC-3.md §XXXII.2)
Bob's proposal: mint against a "sim thumbdrive" -- a scratch blkio_dev from the existing blk_subsys_add_raw_device() mechanism (same one the kernel ramdrive already uses), then hand-transcribe just drive_uuid + DER cert as FORTH literals into a normal .4th capsule block. This dissolves the wrinkle the first pass of Q2 left open: capsule_mint_identity() runs completely unmodified against the scratch device (same live Zuse-signing op, same rng_get_bytes() draw for drive_uuid), so vm_identity_from_cert() needs zero changes -- no origin-aware variant, no content-hash-derived substitute. "Sim" describes only where the bytes were written, invisible to verify. Also corrects an error in the first pass: cert production cannot be "offline" -- capsule_mint_identity() requires Zuse's live in-kernel signing key, so minting is inescapably a two-boot runtime operation (mint on boot N, package by hand, birth on boot N+1). Punch list updated to match. Still design-only, no code. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K |
||
|
|
06e2507ce4 |
Stage E ratified: unattended identity is cert+personality only, no seed (FABRIC-3.md §XXXII.2)
Settles all 4 open questions on paper, no code: - No seed baked into capsules for this pass (Bob's call); runtime-minted seed documented as a future, separately-scoped possibility - Birth via capsule_birth_baby() (unmodified, already generic), not capsule_runcap_birth() -- that path requires a real blkio_dev*+ homeblocks_sig_t* an unattended identity can't provide, and its own header says build-time-baked content is exactly the wrong case for it - Identity population is the same post-birth struct-assignment pattern capsule_wirebind.c already uses live (vm->identity = identity) - §XXIV's block-collision machinery already covers a cert-carrying capsule -- no new risk, no change needed there - ACL gate lives on the attaching human's credentials, since an unattended instance holds no secret to prove anything about itself - xHCI/live-table machinery confirmed irrelevant, independent of the seed question One real wrinkle flagged, not resolved: vm_identity_from_cert()'s serial-number check binds to a physical drive_uuid an unattended identity doesn't have -- needs Bob's call before implementation. Punch list recorded; implementation not started per standing "plan approval is not a start signal" rule. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K |
||
|
|
42d4bf3dad |
Stage D batch 7 (final): mama_forth_words.c groups 5+6 -- dictionary lookup/test words + Stadium primitives; Stage D closed (FABRIC-3.md §XXXII.6)
NAME>XT/RUNCAP-TEST/PAIR-TEST and the six STADIUM-* physics primitives (STADIUM-ADMIT/STADIUM-EVICT/STADIUM-RES-PULL/STADIUM-RES-PUSH/ STADIUM-HEAT@/STADIUM-HEAT!), 12 sites -- closes mama_forth_words.c and the entire kernel-only error-handling audit. All 105 sites from the §XXXII.3 triage now accounted for: 28 already correct, 75 silent sites fixed across repl.c/inference_words.c/ log_words.c/vm_core.c/mama_forth_words.c, 2 special cases resolved by dropping the error per their own documented contract, 1 resolved via console_println() per its own recursion constraint. Verified with awk: zero remaining vm->error=1 sites in mama_forth_words.c lack a diagnostic within the preceding three lines. Three-arch clean qemu acceptance passed. This closes Stage D and the USE/logging/audit thread opened in §XXXII; Stage E (human-vs-unattended identity model) remains open, not started this pass. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K |
||
|
|
90c86c6006 |
Stage D batch 6: mama_forth_words.c group 4 -- identity/crypto words (FABRIC-3.md §XXXII.6)
MINT + mint_pop_string()/ZUSE-ELIGIBILITY-ADD/ZUSE-ELIGIBLE?/ ELEVATE-PUBKEY-UNPACK, 10 sites. mint_pop_string() gained a field_name parameter so its diagnostics name which of MINT's four string arguments failed (phone/email/username/full_name), rather than a generic message that would leave the operator guessing. Diagnostic placement matched each function's own sibling convention where one exists (ZUSE-ELIGIBILITY-ADD), log_message() default otherwise. Three-arch clean qemu acceptance passed. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K |
||
|
|
8b5300fc4f |
Stage D batch 5: mama_forth_words.c group 3 -- cross-VM execution/dispatch, both flagged special cases resolved (FABRIC-3.md §XXXII.6)
VM-STEP/VM-EXEC/VM-CALL (5 sites) gained console_println() diagnostics
matching their own existing sibling guards.
SWITCH-MARK-WORK and VM-HEAT (3 sites) resolved per §XXXII.3's own
recommendation: dropped vm->error entirely rather than diagnosing it,
matching each function's own doc comment ("must never error or spam
the console" / "does not print/error"). SWITCH-MARK-WORK is the same
function §XXVIII.3 already fixed once for an off-by-one that fired
silently on every MSG-SEND in the system -- the guard now genuinely
cannot repeat that by contract, not just by the threshold being right.
VM-HEAT's guards now push 0 and return, matching its own
always-returns-a-value stack effect.
Three-arch clean qemu acceptance passed. Zuse's own WIREBIND attach
(every boot) drives MSG-SEND -> SWITCH-MARK-WORK, so this batch's most
safety-critical fix is exercised by standard acceptance, not just
compiled.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K
|
||
|
|
b42c3b195c |
Stage D batch 4: mama_forth_words.c groups 1+2 -- capsule/lifecycle words + BIRTH/START/KILL (FABRIC-3.md §XXXII.6)
15 of mama_forth_words.c's 48 silent sites fixed, split by functional
grouping per direct instruction: CAPSULE@/CAPSULE-HASH@/CAPSULE-FLAGS@/
CAPSULE-LEN@/CAPSULE-BIRTH/CAPSULE-RUN/EXEC (9), and BIRTH/START/KILL
(6).
Refined the diagnostic-placement rule: match whichever convention that
same function's other already-correct guards use, rather than
defaulting uniformly. BIRTH/START/KILL/EXEC each already had a
console_println() sibling guard ("name too long or empty") -- their
newly-diagnosed guards now match that, same shape as USE's own fix.
The CAPSULE*@ words have no sibling guard to match, so they keep
log_message() (defer_words.c's gold-standard default).
Three-arch clean qemu acceptance passed; BIRTH itself is exercised by
every boot (Hermes/Artemis birth).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K
|
||
|
|
66c2f3e539 |
Stage D batch 3: fix silent error sites in vm_core.c, incl. the two highest-value primitives; two real NULL-deref bugs found and fixed (FABRIC-3.md §XXXII.6)
All 13 silent vm->error=1 sites in vm_core.c now log a diagnostic
first: vm_enter_compile_mode, vm_compile_word, vm_compile_literal,
vm_compile_call, vm_exit_compile_mode, execute_colon_word (2 sites
each/combined), and the four fundamental memory primitives
vm_load_u8/vm_store_u8/vm_load_cell/vm_store_cell.
Found and fixed two real NULL-pointer-dereference risks while adding
the diagnostics: vm_compile_call() and vm_exit_compile_mode() each had
a combined `if (!vm || <cond>) { vm->error = 1; ... }` guard that
dereferenced vm->error even on the !vm branch of its own condition.
Split both, and applied the same defensive split to the four memory
primitives since vm_ptr()/vm_addr_ok() both tolerate vm==NULL
internally.
Live-verified, amd64: 999999999 @ . recovers correctly at Zuse's own
console (Stage A's mechanism holds), but the new vm_load_cell
diagnostic itself didn't print -- traced to memory_words.c's own
redundant, still-silent vm_addr_ok() pre-check in memory_word_fetch()
(and the same shape in memory_word_store()), which intercepts before
ever reaching vm_load_cell(). memory_words.c is vendored, out of this
initiative's scope, spun off to FABRIC-4.md -- recorded as a concrete
cross-reference for that future work rather than left to be
rediscovered.
Three-arch clean qemu acceptance passed, full POST suite included.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K
|
||
|
|
1a8c0e4fcf |
Stage D batch 2: fix silent error sites in log_words.c, resolve the flagged category-iii special case (FABRIC-3.md §XXXII.6)
log_word_set_level(), log_do_emit(), and log_emit_string() (7 sites total) now log a diagnostic via log_message(LOG_ERROR, ...) before setting vm->error, matching this file's own already-correct log_str_emit() sibling and defer_words.c's gold-standard pattern. log_word_append_raw() (backing (LOG-APPEND-RAW), 3 sites) resolved differently per §XXXII.3's own triage note: its doc comment forbids log_message() here (recursion into the log ring it writes to), but the word can also be invoked by hand at the console -- console_println() carries no such recursion risk and matches Stage B's policy for a manual interactive invocation. Added the console.h include this required. Three-arch clean qemu acceptance passed. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K |
||
|
|
5417ffb2bf |
Stage D batch 1: fix silent error sites in repl.c and inference_words.c (FABRIC-3.md §XXXII.6)
repl.c's sk_word_blk_attach_ack() and inference_words.c's array_ptr() helper + infer_word_run()'s allocation guard now log a diagnostic via log_message(LOG_ERROR, ...) before setting vm->error, matching defer_words.c's own gold-standard pattern (§XXXII.3) -- these are internal/background conditions (a malformed message-callback, a bad array reference or allocation failure), not interactive usage mistakes, so the fix keeps the fault and reports it rather than dropping it like USE's own fix did. Batched together (4 sites total, smaller combined than the next file) rather than two separate acceptance cycles for negligible size. Three-arch clean qemu acceptance passed. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K |
||
|
|
09959c6ca2 |
Stage C: primitive error-handling audit triage complete, 105 sites classified, no code changed (FABRIC-3.md §XXXII.3)
Read every one of the 105 vm->error=1 sites across the 8 kernel-only files in context (not sampled): 28 already correct (diagnostic before/ without erroring, e.g. defer_words.c's 15-for-15 gold-standard pattern), 75 silent (the exact USE-defect shape), 2 special cases requiring individual handling rather than a generic fix. Two findings flagged above the rest: SWITCH-MARK-WORK and VM-HEAT (mama_forth_words.c) each set vm->error on their own guard despite their own doc comments explicitly saying they must never error -- SWITCH-MARK-WORK is the same function whose off-by-one already caused a stray silent error to fire on every MSG-SEND once before (§XXVIII.3). And vm_core.c's four fundamental memory primitives (vm_load_u8/ vm_store_u8/vm_load_cell/vm_store_cell) silently fault on any out-of-bounds address -- the highest-reach fix candidates in the audit, hit by far more FORTH words than any single mama_forth_words.c site. Full per-file site list and per-category breakdown recorded for Stage D's own reference. Triage only -- no code changed this stage. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K |
||
|
|
c05b70c8d6 |
Stage B: logging policy documented, level-aware log-ring eviction built, LOG-FLUSH deferred again (FABRIC-3.md §XXXII.4)
Policy decided for the kernel-only audit scope: an interactive command's direct response stays on console_println/console_puts; everything else (state transitions, background diagnostics, audit trails) routes through log_message() at the appropriate level, matching capsule_mint.c's verify_mint() precedent. Documented, not code-swept here -- reclassifying individual sites is Stage D's job. LOG-FLUSH deferred again, explicitly: the per-VM log buffer its own doc comment presumes (vm_log_buffer.h) doesn't exist anywhere in the tree -- building it is real feature work needing its own scoped stage. Level-aware eviction built: log_region_append() now reads the oldest ring slot's own level before evicting it, protecting ERROR/WARN records from being pushed out by INFO/DEBUG churn -- drops the incoming low-priority record instead. Found and fixed an adjacent bug while making this change: the prior two-valued return contract would have made a benign "dropped by design" outcome indistinguishable from a genuine write failure to its one caller, which unconditionally set vm->error on any nonzero return. Changed to a three-valued contract (0 success, 1 dropped by design, -1 genuine failure). Three-arch clean qemu acceptance passed. Eviction path itself not live-exercised (needs 128+ LOG-APPEND calls to fill the ring) -- flagged, matching this project's own precedent for that kind of gap. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K |
||
|
|
5c5896fbc1 |
Stage A: fix USE's silent stack-underflow/bad-address guards, and the Hera fault-scoping gap they exposed (FABRIC-3.md §XXXII.1)
mama_word_use()'s two silent vm->error=1 guards (dsp<1 stack underflow, NULL from vm_ptr()) now print a diagnostic and return, matching the function's other five guards. Live verification of that fix alone surfaced a bigger problem: with the guard no longer silent, the REPL proceeds to interpret the leftover token as an unrecognized word, which independently sets vm->error, and sk_repl_step()/sk_repl_run()'s Hera-branch still hard-halted on that. Investigated kernel_main.c's boot/capsule-load paths directly: they already catch and clear mama->error entirely separately, before sk_repl_run() is ever entered -- so the "no fallthrough surface" halt in these two REPL functions was never protecting a boot-time fault, only an ordinary interactive REPL-turn one. Both functions now recover unconditionally on any VM's error, Hera included, matching how a redirected (WIREBIND/USE'd) identity's session already recovered. Removed the now-fully-unused sk_fault_handler(). Verified live, amd64: USE rajames at Zuse's own console prints the new diagnostic, then "VM fault -- session recovered, resuming", console stays interactive afterward. Three-arch clean qemu acceptance passed before and after the Hera-fault-scoping change. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K |
||
|
|
05159c9f9e |
Scope new initiative: unattended identity model, USE fail-closed-halt root cause, primitive error-handling audit, logging cleanup (FABRIC-3.md §XXXII)
USE crash fully root-caused against current source: mama_word_use() has seven guards, two of which set vm->error silently (stack underflow, bad VM address) while the other five print diagnostics; sk_repl_step() halts the whole kernel only when the faulting VM is Hera; Zuse's console runs directly on Hera's own VM (confirmed in capsule_zuse_boot.c), making a benign typo at her prompt the one deterministic path to the fail-closed halt. Three fix options named, none applied yet. Human-vs-unattended identity/console-birth model scoped: console/user VM pairing is pure name-convention + a per-console-VM VM-NAME-REG call, so after-the-fact console attach needs no new mechanism. Named the invariant an unattended birth path must not violate (never touch the global driving the "no thumbdrive, no prompt" gate). Error-handling audit scoped kernel-only (mama_forth_words.c + 5 kernel-only word_source files + vm_core.c/repl.c): 107 vm->error=1 sites across 8 files, counted directly. The vendored word_source sweep is explicitly out of scope here, spun off as future FABRIC-4.md work. Logging cleanup sequenced ahead of the audit's own fixes, with the two open §XXVII gaps (LOG-FLUSH doesn't exist; FIFO eviction is level-blind) named for an explicit in/out-of-scope call. Staged A-E plan proposed; nothing implemented this pass. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K |
||
|
|
16cc74243c |
std79 DoE campaign rerun post-Stage-4: 81/81, §XIV caught live by rerun's own harness bug (FABRIC-3.md §XXXI)
Reran the established 3x9x3 randomized full-factorial campaign (std79-doe.fth) on all 3 architectures per standing project discipline (any Stadium-adjacent change reruns the whole DoE from the top). First amd64 attempt exposed a real test-harness bug that re-triggered the already-known §XIV concurrent-attach gap: the sequential-attach wait loop checked for any recent WIREBIND-attached line instead of the specific identity requested, firing the next device_add before the kernel finished the current one. Only 2 of 8 identities attached; the campaign itself completed cleanly with graceful "VM-EXEC: VM not found" refusals rather than corrupting anything. Fixed the wait loop, discarded the invalid run's campaign result (its boot log kept for the record), reran clean. Corrected reruns: all 8 identities individually confirmed on all 3 architectures, zero VM-not-found errors, zero faults, 81/81 trials correct against established baseline values. Raw logs archived at experiments/std79-doe/results-20260915-stage4/. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K |
||
|
|
597f5a6cd8 |
Stage 4 verified: reap mechanism proven, real WIREBIND multiuser+multitasking confirmed live (FABRIC-3.md §XXX)
Reap mechanism (increments 2+3): temporary probe using Artemis as a safe stand-in parked identity, 3/3 checks PASS on all 3 architectures (refuse on SWITCHED_OUT, correct post-reap state, switch-signal slot released). Probe reverted, all 3 architectures re-verified clean. Live multiuser verification (increment 4, no code changes): a real previously-unattached identity thumbdrive attached via QMP on a running boot on all 3 architectures. VM-EXEC dispatch into her own live VM computed correctly, tagged with her own name in console output, full Tripod fleet unaffected. EJECT cleanly tore down both her VMs and released the switch-signal slot -- confirms increment 1's per-device fix under a real live attach/detach. Found and flagged, not fixed: interactive USE on a freshly-attached identity halts the kernel outright. Confirmed NOT caused by Stage 4 -- reproduced identically on the commit before any Stage 4 work. Direct VM-EXEC dispatch into the same identity works correctly; this is specific to the USE/BINDSTEP codepath, plausibly never caught before since every prior identity campaign used VM-EXEC, never interactive USE. Stage 4's original scope is complete: multitasking (Tripod) and multiuser (WIREBIND) are now genuinely composed, verified against real hardware-driven identity attach on all 3 architectures. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K |
||
|
|
d9da82b065 |
Stage 4 increments 2+3: WIREBIND VMs as switch-signal participants + mark-and-defer tombstone reap (FABRIC-3.md §XXVIII Stage 4)
Increment 2: WIREBIND user VMs (the ones that actually run FORTH work; console VMs are pure REPL proxies and never participate) register as Stage 3 switch-signal participants at attach, unregister at teardown. Slot table bumped 8 -> 16, matching messaging.4th's own VM-MAX -- a real, already-agreed ceiling, not an invented number. Added sk_vm_switch_signal_unregister() (compaction-based; Tripod VMs never needed removal, WIREBIND VMs cycle constantly and would otherwise exhaust the bounded table). Increment 3: implements the plan's own ratified option (A) for the async-detach UAF risk -- mark-and-defer via a new pending_reap flag on VMRegistryEntry, deliberately not a new VMState (capsule_vm_kill() already treats VM_STATE_DEAD as idempotent success, which would silently swallow a reap attempt; SWITCHED_OUT still accurately describes a tombstoned VM until the moment it's actually freed). unclean_detach() sets it when capsule_vm_kill() refuses a SWITCHED_OUT target; the Stage 3 checkpoint (vm_core.c) checks it before ever attempting to resume a pending switch target, and calls the new capsule_vm_force_reap() instead -- the one caller allowed to bypass capsule_vm_kill()'s own refusal, because it runs at the exact safe cooperative point the switcher itself controls. A new idle-tick sweep cleans up the WIREBIND live-table entry once the reap has actually happened. Verified clean on all 3 architectures (baseline regression -- no WIREBIND attach happens in a plain boot). The reap mechanism's own correctness under a genuinely parked context is verified separately, next, via a temporary deterministic probe. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K |
||
|
|
9f0f33dfc5 |
Stage 4 increment 1: per-device WIREBIND tracking, fixing a real multi-identity detach leak (FABRIC-3.md §XXVIII Stage 4)
capsule_wirebind_unclean_detach()/eject() tracked "the attached identity" as a single global, correct for the console-pairing UX (one physical console) but wrong for detach safety: since §XV/§XVI proved multiple identities genuinely live simultaneously via this same attach path, every attach after the first silently overwrote the singleton, so an unclean detach of any but the most-recently-attached identity was silently ignored -- that VM leaked forever, no trace in the log. Adds a per-device live-identity table, separate from the (unchanged) console-pairing singleton, so unclean-detach resolves any attached device to its own identity. Sized off messaging.4th's own VM-MAX (16) minus Tripod's 3 reserved slots, not an invented number. Corrects the stale "single-USB-device constraint" doc claim in capsule_wirebind.h, false since §XV/§XVI. Groundwork for Stage 4's real deliverable (WIREBIND VMs as switch-signal participants) -- this increment only fixes detach targeting; switch- signal registration is next. Verified clean on all 3 architectures (no WIREBIND attach happens in a plain boot, so this is a regression check on the existing Tripod-only path; live multi-identity verification comes with the switch-signal registration increment). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K |
||
|
|
1c220ad4b4 |
Close §XXVIII: defer 2 known bugs to Stage 4 (FABRIC-3.md §XXIX)
Bob's call after an honest end-to-end status check: multitasking (fixed Tripod fleet) and multiuser (Zuse/WIREBIND) each work for their own tested paths, but two known bugs remain rather than zero -- the §XIV concurrent WIREBIND attach detection gap, and the never-reconciled MSG-TICK/Stage-3-switch dual-ownership rough edge. Both deliberately deferred to be addressed during or at the close of Stage 4, since both bear directly on WIREBIND VMs joining the switch-signal population. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K |
||
|
|
05ae7aa886 |
Fix SWITCH-MARK-WORK off-by-one: was tripping vm->error on every MSG-SEND (FABRIC-3.md §XXVIII.2)
mama_word_switch_mark_work()'s stack-underflow guard checked dsp < 2, requiring 3+ items, when it only ever needs the 2 IDX>NAME leaves it (caddr u). Since dsp is index-based (2 items == dsp 1), this rejected every normal call. MSG-SEND tail-calls SWITCH-MARK-WORK unconditionally, so this fired on every message sent anywhere in the system -- visible only where a caller happened to check the target VM's error flag afterward (mama_word_vm_exec()'s "VM-EXEC: ERROR in Artemis" report). Verified clean on all 3 architectures: boot reaches Startup: Artemis live -> zuse@Hera] ok> with no VM-EXEC: ERROR line at all. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K |
||
|
|
66beae7fd4 |
Stage 3 follow-on: message-arrival eligibility hook + trampoline-blind switch-storm fix (FABRIC-3.md §XXVIII.2)
Implements the message-arrival eligibility signal FABRIC-3.md §XXVIII.1 left open (has_work per-slot flag, set via new SWITCH-MARK-WORK primitive from MSG-SEND) so an idle VM never becomes a switch target purely by waiting out the readiness threshold. Also root-causes and fixes a second, independent switch-storm: the tick's "who is current" check used vm_log_attributed_vm(), which can't see a VM parked in switch.c's own raw trampoline. Replaced with a dedicated g_switch_current_vm tracked by the switch mechanism itself, and moved target-slot eligibility reset to the switch decision point instead of relying on ISR polling to observe a window that can be only a few instructions wide. Verified live on all 3 architectures: clean boot to zuse@Hera] ok>, live cross-VM message dispatch, and (since a quiet log looks identical to a livelocked storm once the DoE probe is gone) confirmed genuine REPL liveness via QMP send-key + screendump on aarch64/riscv64, not log inspection alone. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K |
||
|
|
862d7d9c48 |
Stage 3 follow-on: fix stack-ownership corruption + DoE switch columns (FABRIC-3.md §XXVIII.1)
DoE CSV gained 6 switch-signal columns (switch_count_cumulative, switch_current_slot, switch_*_readiness, switch_ticks_since), and verifying them with a boot-time HB-ON probe surfaced a real livelock: the preemption checkpoint could fire inside a VM-EXEC-nested execute_colon_word() call and switch away from a stack it didn't own, parking a borrowed region of the caller's stack under the wrong VM's saved-context pointer. The trampoline bounce was the visible (safe) half of this; the corruption was the quiet half, live in every prior "clean" Stage 3 boot without ever showing up in the log. Fixed by gating the checkpoint on being at the outermost vm_interpret() call (g_vm_interpret_depth / sk_vm_at_outermost_interpret(), vm_core.c), per Bob's decision. Also fixed two related bugs found in the same pass: g_switch_back_to was a single global stale after first entry, now per-VM state (native_switch_back_to); note_switch_performed() fired on resume instead of switch-out, now called before the switch. Verified on all 3 architectures: steady log growth (no freeze), zero leaked QEMU processes, DoE columns internally consistent, Hermes/Artemis confirmed genuinely executing (not just trampoline-bouncing). Temporary HB-ON boot probe reverted after capture. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016UNhH1mhi52i6Qihh7ZV5S |
||
|
|
986d042aa7 |
Stage 3: timer-driven preemptive switching, live on all 3 arches (FABRIC-3.md §XXVIII)
Fourth stage of the preemptive context-switching plan, and the biggest. LithosAnanke now genuinely, continuously preempts between Hera, Hermes, and Artemis -- timer-driven, running live for the entire remainder of every boot once the Tripod fleet registers, not a bounded probe. A real design fork was resolved before writing code: the naive approach (the timer ISR calling Stage 2's sk_vm_context_switch() directly) is broken -- Stage 0's trap frame lives on whatever stack was active at interrupt time, and jumping to a different stack via Stage 2's own independent swap mid-handler would abandon that trap frame unresumed, guaranteed corruption on the first tick. Chose the safer of two named options: the ISR only ever sets a flag and returns completely normally through its own full epilogue; the actual switch happens moments later, via Stage 2's already-proven mechanism, at a safe cooperative checkpoint on the mainline (execute_colon_word()'s per-word dispatch loop, checked on literally every word, not throttled to the existing 256-word heartbeat-tuning cadence) -- confirmed with the user that word-level granularity is fine-grained enough given the eventual Zynq FPGA target where a word is a mnemonic. New capsule_vm_switch_signal.c/.h: a purpose-built run-readiness signal, deliberately separate from capsule_vm_physics.c's execution-heat engine (that one's own header documents itself as never touched from interrupt context, by design). Slot table sized with headroom (8) rather than hardcoded to today's 3 participants, so extending participation later is another register() call, not a redesign -- per direct request to leave room for swapping the participant set. Simple linear accumulate-then- threshold for this first cut; a fancier law can replace it later without touching the mechanism around it. heartbeat_tick() gains its one deliberate, documented amendment to this file's own top-half/bottom-half discipline -- the first time this codebase reaches into VM-scheduling state from real ISR context. Registration happens only after all three VMs are fully born, right before the REPL starts -- no critical-section protection yet against being switched away mid-birth-setup. Known, flagged rough edge (not reconciled this pass): MSG-TICK's own idle-pump and this new mechanism can still independently move control between the same VMs; not observed to interact badly in verification, but not fully unified either. Verified interactively at the console on all 3 architectures with continuous background preemption running throughout -- amd64 computed `1 1 + .` -> 2, aarch64 computed `1 1 + dup DUP * . CR` -> 4, both correct, REPL fully responsive, zero fault indicators over sustained runtime. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016UNhH1mhi52i6Qihh7ZV5S |
||
|
|
f790d0995e |
Stage 2: cooperative VM context switch primitive, proven on all 3 arches (FABRIC-3.md §XXVIII)
Third stage of the preemptive context-switching plan. The real save/restore switch mechanism now exists -- the first time anything has ever executed on a VM's own native stack (Stage 1 allocated them, unused). New sk_vm_switch_to() (switch.S, one per arch) is an ordinary function call, not an interrupt -- so unlike Stage 0's trap frame, the ABI already covers every caller-saved register; only the callee-saved set needs explicit save/restore (amd64: rbx/rbp/r12-r15, no FP at all since SysV has no callee-saved XMM; aarch64: x19-x28/x29/x30 + d8-d15; riscv64: s0-s11/ra + fs0-fs11, FS-gated like Stage 0 but read once and reused for both halves within one call, since FS is genuine global CPU state, not part of what's switched). A sibling sk_vm_switch_prime() in the same file builds the synthetic first-entry frame, kept in assembly so the layout can never drift out of sync with sk_vm_switch_to() itself. New switch.c/switch.h: sk_vm_context_switch(from, to) handles first-entry priming vs. resuming a parked context, and updates registry state (new VM_STATE_SWITCHED_OUT, distinct from VM_STATE_STOPPED -- STOPPED means no live frame, this means the opposite). sk_vm_switch_entry() is the minimal permanent trampoline every freshly-entered VM lands in: no production behavior defined yet, so it just yields straight back to whoever switched to it, forever. Closes the confirmed unguarded-KILL UAF found during planning: capsule_vm_kill(), mama_word_kill(), and capsule_vm_kill_all_nonmama() all now refuse (or silently leak rather than free, on the cold-restart path where arch_cold_reset() wipes everything immediately after anyway) tearing down a switched-out VM. Side effect found, not built on purpose: the existing MSG-TICK idle-pump already filters on VM_STATE_LIVE, so it automatically stopped dispatching into a switched-out VM with zero changes needed there. Verified via a temporary SWITCH-TEST probe (boot-triggered, since nothing can type interactively into a foreground-only QEMU session) that round-tripped a sentinel through 5 real Hera<->Hermes switches on all 3 architectures: 5/5 rounds, 0 failures, clean continuation to ok>. Probe fully reverted after capture; kernel_main.c shows zero diff. Also: Makefile.starkernel's LOADER_EXTRA_SRCS/LOADER_ASM needed the new files added explicitly (this project's "loader" PE binary is the full running kernel, not a thin bootstrap stage), and aarch64's switch.S needed the same #ifndef _WIN32 guard around .hidden that isr.S already carries (aarch64's loader assembles via clang targeting a PE/COFF target with no .hidden equivalent) -- caught by a build failure, fixed. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016UNhH1mhi52i6Qihh7ZV5S |
||
|
|
57ac3fc304 |
Stage 1: per-VM native stacks, allocated but not yet executed on (FABRIC-3.md §XXVIII)
Second stage of the preemptive context-switching plan. Every VM (Hera, every capsule_birth_baby()-born VM including WIREBIND identities) now gets its own dedicated 2 MiB native C stack at birth -- but nothing runs on it yet, that's Stage 2. Pure allocation-machinery proof. Design correction made before writing code: the plan called for cloning sk_vm_arena_alloc()'s guard-page pattern, but that pattern turns out to be Mama-only -- host_services.c's kernel_alloc() gives every baby VM a plain kmalloc() block for its dictionary arena, not a real guarded PMM allocation. Stacks get the real treatment instead (new sk_vm_native_stack_alloc()/_free() in arena.c): independent pmm_alloc_contiguous() + guard pages for every VM without exception, no singleton, no kmalloc fallback -- a stack overflow is exactly the failure mode guard pages exist for, and a corrupted stack could corrupt whatever saved context Stage 2 trusts. 2 MiB size matches this project's own established kernel-stack convention (g_kernel_stack/g_rpi5_native_stack), not a guess -- that one shared 2 MiB stack today already carries all VMs' combined nested VM-EXEC recursion. Three new VM struct fields, freed in vm_cleanup() alongside the existing call_stack free. Allocation failure is non-fatal to birth. All 3 architectures re-verified clean boot to ok>, no native-stack allocation failures for any Tripod-fleet VM. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016UNhH1mhi52i6Qihh7ZV5S |