Baseline run for StarForth-v4.0.0 as it stands, before any further work.
`make -f kernel/Makefile ARCH=<arch> clean qemu` for amd64, aarch64 and
riscv64, one at a time, disk/artemis.img and
disk/thumbdrives/zuse-thumb-ident.img restored to their committed state
before each.
All three reach `[zuse@Hera] ok>`, with zero UNKNOWN WORD, PARITY:OK,
1050 POST tests run and ALL IMPLEMENTED TESTS PASSED,
stadium_conserved(Artemis)=true.
dict_hash is identical across the three:
Hera (PARITY:M7.1a, word_count=530) 0x08873e0f44b7cb2a
dict_hash 0xe11082140cf86b05
dict_hash 0xb1256603f848e2b9
dict_hash 0xc791409ac1715690
QEMU was asked to quit over QMP once the prompt had appeared and the
serial log had been quiet for 10 seconds.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
D-12 (ruled 2026-10-02): Q.SQRT of a negative value and Q.LOG of zero or
a negative value return 0 and set NODE-ERROR, as Q./ does on division by
zero.
Q.SQRT is v3's q48_sqrt_approx: Newton from x0 = q/2 + 0.25, up to 8
rounds of x' = (x + q/x) / 2, returning x once |x' - x| < 10 ulp;
sqrt(0) = 0 and sqrt(1.0) = 1.0. q, x and the rounds left live in a
5-cell variable (QR). Checked bit for bit against v3's q48_sqrt_approx
(ported into the test) on 19 edge values and 3000 pseudo-random q below
2^48, at both cell widths, optimised and ASan+UBSan. Above 2^48 v3's
q48_div saturates and v4 divides correctly, so there is no parity there.
Bug found and fixed on the way: the "< 10 ulp" test subtracted 10 from
the step's low cell as a signed number, so at 32-bit cells a step of
2^31 or more read as small and the iteration stopped early. It now
tests the low cell's top bit first, as Q.EXP does. The first test run
missed it because `make build/32-test_foundation.c` rebuilds nothing
(the Makefile's rules use absolute paths); the mutation runs, compiled
from source, exposed it. Only `make test` is used from here.
Headroom (data under arg / return): Q.SQRT 3/2.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Q.EXP is v3's q48_exp_approx: the Taylor series to 10 terms on |q|,
stopping below 50 ulp, 1.0 Q./ e^|q| for q < 0, e^0 = 1.0, and for
|q| >= 16.0 either 0 or Q max (v3 returned all ones, which is -ulp read
signed under D-8). x, term, sum, sign and term index live in a new
8-cell variable (QE). Checked bit for bit against v3's q48_exp_approx
(ported into the test) on 23 edge values and 3000 pseudo-random q in
(-16.0, 16.0), at both cell widths, optimised and ASan+UBSan; dropping a
term, moving the 50-ulp stop or skipping the reciprocal all fail.
Q.EXP calls Q./, which calls (UQ/), which calls D2*C, and at first it
left its caller no return entries. Lightened along the chain:
- D2*C takes cin in A and builds 2hi+m with `over + +`: no ROT, no SWAP.
- (UQ/) parks the quotient bit instead of using ROT, stores q without
SWAP, and does its full trial subtraction (with borrow) in line
instead of calling DNEGATE D+.
- Q.EXP keeps its term index in (QE) instead of a FOR count.
Headroom (data under args / return): (UQ/) 4/3 -> 4/4, Q./ 3/2 -> 3/3,
Q.EXP 3/2.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
D-11 (ruled 2026-10-02): Q./ rounds toward zero, saturates to Q max /
Q min on overflow, and on division by zero returns Q max / Q min by the
dividend's sign (0 for 0/0) and sets the new NODE-ERROR register (7),
which VM-ERROR? reads.
- (UQ/): unsigned floor(a * 2^16 / b) for b != 0 and no overflow.
Restoring division, 2N+16 steps, on one shifting register (quotient
above remainder). Quotient register and divisor live in a 5-cell
variable (Q/), like BASE and the hold buffer, so the stacks carry only
the remainder, the quotient bit and the loop count. A first version
kept the divisor on the return stack and overflowed it when called
from Q./ (it never returned).
- D2*C: double shift left with carry in and out.
- Q./: zero-divisor handling, signs (kept in (Q/)), an overflow test
that needs no division (|a| * 2^16 >= |b| * 2^(2N-1), possible only
for |b| < 2^17), then (UQ/) and the sign.
Checked at 32- and 64-bit cells, optimised and ASan+UBSan, against an
independent limb-by-limb long division (including NODE-ERROR): every
pair of 18 edge Q values, 13 width-native edge values (true Q max/min at
either width), the overflow boundary, divisors in [2^(2N-2), 2^(2N-1)),
and pseudo-random cases; and against v3's q48_div for non-negative
a < 2^48. Q.* now also runs on the width-native edge values.
Mutations of the compare paths, the take path, the error flag and the
overflow mask are all caught (the mask matters only for Q min / -1.0).
Headroom (data under args / return): Q./ 3/2.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Q.* now forms its products in the order a1*b1 (low cell), a0*b1, a1*b0,
a0*b0, dropping each input after its last use, with b0 waiting on the
return stack. Each sign correction, (b1<0 ? a0 : 0) and (a1<0 ? b0 : 0),
is folded into cell 2 as soon as its operands are adjacent, using -if
rather than 0< calls. At most four live values sit under any UM* call.
Headroom (data cells under args / return entries): 2/2 -> 3/3.
Same checks as before at 32- and 64-bit cells, optimised and ASan+UBSan;
dropping either sign correction fails about 9900 checks.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
UM* parked its three corrections (c_hi, c_lo, t0) on the return stack
before the +* loop, leaving its caller 4 return entries. Now u1 and u2
wait on the data stack under the loop -- +* touches only T, S and A --
and the corrections are made after it, with -if in place of 0< calls.
The return stack holds the loop count, then at most two temporaries.
The loop count is pushed before s is made, so the data stack never holds
more than four cells: data headroom is unchanged at 6.
Headroom (data cells under args / return entries): 6/4 -> 6/6.
Q.*, which calls UM* four times, goes from 2/1 to 2/2.
Same checks as before at 32- and 64-bit cells, optimised and
ASan+UBSan; breaking any of the four corrections fails 4900+ checks.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Q.* is floor(a*b / 2^16), signed (D-8), cut to two cells. Cells 0..2 of
the product come from the unsigned cell products a0*b0, a0*b1, a1*b0 and
the low cell of a1*b1; reading a1 and b1 as signed takes
(a1<0 ? b0 : 0) + (b1<0 ? a0 : 0) off cell 2. The result is cells 0..2
shifted right 16 with Q.TO-INT twice. Clobbers A.
Checked at 32- and 64-bit cells, optimised and ASan+UBSan, on every pair
of 17 edge Q values plus 20000 pseudo-random pairs:
- against an independent reference (16-bit limbs, magnitudes multiplied
schoolbook, negated, shifted), and
- against v3's q48_mul for non-negative operands.
The first draft applied only a1's sign correction; the reference caught
the missing b1 term (9860 failures at 32-bit cells).
v3 bug found, not fixed: q48_mul's portable branch (used where there is
no __int128) returns (result_hi << 16) | (result_lo >> 16); the high
part must move up 48, so it is wrong whenever the product passes bit
64, e.g. 0.5 * Q max gives 0000ffffffffffff instead of 3fffffffffffffff.
The parity reference here is that branch with the shift corrected,
which matches v3's __int128 branch.
Stack use is tight: Q.* leaves its caller 2 data cells and 1 return
entry, because UM* parks three values on the return stack. Recorded in
5.26; to be improved next.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Executed on the golden model as written, and correct: <, =, D<, 2OVER,
and Q.=, Q.<, Q.0= (the D=, D<, D0= words).
Fixed, because they could not work under D-2 or were wrong:
- DMAX, DMIN: 2OVER 2OVER D< needs 8 data cells plus D<'s 2, the whole
10-deep data stack, so both failed on every case once the caller held
anything. Rewritten on a new call-free helper
(D<) ( d1 d2 -- d1 d2 flag )
which compares copies of the high cells, then the low cells unsigned
on a tie, with in-line sign tests; DMAX and DMIN then drop the loser.
- 2SWAP: ROT and SWAP written in line. Return headroom 2 -> 4, which
also lifts 2OVER 1 -> 3 and Q.> 1 -> 2.
- Q.>: section 5.26 gave SWAP D<, which swaps single cells; it was wrong
in 11702 of 20169 cases. Now 2SWAP D<.
Checked at 32- and 64-bit cells, optimised and ASan+UBSan: every
edge-vector quadruple (50625) for D<, (D<), 2SWAP, 2OVER, DMAX and DMIN;
Q comparisons on every pair of 13 edge Q values plus 20000 pseudo-random
pairs weighted to ties and one-bit differences, signed (D-8) and against
v3's unsigned comparisons where both values have the same sign.
Mutating any (D<) branch or subtraction fails 700+ checks.
The new fatal-UBSan setting caught a test-side array overflow in the
headroom probe (results buffer sized 4, six needed); fixed.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Without -fno-sanitize-recover=all, UBSan prints a report and the run
carries on to "all v4 tests passed". The sanitize build now aborts on
the first report. Verified by rebuilding the test-side shift bug found
in 981f4180 with these flags: the run exits 1 at the report.
Both test binaries now also depend on v4/Makefile, so a flag change
rebuilds them instead of leaving stale binaries in place.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
DECOMPOSITION.md 5.26 described both words in prose only. They are now
written out, call-free, on one observation: +* with S = 0 never adds, so
each step is an exact arithmetic right shift of the double T:A.
: Q.FROM-INT ( n -- q ) push 0 a! 0 pop 15 FOR +* UNEXT push drop a pop ;
T:A = n:0 is n * 2^N; N-16 shifts leave n * 2^16 (count 15 at
32-bit cells, 47 at 64).
: Q.TO-INT ( q -- n ) push a! 0 pop 15 FOR +* UNEXT drop drop a ;
16 shifts at every width; the low cell is left in A. Rounds toward
minus infinity, as v3's arithmetic shift does.
Both clobber A; added to the section 2 list.
Checked at 32- and 64-bit cells, optimised and ASan+UBSan:
- Q.FROM-INT against n * 2^16 for edge values and 20000 pseudo-random n,
and against v3's q48_from_u64 for n >= 0 (low cell at 64-bit, D-10).
- Q.TO-INT against v3's (int64_t)q >> 16 on the edge Q values and 20000
pseudo-random Q values, and against a C double shift on 20000
arbitrary doubles.
- Round trip Q.TO-INT(Q.FROM-INT(n)) = n.
Mutating either shift count or the S = 0 setup fails 40000+ checks.
Note: make sanitize prints UBSan reports but does not fail on them. One
was found here, in the test's own random shift amount (fixed); a grep of
the full sanitize output now shows no runtime errors.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
D-10 (ruled 2026-10-02): a Q48.16 value is two cells at every cell
width, so every Q word is the same double word on 32- and 64-bit nodes.
Recorded in DECOMPOSITION.md section 3 and 5.26; the open note on the
Q.+/Q.- row is resolved by it.
Q.ABS and Q.NEG are the DABS and DNEGATE words. Checked against v3's
q48_abs and 0 - q on the 15 edge Q values and 20000 pseudo-random
values, at 32- and 64-bit cells, optimised and ASan+UBSan:
- 32-bit cells: bit-for-bit v3, including v3's wrap of Q min to itself.
- 64-bit cells: the low cell is v3's; the high cell is the true sign
(so ABS and NEG of Q min give +2^63 rather than wrapping), per D-10.
Breaking DNEGATE's carry fails Q.NEG, Q.ABS and Q.- checks.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
DECOMPOSITION.md 5.26 makes Q.+ and Q.- the D+ and D- words. They are
checked against v3's q48_add/q48_sub (uint64_t a + b, a - b, wrapping)
on every pair of 15 edge Q values (0, ulp, 0.5, 1.0, 1.5, -1.0, -ulp,
values across the 32-bit seam, Q max and min, +-12345.0) and 20000
pseudo-random pairs, 2682 of which overflow Q48.16.
- 32-bit cells: bit-for-bit v3's result, overflow wrap included.
- 64-bit cells, with a Q value as a sign-extended double: the low cell is
v3's result; on overflow the high cell holds the true carry where v3
wraps. Whether a Q value is one cell or two on a 64-bit node is not
ruled; recorded as open in 5.26.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
All four run exactly as written in DECOMPOSITION.md 5.6/5.7 and need no
change now that D+ and DNEGATE are call-free.
Checked against C at 32- and 64-bit cells, optimised and ASan+UBSan:
every edge-vector combination (M+ over all triples, D- and D= over all
50625 quadruples, D0= over all pairs) plus 20000 pseudo-random cases
weighted to equal low or high cells and low sums that wrap to 0.
Mutations of D0= and M+ are caught.
Headroom (data cells under args / return entries under return address):
M+ 6/4, D- 5/4, D0= 7/7, D= 5/3.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
D+ as written in 5.7 was exact but kept two cells on the return stack
while calling U> -> SWAP/U<, leaving its caller one return entry: any
word calling a word that calls D+ (M+, D-, D=, Q.+, Q.-) would have a
return address silently overwritten.
The new D+ makes no calls. Both high cells wait on the return stack; the
carry out of the low-cell add comes from sign tests -- if the low cells'
top bits differ, there is a carry exactly when the sum's top bit is
clear; if they match, exactly when both are set.
Headroom (data cells under args / return entries): 4/1 -> 6/5.
Checked against C over all 50625 edge-vector quadruples and 20000
pseudo-random pairs (weighted to top-bit cases and low sums wrapping to
0), at 32- and 64-bit cells, optimised and ASan+UBSan. Retargeting each
of the four carry branches fails more than 13000 checks.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
SM/REM as written in section 4 never returned correctly for a negative
dividend. It holds three entries on the return stack and then calls
DABS -> DNEGATE -> D+ -> U> -> SWAP/U<, which overflows the 9-deep
circular return stack (D-2). DABS itself could not run: DNEGATE as
written (inv SWAP inv SWAP 1 0 D+) left its caller no return entries.
- DNEGATE: inv over if L1 drop push inv 1 + pop ; L1: drop 1 + ;
i.e. ~d + 1, carrying into the high cell exactly when lo = 0.
- SM/REM: sign tests are native -if (as in 0<) and NEGATE is in line;
the only calls are DNEGATE and UM/MOD, both call-free inside.
Executed on the golden model at 32- and 64-bit cells, optimised and
under ASan+UBSan:
- SM/REM on dividends built as q*n + r with |r| < |n| and r signed as
d: every edge-vector q, n with r = 0 and r = +-(|n|-1), plus 20000
pseudo-random cases.
- /MOD, U>, ABS, S>D, D+ and DABS as written in section 5, against C
(D+ over all 50625 edge-vector quadruples).
Mutations of DNEGATE's carry and of each SM/REM sign branch are caught.
Headroom (data cells under args / return entries under return address):
SM/REM 5/3, /MOD 5/2, DNEGATE 7/7, DABS 7/6, ABS 8/7, U> 6/5.
D+ as written is exact but leaves only 1 return entry; recorded in
DECOMPOSITION.md 5.7 as not yet revised.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The first UM/MOD was exact but called 0<, U<, SWAP and OR inside its
loop, so it left its caller only 3 return-stack entries. SM/REM pushes
two signs before calling it, so under /MOD, M/MOD or */MOD the caller's
return address would be silently overwritten (D-2 circular stacks).
The loop now makes no calls. It branches on hi's top bit with -if, does
the unsigned hi' >= d test as U< does but with in-line sign tests,
subtracts with `inv a + inv`, and sets the quotient bit with `1 +` on an
even lo'. The final SWAP is in line.
Measured on the golden model at 32- and 64-bit cells:
headroom data 3 -> 6 cells under args, return 3 -> 6 entries
speed ~1355-1605 -> ~227-313 instruction words per call
Still exact for every uhi < ud (edge-vector triples and 20000 random
cases, optimised and ASan+UBSan). Retargeting each of the four in-loop
branches to the wrong label fails more than 12000 checks each.
DECOMPOSITION.md: section 4 UM/MOD replaced, with its derivation and
stack limits.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
UM/MOD is assembled exactly as DECOMPOSITION.md section 4 gives it and
needs no change: it is exact for every uhi < ud at 32- and 64-bit cells.
It is checked against q*d + r = uhi:ulo with r < d through v4_umul, so
no second C divider has to be trusted. Coverage: all edge-vector triples
with uhi < ud plus 20000 pseudo-random cases (top-bit, small and
near-maximum divisors), optimised and under ASan+UBSan. Two hand
mutations each fail more than 15000 checks.
New headroom probe: runs a word with marked cells under the canary and
under its return address and reports how many survive, since the D-2
circular stacks overwrite silently instead of faulting. Measured:
UM* 6 data cells under its args, 4 return entries under its return
UM/MOD 3 data cells under its args, 3 return entries under its return
DECOMPOSITION.md: UM/MOD marked executed, with its defined range and
stack limits.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
UM* as first written in DECOMPOSITION.md section 4 was exact only while
u1 <= 2^(n-2): plain +* loses the carry out of T and its shift keeps T's
sign bit, so the loop is exact only while S and T stay in
[-2^(n-2), 2^(n-2)).
The rewrite multiplies by s = u1 2/, which always lies in that range,
starting T at t0 = u2 2/ when u1 is odd, so the loop yields
hi:lo = t0 + s*u2 exactly. Then
u1*u2 = 2*(hi:lo) + c_lo + c_hi*2^n
c_lo = u1 & u2 & 1
c_hi = (u1<0 ? u2 : 0) + (u1 odd and u2<0 ? 1 : 0)
restores the halved-away bits and the unsigned reading of both top bits.
test_foundation.c runs the new definition against v4_umul over every
pair of the edge vectors plus 20000 pseudo-random pairs, at 32- and
64-bit cells, optimised and under ASan+UBSan. The two pinned failing
cases are now ordinary exactness checks. Two hand mutations of the
correction step each fail more than 10000 checks at both widths.
DECOMPOSITION.md: section 4 UM* replaced, D-3 ruling text updated.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
First code for StarForth v4 (JUSTIFICATION.md section 10, step 1): one node
of the 32-instruction core as a C99 model, with cell width as a build
parameter.
- Node: P, A, B, F18 circular stacks (10 and 9 deep, D-2), word-addressed
memory (D-1), 5% guard bands on every bounded list.
- Instruction word: six 5-bit slots in 32 bits at every cell width.
- Executor: all 32 opcodes of DECOMPOSITION.md 1.3. Cell arithmetic wraps
explicitly; no signed overflow or implementation-defined shift.
- Heat: per-opcode and per-call-target counters and the anti-clock, driven
by instruction retirement (1.4, D-6 interim).
- Slot packer and runner for tests, and a reference unsigned multiply in
plain C99 with no 128-bit type.
Tests run at 32- and 64-bit cells, and under ASan and UBSan. They cover
every opcode and execute the first section 4 definitions (NIP SWAP OR
NEGATE ROT 0< 0= 2DUP - U<) against the C operation each stands for.
UM* as written in section 4 is exact only while u1 <= 2^(n-2). Two known
failing cases are pinned in test_foundation.c until it is rewritten.
DECOMPOSITION.md: record D-9, the instruction word is 32 bits at every
cell width (ruled 2026-10-02).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Source tree reorganization:
- Move StarForth v3 engine to v3/ (src/, include/, Makefile)
- Move kernel to kernel/ (src/, include/, linker/, Makefile)
- Create v4/ skeleton for F18-ISA golden model (DECOMPOSITION.md, JUSTIFICATION.md)
- Move FABRIC-0..4.md to docs/fabric/
- Move ONTOLOGY.md and ROADMAP.md to docs/
Board infrastructure:
- Add boards/ser5/, boards/raspi/, boards/milkv/, boards/zynq7020/
- Each board has board.mk (ISA, CPU flags, boot recipe) and README.md
- Root Makefile becomes thin dispatcher: boot_image, all, clean, docs take TARGET
- make boot_image TARGET=SER5|RASPI|MILKV builds one GPT/MBR image per board
- ZYNQ7020 target exists but stops with clear error (ARMv7 port not built yet)
- scripts/mkdiskimage.sh builds disk images for all boards
Docs pipeline:
- docs/book/ with LaTeX master (main.tex) and Makefile
- pandoc converts Markdown to LaTeX at build time
- Two Lua filters: table-widths.lua (wide tables wrap), code-breaks.lua (inline code breaks)
- make docs builds single PDF (754 pages, 0 missing characters)
- make docs TARGET=<board> adds board appendix
- build/docs/<book|board>/meta.tex stamps git commit into PDF
Bug fixes:
- 42 include paths that only worked by accident now use correct relative paths
- clang-18 hardcode replaced with configurable CC variable (fixed aarch64 build)
- Pi 5: kernel_2712.img linked at 0x80000, .bss zeroed, memory reserved
- Doxyfile, .clang-tidy, README.md, Kconfig paths updated
Verified:
- Hosted v3 build passes 1012 tests, 0 failures
- SER5 image boots in QEMU (OVMF), POST passes, K exact (65536 = Q48_ONE)
- Milk-V image boots in QEMU (OpenSBI + U-Boot + bootefi), POST passes
- make clean TARGET=<board> removes only that board and its ISA objects
- make all builds all boards, hosted v3, and docs in one run
Co-authored-by: Junie <junie@jetbrains.com>
D-1 word addressing; D-2 F18 circular stacks (10/9 deep), hidden, so
DEPTH/PICK/ROLL/.S/SP@/SP! are retired everywhere and the DSP register is
dropped; D-3 plain F18 +* (UM* flagged for revision); D-5 host width
matches the host CPU; D-7 moot; D-8 signed Q48.16; D-4 and D-6 deferred
to the hosted-mesh step. Adds a note that every CAP definition must be
re-checked against the 10/9 stack depths.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BY9HMwK5Cetz3caBgHGyds
JUSTIFICATION.md records why v4 exists and the reasoning behind each
major design decision. DECOMPOSITION.md assigns every v3 C primitive a
fate on the 32-instruction F18-derived core.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BY9HMwK5Cetz3caBgHGyds
Distinguishes it clearly from the still-accepted "cloned outside
StarshipOS entirely" limitation this document already settled -- Phase
8 v3 closes a narrower, different threat: cloning block content using
StarshipOS's own console primitives, now refused by MOVE/CMOVE/CMOVE>/
RELOCATE-BLOCK for cross-device copies.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Two research passes confirmed the identity record itself (seed/pubkey/
cert) is already unreachable from any FORTH primitive -- only C-level
read_devblock/write_devblock touch it. But a device's ordinary user block
content CAN be copied between two attached devices today, using only
stock, unpinned words: <src> BLOCK <dst> BUFFER 1024 MOVE UPDATE
SAVE-BUFFERS, or the dedicated RELOCATE-BLOCK word (whose own doc comment
already admits "performs no policy validation of its own"). Checked
whether the existing per-block owner_fp/BLK-ACL-ALLOW@ metadata already
solves this -- it doesn't: owner_fp encodes who (a VM identity pubkey),
never where (physical device), and blk_get_buffer()/blk_update() never
consult acl_allow/acl_ttl at all -- those fields are completely inert.
Small, targeted fix, no rearchitecture:
- blk_subsys_relocate_block() (block_subsystem.c): same-device check.
Its own documented purpose is wear-leveling (relocate on the SAME
device) -- never stated as cross-device, and nothing enforced that
until now.
- New public blk_lbn_device_handle() (block_subsystem.c/.h): the missing
LBN-to-device direction (blk_get_device_range() already goes the other
way). Opaque, stable, == comparable.
- MOVE (memory_words.c) and CMOVE/CMOVE> (string_words.c): refuse when
both addresses are block-window addresses backed by two different
devices -- the exact shape of the composed attack. A copy where either
end is ordinary VM memory (the overwhelming common case: staging text
from PAD, editing a block in place) is untouched.
- blk_vm_check_epoch()/blk_vm_slot_for_addr() exposed (block_words.h) so
the two new call sites share the same window-slot invalidation contract
rather than a second, divergent copy of it.
Verified live on all three architectures, not just boot-clean: same-
device MOVE/RELOCATE-BLOCK still succeed exactly as before; cross-device
MOVE/CMOVE/RELOCATE-BLOCK all refused. Zero UNKNOWN WORD, identical
dict_hash across all three (this change adds no FORTH-visible word, only
internal refusal conditions, as predicted).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Captain Bob rejected a PIN/passphrase second factor firmly and directly
after it was built and live-tested on all three architectures: "nothing
like a pin or a password or secret code or any bullshit... Everybody has
secrets. There's only the drive." All PIN-related code (KDF, XOR
keystream seed encryption, no-echo input, MINT/WIREBIND prompts,
user_identity_seed_t v3 format) was reverted before commit -- none of it
ever landed in git history.
This project's identity model has no knowledge factor, by design:
physical possession of the thumbdrive is the entire credential. A
byte-for-byte clone being equivalent to the real drive is the accepted
model, not a gap needing a fix. Recorded here so future work doesn't
default back to a PIN/password approach.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Found while starting Part A's implementation, before any code was written
against the original design: reading the actual deleted SEND-ELEVATE-REQUEST
source (git show 3e201c8^:capsules/common/messaging.4th, block 5040) shows
it copied the target word's name as literal character bytes into a scratch
buffer, building "S" <name-text>" <pk0> <pk1> <pk2> <pk3> ELEVATE-GRANT"
entirely in the sending VM's own memory, then sent that finished string --
never a raw address -- to Hera. waddr/wu never crossed the VM boundary as
numbers anywhere in this flow. The original write-up reasoned from
ELEVATE-GRANT's own signature alone, without first reading how the caller
actually built its message.
Corrected in place, wrong original text kept struck-through for
traceability rather than deleted, per this series' own convention.
Net effect: Part A (the buffer/message redesign) is not needed -- the
original mechanism was already safe. Part B (capsules/zuse.4th, already
committed and three-arch verified this session, 089ab21) stands on its
own, unaffected.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
ZUSE-ELIGIBILITY-ADD's own doc comment admitted "no authorization check
here or anywhere else... applied later if and when actually needed --
not invented here." That's now: anyone reaching a Hera FORTH prompt
could add their own pubkey to the eligibility list with zero legitimate
identity material -- no minted drive, no WIREBIND, no cert-signature
check involved at all. Once a future caller reaches ELEVATE-GRANT again,
a self-added pubkey would pass zuse_eligibility_is_member() and grant
ACL-ALLOW!/ACL-TTL! on any named word.
Fixed the FORTH-only way, matching this project's own convention (ACL
policy belongs in ACL.4th, never in C; never gate on zuse_session in C --
her power is the absence of ACLs, not a hardcoded session check):
ZUSE-ELIGIBILITY-ADD is now denied by default (capsules/zuse.4th block
4016), granted and pinned only inside ACL-ZUSE-BOOT's already-existing
authenticated branch (block 4017) -- the same gate her own god-mode
already goes through, requiring a real cert-verified Zuse before it opens.
Live-verified on all three architectures, not just boot-clean: after
genesis authentication, ACL-ALLOW@ and ACL-PINNED? both read -1, and
HERE ZUSE-ELIGIBILITY-ADD executes successfully past the ACL gate.
Phase 8 v1 plan: /home/rajames/.claude/plans/jiggly-cuddling-stallman.md
Part A (the ELEVATE-GRANT pointer-confusion fix, FABRIC-3.7.md) is
separate, not yet built.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
The hosted Makefile writes include/version.h with FORCE as a prerequisite
(always regenerates), but Makefile.starkernel's own rule had no
prerequisite at all -- Make only rebuilds a target with no prerequisites
when the file is missing. Since both Makefiles write the same path with
incompatible content (the hosted version has no LITHOS_VERSION/
LITHOS_VERSION_STR at all), running a bare `make -f Makefile.starkernel`
after a hosted `make` build silently reused the wrong file and failed
deep in kernel_main.c with "LITHOS_VERSION_STR undeclared".
Added FORCE (declared .PHONY, matching the hosted Makefile's own existing
pattern) as include/version.h's prerequisite in Makefile.starkernel.
Verified the fix directly: poisoned version.h with a hosted build, then
ran a bare (non-clean) kernel build and confirmed it self-heals.
Full three-architecture acceptance: amd64 (logs/20260922-215333/),
aarch64 (logs/20260922-215803/), riscv64 (logs/20260922-220217/) -- all
three reach [zuse@Hera] ok>, zero UNKNOWN WORD.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
New document, not a reopening of the closed FABRIC-3.5.md/FABRIC-3.6.md --
successor for exactly one topic, per those documents' own close discipline.
Records a security defect found by inspection while auditing Phase 4's
collateral damage to ELEVATE-GRANT: the old SEND-ELEVATE-REQUEST mechanism
passed a raw address (waddr/wu) computed in the sending VM's own memory
space across to Hera, which dereferences it in Hera's own space --
per-VM vaddr_t means those are never the same address space. Whoever
controls waddr controls what dictionary entry NAME>XT resolves to on
Hera, independent of the caller's actual pubkey/eligibility.
Design fix: never cross an address, only ever cross bytes -- generalizes
this session's own payload-aliasing fix (SkHermesMessage.payload_buf) one
level up. Send the target word's name as inline payload bytes, copy them
into a fixed kernel-owned buffer already in Hera's own memory on receipt,
and hand vm_interpret() only kernel-controlled integer literals referencing
that buffer. ELEVATE-GRANT itself is unchanged -- policy logic stays in
FORTH, per ACL.4th's own rule.
No code written or authorized by this document.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
FABRIC-3.5.md's provisional "design phase closed, not yet archival" header
replaced with the real CLOSED/ARCHIVAL form per its own §XXVI.5 spec,
naming v2.1.0 -- prior status headers kept underneath, not deleted.
FABRIC-3.6.md gets the same treatment: a CLOSED/ARCHIVAL banner above the
START HERE section, which stays as historical record rather than being
removed. Neither closure triggers the FABRIC-0 -> -1 -> -2 -> -3
carry-forward chain (both are standalone topic documents) and neither
touches FABRIC-3.md, which remains open for its own topic.
The Tripod/kernel reshuffle is complete: Hermes moved into the kernel as
kernel-Hermes, the Tripod is Hera/Artemis/Hestia, tagged v2.1.0.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
refs/heads/v2.0.1's tip is the exact merge-base with master -- a strict
ancestor, fully subsumed, confirmed by its own commit message ("master
fast-forwarded to v2.0.1, verified"). Stale ref hygiene debt, safe to
delete, not deleted here -- needs Captain Bob's explicit go-ahead.
PR #1 turned out not to be a stray reference at all: it's a PR from this
same branch against master, already closed (not merged) 20 seconds after
it was opened on 2026-09-18, before any of this reshuffle's work existed.
Nothing to resolve beyond recording what it is.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Advisor review of the Phase 5 close-out commit caught two real issues:
- Task 5.1's write-up claimed the Isabelle pass "confirms no regression" --
overstated. Nothing in proof/'s scope changed, so the pass isn't
regression evidence, it's a build-completeness formality; the boundary
argument alone already supports the real conclusion. Sharpened.
- sbom.spdx was regenerated before the 5.4 version bump, so it briefly
understated the engine version. Root cause: the hosted Makefile carries
its own separate hardcoded VERSION (Makefile:17, still 3.1.0), distinct
from Makefile.starkernel's copy that 5.4 bumped -- duplication CLAUDE.md's
own "two independently tracked version strings" note doesn't document.
Bumped to 3.2.0 to match, regenerated (PackageVersion now correct), and
rebuilt the hosted starforth binary so lfs/amd64/starforth reflects it.
Also recorded why 5.1-5.4 landed as one commit (deviation from this
document's per-task-commit discipline) and sharpened the riscv64 log
entry so the FAILED run and the accepted retry aren't ambiguous to a
future reader grepping logs/.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
5.1: Isabelle/HOL pass (52 theories, clean) -- restated the boundary rather
than just citing the green build: proof/ scope was already entirely
outside this reshuffle's footprint (src/starkernel/, capsules/*.4th),
so the boundary is unchanged, not moved.
5.2: Documentation sweep. CLAUDE.md's stale WIP banner and Tripod fleet
description updated now that Phases 0-4 have actually landed (Hera/
Artemis/Hestia, no Hermes). MANIFEST.md rides the strip -- hermes/init.4th's
block table replaced with a deletion note, init.4th/doe-campaign.4th/
hestia/init.4th entries corrected to match the post-strip live files.
Confirmed the TRIPOD.md/0.1 contradiction was already resolved (2026-08-13).
Settled the superseded-docs call explicitly: archive as-is, do not rewrite.
Fixed experiments/bare_metal/README.md's block-size framing (still said
1024-byte budget; real rule is 64 chars x 16 lines). K-qualification
checked clean against the two living documents; full retroactive sweep
of the closed archival FABRIC corpus explicitly declined as disproportionate.
5.3: make sbom. Installed syft (user-local, approved). Found and fixed a
real Makefile bug while at it -- the sbom target hardcoded
--source-name StarForth, so DocumentName was wrong even after regenerating.
5.4: LITHOS_VERSION 2.0.0 -> 2.1.0, engine VERSION 3.1.0 -> 3.2.0 (minor,
per the dictionary-visible-only rule). Replaced the stale version-comment
block in Makefile.starkernel (had the odd/even LTS rule backwards) and
docs/lithosananke/ROADMAP.md's retired versioning-policy section with the
ratified ladder. Verified on all three architectures; riscv64's first pass
hit a transient virtio_blk timeout during boot-time Zuse genesis mint,
reported and confirmed non-reproducing on an immediate clean retry.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
All FORTH-owned message types were cut over to kernel-Hermes in Phase 3
(tasks 3.8-3.10). This removes the now-dead FORTH messaging layer and
the Hermes VM itself: capsules/common/messaging.4th, capsules/hermes/init.4th,
the slot-3 VM-NAME-REG pairing convention, and every load-site/birth-site
reference across capsules/init.4th, artemis/init.4th, hestia/init.4th,
doe-campaign.4th (Artemis-only now), capsule_console.c, capsule_mint.c,
capsule_wirebind.c, capsule_birth.c, and kernel_main.c.
Verified on all three architectures: clean boot, mkcapsule --lint clean
(36 files, 0 violations), zero UNKNOWN WORD, identical dict_hash across
amd64/aarch64/riscv64 for every VM, and a full mint -> WIREBIND-attach ->
USE -> relay round-trip exercising the two highest-risk edits
(capsule_console.c/capsule_mint.c).
Found, not fixed: deleting messaging.4th removes SEND-ELEVATE-REQUEST,
which was the only caller of KH-ELEVATE-SEND and the only path to
ELEVATE-GRANT (zuse-eligibility.4th, still loaded at boot) -- Phase 8
PKI's own elevation entrypoint. Needs a decision before Phase 5 close-out.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Reachability verified live before writing any code, per this
project's own standing rule (grep cannot establish reachability
alone): FIND SEND-ELEVATE-REQUEST / FIND ELEVATE-GRANT / FIND
CH-REQUEST all resolve on a live Hera boot, though grep across
capsules/experiments/docs found zero callers of SEND-ELEVATE-REQUEST
-- a real, complete, directly-callable entrypoint (H.5/H.8's own
design) with no current automatic trigger, not dead code.
Correction to a prior finding, made in the course of this check: task
3.8's write-up claimed "Hera's own pre-existing inability to load
common:messaging.4th" -- false. capsules/init.4th (Hera's own
MAMA_INIT capsule) loads it directly, and SEND-ELEVATE-REQUEST lives
and works in her dictionary right now. Task 3.8's own actual scope is
unaffected by this correction.
Cutover: SEND-ELEVATE-REQUEST (messaging.4th) no longer ends in
CH-REQUEST's COMMON-CH/MSG-SEND path; it now calls KH-ELEVATE-SEND
(repl.c), a new C word wrapping sk_hermes_send_one(), registered
unconditionally for every VM. from/to are derived from the calling VM
and sk_get_mama_vm() directly in C, never taken from the stack -- a
real correctness improvement over CH-REQUEST's own initiator-only
gate, which only existed because a caller COULD pass the wrong from
value; deriving it in C makes that spoof structurally impossible.
SK_HERMES_MSG_TYPE_ELEVATE_REQUEST reuses ELEVATE-REQUEST's own value
(8), same partition-rule reasoning as tasks 3.8/3.9. Delivery is
unchanged task 3.4 machinery. No new static-buffer lifetime caveat --
the payload-aliasing fix landed before this task started.
CH-REQUEST (messaging.4th) is now dead code, its one real caller just
removed -- found, not fixed, per Captain Bob's Law.
Verified live on all three architectures: 0 0 0 0 S" DUP"
SEND-ELEVATE-REQUEST (deliberately-invalid pubkey, so ELEVATE-GRANT
correctly refuses -- the check is the pipeline running, not a grant
succeeding) fires the evidence line and completes cleanly, DUP
unaffected afterward. Zero UNKNOWN WORD, dict_hash identical across
all three architectures (changed uniformly from prior runs -- one new
word registered -- not diverged, matching SXXXIV.6's own rule).
This closes task 3.11 (Phase 3 gate): all of messaging.4th's live
FORTH-owned message types (BLK-ATTACH-EVENT, CONSOLE-CMD-EVENT,
ELEVATE-REQUEST) are now real kernel-Hermes cutovers. Phase 4 may
begin.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Two separate fixes, found and closed together after task 3.9/prompt-
format landed (Captain Bob: "finish the work first then we'll do
cleanup before 3.10 begins").
1. Logging noise (confirmed live via QMP screendump): three sources
were cluttering ordinary interactive console sessions.
- INFERENCE "Output validation failed, ignoring results"
(vm_runtime.c, kernel; vm_time.c, hosted mirror) and xhci "CSW
status = FAILED"/"unit not ready -- retrying" (xhci.c) were
already log_message(LOG_WARN/LOG_ERROR, ...) calls, just visible
at the default runtime LOG_WARN level -- downgraded to LOG_INFO,
all three fire routinely and self-resolve (xhci.c's own existing
comment already documents the retry as expected SCSI UNIT
ATTENTION behavior, not a driver defect).
- Stadium: dispatch cell=... (stadium.c's stadium_dispatch()) was a
genuine defect: an unconditional console_puts()/console_println()
sequence with no level gating at all, printing on every single
dispatch. Rewritten through log_message(LOG_DEBUG, ...).
2. CRLF double-submit (found while investigating why the cleaned-up
noise still didn't look like a normal single-VM session):
sk_console_readline() (repl.c) breaks on '\r' OR '\n' as independent
terminators, so a line sent as both bytes submits twice -- the real
line, then an immediate empty-line submit on the second byte, each
printing its own " ok". Pre-dates Stage D entirely, not an async-
relay artifact. Fixed with a single non-blocking peek-and-discard
for the paired byte right where the line terminates.
Verified live on all three architectures for both fixes: zero UNKNOWN
WORD, dict_hash identical to every prior acceptance run in this
document (both changes are display/interaction-only).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
sk_hermes_send_one() -- the single funnel every sender, including
sk_hermes_publish(), already goes through -- used to store the
caller's own payload_addr pointer as-is. Two sends before either
drains meant both messages pointed at the same caller-owned buffer,
whichever send wrote last silently winning: real, confirmed live (a
second identity thumbdrive attached at boot alongside Zuse's own left
its WIREBIND pairing silently never happening).
SkHermesMessage gains an inline payload_buf[SK_HERMES_CHUNK_MAX_
PAYLOAD] field; sk_hermes_send_one() now memcpy()s the caller's
payload into it and points payload_addr at that copy instead. No
sender or reader call site needed to change -- every existing reader
already only ever reads through payload_addr, which still points at
valid bytes of the same length, now message-owned. g_kh_blk_attach_buf/
g_kh_console_cmd_buf (repl.c, tasks 3.8/3.9) no longer need to survive
past their own send call; their doc comments, which had claimed the
old aliasing shape was benign, are corrected.
Verified the original bug is actually gone: reproduced the exact
original scenario (a real, sequentially-minted rajames identity
attached at boot alongside Zuse's own, two simultaneous BLK-ATTACH-
EVENT sends in one idle-loop pass) -- WIREBIND pairing, USE, and the
Stage D relay all now work where WIREBIND previously silently failed.
Standard three-ISA acceptance also clean: zero UNKNOWN WORD, dict_hash
identical across all three and matching every prior acceptance run in
this document.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
sk_console_user_prefix() (repl.c) previously always returned "zuse"
(or the WIREBIND-attached username) as the left side of the bracket
prefix, regardless of which VM the console was actually pointed at --
"[zuse@rajames]" after USE rajames, always showing the authenticating
superuser rather than the active identity.
Changed on Captain Bob's direct instruction: once the console is
redirected into a WIREBIND identity's own console VM
(console_get_vm_name() != "Hera"), show that same name on both sides
-- "[rajames@rajames]" -- since WIREBIND births the console VM
literally named after the identity, so the identity IS that VM, not a
separate label. At the top level (still on Hera, nothing has
redirected yet), the original zuse_session/WIREBIND-username logic is
unchanged.
An earlier, more ambitious attempt (separate identity/machine tracked
state across every console_set_vm_name() call site) regressed live to
a wrong [zuse@Artemis] prompt and was fully reverted before reaching
any acceptance run -- the landed fix needed none of that new state,
just this one function.
Verified live on all three architectures: [zuse@Hera] at the top
level and after a live WIREBIND attach (before USE), [rajames@rajames]
after USE rajames, with task 3.9's Stage D relay still firing
correctly on top of it. Zero UNKNOWN WORD, dict_hash unaffected
(display-only change).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Send-side cutover: sk_repl_dispatch_line()'s FORTH-string
"CONSOLE-CMD-EVENT 0 3 S\" ...\" 0 MSG-SEND" interpret is replaced with
a direct sk_hermes_send_one() call (SK_HERMES_MSG_TYPE_CONSOLE_CMD,
kernel_hermes.h, deliberately reusing CONSOLE-CMD-EVENT's own value 7,
same partition-rule reasoning as task 3.8's BLK-ATTACH-EVENT cutover).
Real finding along the way: kernel-Hermes's drain only ever runs as a
side effect of vm_interpret() being called on the target VM. Task
3.8's target (Hera) is always being interpreted via the interactive
REPL loop; task 3.9's target is a WIREBIND identity's own ~user VM, a
passive receiver nothing else drives. The existing idle-loop pump only
ticked VMs with the old FORTH MSG-TICK word ACL-allowed -- a VM minted
with the STD79-lockdown personality never has it, so the pump silently
skipped it forever and queued messages never delivered. Fixed by
adding an unconditional, direct sk_hermes_drain_checkpoint() call in
the same pump loop, independent of the MSG-TICK gate.
Verified live on all three architectures: WIREBIND-attach a real
identity, USE into it, type a plain console line, confirm the Stage D
evidence line and correct relayed execution result. Zero UNKNOWN WORD,
dict_hash identical across all three ISAs and matching task 3.8's own
baseline. The FABRIC-3.md-documented USE/BINDSTEP crash did not
reproduce in any of these live sessions (recorded as a finding, not
chased further).
Depends on the zuse_root_pubkey_known fix already landed in ac4d431 --
without it, WIREBIND cannot attach any identity on a fresh boot at
all, which blocked this task's own verification until found and fixed
separately.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
capsule_zuse_boot_load_root_pubkey() is the only writer of
zuse_root_pubkey_known, and it only ever runs once, synchronously, at
Artemis's boot-time virtio-blk attach -- before genesis mint has
happened on a fresh artemis.img, so it finds no marker yet and leaves
the flag 0. Nothing re-triggers it after genesis mint completes.
Silent result: capsule_wirebind_try_attach()'s own
`if (!mama_vm->zuse_root_pubkey_known) return;` gate then refuses
every identity attach for the rest of that boot, with no message at
all -- meaning WIREBIND was silently dead on any boot that reset
artemis.img fresh, which is exactly what this project's own
acceptance convention does before every run.
Found live verifying FABRIC-3.6.md task 3.9's console-session
acceptance. Fixed at the source: install_and_activate() already has
the root pubkey in hand (from genesis mint or an already-Zuse
re-attach) -- activate it directly there instead of relying on a disk
re-read that may never happen in-session.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
While orienting for task 3.9 (CONSOLE-CMD-EVENT cutover, per Captain
Bob's ruling to verify via a real QMP/serial-socket console session),
booting with a second real identity drive attached alongside Zuse's
own exposed a real defect in the already-closed task 3.8 code:
g_kh_blk_attach_buf (repl.c) is one static buffer, and
sk_hermes_send_one() stores payload_addr as a caller-owned pointer,
not a copy. Two real USB-MSC attaches in one sk_repl_idle() pass send
before either drains, so both messages alias the same buffer -- only
one identity ever completed.
Task 3.8's own write-up claimed this "carries the same single-buffer-
reuse shape ATTACH-ACK-BUF itself already had... not a new hazard" --
that claim was wrong and is amended in FABRIC-3.6.md's findings log.
FORTH's own MSG-SEND had the identical pointer-aliasing shape but
never hit the window: MSG-TICK drained from the same sk_repl_idle()
pass that queues attaches. Kernel-Hermes drains at interpret
checkpoints, which don't fire during that pass at all -- the cutover
changed not just how delivery happens but when, opening a window
FORTH's own design never had.
Reported, not fixed here, per Captain Bob's Law: this is a defect in
closed task 3.8 code, found while scoping a different task. A real
fix changes SkHermesMessage's own shape to own its payload bytes
rather than reference a caller's pointer -- bigger than a repl.c
patch, touches every existing sender, needs its own three-ISA
acceptance.
Task 3.9's own send-side cutover (kernel_hermes.h, repl.c) is written
but deliberately left uncommitted -- it would inherit the identical
defect shape if shipped now. logs/20260922-111840/amd64/ is the
reproduction.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
The reply leg (Artemis -> Hera ack) that used to flow through
common:messaging.4th's MSG-SEND/MSG-TICK now goes through
kernel-Hermes's sk_hermes_send_one()/sk_hermes_drain_checkpoint()
instead -- FORTH Hermes never sees a BLK-ATTACH-EVENT message again
(SXXXIV.2's partition rule). The request leg was never real FORTH
messaging traffic to begin with (a direct VM-EXEC, no type tag,
forced by Hera's own inability to load common:messaging.4th), so it
is untouched.
New KH-BLK-ATTACH-SEND (repl.c) wraps sk_hermes_send_one(), reached
from capsules/artemis/init.4th's HERA-BLK-ATTACH-REQ. Delivery reuses
task 3.4's already-wired sk_hermes_drain_checkpoint(); BLK-ATTACH-ACK
itself is unchanged, just reached by a different layer.
SK_HERMES_MSG_TYPE_BLK_ATTACH deliberately reuses BLK-ATTACH-EVENT's
own value (9) to document this as a cutover of the same message, not
a new one.
Two real bugs found on the way, both recorded in FABRIC-3.6.md's
findings log:
- A popped FORTH CREATE-buffer address was raw-cast to a host pointer
instead of going through vm_ptr() -- silently read all-zero memory,
no crash, no error, just a message that arrived and did nothing.
Fixed; the rule and its exception (repl.c's own dev-addr is
legitimately a raw pointer, formatted that way by its own pushing
code) are written up for the next FORTH-facing C word.
- A separate, genuine hang on the very first live exercise of this
path, never reproduced across ten subsequent boots. Reported, not
chased -- not blocking, per the task's own check being otherwise
fully satisfied.
Also found live: log_message() is invisible in this build's actual
serial-log capture at every level -- settled on a single
console_println in the real drain target instead, one line per real
USB attach, not a hot-path.
Final acceptance (logs/20260922-105501, -105758, -110304, disk images
reset before each): dict_hash identical across all three
architectures for every VM, zero UNKNOWN WORD, mkcapsule --lint
clean, real ledger+stadium_conserved(Artemis)=true evidence on every
boot.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Added HERMES-CHANNEL-OPEN? ( req-hi req-lo -- allow? ) at
capsules/ACL.4th block 4008 (default: approve everything) -- the one
word policy authors edit. sk_hermes_channel_open_policy(VM*, VMUuid)
(kernel_hermes.h/.c) is the C-side query that calls it via plain
word-dispatch against the target VM's own dictionary/stack, never
vm_interpret() (avoids task 3.4's input-buffer cursor hazard entirely)
and never decides the answer itself. Fails closed: no policy word,
a policy error, or stack underflow all deny, matching CLAUDE.md's
posture that absence of policy must never mean "always allow."
Two real bugs found and fixed before this was called done:
missing current_executing_entry assignment before calling the word's
func pointer (colon words silently no-op without it, vm_core.c:730 --
no crash, just a wrong answer); and a second FAIL with debug
instrumentation still in place whose precise cause isn't
reconstructable, since no intermediate commit exists for that attempt.
Self-test proves the task's check four ways against the same
unchanged C function: default approve, live redefinition to deny
(zero C change), restore, and a VM with no ACL.4th loaded at all
(fail closed). A fifth check wires the result into task 3.6's
sk_hermes_channel_respond() end to end: a denied policy produces a
NACK and no channel, ledger/stadium_conserved() holding throughout.
Scope, per Captain Bob's ruling: closes with the query built and
proven; sk_hermes_channel_respond() still takes a caller-supplied
approved bool rather than calling the policy internally. Wiring a
real channel-open call site to only this query is deferred to
whichever later task first needs a live decision.
dict_hash identical across amd64/aarch64/riscv64 for every VM, zero
UNKNOWN WORD, mkcapsule --lint clean (38 files, 0 violations).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Extracted sk_hermes_send_one() from sk_hermes_publish()'s own
per-subscriber body -- one code path for both point-to-point and
fan-out delivery, so the ledger can never diverge between them.
Point-to-point addressing turned out to be load-bearing, not
incidental: sk_hermes_publish()'s fan-out sets msg->to to whichever
member it is iterating, so a negotiation message "published" to the
common channel would spuriously reach every member, not just the real
target (checked with advisor() before building the naive version).
"Over the common channel" means every VM is reachable from birth (task
3.2), not that the exchange itself fans out -- messaging.4th's own
CH-REQUEST carried an explicit `to` for the same reason.
sk_hermes_channel_request/respond/close build the mechanics: request ->
grant (creates a private channel, subscribes both parties, sends
CH_GRANT + one ACK) or NACK ("a deny is a NACK", SXLV.1 -- no separate
type); close authorized by membership alone. The grant/deny decision is
a plain caller-supplied `approved` bool -- task 3.7 replaces the call
site that produces it with a real ACL.4th query, not this signature.
Self-test covers the task's own three checks plus a sibling advisor()
flagged: an approved respond() whose channel creation itself fails
(table exhausted) must still fall through to NACK, not a silent false
grant or half-open channel -- verified by exhausting the whole channel
table and confirming the fallback.
Bug found and fixed before this was called done: the first draft
dropped a message via pending_pop() alone, without releasing it first,
leaking its Stadium heat and failing the self-test's own ledger
baseline check (logs/20260922-065946/amd64/, kept as audit trail).
Fixed and re-verified PASS on all three architectures.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
sk_hermes_publish() now enforces SK_HERMES_CHUNK_MAX_PAYLOAD (1024) on
every message's payload_len uniformly, chunked or not -- closing the
gap task 3.3 explicitly parked. A chunk carrier is
[SkHermesChunkHeader][content slice], slice capped at
1024 - sizeof(header) rather than 1024 itself, so every message on the
wire satisfies the same one-block bound vm_interpret()'s own drain
limit already requires -- a future chunk-aware drain never has to
special-case a carrier that can't be handed to vm_interpret() as-is.
Deliberately no chunking-sender API: building one would need
kernel-Hermes to own chunk-buffer memory with a real lifetime it has no
way to track (kept alive until every subscriber drains it). Sending is
a loop pattern a caller writes with sk_hermes_chunk_count() +
sk_hermes_publish(), demonstrated by this task's own self-test.
sk_hermes_reassemble() is pure and memory-agnostic: validates msg_id
agreement, exact seq coverage, and per-chunk slice sizes before a
single memcpy, with the total length computed once and checked against
the caller's buffer once -- never order-dependent on which chunk
happens to overflow.
Verified live on all three architectures: a 1024-byte payload as one
message, a 1025-byte send refused outright with the ledger untouched,
and a 3000-byte payload split into 3 chunks, drained, and reassembled
byte-exact against the original. dict_hash unmoved and identical
across architectures.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
sk_hermes_drain_checkpoint() interprets one queued payload per checkpoint
(ruled: one message per checkpoint), reusing sk_vm_at_outermost_interpret()
and placed before the switch-signal block in vm_core.c's existing
cooperative checkpoint (sk_vm_context_switch() doesn't return until
switched back to, so drain must come first or it silently never runs on
a switching checkpoint).
Amends FABRIC-3.5.md SXLIII.5, caught by advisor() before writing the
naive version: "recursive drain is prevented for free" via
g_vm_interpret_depth is true but only for same-message re-drain -- it
doesn't cover the separate same-VM reentrancy hazard FABRIC-3.md SXX
already named for Hera specifically (VMCallState saves rsp/exit_colon/
ecw_nesting only, never input_buffer/input_length/input_pos). Draining
calls vm_interpret() on the same vm whose own vm_interpret() call is
still paused mid-word at the checkpoint; without saving and restoring
the cursor by hand, the enclosing REPL line or LOAD block would be
silently truncated. sk_hermes_drain_checkpoint() snapshots and restores
input_buffer/input_length/input_pos/mode/error/abort_requested around
the call. Not a divergence from the ruling -- cursor preservation is the
implementer's own obligation inside the ruled mechanism.
Gated behind a system-wide pending-total counter so the common
no-message-in-flight case costs one integer read per word dispatch, not
a stadium_max_vm_count()-sized queue scan (also flagged by advisor() as
a real hot-path cost, not deferred).
Verified live on all three architectures: a self-test publishes a real
payload to Hermes, proves the depth gate via VM-EXEC-ing an existing
harmless colon word into Hermes (genuine nested vm_interpret(), depth 2,
must not drain), then drains directly from genuinely-outermost context
and confirms exactly one clean drain. dict_hash unmoved and identical
across architectures.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>