Q.* now forms its products in the order a1*b1 (low cell), a0*b1, a1*b0,
a0*b0, dropping each input after its last use, with b0 waiting on the
return stack. Each sign correction, (b1<0 ? a0 : 0) and (a1<0 ? b0 : 0),
is folded into cell 2 as soon as its operands are adjacent, using -if
rather than 0< calls. At most four live values sit under any UM* call.
Headroom (data cells under args / return entries): 2/2 -> 3/3.
Same checks as before at 32- and 64-bit cells, optimised and ASan+UBSan;
dropping either sign correction fails about 9900 checks.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
UM* parked its three corrections (c_hi, c_lo, t0) on the return stack
before the +* loop, leaving its caller 4 return entries. Now u1 and u2
wait on the data stack under the loop -- +* touches only T, S and A --
and the corrections are made after it, with -if in place of 0< calls.
The return stack holds the loop count, then at most two temporaries.
The loop count is pushed before s is made, so the data stack never holds
more than four cells: data headroom is unchanged at 6.
Headroom (data cells under args / return entries): 6/4 -> 6/6.
Q.*, which calls UM* four times, goes from 2/1 to 2/2.
Same checks as before at 32- and 64-bit cells, optimised and
ASan+UBSan; breaking any of the four corrections fails 4900+ checks.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Q.* is floor(a*b / 2^16), signed (D-8), cut to two cells. Cells 0..2 of
the product come from the unsigned cell products a0*b0, a0*b1, a1*b0 and
the low cell of a1*b1; reading a1 and b1 as signed takes
(a1<0 ? b0 : 0) + (b1<0 ? a0 : 0) off cell 2. The result is cells 0..2
shifted right 16 with Q.TO-INT twice. Clobbers A.
Checked at 32- and 64-bit cells, optimised and ASan+UBSan, on every pair
of 17 edge Q values plus 20000 pseudo-random pairs:
- against an independent reference (16-bit limbs, magnitudes multiplied
schoolbook, negated, shifted), and
- against v3's q48_mul for non-negative operands.
The first draft applied only a1's sign correction; the reference caught
the missing b1 term (9860 failures at 32-bit cells).
v3 bug found, not fixed: q48_mul's portable branch (used where there is
no __int128) returns (result_hi << 16) | (result_lo >> 16); the high
part must move up 48, so it is wrong whenever the product passes bit
64, e.g. 0.5 * Q max gives 0000ffffffffffff instead of 3fffffffffffffff.
The parity reference here is that branch with the shift corrected,
which matches v3's __int128 branch.
Stack use is tight: Q.* leaves its caller 2 data cells and 1 return
entry, because UM* parks three values on the return stack. Recorded in
5.26; to be improved next.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Executed on the golden model as written, and correct: <, =, D<, 2OVER,
and Q.=, Q.<, Q.0= (the D=, D<, D0= words).
Fixed, because they could not work under D-2 or were wrong:
- DMAX, DMIN: 2OVER 2OVER D< needs 8 data cells plus D<'s 2, the whole
10-deep data stack, so both failed on every case once the caller held
anything. Rewritten on a new call-free helper
(D<) ( d1 d2 -- d1 d2 flag )
which compares copies of the high cells, then the low cells unsigned
on a tie, with in-line sign tests; DMAX and DMIN then drop the loser.
- 2SWAP: ROT and SWAP written in line. Return headroom 2 -> 4, which
also lifts 2OVER 1 -> 3 and Q.> 1 -> 2.
- Q.>: section 5.26 gave SWAP D<, which swaps single cells; it was wrong
in 11702 of 20169 cases. Now 2SWAP D<.
Checked at 32- and 64-bit cells, optimised and ASan+UBSan: every
edge-vector quadruple (50625) for D<, (D<), 2SWAP, 2OVER, DMAX and DMIN;
Q comparisons on every pair of 13 edge Q values plus 20000 pseudo-random
pairs weighted to ties and one-bit differences, signed (D-8) and against
v3's unsigned comparisons where both values have the same sign.
Mutating any (D<) branch or subtraction fails 700+ checks.
The new fatal-UBSan setting caught a test-side array overflow in the
headroom probe (results buffer sized 4, six needed); fixed.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
DECOMPOSITION.md 5.26 described both words in prose only. They are now
written out, call-free, on one observation: +* with S = 0 never adds, so
each step is an exact arithmetic right shift of the double T:A.
: Q.FROM-INT ( n -- q ) push 0 a! 0 pop 15 FOR +* UNEXT push drop a pop ;
T:A = n:0 is n * 2^N; N-16 shifts leave n * 2^16 (count 15 at
32-bit cells, 47 at 64).
: Q.TO-INT ( q -- n ) push a! 0 pop 15 FOR +* UNEXT drop drop a ;
16 shifts at every width; the low cell is left in A. Rounds toward
minus infinity, as v3's arithmetic shift does.
Both clobber A; added to the section 2 list.
Checked at 32- and 64-bit cells, optimised and ASan+UBSan:
- Q.FROM-INT against n * 2^16 for edge values and 20000 pseudo-random n,
and against v3's q48_from_u64 for n >= 0 (low cell at 64-bit, D-10).
- Q.TO-INT against v3's (int64_t)q >> 16 on the edge Q values and 20000
pseudo-random Q values, and against a C double shift on 20000
arbitrary doubles.
- Round trip Q.TO-INT(Q.FROM-INT(n)) = n.
Mutating either shift count or the S = 0 setup fails 40000+ checks.
Note: make sanitize prints UBSan reports but does not fail on them. One
was found here, in the test's own random shift amount (fixed); a grep of
the full sanitize output now shows no runtime errors.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
D-10 (ruled 2026-10-02): a Q48.16 value is two cells at every cell
width, so every Q word is the same double word on 32- and 64-bit nodes.
Recorded in DECOMPOSITION.md section 3 and 5.26; the open note on the
Q.+/Q.- row is resolved by it.
Q.ABS and Q.NEG are the DABS and DNEGATE words. Checked against v3's
q48_abs and 0 - q on the 15 edge Q values and 20000 pseudo-random
values, at 32- and 64-bit cells, optimised and ASan+UBSan:
- 32-bit cells: bit-for-bit v3, including v3's wrap of Q min to itself.
- 64-bit cells: the low cell is v3's; the high cell is the true sign
(so ABS and NEG of Q min give +2^63 rather than wrapping), per D-10.
Breaking DNEGATE's carry fails Q.NEG, Q.ABS and Q.- checks.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
DECOMPOSITION.md 5.26 makes Q.+ and Q.- the D+ and D- words. They are
checked against v3's q48_add/q48_sub (uint64_t a + b, a - b, wrapping)
on every pair of 15 edge Q values (0, ulp, 0.5, 1.0, 1.5, -1.0, -ulp,
values across the 32-bit seam, Q max and min, +-12345.0) and 20000
pseudo-random pairs, 2682 of which overflow Q48.16.
- 32-bit cells: bit-for-bit v3's result, overflow wrap included.
- 64-bit cells, with a Q value as a sign-extended double: the low cell is
v3's result; on overflow the high cell holds the true carry where v3
wraps. Whether a Q value is one cell or two on a 64-bit node is not
ruled; recorded as open in 5.26.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
All four run exactly as written in DECOMPOSITION.md 5.6/5.7 and need no
change now that D+ and DNEGATE are call-free.
Checked against C at 32- and 64-bit cells, optimised and ASan+UBSan:
every edge-vector combination (M+ over all triples, D- and D= over all
50625 quadruples, D0= over all pairs) plus 20000 pseudo-random cases
weighted to equal low or high cells and low sums that wrap to 0.
Mutations of D0= and M+ are caught.
Headroom (data cells under args / return entries under return address):
M+ 6/4, D- 5/4, D0= 7/7, D= 5/3.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
D+ as written in 5.7 was exact but kept two cells on the return stack
while calling U> -> SWAP/U<, leaving its caller one return entry: any
word calling a word that calls D+ (M+, D-, D=, Q.+, Q.-) would have a
return address silently overwritten.
The new D+ makes no calls. Both high cells wait on the return stack; the
carry out of the low-cell add comes from sign tests -- if the low cells'
top bits differ, there is a carry exactly when the sum's top bit is
clear; if they match, exactly when both are set.
Headroom (data cells under args / return entries): 4/1 -> 6/5.
Checked against C over all 50625 edge-vector quadruples and 20000
pseudo-random pairs (weighted to top-bit cases and low sums wrapping to
0), at 32- and 64-bit cells, optimised and ASan+UBSan. Retargeting each
of the four carry branches fails more than 13000 checks.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
SM/REM as written in section 4 never returned correctly for a negative
dividend. It holds three entries on the return stack and then calls
DABS -> DNEGATE -> D+ -> U> -> SWAP/U<, which overflows the 9-deep
circular return stack (D-2). DABS itself could not run: DNEGATE as
written (inv SWAP inv SWAP 1 0 D+) left its caller no return entries.
- DNEGATE: inv over if L1 drop push inv 1 + pop ; L1: drop 1 + ;
i.e. ~d + 1, carrying into the high cell exactly when lo = 0.
- SM/REM: sign tests are native -if (as in 0<) and NEGATE is in line;
the only calls are DNEGATE and UM/MOD, both call-free inside.
Executed on the golden model at 32- and 64-bit cells, optimised and
under ASan+UBSan:
- SM/REM on dividends built as q*n + r with |r| < |n| and r signed as
d: every edge-vector q, n with r = 0 and r = +-(|n|-1), plus 20000
pseudo-random cases.
- /MOD, U>, ABS, S>D, D+ and DABS as written in section 5, against C
(D+ over all 50625 edge-vector quadruples).
Mutations of DNEGATE's carry and of each SM/REM sign branch are caught.
Headroom (data cells under args / return entries under return address):
SM/REM 5/3, /MOD 5/2, DNEGATE 7/7, DABS 7/6, ABS 8/7, U> 6/5.
D+ as written is exact but leaves only 1 return entry; recorded in
DECOMPOSITION.md 5.7 as not yet revised.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The first UM/MOD was exact but called 0<, U<, SWAP and OR inside its
loop, so it left its caller only 3 return-stack entries. SM/REM pushes
two signs before calling it, so under /MOD, M/MOD or */MOD the caller's
return address would be silently overwritten (D-2 circular stacks).
The loop now makes no calls. It branches on hi's top bit with -if, does
the unsigned hi' >= d test as U< does but with in-line sign tests,
subtracts with `inv a + inv`, and sets the quotient bit with `1 +` on an
even lo'. The final SWAP is in line.
Measured on the golden model at 32- and 64-bit cells:
headroom data 3 -> 6 cells under args, return 3 -> 6 entries
speed ~1355-1605 -> ~227-313 instruction words per call
Still exact for every uhi < ud (edge-vector triples and 20000 random
cases, optimised and ASan+UBSan). Retargeting each of the four in-loop
branches to the wrong label fails more than 12000 checks each.
DECOMPOSITION.md: section 4 UM/MOD replaced, with its derivation and
stack limits.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
UM/MOD is assembled exactly as DECOMPOSITION.md section 4 gives it and
needs no change: it is exact for every uhi < ud at 32- and 64-bit cells.
It is checked against q*d + r = uhi:ulo with r < d through v4_umul, so
no second C divider has to be trusted. Coverage: all edge-vector triples
with uhi < ud plus 20000 pseudo-random cases (top-bit, small and
near-maximum divisors), optimised and under ASan+UBSan. Two hand
mutations each fail more than 15000 checks.
New headroom probe: runs a word with marked cells under the canary and
under its return address and reports how many survive, since the D-2
circular stacks overwrite silently instead of faulting. Measured:
UM* 6 data cells under its args, 4 return entries under its return
UM/MOD 3 data cells under its args, 3 return entries under its return
DECOMPOSITION.md: UM/MOD marked executed, with its defined range and
stack limits.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
UM* as first written in DECOMPOSITION.md section 4 was exact only while
u1 <= 2^(n-2): plain +* loses the carry out of T and its shift keeps T's
sign bit, so the loop is exact only while S and T stay in
[-2^(n-2), 2^(n-2)).
The rewrite multiplies by s = u1 2/, which always lies in that range,
starting T at t0 = u2 2/ when u1 is odd, so the loop yields
hi:lo = t0 + s*u2 exactly. Then
u1*u2 = 2*(hi:lo) + c_lo + c_hi*2^n
c_lo = u1 & u2 & 1
c_hi = (u1<0 ? u2 : 0) + (u1 odd and u2<0 ? 1 : 0)
restores the halved-away bits and the unsigned reading of both top bits.
test_foundation.c runs the new definition against v4_umul over every
pair of the edge vectors plus 20000 pseudo-random pairs, at 32- and
64-bit cells, optimised and under ASan+UBSan. The two pinned failing
cases are now ordinary exactness checks. Two hand mutations of the
correction step each fail more than 10000 checks at both widths.
DECOMPOSITION.md: section 4 UM* replaced, D-3 ruling text updated.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
First code for StarForth v4 (JUSTIFICATION.md section 10, step 1): one node
of the 32-instruction core as a C99 model, with cell width as a build
parameter.
- Node: P, A, B, F18 circular stacks (10 and 9 deep, D-2), word-addressed
memory (D-1), 5% guard bands on every bounded list.
- Instruction word: six 5-bit slots in 32 bits at every cell width.
- Executor: all 32 opcodes of DECOMPOSITION.md 1.3. Cell arithmetic wraps
explicitly; no signed overflow or implementation-defined shift.
- Heat: per-opcode and per-call-target counters and the anti-clock, driven
by instruction retirement (1.4, D-6 interim).
- Slot packer and runner for tests, and a reference unsigned multiply in
plain C99 with no 128-bit type.
Tests run at 32- and 64-bit cells, and under ASan and UBSan. They cover
every opcode and execute the first section 4 definitions (NIP SWAP OR
NEGATE ROT 0< 0= 2DUP - U<) against the C operation each stands for.
UM* as written in section 4 is exact only while u1 <= 2^(n-2). Two known
failing cases are pinned in test_foundation.c until it is rewritten.
DECOMPOSITION.md: record D-9, the instruction word is 32 bits at every
cell width (ruled 2026-10-02).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
D-1 word addressing; D-2 F18 circular stacks (10/9 deep), hidden, so
DEPTH/PICK/ROLL/.S/SP@/SP! are retired everywhere and the DSP register is
dropped; D-3 plain F18 +* (UM* flagged for revision); D-5 host width
matches the host CPU; D-7 moot; D-8 signed Q48.16; D-4 and D-6 deferred
to the hosted-mesh step. Adds a note that every CAP definition must be
re-checked against the 10/9 stack depths.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BY9HMwK5Cetz3caBgHGyds
JUSTIFICATION.md records why v4 exists and the reasoning behind each
major design decision. DECOMPOSITION.md assigns every v3 C primitive a
fate on the 32-instruction F18-derived core.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BY9HMwK5Cetz3caBgHGyds