Node memory is a build parameter. Tests named test_host_*.c are now
built with V4_NODE_WORDS=16384, the host node that will hold the
compiler, the dictionary and the text being compiled; every other test
keeps the 1024-word mesh node.
The first such test checks the node at that size and the text
assembler's branch placement, which only matters there: a branch slot
that cannot reach the whole node is not used.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Test support, beside the slot packer: reads definitions in the notation
DECOMPOSITION.md uses -- words, opcodes, literals, constants, labels,
branches, FOR NEXT and FOR UNEXT, in-line macros, comments -- and lays
them down through the slot packer, so they no longer have to be retyped
opcode by opcode in C.
Tested by assembling SWAP, UM/MOD, C@ and C! both ways on two nodes and
comparing memory word for word, by running what is assembled, and on
every error it reports, each with its line number.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
?TERMINAL reads CONSOLE-STATUS; KEY polls it until a character is
pending and then takes it from CONSOLE-RX. Executed on the golden
model at both cell widths, including a transcript of the v3 binary, and
KEY shown still waiting after 5000 instruction words with no input.
v3's KEY returned -1 at the end of its input; here KEY waits.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The console's receive side, standing in for the console node until the
mesh exists, as CONSOLE-TX does for output. A data fetch (@, @+, @b)
from CONSOLE-STATUS gives -1 when a character is pending and 0 when
not; from CONSOLE-RX it gives the next character and takes it, or -1
with none pending. The characters come from a queue the test feeds.
Instruction words and literals are still fetched with v4_node_load, so
code at a register's address is never taken for the register.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
(DO) (?DO) (LOOP) (+LOOP) (LEAVE) I J UNLOOP and (0BRANCH), laid down
in line by hand as the compiler will, executed on the golden model at
both cell widths against C on every pair of 14 loop ends and 11 steps,
and against sequences recorded from the v3 binary.
(LOOP) goes round again while index < limit, signed, as v3 does; the
document's equality test differed whenever start >= limit. (+LOOP) is
now specified. J keeps the outer index in A and needs no extra return
entry.
LEAVE discards the limit and index. v3 left them on its return stack,
which made a LEAVE in an inner loop stop the outer one.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
M- M* M/MOD MOD */ */MOD, D0< D2* D2/ 2ROT, 2DROP and 2>R 2R@ 2R>,
beside UM* and SM/REM in test_foundation.c. Executed on the golden
model at both cell widths against C and results recorded from the v3
binary.
M- widens n before negating it, so the most negative n is right.
*/MOD goes through a full double product. D2/ is one +* step.
M- and M/MOD take the double in the standard order ( lo hi ), as M+
does; v3 took its low cell on top.
The foundation test's node is now full: 958 of the 960 words below its
variables.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
NIP SWAP ROT -ROT ?DUP 2DUP, >R R@ R>, @ ! +! -! 2@ 2! CELLS,
- NEGATE 1+ 1- 2+ 2- MIN MAX, OR NOT 0= 0< 0<> 0> = <> < > <= >= U<
WITHIN TRUE FALSE in a new test, and * and / beside UM* and /MOD in
test_foundation.c. Executed on the golden model at both cell widths
against C on every combination of the edge values and against recorded
transcripts of the v3 binary.
WITHIN is low <= n < high with signed comparisons, which is what v3
computes; the document's circular form differed for low > high.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
FILL ERASE MOVE, COUNT CMOVE CMOVE> BLANK -TRAILING COMPARE SEARCH SCAN
SKIP, as loops over the call-free C@ and C!, executed on the golden
model at both cell widths against C and five recorded transcripts of
the v3 binary.
A negative count reads as 0 where v3 read it so, and elsewhere writes
nothing and sets NODE-ERROR. v3's counted-string auto-detection is not
kept in any of them.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
One case per byte position, like C!: the cell is shifted down with a
2/ loop and masked. It replaces the version built on LSHIFT, RSHIFT
and SWAP calls. C@ now leaves its caller 7 return entries (was 4), and
the words above it gain with it: TYPE 6 (was 3), DUMP 4, Q.PRINT, U.
and U.R 3 (were 2). The signed number words stay at 2: their sign
waits on the return stack.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
BASE and HLD leave their variable's word address; DECIMAL, HEX and
OCTAL store 10, 16 and 8. Executed on the golden model at both cell
widths, including a recorded transcript of the v3 binary.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Built on the number-output words and executed on the golden model at
both cell widths against a C reference and recorded transcripts of the
v3 binary (DUMP byte for byte at 64-bit cells).
Q.PRINT is signed (D-8) and always decimal; DUMP is always hex; both
put BASE back. Each keeps its working state in a variable, (QP) and
(DP), so each leaves its caller 4 data cells and 2 return entries.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The number-printing words, on pictured output and TYPE, executed on the
golden model at both cell widths against a C reference (ten bases,
eleven field widths) and six recorded transcripts of the v3 binary.
Each plain word is its .R word with a width of 0; all six share one
tail. The field width waits in a variable, (W), so the picture runs no
deeper than it does on its own: each word leaves its caller 4 data
cells and 2 return entries.
As v3, the .R words print a trailing space. Unlike v3, D. takes
( lo hi ), prints the whole double, and printing honours BASE.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
: EMIT ( c -- ) CONSOLE-TX b! !b ;
: CR ( -- ) 10 jump EMIT
: SPACE ( -- ) 32 jump EMIT
: TYPE ( baddr u -- )
-if OK drop drop NODE-ERROR b! -1 !b ;
OK: if DONE over C@ EMIT push 1 + pop -1 + jump OK
DONE: drop drop ;
They follow v3 (v3/src/word_source/io_words.c): EMIT prints the low byte
of the cell, CR character 10, SPACE a blank; TYPE prints u bytes,
nothing for u = 0, and for u < 0 nothing, with NODE-ERROR set where v3
raised its error flag. TYPE does not check the address range, which v3
does; out-of-range addressing is still an open question in node.h.
New test v4/tests/test_terminal.c, on a node of its own with the console
attached, 939 checks per width. Two transcripts of the real v3 binary
are recorded as expected output:
65 EMIT 66 EMIT SPACE 67 EMIT CR 68 EMIT 321 EMIT -> "AB C\nDA"
S" Hello, v3" TYPE 91 EMIT <text> 0 TYPE 93 EMIT -> "Hello, v3[]"
Also: EMIT of every byte and of wider values, CR, SPACE, TYPE from every
start within a cell at lengths 0..40, bytes with the top bit set and
zero bytes, negative counts, and that TYPE writes no memory. At 32- and
64-bit cells, optimised and ASan+UBSan (`make test`, `make sanitize`).
Five mutations each fail.
Headroom (data cells under args / return entries): EMIT 8/8, TYPE 4/3.
TYPE's depth is C@'s, which is still as written (C@ -> RSHIFT -> SWAP).
Test results on the amd64 host only. This is a development check, not
acceptance (JUSTIFICATION.md section 16).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
EMIT is a device service: a character sent to the console node. The
mesh and its ports are development step 2 and the memory map is open
(D-4), so the single-node model now stands in for the console with one
memory-mapped register, so that printing words can be run and their
output compared with v3's.
v4_node_console_attach(n, addr): after it, a store to word address
`addr` appends the low 8 bits of the value to a buffer on the node
(V4_CONSOLE_CAP characters, 4096 by default) and does not write memory;
a load from `addr` reads the memory word as before. Characters past the
capacity are counted in console_dropped and discarded. The hook is in
v4_node_store, which all four store opcodes use.
It is off by default: v4_node_reset detaches the console (address -1)
and empties the capture, so the ISA's behaviour and every existing test
are unchanged unless a console is attached. The address is the
caller's choice.
New test v4/tests/test_console.c, 22 checks per width: direct stores,
!b, ! and !+ through the executor printing "Hi!", memory at and around
the register untouched, the low byte only, overflow, detach, a console
at word 0, and reset. At 32- and 64-bit cells, optimised and ASan+UBSan
(`make test`, `make sanitize`); all other v4 tests still pass. Four
mutations of the hook each fail.
DECOMPOSITION.md section 7 gains the CONSOLE-TX row. No printing word is
defined here; EMIT and the words on it are still to do.
Test results on the amd64 host only. This is a development check, not
acceptance (JUSTIFICATION.md section 16).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
`#` held the high quotient on the return stack across its second
UM/MOD, which was the deepest point of pictured output once C! was made
call-free. It now rotates it under the division on the data stack
(-ROT in line: SWAP push SWAP pop).
Headroom, return entries: <# #S #> 2 -> 3, the signed picture 1 -> 2.
Data room is unchanged.
Same 4716 checks per width at 32- and 64-bit cells, optimised and
ASan+UBSan (`make test`, `make sanitize`); the headroom checks now
require the new figures.
Test results on the amd64 host only. This is a development check, not
acceptance (JUSTIFICATION.md section 16).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
C! as written in 5.3 called RSHIFT and LSHIFT (which call SWAP) and OR.
It is now one straight-line case per byte position: the byte is shifted
up with a `2* unext` loop, that byte of the cell cleared with a constant
mask, and the two added. The word address is `2/ 2/` of the byte
address. The upper half of a 64-bit cell is still preserved.
Headroom (data cells under args / return entries): 5/4 -> 7/7.
Same checks as before (every byte of two adjacent words, eight values,
neighbours and the upper half untouched) at 32- and 64-bit cells,
optimised and ASan+UBSan (`make test`, `make sanitize`). Four mutations
fail, one of them only at 64-bit cells, where it differs.
test_pictured.c carries the same C!. The pictured words did not gain
return-stack room from this: <# #S #> still leaves 2 entries and the
signed picture 1. With C! shallow, the deepest point is now inside `#`,
which holds the high quotient on the return stack across its second
UM/MOD. Recorded in 5.8; `#` is unchanged.
Test results on the amd64 host only. This is a development check, not
acceptance (JUSTIFICATION.md section 16).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
DECOMPOSITION.md 5.8 only named these as "standard pictured-output
definitions over UM/MOD and a hold buffer". They are now written out
(5.8) and run on the golden model.
D-13 (ruled 2026-10-03): the hold buffer takes 63 characters, as v3's;
HOLD of a value outside 0-255, or into a full buffer, stores nothing and
sets NODE-ERROR, which is what v3 did with its error flag.
Behaviour follows v3 otherwise: digits 0-9 then A-Z, BASE outside 2..36
reads as 10, SIGN ( n -- ). Stack effects are the standard ones, as 5.8
already ruled (v3 took its double low cell on top). The buffer is 64
bytes, filled backwards from its end through HLD; `#` divides by the
base in two UM/MOD steps. HOLD, SIGN and # end in a jump to the next
word instead of a call, which saves two return-stack entries: with
calls, the signed picture `.` needs overflowed the 9-deep return stack.
New test file v4/tests/test_pictured.c, on a node of its own:
test_foundation.c's hand-assembled words already fill 901 of a node's
1024 words. It re-assembles SWAP, OR, UM/MOD, LSHIFT, RSHIFT and C! as
they are in test_foundation.c.
Checked against a C reference (short division on 16-bit limbs) at 32-
and 64-bit cells, optimised and ASan+UBSan (`make test`, `make
sanitize`), 4716 checks per width: <# #S #>, <# # # 46 HOLD #S #> and
the signed picture, in bases 10, 16, 2, 8, 36, 3 and the invalid 0, 1,
37, -5, on every pair of 15 edge cells; HOLD of 12 values; and a base-2
number that overflows the buffer (63 characters kept, NODE-ERROR set,
nothing outside the buffer written). Six mutations all fail.
Headroom: <# #S #> leaves 3 data cells and 2 return entries; the signed
picture leaves 1 return entry. The depth is C!'s as written (C! ->
LSHIFT -> SWAP), which is unchanged.
Test results on the amd64 host only. This is a development check, not
acceptance (JUSTIFICATION.md section 16).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Both run exactly as written in DECOMPOSITION.md 5.3 and need no change:
four bytes to a cell, little-endian, byte address = 4 * word address +
byte index, at either cell width.
Checked at 32- and 64-bit cells, optimised and ASan+UBSan (`make test`,
`make sanitize`): every byte of two adjacent words, eight values
including ones wider than a byte, with the other bytes of the word, the
upper half of a 64-bit cell and the neighbouring words left untouched.
Breaking C!'s mask or C@'s mask fails.
Headroom (data under args / return): C@ 6/4, C! 5/4.
Test results on the amd64 host only. This is a development check, not
acceptance (JUSTIFICATION.md section 16).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Both run exactly as written in DECOMPOSITION.md 5.5 and need no change.
Checked against C on the edge vectors for every count 0 .. N (a count
of N gives 0), at 32- and 64-bit cells, optimised and ASan+UBSan
(`make test`, `make sanitize`). Removing RSHIFT's sign-bit mask fails.
Headroom (data under args / return): 7/5 each.
They are dependencies of C@ and C!, which the pictured-output hold
buffer needs.
Test results on the amd64 host only. This is a development check, not
acceptance (JUSTIFICATION.md section 16).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
DECOMPOSITION.md 5.26 gives them fate IN: the compiler places two
literals, low cell first. Q.1 and Q.SCALE are `65536 0`, Q.0 is `0 0`;
v3's are 65536, 0 and 65536.
Each expansion is assembled and run, and checked against v3's value;
then in use: Q.1 Q.TO-INT is 1, 1 Q.FROM-INT is Q.1, and q Q.1 Q.* and
q Q.0 Q.+ return q for 2000 pseudo-random q plus 0, Q max and Q min.
At 32- and 64-bit cells, optimised and ASan+UBSan (`make test`,
`make sanitize`). A wrong value in any of the three fails.
Test results on the amd64 host only. This is a development check, not
acceptance (JUSTIFICATION.md section 16).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Q.COS is v3's q48_cos_approx: on |(Q.REDUCE) x|, the even Taylor series
to n = 10. It sets term and sum to 1.0, n to 2 and the sign cell to 0,
then jumps into Q.SIN's loop, so it adds no call level.
Checked bit for bit against v3's q48_cos_approx (ported into the test)
on the same 33 edge angles and 3000 pseudo-random angles as Q.SIN, at
32- and 64-bit cells, optimised and ASan+UBSan (`make test`,
`make sanitize`). Mutating the start index, the start sum, the
alternation or the sign cell each fails 2500+ cosine checks and no sine
checks.
Headroom (data under arg / return): Q.COS 3/2.
Test results on the amd64 host only. This is a development check, not
acceptance (JUSTIFICATION.md section 16).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
(Q.REDUCE) is v3's q48_reduce_angle: the angle as one cell in [-pi, pi].
x / 2pi does not fit a cell, but only the remainder is needed, and
|x| mod 2pi is two UM/MOD steps (high cell first, its remainder leading
the low cell); then the sign of x and one step of 2pi back into range.
Q.SIN is v3's q48_sin_approx on the reduced angle: the odd Taylor series
to n = 11, each term the last times x^2 over n(n-1), stopping below 10
ulp. After the reduction every value fits one cell, so the 6-cell
variable (QT) holds single cells. The loop is laid out for Q.COS to jump
into, so the pair share it without an extra call level.
Checked bit for bit against v3 (both functions ported into the test) on
33 edge angles -- around +-pi/2, +-pi, +-2pi, the 32-bit seam and both
64-bit extremes -- and 3000 pseudo-random angles of every size and sign,
at 32- and 64-bit cells, optimised and ASan+UBSan (`make test`,
`make sanitize`). Mutating either range boundary, the series length or
the sign alternation fails. Moving the 10-ulp stop to 11 fails nothing,
and a scan of all 205888 reduced angles in C shows why: no input's
result depends on it, for sine or cosine.
Headroom (data under arg / return): (Q.REDUCE) 5/5, Q.SIN 3/2.
Test results on the amd64 host only. This is a development check, not
acceptance (JUSTIFICATION.md section 16).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Q.LOG is v3's q48_log_approx: x = 2^k * m with 1.0 <= m < 2.0, ln m by
up to 6 Newton rounds on e^y = m, result y + k * ln 2. As ruled, e^y is
Q.EXP's Taylor loop written in line rather than a call to Q.EXP, so
Q.LOG calls only Q.*, Q./ and UM*. After the reduction m, y, the Taylor
term and its sum each fit one cell, so the 7-cell variable (QL) holds
single cells. x <= 0 returns 0 and sets NODE-ERROR (D-12).
Checked bit for bit against v3's q48_log_approx (ported into the test)
on 22 positive edge values, 2000 pseudo-random x of every magnitude and
a sweep of every 17th reduced m, at 32- and 64-bit cells, optimised and
ASan+UBSan (`make test`, `make sanitize`). Mutating the round count, the
100-ulp stop, ln 2 or the Taylor length all fail.
v3's clamp y = max(y - corr, 0) is kept but cannot fire: a scan of all
65536 values of m in C never reaches it, and never needs more than 4
rounds. So no test covers it.
The test harness's per-call step limit is raised from 100000 to 4000000
instruction words: a 64-bit Q./ takes about 6500, and a mutant forcing
extra Newton rounds ran past the old limit and looked like a width
difference.
Headroom (data under arg / return): Q.LOG 3/2.
Test results on the amd64 host only. This is a development check, not
acceptance (JUSTIFICATION.md section 16).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Both ruled by Captain Bob, 2026-10-03.
Products (README.md, JUSTIFICATION.md section 15): Hosted StarForth F18
(native Linux, amd64/arm64/riscv64, on hardware); FPGA StarForth F18 (a
32-bit build loaded into the FPGA, the gateway and foundation); and
StarshipOS (bare metal, LithosAnanke on the v4 F18 engine). FORTH-79
recomposed on the F18 engine is stored as a capsule, as are the
StarshipOS portions; the tree will be reorganised around this.
Acceptance (JUSTIFICATION.md section 16, v4/README.md): v4 is equivalent
to v3 at any point in time -- same vocabularies, same behaviour, on the
F18-derived engine -- and every ISA, hosted and bare metal, must still
reach its ok prompt. `make -C v4 test` is a development check, not
acceptance. v4 had no acceptance criteria before this.
Documentation only; no code changed.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Baseline run for StarForth-v4.0.0 as it stands, before any further work.
`make -f kernel/Makefile ARCH=<arch> clean qemu` for amd64, aarch64 and
riscv64, one at a time, disk/artemis.img and
disk/thumbdrives/zuse-thumb-ident.img restored to their committed state
before each.
All three reach `[zuse@Hera] ok>`, with zero UNKNOWN WORD, PARITY:OK,
1050 POST tests run and ALL IMPLEMENTED TESTS PASSED,
stadium_conserved(Artemis)=true.
dict_hash is identical across the three:
Hera (PARITY:M7.1a, word_count=530) 0x08873e0f44b7cb2a
dict_hash 0xe11082140cf86b05
dict_hash 0xb1256603f848e2b9
dict_hash 0xc791409ac1715690
QEMU was asked to quit over QMP once the prompt had appeared and the
serial log had been quiet for 10 seconds.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
D-12 (ruled 2026-10-02): Q.SQRT of a negative value and Q.LOG of zero or
a negative value return 0 and set NODE-ERROR, as Q./ does on division by
zero.
Q.SQRT is v3's q48_sqrt_approx: Newton from x0 = q/2 + 0.25, up to 8
rounds of x' = (x + q/x) / 2, returning x once |x' - x| < 10 ulp;
sqrt(0) = 0 and sqrt(1.0) = 1.0. q, x and the rounds left live in a
5-cell variable (QR). Checked bit for bit against v3's q48_sqrt_approx
(ported into the test) on 19 edge values and 3000 pseudo-random q below
2^48, at both cell widths, optimised and ASan+UBSan. Above 2^48 v3's
q48_div saturates and v4 divides correctly, so there is no parity there.
Bug found and fixed on the way: the "< 10 ulp" test subtracted 10 from
the step's low cell as a signed number, so at 32-bit cells a step of
2^31 or more read as small and the iteration stopped early. It now
tests the low cell's top bit first, as Q.EXP does. The first test run
missed it because `make build/32-test_foundation.c` rebuilds nothing
(the Makefile's rules use absolute paths); the mutation runs, compiled
from source, exposed it. Only `make test` is used from here.
Headroom (data under arg / return): Q.SQRT 3/2.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Q.EXP is v3's q48_exp_approx: the Taylor series to 10 terms on |q|,
stopping below 50 ulp, 1.0 Q./ e^|q| for q < 0, e^0 = 1.0, and for
|q| >= 16.0 either 0 or Q max (v3 returned all ones, which is -ulp read
signed under D-8). x, term, sum, sign and term index live in a new
8-cell variable (QE). Checked bit for bit against v3's q48_exp_approx
(ported into the test) on 23 edge values and 3000 pseudo-random q in
(-16.0, 16.0), at both cell widths, optimised and ASan+UBSan; dropping a
term, moving the 50-ulp stop or skipping the reciprocal all fail.
Q.EXP calls Q./, which calls (UQ/), which calls D2*C, and at first it
left its caller no return entries. Lightened along the chain:
- D2*C takes cin in A and builds 2hi+m with `over + +`: no ROT, no SWAP.
- (UQ/) parks the quotient bit instead of using ROT, stores q without
SWAP, and does its full trial subtraction (with borrow) in line
instead of calling DNEGATE D+.
- Q.EXP keeps its term index in (QE) instead of a FOR count.
Headroom (data under args / return): (UQ/) 4/3 -> 4/4, Q./ 3/2 -> 3/3,
Q.EXP 3/2.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
D-11 (ruled 2026-10-02): Q./ rounds toward zero, saturates to Q max /
Q min on overflow, and on division by zero returns Q max / Q min by the
dividend's sign (0 for 0/0) and sets the new NODE-ERROR register (7),
which VM-ERROR? reads.
- (UQ/): unsigned floor(a * 2^16 / b) for b != 0 and no overflow.
Restoring division, 2N+16 steps, on one shifting register (quotient
above remainder). Quotient register and divisor live in a 5-cell
variable (Q/), like BASE and the hold buffer, so the stacks carry only
the remainder, the quotient bit and the loop count. A first version
kept the divisor on the return stack and overflowed it when called
from Q./ (it never returned).
- D2*C: double shift left with carry in and out.
- Q./: zero-divisor handling, signs (kept in (Q/)), an overflow test
that needs no division (|a| * 2^16 >= |b| * 2^(2N-1), possible only
for |b| < 2^17), then (UQ/) and the sign.
Checked at 32- and 64-bit cells, optimised and ASan+UBSan, against an
independent limb-by-limb long division (including NODE-ERROR): every
pair of 18 edge Q values, 13 width-native edge values (true Q max/min at
either width), the overflow boundary, divisors in [2^(2N-2), 2^(2N-1)),
and pseudo-random cases; and against v3's q48_div for non-negative
a < 2^48. Q.* now also runs on the width-native edge values.
Mutations of the compare paths, the take path, the error flag and the
overflow mask are all caught (the mask matters only for Q min / -1.0).
Headroom (data under args / return): Q./ 3/2.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Q.* now forms its products in the order a1*b1 (low cell), a0*b1, a1*b0,
a0*b0, dropping each input after its last use, with b0 waiting on the
return stack. Each sign correction, (b1<0 ? a0 : 0) and (a1<0 ? b0 : 0),
is folded into cell 2 as soon as its operands are adjacent, using -if
rather than 0< calls. At most four live values sit under any UM* call.
Headroom (data cells under args / return entries): 2/2 -> 3/3.
Same checks as before at 32- and 64-bit cells, optimised and ASan+UBSan;
dropping either sign correction fails about 9900 checks.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
UM* parked its three corrections (c_hi, c_lo, t0) on the return stack
before the +* loop, leaving its caller 4 return entries. Now u1 and u2
wait on the data stack under the loop -- +* touches only T, S and A --
and the corrections are made after it, with -if in place of 0< calls.
The return stack holds the loop count, then at most two temporaries.
The loop count is pushed before s is made, so the data stack never holds
more than four cells: data headroom is unchanged at 6.
Headroom (data cells under args / return entries): 6/4 -> 6/6.
Q.*, which calls UM* four times, goes from 2/1 to 2/2.
Same checks as before at 32- and 64-bit cells, optimised and
ASan+UBSan; breaking any of the four corrections fails 4900+ checks.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Q.* is floor(a*b / 2^16), signed (D-8), cut to two cells. Cells 0..2 of
the product come from the unsigned cell products a0*b0, a0*b1, a1*b0 and
the low cell of a1*b1; reading a1 and b1 as signed takes
(a1<0 ? b0 : 0) + (b1<0 ? a0 : 0) off cell 2. The result is cells 0..2
shifted right 16 with Q.TO-INT twice. Clobbers A.
Checked at 32- and 64-bit cells, optimised and ASan+UBSan, on every pair
of 17 edge Q values plus 20000 pseudo-random pairs:
- against an independent reference (16-bit limbs, magnitudes multiplied
schoolbook, negated, shifted), and
- against v3's q48_mul for non-negative operands.
The first draft applied only a1's sign correction; the reference caught
the missing b1 term (9860 failures at 32-bit cells).
v3 bug found, not fixed: q48_mul's portable branch (used where there is
no __int128) returns (result_hi << 16) | (result_lo >> 16); the high
part must move up 48, so it is wrong whenever the product passes bit
64, e.g. 0.5 * Q max gives 0000ffffffffffff instead of 3fffffffffffffff.
The parity reference here is that branch with the shift corrected,
which matches v3's __int128 branch.
Stack use is tight: Q.* leaves its caller 2 data cells and 1 return
entry, because UM* parks three values on the return stack. Recorded in
5.26; to be improved next.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Executed on the golden model as written, and correct: <, =, D<, 2OVER,
and Q.=, Q.<, Q.0= (the D=, D<, D0= words).
Fixed, because they could not work under D-2 or were wrong:
- DMAX, DMIN: 2OVER 2OVER D< needs 8 data cells plus D<'s 2, the whole
10-deep data stack, so both failed on every case once the caller held
anything. Rewritten on a new call-free helper
(D<) ( d1 d2 -- d1 d2 flag )
which compares copies of the high cells, then the low cells unsigned
on a tie, with in-line sign tests; DMAX and DMIN then drop the loser.
- 2SWAP: ROT and SWAP written in line. Return headroom 2 -> 4, which
also lifts 2OVER 1 -> 3 and Q.> 1 -> 2.
- Q.>: section 5.26 gave SWAP D<, which swaps single cells; it was wrong
in 11702 of 20169 cases. Now 2SWAP D<.
Checked at 32- and 64-bit cells, optimised and ASan+UBSan: every
edge-vector quadruple (50625) for D<, (D<), 2SWAP, 2OVER, DMAX and DMIN;
Q comparisons on every pair of 13 edge Q values plus 20000 pseudo-random
pairs weighted to ties and one-bit differences, signed (D-8) and against
v3's unsigned comparisons where both values have the same sign.
Mutating any (D<) branch or subtraction fails 700+ checks.
The new fatal-UBSan setting caught a test-side array overflow in the
headroom probe (results buffer sized 4, six needed); fixed.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Without -fno-sanitize-recover=all, UBSan prints a report and the run
carries on to "all v4 tests passed". The sanitize build now aborts on
the first report. Verified by rebuilding the test-side shift bug found
in 981f4180 with these flags: the run exits 1 at the report.
Both test binaries now also depend on v4/Makefile, so a flag change
rebuilds them instead of leaving stale binaries in place.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
DECOMPOSITION.md 5.26 described both words in prose only. They are now
written out, call-free, on one observation: +* with S = 0 never adds, so
each step is an exact arithmetic right shift of the double T:A.
: Q.FROM-INT ( n -- q ) push 0 a! 0 pop 15 FOR +* UNEXT push drop a pop ;
T:A = n:0 is n * 2^N; N-16 shifts leave n * 2^16 (count 15 at
32-bit cells, 47 at 64).
: Q.TO-INT ( q -- n ) push a! 0 pop 15 FOR +* UNEXT drop drop a ;
16 shifts at every width; the low cell is left in A. Rounds toward
minus infinity, as v3's arithmetic shift does.
Both clobber A; added to the section 2 list.
Checked at 32- and 64-bit cells, optimised and ASan+UBSan:
- Q.FROM-INT against n * 2^16 for edge values and 20000 pseudo-random n,
and against v3's q48_from_u64 for n >= 0 (low cell at 64-bit, D-10).
- Q.TO-INT against v3's (int64_t)q >> 16 on the edge Q values and 20000
pseudo-random Q values, and against a C double shift on 20000
arbitrary doubles.
- Round trip Q.TO-INT(Q.FROM-INT(n)) = n.
Mutating either shift count or the S = 0 setup fails 40000+ checks.
Note: make sanitize prints UBSan reports but does not fail on them. One
was found here, in the test's own random shift amount (fixed); a grep of
the full sanitize output now shows no runtime errors.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
D-10 (ruled 2026-10-02): a Q48.16 value is two cells at every cell
width, so every Q word is the same double word on 32- and 64-bit nodes.
Recorded in DECOMPOSITION.md section 3 and 5.26; the open note on the
Q.+/Q.- row is resolved by it.
Q.ABS and Q.NEG are the DABS and DNEGATE words. Checked against v3's
q48_abs and 0 - q on the 15 edge Q values and 20000 pseudo-random
values, at 32- and 64-bit cells, optimised and ASan+UBSan:
- 32-bit cells: bit-for-bit v3, including v3's wrap of Q min to itself.
- 64-bit cells: the low cell is v3's; the high cell is the true sign
(so ABS and NEG of Q min give +2^63 rather than wrapping), per D-10.
Breaking DNEGATE's carry fails Q.NEG, Q.ABS and Q.- checks.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
DECOMPOSITION.md 5.26 makes Q.+ and Q.- the D+ and D- words. They are
checked against v3's q48_add/q48_sub (uint64_t a + b, a - b, wrapping)
on every pair of 15 edge Q values (0, ulp, 0.5, 1.0, 1.5, -1.0, -ulp,
values across the 32-bit seam, Q max and min, +-12345.0) and 20000
pseudo-random pairs, 2682 of which overflow Q48.16.
- 32-bit cells: bit-for-bit v3's result, overflow wrap included.
- 64-bit cells, with a Q value as a sign-extended double: the low cell is
v3's result; on overflow the high cell holds the true carry where v3
wraps. Whether a Q value is one cell or two on a 64-bit node is not
ruled; recorded as open in 5.26.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
All four run exactly as written in DECOMPOSITION.md 5.6/5.7 and need no
change now that D+ and DNEGATE are call-free.
Checked against C at 32- and 64-bit cells, optimised and ASan+UBSan:
every edge-vector combination (M+ over all triples, D- and D= over all
50625 quadruples, D0= over all pairs) plus 20000 pseudo-random cases
weighted to equal low or high cells and low sums that wrap to 0.
Mutations of D0= and M+ are caught.
Headroom (data cells under args / return entries under return address):
M+ 6/4, D- 5/4, D0= 7/7, D= 5/3.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
D+ as written in 5.7 was exact but kept two cells on the return stack
while calling U> -> SWAP/U<, leaving its caller one return entry: any
word calling a word that calls D+ (M+, D-, D=, Q.+, Q.-) would have a
return address silently overwritten.
The new D+ makes no calls. Both high cells wait on the return stack; the
carry out of the low-cell add comes from sign tests -- if the low cells'
top bits differ, there is a carry exactly when the sum's top bit is
clear; if they match, exactly when both are set.
Headroom (data cells under args / return entries): 4/1 -> 6/5.
Checked against C over all 50625 edge-vector quadruples and 20000
pseudo-random pairs (weighted to top-bit cases and low sums wrapping to
0), at 32- and 64-bit cells, optimised and ASan+UBSan. Retargeting each
of the four carry branches fails more than 13000 checks.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
SM/REM as written in section 4 never returned correctly for a negative
dividend. It holds three entries on the return stack and then calls
DABS -> DNEGATE -> D+ -> U> -> SWAP/U<, which overflows the 9-deep
circular return stack (D-2). DABS itself could not run: DNEGATE as
written (inv SWAP inv SWAP 1 0 D+) left its caller no return entries.
- DNEGATE: inv over if L1 drop push inv 1 + pop ; L1: drop 1 + ;
i.e. ~d + 1, carrying into the high cell exactly when lo = 0.
- SM/REM: sign tests are native -if (as in 0<) and NEGATE is in line;
the only calls are DNEGATE and UM/MOD, both call-free inside.
Executed on the golden model at 32- and 64-bit cells, optimised and
under ASan+UBSan:
- SM/REM on dividends built as q*n + r with |r| < |n| and r signed as
d: every edge-vector q, n with r = 0 and r = +-(|n|-1), plus 20000
pseudo-random cases.
- /MOD, U>, ABS, S>D, D+ and DABS as written in section 5, against C
(D+ over all 50625 edge-vector quadruples).
Mutations of DNEGATE's carry and of each SM/REM sign branch are caught.
Headroom (data cells under args / return entries under return address):
SM/REM 5/3, /MOD 5/2, DNEGATE 7/7, DABS 7/6, ABS 8/7, U> 6/5.
D+ as written is exact but leaves only 1 return entry; recorded in
DECOMPOSITION.md 5.7 as not yet revised.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The first UM/MOD was exact but called 0<, U<, SWAP and OR inside its
loop, so it left its caller only 3 return-stack entries. SM/REM pushes
two signs before calling it, so under /MOD, M/MOD or */MOD the caller's
return address would be silently overwritten (D-2 circular stacks).
The loop now makes no calls. It branches on hi's top bit with -if, does
the unsigned hi' >= d test as U< does but with in-line sign tests,
subtracts with `inv a + inv`, and sets the quotient bit with `1 +` on an
even lo'. The final SWAP is in line.
Measured on the golden model at 32- and 64-bit cells:
headroom data 3 -> 6 cells under args, return 3 -> 6 entries
speed ~1355-1605 -> ~227-313 instruction words per call
Still exact for every uhi < ud (edge-vector triples and 20000 random
cases, optimised and ASan+UBSan). Retargeting each of the four in-loop
branches to the wrong label fails more than 12000 checks each.
DECOMPOSITION.md: section 4 UM/MOD replaced, with its derivation and
stack limits.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
UM/MOD is assembled exactly as DECOMPOSITION.md section 4 gives it and
needs no change: it is exact for every uhi < ud at 32- and 64-bit cells.
It is checked against q*d + r = uhi:ulo with r < d through v4_umul, so
no second C divider has to be trusted. Coverage: all edge-vector triples
with uhi < ud plus 20000 pseudo-random cases (top-bit, small and
near-maximum divisors), optimised and under ASan+UBSan. Two hand
mutations each fail more than 15000 checks.
New headroom probe: runs a word with marked cells under the canary and
under its return address and reports how many survive, since the D-2
circular stacks overwrite silently instead of faulting. Measured:
UM* 6 data cells under its args, 4 return entries under its return
UM/MOD 3 data cells under its args, 3 return entries under its return
DECOMPOSITION.md: UM/MOD marked executed, with its defined range and
stack limits.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
UM* as first written in DECOMPOSITION.md section 4 was exact only while
u1 <= 2^(n-2): plain +* loses the carry out of T and its shift keeps T's
sign bit, so the loop is exact only while S and T stay in
[-2^(n-2), 2^(n-2)).
The rewrite multiplies by s = u1 2/, which always lies in that range,
starting T at t0 = u2 2/ when u1 is odd, so the loop yields
hi:lo = t0 + s*u2 exactly. Then
u1*u2 = 2*(hi:lo) + c_lo + c_hi*2^n
c_lo = u1 & u2 & 1
c_hi = (u1<0 ? u2 : 0) + (u1 odd and u2<0 ? 1 : 0)
restores the halved-away bits and the unsigned reading of both top bits.
test_foundation.c runs the new definition against v4_umul over every
pair of the edge vectors plus 20000 pseudo-random pairs, at 32- and
64-bit cells, optimised and under ASan+UBSan. The two pinned failing
cases are now ordinary exactness checks. Two hand mutations of the
correction step each fail more than 10000 checks at both widths.
DECOMPOSITION.md: section 4 UM* replaced, D-3 ruling text updated.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
First code for StarForth v4 (JUSTIFICATION.md section 10, step 1): one node
of the 32-instruction core as a C99 model, with cell width as a build
parameter.
- Node: P, A, B, F18 circular stacks (10 and 9 deep, D-2), word-addressed
memory (D-1), 5% guard bands on every bounded list.
- Instruction word: six 5-bit slots in 32 bits at every cell width.
- Executor: all 32 opcodes of DECOMPOSITION.md 1.3. Cell arithmetic wraps
explicitly; no signed overflow or implementation-defined shift.
- Heat: per-opcode and per-call-target counters and the anti-clock, driven
by instruction retirement (1.4, D-6 interim).
- Slot packer and runner for tests, and a reference unsigned multiply in
plain C99 with no 128-bit type.
Tests run at 32- and 64-bit cells, and under ASan and UBSan. They cover
every opcode and execute the first section 4 definitions (NIP SWAP OR
NEGATE ROT 0< 0= 2DUP - U<) against the C operation each stands for.
UM* as written in section 4 is exact only while u1 <= 2^(n-2). Two known
failing cases are pinned in test_foundation.c until it is rewritten.
DECOMPOSITION.md: record D-9, the instruction word is 32 bits at every
cell width (ruled 2026-10-02).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Source tree reorganization:
- Move StarForth v3 engine to v3/ (src/, include/, Makefile)
- Move kernel to kernel/ (src/, include/, linker/, Makefile)
- Create v4/ skeleton for F18-ISA golden model (DECOMPOSITION.md, JUSTIFICATION.md)
- Move FABRIC-0..4.md to docs/fabric/
- Move ONTOLOGY.md and ROADMAP.md to docs/
Board infrastructure:
- Add boards/ser5/, boards/raspi/, boards/milkv/, boards/zynq7020/
- Each board has board.mk (ISA, CPU flags, boot recipe) and README.md
- Root Makefile becomes thin dispatcher: boot_image, all, clean, docs take TARGET
- make boot_image TARGET=SER5|RASPI|MILKV builds one GPT/MBR image per board
- ZYNQ7020 target exists but stops with clear error (ARMv7 port not built yet)
- scripts/mkdiskimage.sh builds disk images for all boards
Docs pipeline:
- docs/book/ with LaTeX master (main.tex) and Makefile
- pandoc converts Markdown to LaTeX at build time
- Two Lua filters: table-widths.lua (wide tables wrap), code-breaks.lua (inline code breaks)
- make docs builds single PDF (754 pages, 0 missing characters)
- make docs TARGET=<board> adds board appendix
- build/docs/<book|board>/meta.tex stamps git commit into PDF
Bug fixes:
- 42 include paths that only worked by accident now use correct relative paths
- clang-18 hardcode replaced with configurable CC variable (fixed aarch64 build)
- Pi 5: kernel_2712.img linked at 0x80000, .bss zeroed, memory reserved
- Doxyfile, .clang-tidy, README.md, Kconfig paths updated
Verified:
- Hosted v3 build passes 1012 tests, 0 failures
- SER5 image boots in QEMU (OVMF), POST passes, K exact (65536 = Q48_ONE)
- Milk-V image boots in QEMU (OpenSBI + U-Boot + bootefi), POST passes
- make clean TARGET=<board> removes only that board and its ISA objects
- make all builds all boards, hosted v3, and docs in one run
Co-authored-by: Junie <junie@jetbrains.com>
D-1 word addressing; D-2 F18 circular stacks (10/9 deep), hidden, so
DEPTH/PICK/ROLL/.S/SP@/SP! are retired everywhere and the DSP register is
dropped; D-3 plain F18 +* (UM* flagged for revision); D-5 host width
matches the host CPU; D-7 moot; D-8 signed Q48.16; D-4 and D-6 deferred
to the hosted-mesh step. Adds a note that every CAP definition must be
re-checked against the 10/9 stack depths.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BY9HMwK5Cetz3caBgHGyds
JUSTIFICATION.md records why v4 exists and the reasoning behind each
major design decision. DECOMPOSITION.md assigns every v3 C primitive a
fate on the 32-instruction F18-derived core.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BY9HMwK5Cetz3caBgHGyds