Files
LithosAnanake/include/starkernel/vm/kernel_hermes.h
T
Robert Allan JamesandClaude Sonnet 5 205a49ecd0 ACK/NACK and private-channel negotiation -- FABRIC-3.6.md task 3.6
Extracted sk_hermes_send_one() from sk_hermes_publish()'s own
per-subscriber body -- one code path for both point-to-point and
fan-out delivery, so the ledger can never diverge between them.
Point-to-point addressing turned out to be load-bearing, not
incidental: sk_hermes_publish()'s fan-out sets msg->to to whichever
member it is iterating, so a negotiation message "published" to the
common channel would spuriously reach every member, not just the real
target (checked with advisor() before building the naive version).
"Over the common channel" means every VM is reachable from birth (task
3.2), not that the exchange itself fans out -- messaging.4th's own
CH-REQUEST carried an explicit `to` for the same reason.

sk_hermes_channel_request/respond/close build the mechanics: request ->
grant (creates a private channel, subscribes both parties, sends
CH_GRANT + one ACK) or NACK ("a deny is a NACK", SXLV.1 -- no separate
type); close authorized by membership alone. The grant/deny decision is
a plain caller-supplied `approved` bool -- task 3.7 replaces the call
site that produces it with a real ACL.4th query, not this signature.

Self-test covers the task's own three checks plus a sibling advisor()
flagged: an approved respond() whose channel creation itself fails
(table exhausted) must still fall through to NACK, not a silent false
grant or half-open channel -- verified by exhausting the whole channel
table and confirming the fallback.

Bug found and fixed before this was called done: the first draft
dropped a message via pending_pop() alone, without releasing it first,
leaking its Stadium heat and failing the self-test's own ledger
baseline check (logs/20260922-065946/amd64/, kept as audit trail).
Fixed and re-verified PASS on all three architectures.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-22 07:25:51 -04:00

733 lines
37 KiB
C
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
/*
StarForth — Steady-State Virtual Machine Runtime
Copyright (c) 2023–2025 Robert A. James
All rights reserved.
This file is part of the StarForth project.
Licensed under the StarForth License, Version 1.0 (the "License");
you may not use this file except in compliance with the License.
You may obtain a copy of the License at:
https://github.com/star.4th@proton.me/StarForth/LICENSE.txt
This software is provided "AS IS", WITHOUT WARRANTY OF ANY KIND,
express or implied, including but not limited to the warranties of
merchantability, fitness for a particular purpose, and noninfringement.
See the License for the specific language governing permissions and
limitations under the License.
*/
/**
* kernel_hermes.h - Kernel-resident Hermes: message and membership
* structures (FABRIC-3.6.md task 2.1, item 28; design: FABRIC-3.5.md
* SIII/SXXXIII/SXXXIV).
*
* Phase 2, task 2.1 ONLY: type definitions, wired to nothing, drawing no
* heat. No allocator, no send/deliver/reap logic, no registration
* anywhere -- those are tasks 2.2 onward, each its own commit. This file
* existing and compiling changes no VM's dictionary and no runtime
* behaviour; that is deliberate (FABRIC-3.5.md SXXII.4: Phase 2 structures
* come first and prove nothing until the allocator is built on top, SXL.4
* item 41).
*
* SkHermesMessage mirrors capsules/common/messaging.4th's live MSG-CELLS
* layout (9 cells: MSG-TYPE@/FROM@/TO@/PADDR@/PLEN@/STADIUM-CELL@/SEQ@/
* CH@/ORIG-TYPE@) field-for-field, per FABRIC-3.5.md SXXXIII.4 item 1 --
* "roughly half the file is accessors that become struct fields." The
* Stadium-cell field is the heat coupling itself: a message's heat is not
* a field of its own, it IS the Stadium cell it occupies (SXL.4's
* consumption model, SXXXIX.4's per-VM invariant) -- there is deliberately
* no separate heat field here to keep that single-source-of-truth.
*
* SkHermesMembership is the "one broadcast membership list" SXXXIII.4
* item 3 and SXXXIII.5 recommend in place of messaging.4th's 28-word
* channel abstraction (CH-REQUEST/ACCEPT/CONFIRM/CLOSE/MINT-ID and the
* CH-NEGOTIATING/OPEN/CLOSING state machine) -- traced to have exactly one
* live caller, CH-ADD-MBR, everything else channel-shaped is unexercised.
* Whether kernel-Hermes ever adds negotiation on top is item 27, an open
* Phase 3 ruling (FABRIC-3.6.md B1) -- this structure does not answer
* that question, it only holds a flat list, which is correct either way.
*/
#ifndef STARKERNEL_VM_KERNEL_HERMES_H
#define STARKERNEL_VM_KERNEL_HERMES_H
#ifdef __STARKERNEL__
#include <stddef.h>
#include <stdint.h>
#include "starkernel/vm_uuid.h" /* VMUuid */
#include "starkernel/q48_16.h" /* Q48_ONE */
/*
* SK_HERMES_MSG_MAX / SK_HERMES_MEMBER_MAX - sizing. Mirrors
* messaging.4th's own MSG-MAX (32) and MBR-MAX (64) as a starting point --
* kernel-Hermes is a single central pool rather than N per-VM arenas, so
* these may need revisiting once real traffic exists to size against. Not
* a ruling, just where the FORTH precedent already was.
*/
#define SK_HERMES_MSG_MAX 32
#define SK_HERMES_MEMBER_MAX 64
/*
* SK_HERMES_Q_SLOT - per-message admission heat, task 2.2 (item 28).
* messaging.4th:44-46 derives Q.SLOT as the reservoir remaining after
* COMMON-CH's own Q.1/3 floor, split across MSG-MAX + (CH-MAX-1) slots --
* a floor that exists to reserve heat for the one live channel object
* itself. Kernel-Hermes has no such object: SXXXIII.4 item 3 and
* SXXXIII.5 replace the channel abstraction with a flat membership list
* that carries no Stadium heat of its own, so there is nothing left for a
* floor to protect. Deliberately simpler here rather than carrying the
* old formula's now-unmotivated term forward: reservoir split evenly
* across message slots only.
*/
#define SK_HERMES_Q_SLOT ((uint64_t)Q48_ONE / SK_HERMES_MSG_MAX)
/*
* SkHermesMessage - one message slot, field-for-field mirror of
* messaging.4th's 9-cell MSG layout.
*
* @field type Message type code (SPAWN/PAUSE/RESUME/KILL-style
* codes are Category A/dead per task 0.1/0.4; live
* types today are CONSOLE-CMD-EVENT(7),
* ELEVATE-REQUEST(8), BLK-ATTACH-EVENT(9), messaging.4th
* reserved sentinels MSG-NACKED(253)/MSG-DELIVERED(255)).
* @field from Sending VM (mirrors MSG-FROM@).
* @field to Target VM (mirrors MSG-TO@).
* @field payload_addr Out-of-line payload address (mirrors MSG-PADDR@).
* FABRIC-3.5.md SXLIII.6.1 (item 44, still open): a
* payload above INPUT_BUFFER_SIZE-1 (1024) bytes cannot
* be drained in one interpret call -- bound or chunk it
* before Phase 3, not here.
* @field payload_len Payload length in bytes (mirrors MSG-PLEN@).
* @field stadium_cell Index into the Stadium cell array this message's
* heat currently occupies, or a sentinel meaning "none"
* (mirrors MSG-STADIUM-CELL@) -- the heat coupling
* itself; see this file's own top comment.
* @field seq Monotonic send sequence (mirrors MSG-SEQ@).
* @field channel Broadcast/channel marker; unused today (mirrors
* MSG-CH@) -- item 27 territory, not decided here.
* @field orig_type Original type before a NACK/redeliver rewrite
* (mirrors MSG-ORIG-TYPE@).
* @field in_use Free-list occupancy flag for task 2.2's allocator.
* Not present in the FORTH layout (which uses
* MSG-TYPE@ 0<> as its own live/free test) -- kept
* explicit here rather than overloading `type == 0`,
* since kernel-Hermes's held/pulled/returned/consumed
* ledger (task 2.4, SXL.4) needs an unambiguous
* occupancy bit independent of the type field's value.
*/
typedef struct {
uint32_t type;
VMUuid from;
VMUuid to;
void *payload_addr;
uint32_t payload_len;
int32_t stadium_cell;
uint32_t seq;
uint32_t channel;
uint32_t orig_type;
int in_use;
VMUuid owner; /* VM whose reservoir funded this message (task 2.7) */
} SkHermesMessage;
/*
* SkHermesMembership - one flat broadcast membership list: every VM that
* has joined, no per-member state beyond identity. Replaces the 28-word
* channel abstraction; see this file's own top comment for why.
*
* @field members Member VM identities, valid for indices < count.
* @field count Number of valid entries in members[].
*/
typedef struct {
VMUuid members[SK_HERMES_MEMBER_MAX];
size_t count;
} SkHermesMembership;
/*
* sk_hermes_alloc - Heat-coupled allocate (task 2.2, item 28): pull
* SK_HERMES_Q_SLOT from vm_id's own Stadium reservoir, admit a real
* Stadium-floor patron with that heat (behaviour DELIVER, matching
* messaging.4th's own `SB-DELIVER STADIUM-ADMIT`), and claim a free
* message slot recording the admitted cell. Refuses cleanly, rolling
* back anything already pulled/admitted, if reservoir, arena, or the
* Stadium floor itself refuses.
*
* CORRECTED in task 2.3 (see this file's .c counterpart's own top
* comment): the first cut of this function never admitted a real
* Stadium patron, which left held message heat invisible to SXL.4's
* conservation invariant while allocated. It now does, which is what
* makes sk_hermes_release()'s eviction meaningful.
*
* Wired to nothing outside this file's own self-test (kernel_main.c) as
* of task 2.2/2.3 -- no FORTH word, no capsule interaction, no protocol
* logic (send/deliver/reap are later tasks). Exercised only by a
* synthetic-VM self-test; does not change any real VM's dictionary.
*
* @param vm_id Caller whose reservoir is charged.
* @param out_msg On success, set to the claimed slot. Untouched on
* refusal.
* @return 0 on success, -1 on refusal (insufficient reservoir, no free
* slot, or the Stadium floor itself refused admission -- this
* function does not distinguish the three in the return value;
* all three leave all state exactly as it was).
*/
int sk_hermes_alloc(VMUuid vm_id, SkHermesMessage **out_msg);
/*
* sk_hermes_release - Return a message's heat via the Stadium eviction
* path (task 2.3, item 28), mirroring MSG-FREE-NODE
* ("DUP 5 CELLS + @ STADIUM-EVICT DROP"). stadium_evict() itself returns
* the departing patron's remaining heat to its owning VM's reservoir --
* this function does not touch the reservoir directly, matching the
* FORTH shape exactly. Clears the slot (in_use = 0) whether or not
* eviction succeeds, since a message this function was asked to release
* should not remain allocated either way.
*
* @param msg A slot previously returned by sk_hermes_alloc(). Refused
* (returns -1, no effect) if NULL or already released.
* @return 0 on successful eviction, -1 if msg was invalid or eviction
* itself was refused by the Stadium floor.
*/
int sk_hermes_release(SkHermesMessage *msg);
/*
* sk_hermes_ledger - Task 2.4 (item 28), the four counters FABRIC-3.5.md
* SXXXVII.3/SXL.4 name: held, pulled, returned, consumed. Each is touched
* at exactly one call site:
*
* held += pulled amount -- sk_hermes_alloc(), on success only
* held -= returned amount -- sk_hermes_release(), on success only
* pulled += pulled amount -- sk_hermes_alloc(), same site as held's
* returned += returned amount -- sk_hermes_release(), same site as held's
* held -= decay delta -- sk_hermes_decay(), same site as consumed's
* consumed += decay delta -- sk_hermes_decay() (task 2.5)
*
* The audit invariant this ledger exists to make checkable (task 2.6,
* epsilon zero): held == pulled - returned - consumed. All four are
* exact integers (SXXXVII.3: "Nothing here is a measurement. There is no
* noise to tolerate.") -- this accessor is how task 2.6's audit, task
* 2.7's Stage B proof, and task 2.8's scan cross-check all read the same
* four numbers, rather than each keeping its own copy.
*
* @param held Set to the current live total (may be NULL).
* @param pulled Set to the cumulative total ever pulled (may be NULL).
* @param returned Set to the cumulative total ever returned (may be
* NULL).
* @param consumed Set to the cumulative total ever consumed by decay
* (may be NULL).
*/
void sk_hermes_ledger(uint64_t *held, uint64_t *pulled, uint64_t *returned, uint64_t *consumed);
/*
* SK_HERMES_Q_DECAY - per-application decay factor, Q48.16. Identical to
* messaging.4th's `65208 CONSTANT Q-DECAY` (65208/65536 ~= 0.99499):
* FABRIC-3.5.md SXL.4 rules that the consumption economy is preserved
* exactly, not re-tuned.
*/
#define SK_HERMES_Q_DECAY ((uint64_t)65208)
/*
* sk_hermes_decay - Apply one decay step to a live message (task 2.5,
* item 28), mirroring MSG-COOL-ONE (`MSG-HEAT@ Q-DECAY Q.* MSG-HEAT!`):
* the message's Stadium-cell heat becomes q48_mul(heat, Q_DECAY). Unlike
* MSG-COOL-ONE, the destroyed difference is RECORDED (SXXXVII.3): the
* delta heat_before - heat_after is added to `consumed` and subtracted
* from `held`, so held == pulled - returned - consumed stays exact.
*
* @param msg A live slot from sk_hermes_alloc(). Refused (returns -1, no
* effect) if NULL, not in use, or holding no Stadium cell.
* @return 0 on success, -1 if refused.
*/
int sk_hermes_decay(SkHermesMessage *msg);
/*
* Task 2.6 (item 28): the self-audit, epsilon zero (SXXXVII.3).
*
* sk_hermes_audit_values - pure predicate: nonzero iff
* held == pulled - returned - consumed exactly. Separate from the live
* ledger so the check can be exercised with deliberately corrupted values
* without mutating real state.
*
* sk_hermes_audit - checks the live ledger (O(1), SXXXVII.4); on failure
* prints a console line and bumps a failure count. Called at the end of
* every successful mutation: alloc, release, decay.
*
* sk_hermes_audit_failure_count - cumulative failures seen by the live
* audit; must be 0 in a healthy kernel.
*/
int sk_hermes_audit_values(uint64_t held, uint64_t pulled, uint64_t returned, uint64_t consumed);
int sk_hermes_audit(void);
uint64_t sk_hermes_audit_failure_count(void);
/*
* Task 2.8 (item 28), SXXXVII.4: the scan that verifies the counters
* themselves. DIAGNOSTIC ONLY -- O(SK_HERMES_MSG_MAX), never called from
* alloc/release/decay.
*
* sk_hermes_scan_held - walks the message arena and sums the live Stadium
* cell heat of every in-use message (the ground truth `held` claims to
* track). Sets *live_count (may be NULL) to the messages counted.
*
* sk_hermes_scan_check - nonzero iff the scan sum equals the `held`
* counter exactly.
*/
uint64_t sk_hermes_scan_held(size_t *live_count);
int sk_hermes_scan_check(void);
/*
* Channel table (FABRIC-3.6.md task 3.2, B1 / FABRIC-3.5.md SXLV.1): B1
* overrules SXXXIII.5's single flat broadcast list -- messaging is
* publish/subscribe, one permanent common channel every VM joins at
* birth, private channels created by request/grant/deny (task 3.6, not
* this one). SkHermesMembership (task 2.1) becomes per-channel: one
* instance per SkHermesChannel table entry instead of one global list.
*
* INERT, same posture as task 2.1: this task wires channel existence and
* membership bookkeeping only. No publish, no dispatch, no ACK/NACK, no
* ACL hook (tasks 3.3, 3.6, 3.7) -- creating/destroying/subscribing a
* channel here changes no VM's dictionary and delivers no message.
*
* Sizing: SXLV.1 says the channel table has "no fixed channel maximum,
* same reasoning as SXLV.2" -- SXLV.2/task 3.1's own ruling is boot-time,
* RAM-derived sizing (Stadium's stadium_max_vm_count_val pattern), not a
* literal unbounded/growable table. No separate numeric sizing rule was
* ruled for the channel table specifically (task 3.0's five sub-items
* covered ACK cadence/ACL-hook-location/switch-table-sizing/chunk-
* framing/drain-cadence, not this). Extending task 3.1's already-ruled
* pattern directly -- same stadium_max_vm_count() bound, same
* kmalloc-at-boot shape -- is the smallest choice consistent with what
* was ruled, flagged here as an engineering extrapolation, not restated
* as a separate Captain Bob ruling.
*/
/* SK_HERMES_CHANNEL_COMMON - the permanent common channel's fixed index.
* Exists (in_use, empty membership) from sk_hermes_channels_boot_init()
* itself, before any VM is born -- every VM joins it at birth (task 3.2's
* own kernel_main.c/capsule_birth.c wiring), never destroyed. */
#define SK_HERMES_CHANNEL_COMMON 0
/*
* SkHermesChannel - one channel/topic: its membership plus an occupancy
* flag for the dynamic table below. No name field -- messaging.4th's own
* channel abstraction (CH-ARENA) has none either; a channel is identified
* by its table index, the same way a Stadium quota slot or a switch-
* signal slot is identified by index, not by string.
*/
typedef struct {
SkHermesMembership membership;
int in_use;
} SkHermesChannel;
/* Boot-time allocation: kmalloc's the channel table to
* stadium_max_vm_count() entries (see this section's own sizing note
* above) and creates the common channel (index SK_HERMES_CHANNEL_COMMON,
* in_use, empty membership -- members join at birth, not pre-populated
* here). Must run after stadium_boot_init() (that bound is 0, and this
* fails, until Stadium has computed it). Soft failure -- returns -1 and
* leaves the table unallocated (capacity 0, so every call below simply
* refuses) rather than halting boot, same posture as
* stadium_boot_init()/session_boot_init()/sk_vm_switch_signal_boot_init().
* Idempotent-unsafe: calling twice leaks the first allocation, so callers
* must call it exactly once. */
int sk_hermes_channels_boot_init(void);
/* sk_hermes_channel_create - allocate a new, empty, inert channel.
* Returns its table index, or -1 if the table is full or
* sk_hermes_channels_boot_init() was never called/failed. */
int sk_hermes_channel_create(void);
/* sk_hermes_channel_destroy - tear down a channel created by
* sk_hermes_channel_create(). Refuses (-1, no effect) if channel_id is
* SK_HERMES_CHANNEL_COMMON (the common channel is permanent, SXLV.1),
* out of range, or not in use. Clears membership. */
int sk_hermes_channel_destroy(int channel_id);
/* sk_hermes_channel_subscribe - add vm_id to channel_id's membership.
* Refuses (-1, no effect) if channel_id is invalid/not in use, vm_id is
* already a member, or the channel's membership is already at
* SK_HERMES_MEMBER_MAX. Idempotent in effect (a second call with the
* same args refuses rather than duplicating), not in return value. */
int sk_hermes_channel_subscribe(int channel_id, VMUuid vm_id);
/* sk_hermes_channel_unsubscribe - remove vm_id from channel_id's
* membership (compacts the list, same shape as
* sk_vm_switch_signal_unregister()). Refuses (-1, no effect) if
* channel_id is invalid/not in use or vm_id is not a member. */
int sk_hermes_channel_unsubscribe(int channel_id, VMUuid vm_id);
/* sk_hermes_channel_is_member - nonzero iff vm_id is currently a member
* of channel_id. Zero (not an error signal) if channel_id is invalid/not
* in use. */
int sk_hermes_channel_is_member(int channel_id, VMUuid vm_id);
/* sk_hermes_channel_member_count - current membership size of channel_id,
* or -1 if channel_id is invalid/not in use. */
int sk_hermes_channel_member_count(int channel_id);
/* sk_hermes_channel_capacity - the table's boot-time-computed capacity
* (0 if sk_hermes_channels_boot_init() was never called or failed) --
* DoE/test observability, mirrors sk_vm_switch_signal_slot_capacity(). */
int sk_hermes_channel_capacity(void);
/* Boot-time allocation for the per-subscriber pending-queue table
* (FABRIC-3.6.md task 3.3), kmalloc'd to stadium_max_vm_count() entries --
* same sizing pattern as the channel table (task 3.2) and switch table
* (task 3.1); see kernel_hermes.c's own comment on this call for why. Must
* run after stadium_boot_init(). Soft failure -- returns -1 and leaves the
* table unallocated (every queue lookup then finds nothing) rather than
* halting boot. Idempotent-unsafe: call exactly once. */
int sk_hermes_queues_boot_init(void);
/*
* sk_hermes_send_one - point-to-point delivery to exactly one VM
* (FABRIC-3.6.md task 3.6, extracted from sk_hermes_publish()'s own
* per-subscriber body, which now calls this once per channel member).
* One allocation, one queue push, same rollback-on-refusal discipline
* sk_hermes_publish() already had -- one code path, so the ledger can
* never diverge between the two callers.
*
* Exists because task 3.6's negotiation messages (request/grant/deny/
* close/ACK/NACK) are inherently addressed to ONE specific VM, and
* sk_hermes_publish()'s fan-out cannot express that: it sets msg->to to
* whichever member it is currently iterating, so a negotiation message
* "published" to the common channel would be delivered to (and drawn
* heat against) every common-channel member, each believing it was the
* addressee, not just the real target. `messaging.4th`'s own
* `CH-REQUEST ( type from to paddr plen -- )` carried an explicit `to`
* for exactly this reason -- point-to-point addressing was in the
* original protocol from the start.
*
* @param from Sender, whose reservoir funds the allocation.
* @param to The one recipient.
* @param type Message type code.
* @param channel_id Tag copied into the message's own `channel`
* field -- purely informational at this level (no
* membership is consulted or required), letting a
* response encode context (e.g. task 3.6's GRANT
* uses this to tell the requester which private
* channel was just created).
* @param payload_addr Out-of-line payload address, passed through
* unchanged.
* @param payload_len Refused (-1) if it exceeds
* SK_HERMES_CHUNK_MAX_PAYLOAD.
* @return 0 on success, -1 on any refusal (bound, reservoir, arena, or
* destination queue full -- all leave state exactly as found).
*/
int sk_hermes_send_one(VMUuid from, VMUuid to, uint32_t type, uint32_t channel_id,
void *payload_addr, uint32_t payload_len);
/*
* Publish path, no dispatch (FABRIC-3.6.md task 3.3, SXLIII.3; heat-cost
* ruling 2026-09-21: one message per subscriber, separate heat draw each --
* matches the existing heat-coupled allocator 1:1, no refcount machinery).
*
* sk_hermes_publish() loops sk_hermes_send_one() (task 3.6 extraction,
* above) once per channel member, funded by the publisher's own
* reservoir each time. It does not interpret, deliver, or otherwise
* dispatch anything -- draining a queue at a VM's own outermost
* interpret checkpoint is task 3.4's scope, not this one's.
*
* Best-effort, not atomic across subscribers: if a given subscriber's
* send_one() call is refused (reservoir exhausted, message arena full,
* or that subscriber's own pending queue full), that one subscriber is
* skipped -- send_one() has already rolled back its own allocation
* internally, so nothing here needs to -- but delivery to every OTHER
* subscriber already queued is not undone. This was not a
* separate Captain Bob ruling; it is the natural reading of "ledger audit
* and stadium_conserved() hold across N publishes to M subscribers" (task
* 3.3's own check) -- those invariants hold under partial delivery just
* as well as under all-or-nothing, and requiring atomicity across M
* independent reservoir-funded allocations would need a two-phase
* commit/rollback this task's inert scope does not call for.
*
* @param from Publisher, whose reservoir funds every allocation.
* @param channel_id Target channel (SK_HERMES_CHANNEL_COMMON or a
* channel from sk_hermes_channel_create()). Refused
* (-1) if invalid/not in use.
* @param type Message type code, passed through unchanged.
* @param payload_addr Out-of-line payload address, passed through
* unchanged.
* @param payload_len Payload length in bytes. Refused (-1, no
* allocation attempted) if it exceeds
* SK_HERMES_CHUNK_MAX_PAYLOAD (task 3.5, added here
* since task 3.3 deliberately parked this check --
* see that section's own doc comment below).
* @return Count of subscribers successfully enqueued to (0..member count),
* or -1 if channel_id itself was invalid or payload_len exceeded
* the bound.
*/
int sk_hermes_publish(VMUuid from, int channel_id, uint32_t type,
void *payload_addr, uint32_t payload_len);
/* SK_HERMES_PENDING_MAX - per-subscriber pending-queue depth. The global
* message arena (SK_HERMES_MSG_MAX) is the real ceiling on how many
* messages can ever be in flight system-wide, so sizing each VM's own
* queue to that same bound is a safe, simple upper limit rather than a
* new number to justify. */
#define SK_HERMES_PENDING_MAX SK_HERMES_MSG_MAX
/* sk_hermes_pending_count - number of messages currently queued for
* vm_id (0 if vm_id has no queue yet -- never having received a message
* is not an error). */
int sk_hermes_pending_count(VMUuid vm_id);
/* sk_hermes_pending_peek - the oldest still-queued message for vm_id, or
* NULL if vm_id has no queue or an empty one. Does not remove it -- task
* 3.4's drain logic is expected to peek, interpret, then pop. */
SkHermesMessage *sk_hermes_pending_peek(VMUuid vm_id);
/* sk_hermes_pending_pop - removes (does not release/interpret) the
* oldest queued entry for vm_id. Callers that also want the message's
* heat returned must call sk_hermes_release() on the value
* sk_hermes_pending_peek() returned, themselves, before or after popping
* -- this function only advances the queue. Refused (-1, no effect) if
* vm_id has no queue or an empty one. */
int sk_hermes_pending_pop(VMUuid vm_id);
/* Forward declaration only -- kernel_hermes.h deliberately does not
* include vm.h (kept decoupled from the full VM struct, same posture as
* every other type in this file), but sk_hermes_drain_checkpoint() below
* needs a VM* parameter. Matches vm.h's own `typedef struct VM VM;`
* shape exactly, so no redefinition conflict. */
typedef struct VM VM;
/*
* Drain at the outermost checkpoint (FABRIC-3.6.md task 3.4, SXLIII.3-.5;
* one message per checkpoint, ruled 2026-09-21). Called from
* execute_colon_word()'s existing cooperative checkpoint in vm_core.c,
* gated the same way the switch-signal checkpoint already is
* (sk_vm_at_outermost_interpret()).
*
* AMENDS SXLIII.5's own claim: "recursive drain is prevented for free"
* via g_vm_interpret_depth is true (the SAME message cannot drain twice),
* but that counter says nothing about a SEPARATE, real hazard SXLIII.5
* never named -- calling vm_interpret() on this same vm, from inside its
* own currently-running vm_interpret() call, overwrites vm->input_buffer/
* input_length/input_pos with the drained payload's own state.
* VMCallState (vm_state_push()/vm_state_pop(), mama_forth_words.c) saves
* only rsp/exit_colon/ecw_nesting -- never these three fields (the exact
* gap FABRIC-3.md SXX documents: "the same class... for input_buffer/
* input_pos not being saved by vm_state_push/pop"). Left unaddressed,
* the enclosing vm_interpret() call's own while loop would silently lose
* the rest of its input line/block the moment a drain fires mid-line --
* the same failure shape as the INPUT_BUFFER_SIZE 256 defect
* .claude/CLAUDE.md calls non-negotiable, and trap #1 in this document's
* own START HERE. Not a divergence from SXLIII.3's ruling (the mechanism
* is exactly as ruled); cursor/mode/error/abort preservation is this
* function's own implementation obligation, not a new design question.
*
* sk_hermes_drain_checkpoint() therefore snapshots vm->input_buffer/
* input_length/input_pos/mode/error/abort_requested before calling
* vm_interpret() on the queued payload, forces vm->mode to
* MODE_INTERPRET for the duration (a checkpoint reached mid-colon-
* definition must not let the payload's words compile into the
* enclosing definition), and restores all six afterward -- fully
* isolating the drain from whatever the enclosing execution was doing.
*
* Payload convention: payload_addr is trusted to point at a
* NUL-terminated C string (vm_interpret()'s own signature takes no
* length) -- payload_len is not consulted here. Bounding/validating that
* is task 3.5's scope (payload bound and chunking), not this one's.
*
* Fast path: sk_hermes_publish()/sk_hermes_pending_pop() maintain a
* single system-wide pending-total counter; this function reads it
* first and returns immediately if it is 0, so the common case (nothing
* in flight) costs one integer read on every word dispatch, not a
* stadium_max_vm_count()-sized queue-table scan.
*
* @param vm The VM at its own outermost checkpoint. NULL is refused.
* @return 1 if a message was drained cleanly, 0 if nothing was pending
* for this vm, -1 if a message was drained but interpreting its
* payload set vm->error (restored to its pre-drain value either
* way -- a bad message must not abort the enclosing execution).
*/
int sk_hermes_drain_checkpoint(VM *vm);
/*
* Payload bound and chunking (FABRIC-3.6.md task 3.5, B4 / FABRIC-3.5.md
* SXLV.3; chunk framing ruled 2026-09-21: each chunk carries (msg_id,
* seq, is_last) ahead of the real content).
*
* Sizing decision (not itself ruled -- see SXLV.4's own note that the
* concrete framing was a task-writing prerequisite, not a ruling):
* SK_HERMES_CHUNK_MAX_PAYLOAD (1024, one block, SXLIV.1) bounds a
* message's payload_len UNIFORMLY -- chunked or not. A chunk carrier's
* own payload is [SkHermesChunkHeader][content slice], so the slice
* itself is capped at 1024 - sizeof(SkHermesChunkHeader), not 1024,
* keeping every message on the wire under the same one-block bound
* SXLIII.6.1 already requires for vm_interpret()'s own drain limit.
* (The alternative -- slice up to 1024, carrier up to
* 1024+sizeof(header) -- was considered and rejected: it would mean a
* chunk carrier can never be handed to vm_interpret() as-is, making a
* later chunk-aware drain a special case instead of the same drain path
* every other message already uses.)
*
* sk_hermes_publish() (task 3.3, above) enforces this bound directly --
* refuses (-1) any single payload_len over SK_HERMES_CHUNK_MAX_PAYLOAD,
* whether or not it's a chunk carrier, since both cases hold the same
* invariant now.
*
* SENDING a payload over SK_HERMES_CHUNK_MAX_SLICE bytes is deliberately
* NOT a kernel-Hermes API here -- it is a loop a caller writes with the
* primitives below plus sk_hermes_publish(), the same way this task's
* own self-test proves reassembly (kernel_main.c: build N static chunk
* buffers, sk_hermes_publish() each). Building a chunking sender would
* need kernel-Hermes to own chunk-buffer memory with a real lifetime
* (kept alive until every subscriber has drained it, freed only once
* kernel-Hermes has no way to know that) -- a genuine new question this
* task's inert scope does not call for. A real multi-chunk sender is a
* later task's problem, when a real message type actually needs one
* (3.8+, the same boundary chunk-aware drain is already deferred past).
*/
/* One block = 1024 bytes (64x16, SXLIV.1) -- the uniform bound every
* message's payload_len must satisfy, chunked or not (see this
* section's own sizing note above). */
#define SK_HERMES_CHUNK_MAX_PAYLOAD 1024
/* SkHermesChunkHeader - prepended to a chunk carrier's own payload,
* ahead of up to SK_HERMES_CHUNK_MAX_SLICE bytes of real content.
* msg_id is caller-chosen and must be unique per logical multi-chunk
* send from a given sender -- reassembly groups chunks by it. seq is
* 0-based, ascending, no gaps, {0..n_chunks-1}. is_last is set on
* exactly the chunk whose seq is n_chunks-1, no other. */
typedef struct {
uint32_t msg_id;
uint32_t seq;
int is_last;
} SkHermesChunkHeader;
/* Content bytes per chunk -- the one-block bound minus header overhead. */
#define SK_HERMES_CHUNK_MAX_SLICE (SK_HERMES_CHUNK_MAX_PAYLOAD - (uint32_t)sizeof(SkHermesChunkHeader))
/* sk_hermes_chunk_count - number of SK_HERMES_CHUNK_MAX_SLICE-sized
* chunks `len` bytes needs (ceiling division). 0 bytes needs 0 chunks;
* 1..SK_HERMES_CHUNK_MAX_SLICE needs 1; and so on. Pure arithmetic, no
* side effects -- a caller sizing its own chunk-send loop calls this
* first. */
uint32_t sk_hermes_chunk_count(uint32_t len);
/* sk_hermes_reassemble - concatenates n_chunks chunk carrier messages
* (already drained off a subscriber's own pending queue via
* sk_hermes_pending_peek()/sk_hermes_pending_pop(), task 3.3) sharing
* one msg_id back into one byte-exact buffer at out_buf, in ascending
* seq order regardless of the order chunks[] itself is passed in.
*
* Refuses (-1, *out_len untouched, no partial copy ever lands in
* out_buf) if: n_chunks is 0 or exceeds SK_HERMES_MSG_MAX (the system-
* wide message arena size -- chunking can never need more chunk-carrier
* messages than the whole arena holds); any chunk is NULL, has no
* payload, or its payload is shorter than sizeof(SkHermesChunkHeader);
* the chunks' msg_ids disagree; seq values are not exactly
* {0..n_chunks-1} (a duplicate or a gap); is_last is not set on exactly
* the seq==n_chunks-1 chunk and no other; a non-final chunk's content
* slice is not exactly SK_HERMES_CHUNK_MAX_SLICE bytes, or the final
* chunk's is 0 or over that bound; or the total reassembled length
* would exceed out_buf_cap (checked once, from validated seq
* completeness, before any memcpy -- never order-dependent on which
* chunk happens to overflow first).
*
* @param chunks Array of n_chunks chunk carrier message pointers.
* @param n_chunks Chunk count (from sk_hermes_chunk_count() at send
* time, or simply len(chunks[])).
* @param out_buf Caller-owned destination buffer.
* @param out_buf_cap Capacity of out_buf in bytes.
* @param out_len Set to the reassembled length on success only.
* @return 0 on success, -1 on any refusal above.
*/
int sk_hermes_reassemble(SkHermesMessage **chunks, int n_chunks,
uint8_t *out_buf, uint32_t out_buf_cap,
uint32_t *out_len);
/*
* ACK/NACK and private-channel negotiation (FABRIC-3.6.md task 3.6, B1 /
* FABRIC-3.5.md SXLV.1). "A private channel is created by request ->
* grant/deny over the common channel" -- read as: every VM is a
* common-channel member from birth (task 3.2), so a requester can
* always REACH a target without a prior private channel; the exchange
* itself is point-to-point (sk_hermes_send_one(), task 3.6's own
* extraction above), not a fan-out to the whole common-channel
* membership -- see sk_hermes_send_one()'s own doc comment for why
* fan-out is structurally wrong for an addressed request.
* `messaging.4th`'s own `CH-REQUEST` carried an explicit `to` for the
* same reason; this is the same shape, not a new one.
*
* "A deny is a NACK" (SXLV.1) -- there is no separate CH_DENY type;
* denial IS SK_HERMES_MSG_TYPE_NACK. ACK is sent once, for the
* channel-open + delivery moment (the ruled cadence, 2026-09-21), not
* for every message that follows on the new channel.
*
* The GRANT/DENY *decision* is task 3.7's scope -- the ACL policy hook,
* "kernel-Hermes asks ACL.4th, never decides in C, never gates on
* zuse_session" (CLAUDE.md, SXLV.1). sk_hermes_channel_respond() below
* takes that decision as an explicit caller-supplied `approved` flag
* (named for what it is, not `allow`, so it reads as a caller decision
* passed in, never as policy living in C) -- task 3.7 replaces the CALL
* SITE that produces this flag with a real ACL.4th query; this
* function's own signature does not change.
*/
#define SK_HERMES_MSG_TYPE_CH_REQUEST 20
#define SK_HERMES_MSG_TYPE_CH_GRANT 21
#define SK_HERMES_MSG_TYPE_ACK 22
#define SK_HERMES_MSG_TYPE_NACK 23
#define SK_HERMES_MSG_TYPE_CH_CLOSE 24
/* Chosen clear of messaging.4th's own live/reserved type space
* (PAUSE-EVENT=2, RESUME-EVENT=3, KILL-EVENT=4, CONSOLE-CMD-EVENT=7,
* ELEVATE-REQUEST=8, BLK-ATTACH-EVENT=9, MSG-NACKED=253,
* MSG-DELIVERED=255) -- kernel-Hermes remains its own inert, parallel
* type space until a real cutover (task 3.8+) makes the two coincide;
* not itself a ruling, just room left deliberately. */
/* sk_hermes_channel_request - requester asks target for a new private
* channel. A point-to-point CH_REQUEST (sk_hermes_send_one(), tagged
* with SK_HERMES_CHANNEL_COMMON purely for reachability context),
* funded by requester's own reservoir. Creates nothing -- the responder
* decides via sk_hermes_channel_respond() below.
* @return 0 on success, -1 if the send itself was refused. */
int sk_hermes_channel_request(VMUuid requester, VMUuid target);
/* sk_hermes_channel_respond - target answers a pending request from
* requester, `approved` supplied by the caller (see this section's own
* top comment on why that is task 3.7's future call site, not a policy
* decision living here).
*
* Approved: creates a new private channel, subscribes both requester
* and target, sends CH_GRANT to requester (msg->channel = the new
* channel id -- how the requester learns which channel to use) followed
* by one ACK (the ruled channel-open+delivery cadence). If channel
* creation fails (table full) or either subscription fails, the
* partial channel is destroyed and this call falls through to the
* denied path below instead of leaving a half-open channel or a silent
* false grant -- the same "leave no half-state" discipline
* sk_hermes_publish()'s own per-subscriber rollback already uses.
*
* Denied (approved==0, or a failed grant attempt as above): sends NACK
* to requester ("a deny is a NACK", SXLV.1 -- no separate type). No
* channel exists afterward either way.
*
* @return The new channel id (>= 0) on a real grant, or -1 on any deny
* -- both a policy deny and a failed grant attempt leave no
* channel behind, which is what this task's own check asks for.
*/
int sk_hermes_channel_respond(VMUuid target, VMUuid requester, int approved);
/* sk_hermes_channel_close - tears down a private channel. Authority is
* membership alone: either party may close a channel it belongs to
* (narrowest defensible rule given both are already trusted members of
* it -- not gated on anything beyond that, and not itself a ruling).
* Refuses (-1, no effect) if closer is not a member of channel_id, or
* channel_id is SK_HERMES_CHANNEL_COMMON (sk_hermes_channel_destroy()
* itself already refuses the common channel).
* @return 0 on success, -1 on refusal. */
int sk_hermes_channel_close(VMUuid closer, int channel_id);
#endif /* __STARKERNEL__ */
#endif /* STARKERNEL_VM_KERNEL_HERMES_H */