ACK/NACK and private-channel negotiation -- FABRIC-3.6.md task 3.6

Extracted sk_hermes_send_one() from sk_hermes_publish()'s own
per-subscriber body -- one code path for both point-to-point and
fan-out delivery, so the ledger can never diverge between them.
Point-to-point addressing turned out to be load-bearing, not
incidental: sk_hermes_publish()'s fan-out sets msg->to to whichever
member it is iterating, so a negotiation message "published" to the
common channel would spuriously reach every member, not just the real
target (checked with advisor() before building the naive version).
"Over the common channel" means every VM is reachable from birth (task
3.2), not that the exchange itself fans out -- messaging.4th's own
CH-REQUEST carried an explicit `to` for the same reason.

sk_hermes_channel_request/respond/close build the mechanics: request ->
grant (creates a private channel, subscribes both parties, sends
CH_GRANT + one ACK) or NACK ("a deny is a NACK", SXLV.1 -- no separate
type); close authorized by membership alone. The grant/deny decision is
a plain caller-supplied `approved` bool -- task 3.7 replaces the call
site that produces it with a real ACL.4th query, not this signature.

Self-test covers the task's own three checks plus a sibling advisor()
flagged: an approved respond() whose channel creation itself fails
(table exhausted) must still fall through to NACK, not a silent false
grant or half-open channel -- verified by exhausting the whole channel
table and confirming the fallback.

Bug found and fixed before this was called done: the first draft
dropped a message via pending_pop() alone, without releasing it first,
leaking its Stadium heat and failing the self-test's own ledger
baseline check (logs/20260922-065946/amd64/, kept as audit trail).
Fixed and re-verified PASS on all three architectures.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This commit is contained in:
Robert Allan James
2026-09-22 07:25:51 -04:00
co-authored by Claude Sonnet 5
parent 8a2ee0fdad
commit 205a49ecd0
8 changed files with 37482 additions and 44 deletions
+130 -12
View File
@@ -383,26 +383,61 @@ int sk_hermes_channel_capacity(void);
* halting boot. Idempotent-unsafe: call exactly once. */
int sk_hermes_queues_boot_init(void);
/*
* sk_hermes_send_one - point-to-point delivery to exactly one VM
* (FABRIC-3.6.md task 3.6, extracted from sk_hermes_publish()'s own
* per-subscriber body, which now calls this once per channel member).
* One allocation, one queue push, same rollback-on-refusal discipline
* sk_hermes_publish() already had -- one code path, so the ledger can
* never diverge between the two callers.
*
* Exists because task 3.6's negotiation messages (request/grant/deny/
* close/ACK/NACK) are inherently addressed to ONE specific VM, and
* sk_hermes_publish()'s fan-out cannot express that: it sets msg->to to
* whichever member it is currently iterating, so a negotiation message
* "published" to the common channel would be delivered to (and drawn
* heat against) every common-channel member, each believing it was the
* addressee, not just the real target. `messaging.4th`'s own
* `CH-REQUEST ( type from to paddr plen -- )` carried an explicit `to`
* for exactly this reason -- point-to-point addressing was in the
* original protocol from the start.
*
* @param from Sender, whose reservoir funds the allocation.
* @param to The one recipient.
* @param type Message type code.
* @param channel_id Tag copied into the message's own `channel`
* field -- purely informational at this level (no
* membership is consulted or required), letting a
* response encode context (e.g. task 3.6's GRANT
* uses this to tell the requester which private
* channel was just created).
* @param payload_addr Out-of-line payload address, passed through
* unchanged.
* @param payload_len Refused (-1) if it exceeds
* SK_HERMES_CHUNK_MAX_PAYLOAD.
* @return 0 on success, -1 on any refusal (bound, reservoir, arena, or
* destination queue full -- all leave state exactly as found).
*/
int sk_hermes_send_one(VMUuid from, VMUuid to, uint32_t type, uint32_t channel_id,
void *payload_addr, uint32_t payload_len);
/*
* Publish path, no dispatch (FABRIC-3.6.md task 3.3, SXLIII.3; heat-cost
* ruling 2026-09-21: one message per subscriber, separate heat draw each --
* matches the existing heat-coupled allocator 1:1, no refcount machinery).
*
* sk_hermes_publish() allocates one SkHermesMessage per channel member
* (via sk_hermes_alloc(), funded by the publisher's own reservoir) and
* enqueues each onto that member's own pending queue -- FIFO, one queue
* per subscriber VM, found-or-created lazily on first use (same
* find-or-create-by-vm_id shape stadium.c's quota_slot_for_vm() and
* session.c's session_find() already use). It does not interpret,
* deliver, or otherwise dispatch anything -- draining a queue at a VM's
* own outermost interpret checkpoint is task 3.4's scope, not this one's.
* sk_hermes_publish() loops sk_hermes_send_one() (task 3.6 extraction,
* above) once per channel member, funded by the publisher's own
* reservoir each time. It does not interpret, deliver, or otherwise
* dispatch anything -- draining a queue at a VM's own outermost
* interpret checkpoint is task 3.4's scope, not this one's.
*
* Best-effort, not atomic across subscribers: if a given subscriber's
* allocation or enqueue fails (reservoir exhausted, message arena full,
* send_one() call is refused (reservoir exhausted, message arena full,
* or that subscriber's own pending queue full), that one subscriber is
* skipped -- the message already allocated for a failed enqueue is
* released back (rolled back) rather than left orphaned, but delivery to
* every OTHER subscriber already queued is not undone. This was not a
* skipped -- send_one() has already rolled back its own allocation
* internally, so nothing here needs to -- but delivery to every OTHER
* subscriber already queued is not undone. This was not a
* separate Captain Bob ruling; it is the natural reading of "ledger audit
* and stadium_conserved() hold across N publishes to M subscribers" (task
* 3.3's own check) -- those invariants hold under partial delivery just
@@ -609,6 +644,89 @@ int sk_hermes_reassemble(SkHermesMessage **chunks, int n_chunks,
uint8_t *out_buf, uint32_t out_buf_cap,
uint32_t *out_len);
/*
* ACK/NACK and private-channel negotiation (FABRIC-3.6.md task 3.6, B1 /
* FABRIC-3.5.md SXLV.1). "A private channel is created by request ->
* grant/deny over the common channel" -- read as: every VM is a
* common-channel member from birth (task 3.2), so a requester can
* always REACH a target without a prior private channel; the exchange
* itself is point-to-point (sk_hermes_send_one(), task 3.6's own
* extraction above), not a fan-out to the whole common-channel
* membership -- see sk_hermes_send_one()'s own doc comment for why
* fan-out is structurally wrong for an addressed request.
* `messaging.4th`'s own `CH-REQUEST` carried an explicit `to` for the
* same reason; this is the same shape, not a new one.
*
* "A deny is a NACK" (SXLV.1) -- there is no separate CH_DENY type;
* denial IS SK_HERMES_MSG_TYPE_NACK. ACK is sent once, for the
* channel-open + delivery moment (the ruled cadence, 2026-09-21), not
* for every message that follows on the new channel.
*
* The GRANT/DENY *decision* is task 3.7's scope -- the ACL policy hook,
* "kernel-Hermes asks ACL.4th, never decides in C, never gates on
* zuse_session" (CLAUDE.md, SXLV.1). sk_hermes_channel_respond() below
* takes that decision as an explicit caller-supplied `approved` flag
* (named for what it is, not `allow`, so it reads as a caller decision
* passed in, never as policy living in C) -- task 3.7 replaces the CALL
* SITE that produces this flag with a real ACL.4th query; this
* function's own signature does not change.
*/
#define SK_HERMES_MSG_TYPE_CH_REQUEST 20
#define SK_HERMES_MSG_TYPE_CH_GRANT 21
#define SK_HERMES_MSG_TYPE_ACK 22
#define SK_HERMES_MSG_TYPE_NACK 23
#define SK_HERMES_MSG_TYPE_CH_CLOSE 24
/* Chosen clear of messaging.4th's own live/reserved type space
* (PAUSE-EVENT=2, RESUME-EVENT=3, KILL-EVENT=4, CONSOLE-CMD-EVENT=7,
* ELEVATE-REQUEST=8, BLK-ATTACH-EVENT=9, MSG-NACKED=253,
* MSG-DELIVERED=255) -- kernel-Hermes remains its own inert, parallel
* type space until a real cutover (task 3.8+) makes the two coincide;
* not itself a ruling, just room left deliberately. */
/* sk_hermes_channel_request - requester asks target for a new private
* channel. A point-to-point CH_REQUEST (sk_hermes_send_one(), tagged
* with SK_HERMES_CHANNEL_COMMON purely for reachability context),
* funded by requester's own reservoir. Creates nothing -- the responder
* decides via sk_hermes_channel_respond() below.
* @return 0 on success, -1 if the send itself was refused. */
int sk_hermes_channel_request(VMUuid requester, VMUuid target);
/* sk_hermes_channel_respond - target answers a pending request from
* requester, `approved` supplied by the caller (see this section's own
* top comment on why that is task 3.7's future call site, not a policy
* decision living here).
*
* Approved: creates a new private channel, subscribes both requester
* and target, sends CH_GRANT to requester (msg->channel = the new
* channel id -- how the requester learns which channel to use) followed
* by one ACK (the ruled channel-open+delivery cadence). If channel
* creation fails (table full) or either subscription fails, the
* partial channel is destroyed and this call falls through to the
* denied path below instead of leaving a half-open channel or a silent
* false grant -- the same "leave no half-state" discipline
* sk_hermes_publish()'s own per-subscriber rollback already uses.
*
* Denied (approved==0, or a failed grant attempt as above): sends NACK
* to requester ("a deny is a NACK", SXLV.1 -- no separate type). No
* channel exists afterward either way.
*
* @return The new channel id (>= 0) on a real grant, or -1 on any deny
* -- both a policy deny and a failed grant attempt leave no
* channel behind, which is what this task's own check asks for.
*/
int sk_hermes_channel_respond(VMUuid target, VMUuid requester, int approved);
/* sk_hermes_channel_close - tears down a private channel. Authority is
* membership alone: either party may close a channel it belongs to
* (narrowest defensible rule given both are already trusted members of
* it -- not gated on anything beyond that, and not itself a ruling).
* Refuses (-1, no effect) if closer is not a member of channel_id, or
* channel_id is SK_HERMES_CHANNEL_COMMON (sk_hermes_channel_destroy()
* itself already refuses the common channel).
* @return 0 on success, -1 on refusal. */
int sk_hermes_channel_close(VMUuid closer, int channel_id);
#endif /* __STARKERNEL__ */
#endif /* STARKERNEL_VM_KERNEL_HERMES_H */