Skip to content

Optimization and bug notes

Each entry records the mechanism, result or correctness effect, and a reference for further detail. Measurements describe the stated workload and hardware. Current capabilities are in implemented surface; remaining work is in the roadmap.

Bounded replication reads and preserved progress

Section titled “Bounded replication reads and preserved progress”

Owner replication reads now stop after the useful byte-bounded prefix and one lookahead record, preserving snapshot deferral and ordered event gating. Follower progress keeps a pending owner read alive, avoiding discarded blocking scans, while reset and stop still cancel it. A three-node tmpfs comparison at 4 KiB measured lower CPU use and higher throughput, with cache, disk, oversized record and snapshot-boundary regression coverage.

Owner streams now adopt runtime read limits at the next batch, and oversized batches retain credit debt until durable application returns those bytes. Previously, saturating subtraction could let repeated oversized records grow the window. Tests cover live settings adoption and repeated oversized records, while the configuration reference separates queued frame count from outstanding byte credit.

Follower retention and non-destructive repair

Section titled “Follower retention and non-destructive repair”

Snapshot compaction could overtake admitted followers during sustained publishing, and the legacy checkpoint fallback correctly refused to replace their accepted history. Owner compaction now protects live follower cursors, and exhausted history starts consensus-authorized repair in a separate generation while preserving the original recovery evidence. Storage crash-boundary tests and a three-node repair test cover the switch, continued majority confirms and stale-session rejection.

Opt-in background agreement creates durable same-cut queue snapshots across all admitted replicas and advances retention only after consensus publication. Recovery verifies the surviving snapshot and suffix while preserving live payload checks; large live backlogs still require payload reads. Crash, partial-installation and confirmed-suffix tests cover the publication and recovery boundaries.

The Settings page compares saved values with the configurations installed by the local broker and connection updater. Independent cluster and cache revision counters are kept separate, and startup now installs the latest subscribed snapshot before waiting for changes. Existing operations can retain earlier snapshots, so installation does not imply adoption by every worker.

Inspection, replay and staging failures now carry numeric limits, accepted work and refused work into the bounded recovery timeline. An unchanged retry is identified as insufficient for these fixed limits, while unknown errors remain unclassified. This adds observability without changing source proof, budgets or activation requirements.

Boosted navigation releases page interaction listeners, timers, visibility callbacks and event streams. Disposed polling and queued stream callbacks cannot restart polling or restore an old page’s scroll position.

Normal ordered-application progress now wakes only eligible offset ranges, while reset and failure still wake invalidated operations. This removes the wake-up storm reproduced by the historical ten-connection workload, with benchmark results and storage, checkpoint and recovery regressions covering the change.

An isolated early-replication experiment passes 48 schedules covering both logs, partial writes, fsync errors, follower timing and abrupt process exit. No failed publish produced a success receipt or became owner-deliverable, and confirmed prefixes survived reopen. Eight additional TCP cluster cases preserve majority and three-copy requirements through owner failure, recovery and readmission, with the remaining adoption gates tracked in the recovery continuation plan.

Async page requests check their original page lifetime before updating controls, including delayed response bodies and error handlers. Boosted navigation accepts only the latest request and shares pending script loads, preventing older responses or initializers from replacing a newer page. Regression tests complete requests out of order and cover revisiting Settings while a save is pending.

Observer error isolated from the dashboard

Section titled “Observer error isolated from the dashboard”

The intermittent MutationObserver.observe exception reproduces in the embedded browser on script-free iframe controls with both lazy and eager loading. Those controls contain no Fibril assets or application observer calls, isolating the failure from the shipped dashboard. Tracing also identified observer initialization in the browser automation layer, where further investigation belongs.

The existing recovery stage guards now also populate the Cluster dashboard with process-local attempts, peer timings, outcomes and overlapping work, using the same renderer in the docs demo. Retention is bounded to 32 attempts and 256 stages per attempt, with omission counts and restart scope visible; failure detection and client reconnect remain separate measurements (dashboard details).

Zero-grace eager mode monitors idle Raft connections on the active controller and immediately verifies explicit connection loss; successful RPCs clear suspicion and silent failures retain heartbeat expiry. Four same-host, three-replica SATA process-kill screens recorded suspicion 26–34 ms after injection, with majority-durable history checks and old-owner rejoin passing. Recovery remains separately fenced and validated; the default policy stays disabled with a one-second grace (settings and scope).

A copying target reads one page ahead while appending the current page; both operations finish before progress is accepted, and installation still waits for every source reader. A same-host SATA release ABBA test with a missing 64 MiB suffix measured copying at 0.95–1.13 seconds sequentially and 0.75–0.76 seconds with prefetch; complete attempts fell from 1.85–2.08 to 1.70 seconds. The tradeoff is one additional page of at most 16 MiB per active copying target, with concurrency capped at two.

Each of at most two replica pipelines now starts inspection immediately after its own seal succeeds. Source selection waits for all collected evidence, and every source read finishes before installation can replace storage. Repeated recovery, checkpoint, learner-restart and three-replica release gates cover those barriers.

Reusing compatible retained message segments

Section titled “Reusing compatible retained message segments”

Recovery targets can reuse their exact sealed payload range after full CRC and digest checks, and installation shares the completed stage’s closed segments with a private writable tail. The 100k-message, three-replica SATA readiness gate improved from a fresh 23.99-second baseline to 9.31–10.34 seconds; the 14.48-second payload-transfer phase disappeared for matching replicas. Repair privatizes shared files before mutation, and the new manifest format fences older binaries (storage and recovery boundaries).

Inspection now reads each frozen log once per traversal, checks record CRCs and the complete sealed digest, and creates evidence only after full receiver verification. Bounded sessions preserve authorization, cancellation ownership and deadlines; strict target-copy reads remain unchanged. Single-run three-replica SATA readiness improved from 50.72 to 24.11 seconds for 100k outstanding 1 KiB messages and from 46.87 to 18.97 seconds for 100k settled plus one outstanding; retained-data copying was the main cost after this increment (internals).

Reusing authenticated inspection connections

Section titled “Reusing authenticated inspection connections”

Recovery inspection reuses one authenticated connection per source within a bounded inspection; each page still receives fresh authorization and identity validation. Requests own their socket while in flight, so timeout, cancellation and invalid replies discard it before a retry. The 100k-backlog SATA gate passed at 42.17 seconds after artifact reuse measured 42.50 seconds, a difference too small to establish a speedup from these single runs.

The recovery driver keeps the selected replay artifact through plan persistence and verifies it against the committed plan before use, avoiding a second reconstruction in the same attempt. Completed target snapshots remain preferred, and resumed attempts reconstruct the source when no local artifact exists. The 100k-backlog SATA gate passed at 42.50 seconds to readiness, compared with 50.72 seconds before this change; repeated recovery and restart tests also passed.

A subscriber connected only to its queue owner stayed pending after that owner died because topology advertised no surviving address. All five clients now accept fallback discovery addresses and retry temporary recovery responses during supervised reattachment. A three-replica SATA warm-traffic test delivered fresh probes on the original subscription 8.13 seconds after SIGKILL, with all 16,959 confirmed IDs preserved and replicas converged after restart.

Bounded recovery pages and buffered history scans

Section titled “Bounded recovery pages and buffered history scans”

Recovery inspection now uses the existing 4096-record / 16 MiB page limits, and Keratin buffers frozen reads while reusing bounded record scratch space. On the three-replica SATA acceptance host, 10k-backlog readiness fell from 15.12–15.22 s to 4.29 s; previously timing-out 100k backlog and settled-history cases completed in 50.72 s and 46.87 s with full ID, durable-write and restart/convergence checks. Full-history verification and deadlines remain unchanged; subsequent artifact, connection and sequential-inspection improvements are recorded above.

Raft sockets after cancellation and peer restart

Section titled “Raft sockets after cancellation and peer restart”

A stopped embedded Raft core could keep answering fatal errors on a cached socket, preventing retries from reaching its replacement. Requests now own their socket until a complete response arrives; cancellation, failed exchanges and fatal replies retire it. Two deterministic regressions and the repeated learner-restart acceptance test cover the repair (Ganglion 7e96cf0).

Small recovery operations repeatedly spent around 170 ms waiting on quorum-backed metadata authorization. Enabling TCP_NODELAY on Ganglion’s accepted/dialed TCP connections and forwarded writes reduced the measured recovery attempt from 6.75–6.96 s to 0.81–1.22 s across two controls and three changed single-host, three-replica SATA runs with one confirmed 1 KiB message. Durability and recovery authority checks are unchanged; larger histories, loaded clusters and pending-subscription recovery remain separate acceptance work.

Storage preallocation can now be changed from the serving node’s dashboard, with durable overrides and revision checks. Logs sample the shared policy at segment creation/reopen, keep existing segments unchanged, and report pending adoption or filesystem allocation fallback. Tests cover concurrent edits, caller cancellation, restart/reset, failed persistence, node isolation, rollover and recovery staging.

Saving an unrelated runtime setting rebuilt an incomplete document, resetting omitted connection, replication and Plexus stream settings to deserialization defaults. The form now edits a copy of the full loaded document, exposes every current runtime field and retains the version check; regression tests cover preservation, reloads, optional values, locks and duration conversion. A template coverage test compares controls against the serialized Rust settings model so newly added fields cannot silently miss the dashboard.

A metadata election could temporarily leave only one broker marked live, causing majority placement to propose a lower confirmation threshold and then a second recovery as heartbeats returned. The controller now holds that proposal until it can preserve the existing write requirement; confirmed process-loss checks exercise recovery onto two survivors. The fix and regression are in 5c6557d.

A recovered or newly admitted follower could start its ordinary replication worker at zero despite having installed a checkpoint beyond the retained event head. Accepted-history workers now resume from the admitted storage’s complete, idle applied boundary, with current identity rechecked before adoption. A learner regression confirms a new publication using that follower after checkpoint admission; see 5c6557d.

The active metadata controller can now use repeated explicit Raft transport failures and a reconnect grace to begin recovery before heartbeat expiry. The policy is off by default and runtime configurable; successful contact or a fresh heartbeat resets suspicion, while recovery proof still gates serving. Process-kill, brief-pause and stale-owner fencing checks cover the initial path; packet-level partitions and broader disruption measurements remain in the failover plan.

Message and event log staging can now grow on demand, shed empty capacity and release idle allocations through the existing writer loop. Broker measurements found comparable latency and workload-dependent CPU/RSS effects; capacity tracing and forced collection on storage writer threads linked much of the idle RSS excess after bursts to allocator retention after staging capacity was released. Fibril enables the policy by default, with adaptive_staging = false as the retained-buffer opt-out; configuration and retention details are in configuration and Keratin’s experiments/ADAPTIVE_STAGING.md.

Follower transition before storage admission

Section titled “Follower transition before storage admission”

Initial follower setup could fail while history admission was pending, while the routing cache made subsequent reconciliation treat the role change as complete. The broker now retains that rejected transition’s predecessor and retries against current metadata after local admission; removed assignments cancel the pending work. A regression holds metadata unchanged across admission, and a second checks assignment withdrawal.

Authenticated tests now race learner admission against recovery and verify the resulting witness requirement at the unchanged write threshold. An owner whose metadata runtime has stopped cannot confirm further writes after its eligible follower is sealed; repeated learner broker/metadata restarts also preserve incomplete checkpoint backfill and reject stale receipts. These tests complement native process-kill coverage and leave packet-level partitions and whole-process interruption during transfer as separate acceptance work.

An empty replication read panicked an OpenRaft 0.9.24 background task, sometimes while the surrounding test still passed. The 0.9.25 heartbeat fallback also fails at index zero; Ganglion carries a narrow retry patch that reports no progress for unread data and retires old streams on step-down, with subprocess checks for panics, eventual replication and cancellation. Dependency provenance and removal conditions are recorded in Ganglion’s vendor/openraft/PATCHES.md; the scheduling cause of the original intermittent empty read remains unconfirmed.

Assigned replicas excluded from an activated quorum can now copy and backfill the current history while the owner continues confirming writes. A durable, fully applied cut and complete payload dependencies gate exact-instance admission; unchanged sessions keep their progress and the next recovery uses the enlarged witness set. Restart and process-kill tests cover checkpoint backfill and publication of a new local generation with old evidence retained.

Storage normally materialized a queue with an owner role before coordination adjusted it. Learner generations now retain a follower marker across eviction and restart, preventing background timer maintenance from treating a returning copy as an owner during catch-up. Explicit promotion remains a separate operation.

A persisted recovery could stall indefinitely after its proposed owner disappeared, even with sufficient surviving evidence and completed transfers. The controller now selects another available member of the fixed replica set without changing seals, staged data or quorum requirements; activation binds the selected owner through fresh consensus. Three-node tests cover a stopped candidate, preserved staged data, stale-candidate rejection and resumed replicated confirms.

A checkpoint reset could return before older fsync completions were consumed, allowing one to restore a stale durable frontier after the cut. The writer now drains those completions before replacing files and resetting state. A queued-fsync regression reproduced a frontier of one after reset to zero and now verifies the reset remains at zero.

An empty fsync job used the same inclusive offset as a job covering message zero. When its completion updated the manifest, it could invent a one-record tail; completion after a later append could also overstate durability. Fsync jobs and the log’s internal durable boundary now use exclusive counts, with a deterministic regression for both sequences (Keratin d8902a7).

Automatic queue enrollment and interrupted preparation

Section titled “Automatic queue enrollment and interrupted preparation”

Fresh Unix cluster queues now prepare and activate their initial write quorum automatically. A guarded renewal preserves origin IDs across owner replacement before activation, while durable storage intent makes interrupted pristine preparation resumable; activated histories stay on the normal recovery path. Tests cover preparation SIGKILL boundaries, full broker/metadata restarts, and independent follower admission after owner shutdown.

Recovery installation and quorum activation

Section titled “Recovery installation and quorum activation”

Recovery builds a verified storage generation and atomically publishes its route while retaining the old sealed source. Exact process/storage receipts gate new-quorum activation and local admission; bounded authenticated transfer can continue from a completed stage after source loss. Linux crash tests, resumed majority confirmations and consecutive recoveries cover the installed path; ordinary enrollment remains a rollout gate.

The first in-process installation attempt registered the replacement as a new queue, so the actor skipped its saved snapshot despite retaining the payload log. Installation now publishes an existing-storage registry entry under lifecycle serialization. An end-to-end regression performs two recoveries, resumes publishing and verifies delivery of every confirmed message.

Selected queue state and payloads can now be staged in a separate native log while the old sealed source stays readable. Exact page retries, immutable metadata and sequential disk verification preserve the selected baseline through restart; four SIGKILL boundaries and authenticated transfer are covered. A fresh consensus plan check authorizes staging, with replacement and writer admission remaining separate gates.

A verified source selection can now be fixed in consensus before replacement storage work begins. The plan binds the exact state, witnesses, continuation offsets and new history/session IDs; retries retain the first decision while replicas remain fenced. Regressions cover zero boundaries, substituted evidence, stale transitions and consensus retries across three metadata nodes.

Sealed queue replay can now produce deterministic snapshot artifacts that release old leases while preserving logical state and verified live-payload identity. Explicit source selection requires accepted origin authority, the old witness threshold, complete source coverage and consistent comparisons; it leaves installation and writer admission closed. Tests cover partial replay, divergence, crossed tails, stale transitions, older timer semantics and a single intersecting witness over authenticated TCP.

An owner could activate a delayed publish locally, deliver it and record a NACK, while restart or follower replay missed the activation and lost the retry increment. Due timers now move to ready through a bounded, ordered event that consumes their heap entries; an elapsed deadline at broker admission becomes an ordinary enqueue. Regressions cover restart, follower application, different checkpoint starts, idle polls and retry deadlines; delayed retry also removes follower-ready state before waiting.

Sealed history authority and process replacement

Section titled “Sealed history authority and process replacement”

Version-two seals retain the original durable storage receipt, allowing recovery to check the exact accepted history and replica instance after restart. Recovery witness thresholds use the eligible write set while confirmations retain the configured write threshold; a changed registered process requests recovery even when placement is unchanged. Three-node authenticated TCP tests cover sealed receipt admission after majority-confirmed replication.

The recovery scanner counted bytes from the next read position instead of the start of a buffered partial record, overstating the valid end when records crossed read chunks. Keratin d45757f corrects the boundary; tests cover chunk-spanning records, padding and append after repeated reopen. Bound histories also refuse local repair that could discard records and reuse offsets under an unchanged writer session.

Exact initial activation and live replication

Section titled “Exact initial activation and live replication”

Explicit activation now admits only the prepared provider and storage instances, while live replication carries the accepted history through pull and streaming paths. A three-node TCP regression verifies majority confirmation, stale-identity rejection and replacement-storage fencing; automatic recovery readmission remains in the failover plan.

Replication stream rejection and cancellation

Section titled “Replication stream rejection and cancellation”

A rejected stream-start request previously left the follower waiting because the transport ignored ordinary error responses. The reader now reports that rejection, and both transport tasks abort when their parent is dropped; TCP tests verify prompt exit and socket closure.

Remote preparation now verifies the exact target and committed decision through consensus before creating a non-writable storage baseline. Requests use fresh node authentication, bounded frames and deadlines; identical calls share one broker operation that survives caller cancellation. Real three-node TCP tests cover majority receipt collection, retries, stale decisions, wrong targets and reopened storage (initial preparation).

Creation-time enrollment and durable preparation evidence

Section titled “Creation-time enrollment and durable preparation evidence”

Explicit enrollment now blocks ordinary serving from catalogue creation while preserving replica placement, and preparation requires that enrollment before it can certify an empty origin. Deletion retains a retired marker so the remaining assignment cannot accidentally regain legacy serving; exact quorum receipts persist across metadata restart without activating writers. Regressions cover placement stability, retirement/recreation, conflicting receipts and owner replacement (initial preparation).

Consensus preparation without writer admission

Section titled “Consensus preparation without writer admission”

Initial-history decisions persist fixed history/session IDs bound to the owner instance and exact assignment. Replica preparation leaves storage non-writable; receipt collection deduplicates replies, requires the owner and blocks contradictory replica identities. A durable metadata restart regression covers rejection of the old owner grant; automatic enrollment and recovery readmission remain pending (initial preparation).

An explicit initialization primitive now persists incarnation, history and writer-session IDs before admitting pristine storage. Reopening that store blocks ordinary access even at the same assignment epoch, while recovery sealing remains available; conflicting IDs and ordinary checkpoint replacement cannot relabel it. Cancellation, corruption and Linux SIGKILL regressions cover the boundary, with automatic enrollment and recovery readmission still pending (storage binding).

New queue and stream declarations now register their incarnation ID atomically, so concurrent declarations retain one identity and recreation after retirement receives a new one. Conditional deletion protects identified replacements from a delayed old delete, and recovery seal authorization rejects changed or missing incarnation metadata. Consensus snapshot restart and transition tests cover these cases; storage lineage and writer admission remain in the failover plan.

Queue events now finish actor application in log order, allowing snapshots to capture state with its actual exclusive event boundary. Interrupted application cannot be cleared by changing roles, and verified repeated follower batches skip already-applied events so NACK retries do not increment twice. Tests cover event zero, partial batches, cancelled captures, failed writes and restart; checkpoint internals describes the boundary.

Explicit sealed inspection now verifies each snapshot and replays its event suffix to a common target, then compares live payload identities separately from state. Owner-local leases have a separate normalized digest; retry counts, delays, TTL and DLQ state retain their meaning. Authenticated TCP tests cover unequal checkpoint starts, while source authority and automatic activation remain tracked in the failover plan.

Checkpoint rejection and cancelled captures

Section titled “Checkpoint rejection and cancelled captures”

Malformed queue checkpoints could panic on invalid ranges or a truncated custom-DLQ group, and a rejected load could partially replace existing actor state. Decoding now validates into isolated state before replacement, and a pause acquired for checkpoint export is released even when its caller is cancelled while draining active work. Owner capture also serializes with role changes and recovery sealing; malformed-field, cancellation and real broker checkpoint tests cover these paths.

Restartable suffix repair preserves valid events

Section titled “Restartable suffix repair preserves valid events”

The recovery helper used a whole-log checkpoint reset when it intended to remove only a bad suffix, leaving the valid prefix in memory but removing its durable records. Suffix repair now journals the cut, preserves earlier records and resumes before normal log opening; repeated restarts, I/O faults and six SIGKILL boundaries cover the fix. A dangling enqueue also no longer masks later corruption or unexplained missing settlement payloads.

Explicit inspection can reconstruct fully retained queue histories at a common exclusive event frontier and hash their complete canonical state. Payload and input-history digests remain separate, while operation budgets and bounded CPU workers keep this work outside live actor scheduling. Equal state supports comparison but still requires resource lineage and checkpoint authority (recovery sealing).

Sealed-history comparison and dependency diagnostics

Section titled “Sealed-history comparison and dependency diagnostics”

Explicit pair inspection now compares shared offsets across differently retained histories and verifies the complete transferred log digests, retaining only a page of unmatched record IDs. Whole-event reference checks expose incomplete payload batches and preserve explicit replay/checkpoint gaps; timeout, source loss and exhausted budgets discard partial results. Matching overlap remains subject to common-origin and state proof before source selection (recovery sealing).

Recovery can now read bounded pages from live or restarted sealed replicas without opening queue actors, repairing files or lifting the seal. Every page verifies the complete retained logs and snapshot, while authentication, transition checks and cancellation-safe locks preserve the recovery boundary. Full rescans are deliberately confined to explicit recovery calls; compatible-history proof and efficient bulk installation remain pending (recovery sealing).

Explicit seal calls now use fresh authenticated connections with whole-call deadlines, and witness admission binds each reply to the contacted old replica and exact pending transition. Duplicate retries count once, a smaller proposed configuration cannot lower the old threshold, and contradictory evidence from one sealed replica blocks the collection. Threshold completion remains AwaitingHistoryValidation; automatic dispatch, compatible-history proof and activation are still pending.

Internal replication handlers accepted ordinary authenticated clients, and unauthenticated clients when authentication was disabled; a real TCP regression reproduced a read of retained payloads. All replication and recovery controls now require the node principal on the physical connection before decoding or changing state. Tests cover refused reads/writes, every control opcode, resumed sessions, authenticated catch-up/checkpoints, TLS, streaming confirms and reauthentication after transport loss.

Authenticated recovery requests and retained history

Section titled “Authenticated recovery requests and retained history”

Explicit seal requests now require node authentication on the current transport and consensus authorization of the exact pending transition, with concurrent identical requests sharing one operation. Storage persists a fingerprint of the frozen records and snapshot, independent of fencing epochs and replica-local append times; retry reads bypass the cache so disk changes cannot hide behind cached contents. Authentication, cancellation, stale-transition and changed-record regressions cover this increment; automatic dispatch and compatible-history selection remain pending (recovery sealing).

Recoverable follower checkpoint installation

Section titled “Recoverable follower checkpoint installation”

A crash between the two log resets could leave old queue state referring to removed message bodies. A durable installation journal now resumes replacement before ordinary recovery, and a completion receipt prevents retries from erasing later backfill; snapshot fencing rejects stale writes after installation or eviction. Linux fault, cancellation and six-boundary SIGKILL tests cover Keratin 76f1469; coordinated history selection and activation remain pending.

Manifest storage removed the old file before renaming its replacement, allowing a failed rename or crash to lose the persisted epoch. Direct replacement now preserves the previous manifest on rename failure and synchronizes both affected directories on Unix; an injected-failure regression covers Keratin 3435e3d. Windows metadata durability still requires separate implementation and testing.

Replication wakeups independent of confirmation waits

Section titled “Replication wakeups independent of confirmation waits”

Each locally durable queue batch now wakes replication after its payload/enqueue dependency is registered. A previous publication waiting for remote confirmation no longer delays that notification; the notification runs once per batch and skips a cancelled owner runtime. A regression parks a follower beyond batch A, leaves A unconfirmed, and verifies that durable batch B wakes the follower while both confirmations still wait for replica progress.

Confirmed history during follower promotion — unresolved

Section titled “Confirmed history during follower promotion — unresolved”

A real-storage diagnostic confirmed a batch on owner A and follower B, then successfully promoted empty follower C after A stopped. Placement uses advisory heartbeat tails, while promotion establishes local completeness; neither establishes that the selected candidate contains every previously confirmed batch. Preserving that history across assignment changes is a recovery gate for clustered HA and for replication before owner durability.

A persisted seal keeps queue and stream evidence intact across restart, eviction and interrupted fencing of the two logs. The storage primitive blocks ordinary reopening and cleanup and supports identical-request retries; it remains disconnected from automatic failover until history validation, installation and activation are implemented. Failure tests cover caller cancellation, a failed writer, partial fencing, damaged markers and startup suffix repair in Keratin 617e1e8.

An epoch update changed the in-memory manifest before its disk write succeeded, allowing a retry to return success without retrying the failed write. The update now publishes the new in-memory epoch only after persistence succeeds; a filesystem fault test verifies repeated failure followed by successful durable retry in Keratin 71b2f56.

Pending recovery before replicated assignment replacement

Section titled “Pending recovery before replicated assignment replacement”

The controller retains the previous assignment and persists a proposed replacement when ownership, replica membership or policy changes affect replicated confirmation. The request survives metadata restart and is exposed in controller status, preventing followers from switching to an unproven source. Fresh sealing, recovery and activation remain pending, so this first barrier pauses replicated failover while preserving its recovery evidence.

The log reader now accepts a cache hit that contains every record up to the captured durable frontier, even when the requested batch is larger. Incomplete coverage, decode failures and offset discontinuities fall back to the file reader; rollover, eviction and reopen parity are covered by regression tests. The initial three-node SATA/NVMe screen found overlapping throughput and latency ranges, so this change carries no end-to-end speedup claim; see Keratin dd7943e.

Queue replication dependencies and follower persistence

Section titled “Queue replication dependencies and follower persistence”

Queue confirmation and delivery visibility require the same counted follower to cover the payload batch and its exact enqueue-event frontier, including the whole payload group referenced by an indivisible enqueue record. Progress is fenced by assignment epoch and transport session, and replacing an assignment clears its previous proof. Follower message and event writes overlap; both completions drain before state application, while an interrupted or failed apply blocks promotion until recovery or checkpoint resync.

Promotion and direct owner activation check the payload frontier required by ready, delayed, inflight, settled and dead-letter state; recovery keeps a checkpoint with missing payload backfill in the follower role. Client admission cannot perform the watcher’s follower-to-owner transition. Recovery also replays retained event zero when a legacy snapshot’s inclusive zero could denote an empty checkpoint; interrupted replacement of both logs and checkpoint state is covered by the recoverable installation journal described above.

Consecutive batches could reach queue state in reverse order when the earlier durability continuation was delayed. Keratin now reserves an application turn under the append-order lock, allowing persistence to overlap while preserving state submission order; failed or abandoned turns fail the chain closed until recovery. A deterministic overtaking test and the Stroma regression suite cover Keratin 0944dd6.

An ACK from a different consumer could remove the rightful consumer’s delivery tag, and ignored or duplicate requests could leave settlement-drain accounting nonzero. Ownership validation and removal now share the same map lock, while ignored requests release their accounting without changing consumer credit. The ordinary-broker regression covers wrong-consumer ACK/NACK/reject requests, duplicate ACKs and graceful drain in Fibril d64577d.

The Rust confirmation handle implements Future, permitting direct polling without a per-confirmation wrapper. The steady workload offers --ordered-confirmations for FIFO polling on an ordered single-partition workload; the default and cross-broker harness retain concurrent completion collection. This reduces benchmark bookkeeping and does not change broker confirmation guarantees or optimize the other SDKs’ transports.

The shared Rust harness records message identity, offered load, completion latency and resource samples, with isolated single-node Docker provisioning. Existing-server modes also cover pipelined request/reply and multiple connections; portable cluster/fault orchestration remains pending. Raw competitor measurements remain separate from public capacity claims.

Transparent huge pages — deployment option

Section titled “Transparent huge pages — deployment option”

Allocation tracing found that retained log staging reserves two 16 MiB write buffers and two 256 KiB index buffers per materialized queue; their resident cost can grow substantially under THP even when lightly used. Mirrored local tests of the allocator startup option found lower RSS with THP disabled across paced, saturated and three-node request/reply workloads, with workload-dependent CPU and latency costs. The option remains a candidate for memory-sensitive deployments; production defaults are unchanged, and memory pressure and large active working sets require further validation.

Storage writer channel sizing — configurable, defaults unchanged

Section titled “Storage writer channel sizing — configurable, defaults unchanged”

Crossbeam writer and notification channels allocate their full slot arrays when each log opens. In a local 32-queue probe, reducing their capacities from 8,192 to 64 slots reduced post-declaration broker RSS from about 347 MiB to 248 MiB; two short 1 KiB publish/delivery runs per storage type kept throughput within roughly 1% of the default mean on tmpfs and SATA. The startup writer buffer factor exposes this tradeoff while retaining the existing default; shrinking the separate async command pipelines caused substantial throughput losses in the same investigation.

Targeted reads for queue routing — adopted

Section titled “Targeted reads for queue routing — adopted”

Queue routing reads the committed metadata it needs through a targeted path, reducing repeated construction and traversal of the cluster view. Native replication and live-repartition checks covered routing equivalence and stale partition versions; the implementation is in fibril 94cc1e8.

The client buffers individual ACK frames through a bounded command drain and flushes the tail immediately when that drain ends, reducing socket writes while preserving frame order and settlement history. In localhost tmpfs delivery-only runs with 128-byte payloads, four readers, prefetch 16384 and 20 million preloaded messages, mean throughput rose from about 529k to 1.03M messages/s across two runs per variant. fibril e713600 includes ordering, boundary, small-window and failed-flush coverage.

Python and TypeScript publish buffering — adopted

Section titled “Python and TypeScript publish buffering — adopted”

Pipelined confirmed publishes share bounded socket writes while retaining their individual frames and confirmation results. Both clients expose count, byte and time limits plus an immediate-write option; Python’s ordinary awaited publish keeps its immediate-send path. See fibril e713600.

The broker flushes buffered output when both outgoing queues drain, retaining its existing busy-batch bounds and durability gates. In same-host SATA ext4 runs with 1 KiB messages and one outstanding confirmation, mean per-run median confirmation latency fell from about 6.0 ms to 1.4 ms; the higher-concurrency Rust screen showed roughly 3–6% more broker CPU per message and small throughput changes. fibril dc878c9 includes writer and replication-gate tests.

Larger delivery cache — experiment retained, defaults unchanged

Section titled “Larger delivery cache — experiment retained, defaults unchanged”

Increasing the log tail cache from 64 MiB to 4 GiB raised cache hits to 100% in the measured tmpfs window, while delivery stayed around 419–422k messages/s and broker RSS rose from about 0.44 to 4.87 GiB. Faster reads shifted waiting toward consumer-channel sends, so the experiment did not justify a larger default.

Settlement-history removal — not adopted

Section titled “Settlement-history removal — not adopted”

Isolated ACK experiments reduced queue-state work by avoiding full settled history, but the broker comparisons did not establish a useful delivery gain. Full range-compressed settlement history remains in place; later component isolation identified client ACK writes as a stronger delivery limit.

Benchmark methods and earlier detailed experiments are in the optimization log.

Remaining follower kept its previous owner — fixed

Section titled “Remaining follower kept its previous owner — fixed”

After an owner change, a node that stayed a follower could receive a no-op assignment transition and keep pulling from the previous owner. The planner now refreshes that follower when its source or epoch changes, and the worker retains its continuation offsets while retargeting: fibril 3fc4253.

Retired partitions reopened during cleanup — fixed

Section titled “Retired partitions reopened during cleanup — fixed”

Late cleanup or a stale reconnect could recreate a partition after shrink had deleted it. Workers now stop before deletion, storage cleanup avoids creating missing partitions, and a local retirement fence rejects stale access until a fresh assignment admits the partition again: fibril 8edb75f and keratin 470d9e4.

Confirmation waiter missed the final progress update — fixed

Section titled “Confirmation waiter missed the final progress update — fixed”

A follower could report sufficient durable progress between the confirmation waiter’s progress check and notification registration, leaving the waiter asleep until timeout. Registration now precedes the check, with a paused-time test that forces the old race: fibril 7f40f92.

Independent declarations produced conflicting event histories — fixed

Section titled “Independent declarations produced conflicting event histories — fixed”

A broker receiving a declaration could append its own event before placement, leaving a different record at an offset later supplied by the assigned owner. Declarations now carry settings through coordination and only assigned owners append them; followers receive those events through replication: fibril 990765e and keratin 0b75fd5.

Checkpoint reset discarded source epochs — fixed

Section titled “Checkpoint reset discarded source epochs — fixed”

Checkpoint installation dropped the source epochs before destructive log resets, allowing a stale checkpoint to pass an already-advanced local fence. Both epochs now reach storage, are validated before reset submission, and are checked again by each writer in command order: fibril dc878c9 and keratin 11ef6a9.

Replication conflicts lacked preceding context — diagnostics added

Section titled “Replication conflicts lacked preceding context — diagnostics added”

Overlap reports now include bounded control history, offsets and effective record identities, with payloads and headers excluded. An isolated reintroduction of the declaration bug produced a report showing the local declaration and the incoming enqueue at event offset zero; see the diagnostic format and keratin 11ef6a9.

Interrupted checkpoints and early promotion — open

Section titled “Interrupted checkpoints and early promotion — open”

Isolated storage tests reproduced reopening after a message-log reset with snapshot references to missing bodies, and local-tail promotion before checkpoint-required message backfill completed. The recovery plan tracks consistent installation and promotion eligibility through restart; repair implementation and broker-level process-kill validation remain pending. Evidence and test scenarios are recorded in fibril b14a91a.

Writer/fsync-worker deadlock under saturated storage — fixed

Section titled “Writer/fsync-worker deadlock under saturated storage — fixed”

Scheduled commits could exceed the inflight limit and block the writer while the fsync worker was blocked returning completions, producing a circular wait. Every fsync handoff now waits for capacity by draining completions first; the stress reproduction and fix are recorded in keratin 3cb1218.

Recovery wakeups, bounded concurrency and checkpoint cadence

Section titled “Recovery wakeups, bounded concurrency and checkpoint cadence”

Recovery now wakes on committed metadata, and completed local admission wakes assignment retries so follower durability reporting can resume promptly. At most two replicas are prepared concurrently, with all source reads completed before installation; the settled-history SATA screen improved fresh delivery from 4.15 s to 2.00–2.02 s. Checkpoint event/append-byte triggers complement the periodic policy, with shared dashboard age and retention diagnostics; see the current recovery screen.

The receive loop matched only Some(frame), disabling its read branch on EOF while heartbeat and command branches stayed live. Explicit EOF handling now closes the engine and fails pending requests immediately; read-side transport errors remain retryable, and all five SDKs pass a pending-topology EOF check. In the single-host SATA warm-traffic screen, the original subscriber’s fresh delivery moved from roughly 5 s to 1.06 s with the same immediate-mode server.