Failure modes and operations
This page is the operator-facing view of what happens when things go wrong: what survives which failure, and what to do about it. It synthesizes the reliability semantics, replication, recovery quarantine, and reconnects pages into one incident-time reference.
Failure semantics are version-specific. This page describes the current main branch; consult the matching documentation for a deployed version. Single-node durability is a real, tested part of the broker. The cluster path (replication and failover) is experimental and needs more failure testing before it is production-grade high availability.
What survives what
Section titled “What survives what”The delivery foundation is at-least-once: a confirmed publish is durable per the tier below, and a message is redelivered until it is settled (acked or dead-lettered). Fibril does not claim exactly-once.
Two distinct failures matter, and they are not the same:
- Process restart - the broker process stops and starts again on the same machine, with its data directory intact.
- Node loss - the machine (or its disk) is gone, so the local data is not coming back.
| Tier | Survives process restart | Survives node loss | Notes |
|---|---|---|---|
Queue, single owner (local_durable) | Yes (fsync per durable write) | No | The partition is unavailable while the node is down. This is the default and the only mode for a standalone broker. |
Queue, replicated (replica_durable / majority_durable) | Yes | Copies survive; automatic failover paused | Confirmation waits for the required replicas. The controller retains the old assignment while a pending recovery request awaits confirmed-history proof. Requires Ganglion coordination and assigned followers. |
Stream, durable owner-only | Yes (fsync before deliver/confirm) | No | Survives restart, not node loss. Records and cursors are on the one owner. |
Stream, durable replicated (stream_replication_factor >= 2) | Yes | Copies survive; automatic failover paused | Record and cursor logs replicate; replicated ownership changes share the pending-recovery barrier of queues. |
Stream, speculative | Partial | No | Delivers off the staged offset and defers the producer confirm until durable, so a confirmed record is durable, but unconfirmed records in flight at a crash can be lost. |
Stream, ephemeral | Partial | No | Lowest latency, AfterWrite (no per-record fsync). Recent records can be lost on a crash. Intended for where freshness beats durability. |
Cross-cutting guarantees that hold across a restart or a clean failover:
- Assignment fencing. Ownership changes bump an epoch. Nodes that have advanced to it reject older-epoch replication. Cluster-wide preservation of confirmed history during ownership changes remains an acceptance gate.
- Per-partition ordering is preserved; there is no global order across partitions.
- Leased-but-unacked work survives a crash. Inflight (offset, deadline) pairs are persisted, so a message that was leased to a consumer but not yet acked is not lost across a broker restart - it becomes deliverable again.
- Streams resume from the durable cursor, not a client-held offset, so a reconnecting or failed-over reader continues where its committed cursor left off.
Operator runbook
Section titled “Operator runbook”Planned restart or rolling upgrade
Section titled “Planned restart or rolling upgrade”Use drain so in-flight work is not dropped:
POST /admin/api/drain(optionally{ "grace_ms": 5000, "message": "upgrade" }). The broker pushes a going-away notice to every connected client and, in coordinated mode, marks itself draining so the controller hands its partition ownership to caught-up followers. The call returns once ownership has moved orconnection.drain_handoff_timeout_ms(default 30s) elapses, reportingowned_partitions_remainingandhandoff_completeso a restart script knows whether stopping now is gap-free.- Clients surface the notice (so apps can stop producing or finish in-flight work); ownership moves push topology updates that redirect traffic to the new owners while this broker keeps serving whatever it still owns.
- Restart the broker once the drain call returns.
Reconnect grace (on by default, connection.reconnect_grace_ms, 5s) keeps a
returning client’s session and inflight work intact across the brief break.
Drain uses the same recovery barrier as failover for replicated assignments.
Those assignments remain on the old owner while recovery is pending, so drain
can time out with handoff_complete = false. Preserve that owner’s data and
inspect the pending recovery before stopping it. A draining node receives no new
placements.
An owner broker dies (cluster)
Section titled “An owner broker dies (cluster)”For enrolled queue histories, automatic ownership replacement uses the recovery barrier until verified history and an installed write quorum are available:
- The controller persists the previous and proposed assignments, preserving the old replica set and replication source. The proposed owner cannot serve through that unapproved assignment.
- The recovery worker seals eligible survivors, verifies a complete source, transfers its state and activates the installed quorum. If its candidate becomes unavailable, the controller can choose another member of the fixed proposed set.
- Inspect
consensus.controller.pending_recoveriesin the admin topology response and retain every surviving replica’s data. These records preserve the original proposal; candidate-replacement logs identify a subsequent handoff. Returning replicas can supply missing evidence. Serving resumes only after verified activation and exact local admission. - Check the admin queues page for owner and in-sync replica status.
A local_durable single-owner partition has no follower to promote, so it is
unavailable until the node returns with its data dir. This is expected - that
mode trades node-loss survival for lower latency.
Replica-durable publishes start failing fast
Section titled “Replica-durable publishes start failing fast”A replica_durable / majority_durable confirmed publish fails fast (rather
than blocking to the timeout) when fewer replicas are in sync than
min_in_sync_replicas. This is the ISR floor protecting you from acknowledging
writes that are not adequately replicated. Check follower health and lag on the
admin queues page; the partition keeps accepting writes again once enough
followers are back in sync. Set the floor to 1 to disable it.
A damaged log is found on restart
Section titled “A damaged log is found on restart”If a power loss or hardware fault left a queue’s event log unreadable, the broker
does not blindly replay it (which could rebuild wrong state or crash the whole
process). It isolates the partition per the recovery.on_mismatch policy and
surfaces a quarantine banner. Repair truncates the partition to its last valid
point via POST /admin/api/quarantine/repair; a follower then re-fetches the
dropped suffix on its next catch-up. See
recovery quarantine.
On-disk data for a partition this node no longer owns
Section titled “On-disk data for a partition this node no longer owns”After a cold restart, a partition that was reassigned to another node while this node was down is left on disk as inert cold storage: it is not served and not materialized (serving is ownership-gated), so it cannot leak stale data. Its disk is not reclaimed automatically yet - reclaiming it is a manual step for now.
Growing or shrinking a queue’s partitions
Section titled “Growing or shrinking a queue’s partitions”Live repartitioning (Ganglion mode) is driven from the admin topology page. A grow adds partitions; a shrink drains the retiring partitions before retiring them, and the cutover is fenced on cluster-wide client adoption of the new routing, with publish version-fencing as the correctness backstop. Watch the topology page through the transition.
Tuning reconnect grace
Section titled “Tuning reconnect grace”connection.reconnect_grace_ms controls how long a dropped resumable client is
held before its subscriptions are cleaned up and its unsettled inflight is
requeued. It is on by default (5s). Lower it for faster redelivery on a genuinely
dead consumer; raise it for flakier networks; set 0 to disable and clean up
immediately on disconnect.
See also
Section titled “See also”- Reliability semantics - the formal delivery and durability guarantees.
- Replication - durability levels, in-sync replicas, and failover.
- Recovery quarantine - damaged-log handling.
- Reconnects - reconnect grace and resume.
- Configuration - the durability, replication, and connection settings referenced here.