Skip to content

Project status

Fibril is pre-alpha infrastructure. This page summarizes the current main-branch implementation and its operational limits. Cluster operation remains experimental.

For a more detailed checklist of what is wired and what conditions apply, see implemented surface.

FeatureStatusNotes
Durable queuesAvailableDurable message storage, queue state, snapshots, and replay
Publish and subscribeAvailableCustom TCP protocol with Rust, TypeScript, Python, Go, and C# clients
Explicit settlementAvailableACK, fail, immediate retry, and delayed retry paths
LeasingAvailableExpired leases can return to ready
BackpressureAvailablePull-based delivery and bounded prefetch
Delayed publishAvailableBroker path and Rust, TypeScript, Python, Go, and C# client methods are wired
Message TTLAvailablePer-message or per-queue default; expired messages drop or dead-letter. Not queue expiration
Dead letteringAvailableGlobal and per-queue policy are configurable, replay tooling is still early
Sparse queuesAvailableLazy loading, idle cleanup and default adaptive storage buffers reduce idle allocations; memory under busy multi-queue loads remains a measurement focus
Message inspectionAvailableBrowse active queue messages from admin tooling, with optional settled offsets and payload previews
Partitioned queuesAvailableDeclared queues can have multiple partitions, with client-side key routing and transparent fan-in
Plexus streamsAvailableFan-out channel type: every subscriber receives every record, durable named cursors, per-stream durability tiers, partitioning with client-side fan-in, and header filters. Rust, TypeScript, Python, Go, and C# clients
Wildcard subscribeAvailableOpt-in client.routing() subscribes to every queue or stream whose topic matches a *-glob and auto-attaches channels that start matching later, driven by a live cluster catalogue. Client-side only. Rust, TypeScript, Python, Go, and C# clients
Partition ownershipExperimentalEmbedded coordination can assign queue ownership, fence stale owners, and redirect clients to the current owner
Queue replicationExperimentalAuthenticated pull/streaming replication, ordered event application and replica-durable confirms bound to exact history and replica identities
Automatic queue recoveryExperimentalEnrolled Unix queues fence ownership changes, verify surviving history and activate an installed quorum. Compatible retained payloads can be reused after full verification; contradictory or unsupported histories remain fenced
Background replica catch-upExperimentalExcluded assigned replicas catch up while the owner continues serving; durable/applied state and complete payloads gate admission
Eager failure detectionExperimental, opt-inRuntime policy reacts to repeated explicit peer transport failures; zero grace adds idle-disconnect monitoring and immediate reconnect verification. Recovery proof still gates service; silent failures use heartbeat expiry
Stream replicationExperimentalDurable-tier streams replicate record and cursor logs to followers with replica-durable confirms and caught-up failover, placed and owned through embedded coordination. Express tiers stay owner-only
Live repartitioningExperimentalGrow or shrink a queue’s partition count in coordinated mode, from the admin topology page
Reconnect and restart resumeAvailableTyped subscription closure, broker-local persisted resume sessions and stale-delivery settlement across all five clients. Explicit fallback discovery endpoints and temporary-recovery retries support owner-loss reattachment
Checkpoint installationExperimental, UnixJournaled installation resumes across interruption; required payload backfill blocks promotion, including after restart. Linux fault and process-kill tests cover the path
Agreed recovery checkpointsExperimental, opt-in on UnixDurable same-cut capsules and all-admitted replica receipts publish through consensus. Retention preserves the accepted base and later suffix; recovery verifies retained bytes and keeps the existing quorum/fencing requirements. Event/append-byte triggers and shared dashboard diagnostics complement periodic cadence. Live backlog still requires payload validation
Speculative queue deliveryExperimental prototype, unmergedLocal-queue results remain separate from main. Correctness reconciliation and replicated speculation are pending; Plexus stream tiers have their own existing contract
Recovery quarantineAvailableA damaged queue log is detected on recovery and isolated per the recovery.on_mismatch policy, with operator repair
TLS in transitAvailableThe broker listener serves TLS from operator PEMs or generated per-deployment material, mismatches are named in both directions, the clients connect with CA-file, fingerprint-pin, or OS-roots trust, the dashboard serves HTTPS from the same material, and first-boot setup mode offers generate/supply/skip before the broker ever serves. Inter-broker replication and coordination traffic is encrypted too (tls.inter_broker, shared-CA lane for generated material), the serving certificate rotates live via fibrilctl admin reload-tls, and tls.client_auth turns client certificates into credentials (a verified identity authenticates as the matching user with no password, require closes the handshake to certless peers)
Broker authenticationAvailableArgon2 user store seeded from config, managed from the dashboard and fibrilctl, replicated across the cluster. Built-in fibril/fibril credentials work from loopback only. Node-to-node connections authenticate with a cluster shared secret, never a user account
Admin dashboard and demoAvailableShared production views cover queues, messages, topology, resources and settings; the docs include a read-only simulated demo
Runtime and node-local settingsAvailable, scopedRuntime fields have version-checked dashboard controls. Node-local segment preallocation persists separately and reports adoption at segment rollover/reopen; other storage tuning remains startup configuration
Recovery stage diagnosticsAvailableBounded process-local worker timelines in the shared Cluster dashboard and demo; stage outcomes, peers and overlap. Timing excludes failure detection and client reconnect; history resets on restart
Prometheus metricsAvailableGET /metrics on the admin listener behind the same auth and HTTPS as the dashboard: node-level aggregates always, per-channel series from materialized channels gated by admin.metrics_per_channel
Exclusive consumer groupsPartialRust, TypeScript, Python, Go, and C# client opt-in for one active consumer per partition, with sticky assignment and cross-broker coordinator wiring
TransactionsOut of scopeNot planned. Transactional publish/consume workflows are intentionally excluded

Valid histories can exceed current automatic recovery budgets and remain fenced. Large-history recovery availability is a high-priority gap.

Cluster recovery still needs broader partition, power-loss and sustained-load acceptance, authoritative stream-history recovery and safe reclamation of retained recovery data. There is no history-independent failover latency guarantee; see the failover plan for the remaining gates.

Informal internal measurements on a Ryzen 5950X system have observed roughly 250k+ messages/sec ingress and egress with 1KB payloads on the durable path. Plexus stream fan-out reaches roughly 1.5M delivered records/sec across sixteen readers on a single partition, and one partition fans out to hundreds of readers when delivery throughput is not the bottleneck.

These numbers are architecture sanity checks, not a rigorous benchmark suite. Hardware, durability settings, batching, queue depth, workload, and storage behavior all matter. See benchmarks for the fuller queue and stream tables.