Project status
Fibril is pre-alpha infrastructure. This page summarizes the current main-branch implementation and its operational limits. Cluster operation remains experimental.
For a more detailed checklist of what is wired and what conditions apply, see implemented surface.
| Feature | Status | Notes |
|---|---|---|
| Durable queues | Available | Durable message storage, queue state, snapshots, and replay |
| Publish and subscribe | Available | Custom TCP protocol with Rust, TypeScript, Python, Go, and C# clients |
| Explicit settlement | Available | ACK, fail, immediate retry, and delayed retry paths |
| Leasing | Available | Expired leases can return to ready |
| Backpressure | Available | Pull-based delivery and bounded prefetch |
| Delayed publish | Available | Broker path and Rust, TypeScript, Python, Go, and C# client methods are wired |
| Message TTL | Available | Per-message or per-queue default; expired messages drop or dead-letter. Not queue expiration |
| Dead lettering | Available | Global and per-queue policy are configurable, replay tooling is still early |
| Sparse queues | Available | Lazy loading, idle cleanup and default adaptive storage buffers reduce idle allocations; memory under busy multi-queue loads remains a measurement focus |
| Message inspection | Available | Browse active queue messages from admin tooling, with optional settled offsets and payload previews |
| Partitioned queues | Available | Declared queues can have multiple partitions, with client-side key routing and transparent fan-in |
| Plexus streams | Available | Fan-out channel type: every subscriber receives every record, durable named cursors, per-stream durability tiers, partitioning with client-side fan-in, and header filters. Rust, TypeScript, Python, Go, and C# clients |
| Wildcard subscribe | Available | Opt-in client.routing() subscribes to every queue or stream whose topic matches a *-glob and auto-attaches channels that start matching later, driven by a live cluster catalogue. Client-side only. Rust, TypeScript, Python, Go, and C# clients |
| Partition ownership | Experimental | Embedded coordination can assign queue ownership, fence stale owners, and redirect clients to the current owner |
| Queue replication | Experimental | Authenticated pull/streaming replication, ordered event application and replica-durable confirms bound to exact history and replica identities |
| Automatic queue recovery | Experimental | Enrolled Unix queues fence ownership changes, verify surviving history and activate an installed quorum. Compatible retained payloads can be reused after full verification; contradictory or unsupported histories remain fenced |
| Background replica catch-up | Experimental | Excluded assigned replicas catch up while the owner continues serving; durable/applied state and complete payloads gate admission |
| Eager failure detection | Experimental, opt-in | Runtime policy reacts to repeated explicit peer transport failures; zero grace adds idle-disconnect monitoring and immediate reconnect verification. Recovery proof still gates service; silent failures use heartbeat expiry |
| Stream replication | Experimental | Durable-tier streams replicate record and cursor logs to followers with replica-durable confirms and caught-up failover, placed and owned through embedded coordination. Express tiers stay owner-only |
| Live repartitioning | Experimental | Grow or shrink a queue’s partition count in coordinated mode, from the admin topology page |
| Reconnect and restart resume | Available | Typed subscription closure, broker-local persisted resume sessions and stale-delivery settlement across all five clients. Explicit fallback discovery endpoints and temporary-recovery retries support owner-loss reattachment |
| Checkpoint installation | Experimental, Unix | Journaled installation resumes across interruption; required payload backfill blocks promotion, including after restart. Linux fault and process-kill tests cover the path |
| Agreed recovery checkpoints | Experimental, opt-in on Unix | Durable same-cut capsules and all-admitted replica receipts publish through consensus. Retention preserves the accepted base and later suffix; recovery verifies retained bytes and keeps the existing quorum/fencing requirements. Event/append-byte triggers and shared dashboard diagnostics complement periodic cadence. Live backlog still requires payload validation |
| Speculative queue delivery | Experimental prototype, unmerged | Local-queue results remain separate from main. Correctness reconciliation and replicated speculation are pending; Plexus stream tiers have their own existing contract |
| Recovery quarantine | Available | A damaged queue log is detected on recovery and isolated per the recovery.on_mismatch policy, with operator repair |
| TLS in transit | Available | The broker listener serves TLS from operator PEMs or generated per-deployment material, mismatches are named in both directions, the clients connect with CA-file, fingerprint-pin, or OS-roots trust, the dashboard serves HTTPS from the same material, and first-boot setup mode offers generate/supply/skip before the broker ever serves. Inter-broker replication and coordination traffic is encrypted too (tls.inter_broker, shared-CA lane for generated material), the serving certificate rotates live via fibrilctl admin reload-tls, and tls.client_auth turns client certificates into credentials (a verified identity authenticates as the matching user with no password, require closes the handshake to certless peers) |
| Broker authentication | Available | Argon2 user store seeded from config, managed from the dashboard and fibrilctl, replicated across the cluster. Built-in fibril/fibril credentials work from loopback only. Node-to-node connections authenticate with a cluster shared secret, never a user account |
| Admin dashboard and demo | Available | Shared production views cover queues, messages, topology, resources and settings; the docs include a read-only simulated demo |
| Runtime and node-local settings | Available, scoped | Runtime fields have version-checked dashboard controls. Node-local segment preallocation persists separately and reports adoption at segment rollover/reopen; other storage tuning remains startup configuration |
| Recovery stage diagnostics | Available | Bounded process-local worker timelines in the shared Cluster dashboard and demo; stage outcomes, peers and overlap. Timing excludes failure detection and client reconnect; history resets on restart |
| Prometheus metrics | Available | GET /metrics on the admin listener behind the same auth and HTTPS as the dashboard: node-level aggregates always, per-channel series from materialized channels gated by admin.metrics_per_channel |
| Exclusive consumer groups | Partial | Rust, TypeScript, Python, Go, and C# client opt-in for one active consumer per partition, with sticky assignment and cross-broker coordinator wiring |
| Transactions | Out of scope | Not planned. Transactional publish/consume workflows are intentionally excluded |
Valid histories can exceed current automatic recovery budgets and remain fenced. Large-history recovery availability is a high-priority gap.
Cluster recovery still needs broader partition, power-loss and sustained-load acceptance, authoritative stream-history recovery and safe reclamation of retained recovery data. There is no history-independent failover latency guarantee; see the failover plan for the remaining gates.
Early performance observations
Section titled “Early performance observations”Informal internal measurements on a Ryzen 5950X system have observed roughly 250k+ messages/sec ingress and egress with 1KB payloads on the durable path. Plexus stream fan-out reaches roughly 1.5M delivered records/sec across sixteen readers on a single partition, and one partition fans out to hundreds of readers when delivery throughput is not the bottleneck.
These numbers are architecture sanity checks, not a rigorous benchmark suite. Hardware, durability settings, batching, queue depth, workload, and storage behavior all matter. See benchmarks for the fuller queue and stream tables.