Admin Dashboard
The admin dashboard is for operators. It shows broker state, active connections, queues, runtime settings, and message inspection tools.
Explore the read-only web demo without installing a broker. It uses the same templates, scripts and styles as the real dashboard, with illustrative sample data. Browsing, filtering, charts and themes work; broker mutations are disabled. Embedded examples stay on their assigned view; use Open full dashboard above a frame to browse other pages. Sample values are examples, not benchmark results.
The UI is a sidebar shell with a command palette (Ctrl/Cmd-K jumps to any page), dark and light modes, and optional named flavors that re-key the accent (neuronic, chlorophyll, crimson, eosin, azure, iris, and carotene - the original design mockup’s burnt-orange scheme, named for the pigment). The fibril ring mascot greets you on the login page, the 404 page, and the browser tab - where its face reacts: X-eyes while the broker is unreachable, a strained face while it runs flat out. It stays out of the data views. The UI vendors only the small icon set and mascot art it needs and makes no external requests. In cluster mode the top bar shows which broker you are on and switches to another broker’s admin on the same page.
To see the whole dashboard busy without inventing traffic by hand, run the
demo world against any broker: cargo run --release -p fibril-demo. It
drives three small businesses (orders through kitchens into deliveries and
payments, XML freight manifests, hotel bookings with rare tiered-refund
cancellations) with rush hours, a napping consumer that raises and resolves
backlog attention, and a flaky one that feeds the dead-letter flow.
The dashboard is not meant to be a high-frequency monitoring feed. Use it to
answer specific operational questions, and point Prometheus at the /metrics
endpoint on the same listener (same auth, same HTTPS) for continuous
monitoring. See monitoring.
When broker TLS is enabled (tls.enabled = true) the dashboard serves HTTPS
from the same certificate material. A plain http:// request to the TLS
listener gets redirected to the same URL under https://, so a bookmark or
habit-typed address still lands on the dashboard. Set tls.admin_enabled = false to keep it on plain HTTP behind a reverse proxy that terminates TLS. See
configuration for the tls section. The settings page
startup summary shows the current TLS state, including the CA fingerprint for
generated material.
With setup.mode = true and a fresh data dir, the server boots into a
first-boot setup page on 127.0.0.1:<admin port> instead of serving: choose
generated TLS material, paste a certificate and key, or explicitly continue
without TLS, and the broker starts once the choice is applied. See
configuration for the mechanics.
Overview And Diagnostics
Section titled “Overview And Diagnostics”The overview page is intentionally curated. It leads with a needs-attention panel - conditions the broker itself flags, like a backlog with no consumer reading it (or one growing despite consumers), a certificate near expiry, a failed settings load, a quarantined partition, low disk on the data directory, a stalled replication follower, or a broker left draining, each linking to where it is fixed - then live throughput and backlog-over-time charts, state-colored stat cards, process resource use, reconnect outcomes, and a small set of storage health signals. The disk card carries a segmented breakdown of which queues the bytes belong to, largest first, so growth points at its cause - folded by default behind a small toggle so the stat row stays compact, and the choice sticks per browser. In cluster mode the overview also lists the nodes: each registered broker with its address, liveness, owned partitions, and live publish and delivery rates from the coordination heartbeat, with the consensus leader starred - and the top bar carries a live/registered nodes chip beside the stream-health pill and the broker switcher. Chart history is sampled in memory on the broker (the last 30 minutes) and resets on restart; Prometheus stays the durable history.
Pages update live: each open page holds one server-sent-events stream and the broker pushes its data every couple of seconds, serializing each data family once no matter how many pages watch it. An idle dashboard costs the broker nothing, and pages fall back to polling automatically if the stream is unavailable. The live pill in the top bar tracks the stream’s health, and it only declares the broker unreachable on an observed failure - updates merely paused (a background tab, active interaction on the page) show as stale. A passive probe keeps checking while polls pause, so a broker dying behind an inactive tab still turns that tab’s favicon to X-eyes within a couple of minutes (a tab the browser has put to sleep runs nothing and keeps its last face).
The overview also carries a collapsed resources panel - memory, CPU, and disk over the last 30 minutes - for when a resource question needs a shape, not a number. The settings page can enable desktop notifications for new attention conditions (critical only, or critical and warning), a per-browser choice.
The Storage retention panel separates active data and snapshots, agreed checkpoint artifacts, retained or inactive generations, recovery staging and other files under the node’s Stroma root. Unix samples use allocated file blocks and count each hard-linked inode once. Shared bytes are attributed to active data first. Current logs can include segments pinned by an accepted checkpoint. On other hosts the panel labels file lengths, which can count links separately. Directory blocks, filesystem metadata, reflink sharing and files outside that root are excluded.
A request starts a background sample when needed and immediately returns the last result. Samples refresh at most every 30 seconds while the view is in use. One blocking scan runs at a time per process, with a 100,000-entry and five-second traversal budget checked between filesystem calls. The panel shows sample age, skipped entries and errors. An unsuccessful scan retains the last completed sample. These approximate counters grant no permission to reclaim recovery data. The existing Active logs number and history chart still report message and event log file lengths, so they use a different basis from allocated blocks.
The Activity page is the broker’s own diary: operator actions (declares, deletes, drains, test publishes), attention conditions raised and resolved, cluster membership changes, and stream subscribers falling behind into lag recovery, newest first with severity-colored entries and a filter. It holds the newest 512 entries in memory and resets on restart - the broker log remains the durable trail.
The diagnostics page shows lower-level storage and queue metrics, including command-lane depths and timings, command-kind counters, append stats, snapshot cost, and recovery counters. Use it when the overview suggests pressure and you need the next level of detail.
Security
Section titled “Security”The security page shows the served certificate (subject, expiry, and the SHA-256 fingerprint clients pin, with a copy button) and reloads certificate material from disk without a restart - new handshakes get the new leaf while existing sessions keep serving. It also manages the admin users described below.
The security page manages users (the settings page retains a copy of the
section): create or rotate a user (argon2-hashed, never shown), remove a
user, and see the current list with timestamps. The
built-in fibril/fibril pair works from loopback only, so create a user for
remote access. In cluster mode edits replicate to every node. The same
operations are available from fibrilctl user add/passwd/remove/list.
Auth State
Section titled “Auth State”When admin authentication is configured, dashboard pages redirect unauthenticated
requests to /login, and the header shows a logout action.
When authentication is disabled, pages are accessible directly and the header
shows Auth disabled. Treat that mode as local-only or otherwise protected by
network boundaries.
Queues
Section titled “Queues”The queues page lists known queues with ready, inflight, and settled offset information. Use Inspect messages from a queue row when you want to inspect that specific queue.
You can create a queue from the page (partition count, an optional dead-letter policy, and an optional default message TTL) and delete a queue from its row. Delete is single-node: it is refused while a partition still has inflight work, and refused in cluster mode pending coordinated teardown. A hide-inactive toggle and a search filter help when the list is long. Filters persist in the page URL (here and on the streams, connections, and subscriptions pages), so a filtered view survives reload and can be shared by copying the address.
For sparse workloads, the page also shows whether each queue is currently loaded in memory, only indexed on disk, or recently unloaded after being idle. It shows active publisher/subscriber counts, idle time when known, last used time for the current process, and the most recent idle-cleanup result or skip reason.
Partitioned queues are shown as one row per topic with their partition and
loaded counts, a 30-minute depth trend sparkline, live in/s and out/s rates
(derived in the page between refreshes), and the queue’s dead-letter
policy. Clicking the sparkline opens a drilldown with the depth chart, the
rates, and the age of the oldest ready message. Expand a topic to see each partition’s own state, or open Detail
for the full picture of one queue: its depth and leased charts, per-partition
cards, live consumers with their settlement mode, and its declared
configuration. The detail page can also publish a test message through the
broker’s real publish path - durable confirm and delivery included - so you can
verify a queue end to end without a client. Test messages carry a reserved
fibril.test: admin header consumers can recognize and filter.
When this broker is replicating queues from their owners, the page also shows a follower-replication section: which partitions this broker follows, each follower’s status (caught up, pending retry, or checkpoint required), how far it has pulled, and when it last made progress. See replication.
The message inspection link starts near the queue’s settled offset by default so
you do not begin at offset 0 on large queues unless you choose to.
Streams
Section titled “Streams”The streams page lists the Plexus stream channels this broker is currently hosting, grouped by topic. Each partition row shows its head and tail offsets, how many records are retained, its live subscription count, and how often a subscriber overflowed its live buffer and went through lag recovery. Streams with durable subscribers also list their cursors: each named cursor’s partition, whether it sits at the tail or is catching up, how far behind it is, its read rate, and when it last advanced - a parked consumer is visible immediately. The topic heading shows the declared durability tier and retention bound, and a publish test record button that sends one marked record through the real publish path, the same end-to-end check the queue detail page offers.
You can create a stream from the page: a topic, a partition count, a durability tier, and optional retention bounds (records, bytes, or age). A topic is one channel kind for its lifetime, so declaring a stream over an existing queue (or the reverse) is refused.
Message Inspection
Section titled “Message Inspection”Message inspection reads queue state and persisted message data on demand. Use it for debugging and operations, not as a live polling view. The topic and group fields suggest existing names as you type (as do the group and dead-letter fields when declaring a queue); typing a new name works as usual.
Inspecting a queue can load it into memory. If idle queue cleanup is enabled and no publisher or subscriber keeps that queue active, cleanup can unload it again after the idle window.
By default, inspection shows active queue state:
- ready messages
- inflight messages
- delayed messages
- pending DLQ messages
Enable Include settled offsets when you also need persisted records that are no longer active in queue state. Use the status filter when you only care about one status, such as pending DLQ messages.
Payload previews are optional. They are base64 over the admin API and shown as a short preview in the table. Use the payload modal for a larger preview. Large page sizes and large payload previews ask for confirmation because they can read a lot of persisted data.
Dead Letters
Section titled “Dead Letters”The dead-letters page gathers the failure lane in one place: the global dead-letter target with its backlog now and over time, and every queue that declares a dead-letter policy with its depth and an inspector link. Browsing and replaying individual dead letters happens in message inspection.
DLQ Replay
Section titled “DLQ Replay”When inspecting a DLQ queue, select specific offsets and use Replay selected to source. Replay copies the payload and user headers back to the recorded source queue.
Replay does not remove or acknowledge the DLQ message. The result table reports which offsets were replayed and which were skipped.
Runtime Settings
Section titled “Runtime Settings”The runtime form labels the saved document’s authority and revision. Each field also shows the startup seed from this node’s resolved configuration. These seeds initialize a missing settings document, and saved settings take precedence. Keep startup seeds consistent across cluster nodes.
The settings page exposes delivery, connection lifecycle, replication, partitioning, consumer-group defaults, idle queue cleanup and Plexus stream settings. The startup summary shows this node’s storage configuration, including batching, tail-cache budget, writer buffers, adaptive staging and segment sizing; these values remain read-only startup values. The node-local section below can override preallocation; other storage controls require a restart.
Broker runtime settings are node-local in standalone mode and cluster-authoritative in coordinated mode. Saving preserves the complete loaded settings document, including fields not represented by the current form. A version conflict reloads the current document for review before another save. Startup-locked settings are shown as locked and cannot be edited through the dashboard.
A successful save persists the requested settings. Workers adopt them at their supported boundaries: cursor batching reads the next batch’s settings, replication buffer depth applies to the next replication stream, and an active drain keeps its original handoff timeout. The page separately reports whether the broker and connection runtime values installed on this node match the saved document. This comparison uses values because the cluster and local cache have independent revision counters. It does not certify adoption by every in-flight operation. Controller checkpoint policy, eager failover policy and drain handoff timing are not tracked by this observation. Reload the form to refresh the status.
The Node-local Storage section identifies the serving node and controls segment preallocation. Overrides survive restart; clearing the override restores the startup value. The table shows the accepted revision, each open log’s adopted revision, pending rollover/reopen and filesystem allocation fallback. Use Refresh status to check adoption; saving does not force a rollover. Open the other node’s dashboard to edit that node. See node-local configuration.
Global DLQ target changes are persisted in storage-owned state and survive restart.
Connections
Section titled “Connections”The connections page lists every open client connection with its publish count, subscriptions, auth state, and uptime. Its diagram view renders the broker as the fibril ring with clients plugged in: publisher connections run blue strands into the left side, subscribers violet strands out of the right, pulses travel along each strand at a cadence following that connection’s live rate (derived in the page from successive data ticks), and idle connections gather as a dim bundle under the ring. Hovering an endpoint spotlights its strand and its table row. In a cluster the smart clients follow partition ownership, so a connection appears on the broker it actually talks to.
Subscriptions and Cohorts
Section titled “Subscriptions and Cohorts”The subscriptions page lists active subscriptions. When exclusive consumer groups (cohorts) are in use, it also shows this broker’s view of each cohort: the topic, group, and the members with their per-consumer targets and the partitions each member’s live subscription covers. The queue detail page uses the same data to name the covering member on every partition card and to flag a partition no cohort member covers. Cohort assignment is broker-local runtime state, so this is a per-node view rather than a single cluster-wide table.
Topology
Section titled “Topology”When the broker runs in coordinated (Ganglion) mode, the topology page shows the cluster: registered brokers, per-partition ownership with fencing epochs and followers, and the consensus block (leader and voters). See clustering.
The Recovery timeline shows recent recovery worker attempts on the serving node. Select an attempt to see its resource, epoch, transition identity, peers, individual stage durations and overlapping work. Failed and cancelled stages remain visible alongside completed attempts. Failures are red, cancellations use amber striped bars, and phases without recorded stages are grey. Retries can reuse work, so an unobserved phase does not imply a failure. Timing begins at worker entry; failure detection and client reconnect are outside that interval. A completed attempt has activated its assignment, but replica admission can still retry.
Recognized recovery budget failures identify the limit, work accepted before the
check and the additional work refused. These counts describe the failing operation,
not complete recovery verification. An unchanged retry cannot clear the same
fixed limit. A different available source, changed history or supported budget may
allow progress. Legacy errors without numeric diagnostics retain an unknown cause.
The demo includes an archive.import attempt that reaches the inspection record limit.
The shared topology API exposes these observations under
consensus.recovery_timeline. The store holds 32 attempts with at most 256 stages
each, preferring to evict completed attempts; omission and eviction counts are
visible. Labels are capped at 128 characters, and message payloads and raw error
bodies are excluded. History resets when the node restarts. Switch to another
broker for its local attempts, and use the transition identity to correlate the
existing fibril::recovery_timing log events.
The diagram view renders each broker as the fibril ring: active brokers blink more often than idle ones, and one that drops out of the cluster lingers briefly as a dimmed X-eyed ghost so the gap is visible. Replication links draw as fiber strands whose density follows how many partitions ride the link, and a standalone broker sprouts a strand per declared queue. The rings also show real load: each broker reports coarse publish and delivery rates on its coordination heartbeat, and the diagram turns them into signal pulses along the strands (an idle ring shimmers rarely, a busy one fires constantly), faster band-light cycling, a msg/s figure under the broker name, and - flat out - a strained face. The gold consensus mesh hides behind a Consensus toggle (membership already implies a leader connection), and a Fibers toggle turns the animated layer off entirely, defaulting off past twelve rings. Honoring reduced-motion preferences stills all of it.
The list view shows the same cluster as broker cards - liveness, the
consensus leader starred, address, owned and followed partitions, version
and uptime, TLS state with certificate expiry, and live rates, with an
open-its-admin link on every other node - above a placement matrix of
owned partitions per broker per topic. The view choice sticks per browser
and answers to ?view=list in the URL.
The page also exposes three operator actions, each with a confirmation:
- Drain this broker: clients are told to move and, in coordinated mode, partition ownership hands off to caught-up followers before the call returns - zero partitions remaining means stopping the process is gap-free.
- Repartition a queue by setting its partition count.
- Add or remove a consensus voting member.
Use these deliberately. Repartitioning changes placement, and voting-membership changes affect quorum and availability.
Health And Quarantine
Section titled “Health And Quarantine”/healthz is a simple liveness check. /readyz reflects readiness, including
whether any partition is quarantined after a failed recovery.
If a partition was quarantined because its log failed recovery verification, a banner appears across the dashboard. From the banner you can repair the affected partition, which truncates its log to the last valid record and clears the quarantine. See recovery quarantine.