Configuration
Fibril has two configuration layers:
- Startup config decides how the server process starts.
- Runtime settings decide live broker behavior and are persisted after first boot.
That split matters. Startup config is for things the process needs before it can run, such as bind addresses and the data directory. Runtime settings are for behavior that can be changed while the broker is running, such as delivery timing and idle queue cleanup.
Config File
Section titled “Config File”The config file format is TOML. The repository includes fibril.example.toml:
[server]data_dir = "server_data"
[broker.listener]bind = "0.0.0.0:9876"
[admin.listener]bind = "0.0.0.0:8081"
[admin.auth]enabled = falseusername = "fibril"# password = "change-me"
[setup]mode = false
[tls]enabled = false# Supply your own PEM files:# cert_path = "/etc/fibril/tls/server.pem"# key_path = "/etc/fibril/tls/server.key"# Or generate per-deployment material under <data_dir>/tls on first boot:# auto_self_signed = true# The admin dashboard serves HTTPS from the same material when TLS is# enabled. Opt it out when a reverse proxy terminates TLS for the dashboard:# admin_enabled = false# Follower-to-owner replication follows `enabled` too. Opt it out when a# service mesh or tunnel already encrypts inter-broker traffic:# inter_broker = false# CA that peer certificates chain to. Unset falls back to the generated# <data_dir>/tls/ca.pem, then OS roots:# peer_ca_path = "/etc/fibril/tls/ca.pem"
[storage.keratin]fsync_interval_ms = 5# Floor between storage commits while the fsync worker is idle. 0 self-clocks# group commit on fsync completions, the best setting for fast storage (NVMe,# tmpfs). On slow-fsync storage such as SATA SSDs a floor around the fsync# interval (5) gives the drive breathing room between write barriers.min_fsync_interval_ms = 0# Writer and notification channels each have 64 * factor slots per log.# 128 preserves the default 8192 slots; smaller values reduce eager allocation.writer_buffer_factor = 128adaptive_staging = truestaging_decay_secs = 10staging_idle_release_secs = 60
[storage.keratin.message_log]segment_max_bytes = 268435456
[storage.keratin.event_log]segment_max_bytes = 33554432
[runtime_seed.delivery]inflight_ttl_ms = 30000expiry_poll_min_ms = 15000expiry_batch_max = 8192delivery_poll_max_ms = 5000
[runtime_seed.idle_queue_cleanup]enabled = falseevict_after_ms = 600000sweep_interval_ms = 60000# Set this for sparse workloads with long-lived publishing connections.# publisher_idle_timeout_ms = 600000
[runtime_seed.connection]# Reconnect grace is on by default (5000 ms). Set to 0 to disable it.# reconnect_grace_ms = 5000# How stale a persisted session skeleton a restarted broker will honor, so a# fast restart lets clients resume and reconcile (60000 ms). Set to 0 to# disable restart resume.# resume_session_restart_ttl_ms = 60000
[runtime_seed.replication]confirm_timeout_ms = 5000caught_up_poll_ms = 1000retry_poll_ms = 100checkpoint_retry_poll_ms = 5000max_messages_per_read = 2048max_events_per_read = 2048max_bytes_per_read = 8388608max_iterations_per_tick = 8min_in_sync_replicas = 1isr_timeout_ms = 10000read_timeout_slack_ms = 10000owner_connect_timeout_ms = 5000
[runtime_seed.partitioning]default_partition_count = 1
[runtime_seed.consumer_groups]# Blank or omit to disable the under-provisioned signal.# default_target_per_consumer = 4
[coordination.ganglion]heartbeat_interval_ms = 3000liveness_ttl_ms = 9000
[runtime_locks]idle_queue_cleanup = falseRun with a config file:
cargo run --release --bin fibril-server -- --config fibril.tomlor:
FIBRIL_CONFIG=fibril.toml cargo run --release --bin fibril-serverPrecedence
Section titled “Precedence”Startup config is resolved in this order:
compiled defaults < TOML config file < environment variables < CLI argumentsThis precedence applies to startup fields and first-boot runtime seeds. It does not mean environment variables keep overriding persisted runtime settings after runtime state exists.
Startup Fields
Section titled “Startup Fields”These fields are read on process start.
| TOML field | Env var | CLI flag | Default |
|---|---|---|---|
server.data_dir | FIBRIL_DATA_DIR | --data-dir | server_data |
broker.listener.bind | FIBRIL_BROKER_BIND | --broker-bind | 0.0.0.0:9876 |
broker.listener.advertise | FIBRIL_BROKER_ADVERTISE | none | derived (see below) |
admin.listener.bind | FIBRIL_ADMIN_BIND | --admin-bind | 0.0.0.0:8081 |
admin.listener.advertise | FIBRIL_ADMIN_ADVERTISE | none | derived (see below) |
admin.auth.enabled | FIBRIL_ADMIN_AUTH_ENABLED | --admin-auth-enabled | false |
admin.auth.username | FIBRIL_ADMIN_USERNAME | --admin-username | fibril |
admin.auth.password | FIBRIL_ADMIN_PASSWORD | --admin-password | unset |
admin.metrics_per_channel | FIBRIL_ADMIN_METRICS_PER_CHANNEL | none | true |
tls.enabled | FIBRIL_TLS_ENABLED | none | false |
tls.cert_path | FIBRIL_TLS_CERT_PATH | none | unset |
tls.key_path | FIBRIL_TLS_KEY_PATH | none | unset |
tls.auto_self_signed | FIBRIL_TLS_AUTO_SELF_SIGNED | none | false |
tls.admin_enabled | FIBRIL_TLS_ADMIN_ENABLED | none | follows tls.enabled |
tls.inter_broker | FIBRIL_TLS_INTER_BROKER | none | follows tls.enabled |
tls.peer_ca_path | FIBRIL_TLS_PEER_CA_PATH | none | generated CA, then OS roots |
tls.client_auth | FIBRIL_TLS_CLIENT_AUTH | none | off |
tls.client_ca_path | FIBRIL_TLS_CLIENT_CA_PATH | none | generated CA |
setup.mode | FIBRIL_SETUP_MODE | none | false |
auth.allow_default_loopback | none | none | true |
auth.seed_users | FIBRIL_AUTH_USERNAME + FIBRIL_AUTH_PASSWORD (one entry) | none | empty |
coordination.secret_path | FIBRIL_CLUSTER_SECRET_PATH (or FIBRIL_CLUSTER_SECRET for the value) | none | unset |
storage.keratin.fsync_interval_ms | FIBRIL_KERATIN_FSYNC_INTERVAL_MS | --keratin-fsync-interval-ms | 5 |
storage.keratin.min_fsync_interval_ms | FIBRIL_KERATIN_MIN_FSYNC_INTERVAL_MS | --keratin-min-fsync-interval-ms | 0 |
storage.keratin.writer_buffer_factor | FIBRIL_KERATIN_WRITER_BUFFER_FACTOR | none | 128 |
storage.keratin.adaptive_staging | FIBRIL_KERATIN_ADAPTIVE_STAGING | none | true |
storage.keratin.staging_decay_secs | FIBRIL_KERATIN_STAGING_DECAY_SECS | none | 10 |
storage.keratin.staging_idle_release_secs | FIBRIL_KERATIN_STAGING_IDLE_RELEASE_SECS | none | 60 |
storage.keratin.message_log.segment_max_bytes | FIBRIL_KERATIN_MESSAGE_LOG_SEGMENT_MAX_BYTES | --keratin-message-log-segment-max-bytes | 268435456 |
storage.keratin.event_log.segment_max_bytes | FIBRIL_KERATIN_EVENT_LOG_SEGMENT_MAX_BYTES | --keratin-event-log-segment-max-bytes | 33554432 |
coordination.mode | FIBRIL_COORDINATION_MODE | none | static |
coordination.ganglion.heartbeat_interval_ms | FIBRIL_COORDINATION_HEARTBEAT_INTERVAL_MS | none | 3000 |
coordination.ganglion.liveness_ttl_ms | FIBRIL_COORDINATION_LIVENESS_TTL_MS | none | 9000 |
coordination.ganglion.target_followers | none | none | 1 |
coordination.ganglion.stream_replication_factor | none | none | 1 |
coordination.ganglion.repartition_adoption_timeout_ms | none | none | 30000 |
coordination.ganglion.assignment_durability | FIBRIL_COORDINATION_ASSIGNMENT_DURABILITY | none | local_durable |
recovery.on_mismatch | FIBRIL_RECOVERY_ON_MISMATCH | none | quarantine |
Changing these generally requires restarting the server.
broker.listener.advertise is the address (or addresses) the broker tells peers
and clients to reach it on, separate from bind. This matters when bind is not
itself dialable - the common case is binding 0.0.0.0 in a container, which a
peer cannot connect back to. Give it a routable host:port (a service name is
fine, it is resolved at connect time), or several comma-separated entries in
FIBRIL_BROKER_ADVERTISE in priority order. When unset it is derived in
ganglion mode from this node’s coordination peer host plus the broker port, and
otherwise falls back to bind. Standalone single-broker deployments do not need
it (clients connect to the broker directly). Only the first entry is dialed today;
the rest are carried for forward compatibility.
admin.listener.advertise is the same idea for the dashboard: the address a
node registers with the cluster so the Cluster page’s “open its admin” links
and the broker switcher can point a browser at it. When unset and the admin
bind host is unspecified (0.0.0.0), the node reuses the broker advertise
host with the admin port; a raw unspecified address is never linked (the
dashboard shows “no reachable admin address registered” instead). In the
Docker cluster example each node advertises its host-mapped admin port.
coordination.mode is static for a standalone single-broker deployment (the
default) or ganglion to run the embedded coordinator and form a cluster. The
coordination.ganglion.* settings only apply in ganglion mode. See
clustering and
replication.
coordination.ganglion.target_followers is the desired follower count per queue
partition. coordination.ganglion.stream_replication_factor is the equivalent for
DURABLE Plexus stream partitions — tuned separately so stream and queue fault
tolerance can differ; only the durable tier replicates, the express tiers stay
owner-only. A value of one keeps a durable stream available across a single node
loss; zero makes durable streams owner-only (durable on disk, not HA). See
Plexus streams.
coordination.ganglion.assignment_durability is the default durability
policy for new assignments (local_durable, replica_accepted, replica_durable,
or majority_durable).
recovery.on_mismatch controls what happens when recovery finds a damaged queue
log: quarantine (default) isolates the partition, refuse reports not ready,
and ignore truncates to the last valid record. See
recovery quarantine.
admin.auth.enabled = true requires both admin.auth.username and admin.auth.password.
The admin password is intentionally not shown in the dashboard startup summary.
admin.metrics_per_channel controls whether the Prometheus /metrics
endpoint on the admin listener includes per-channel series (queue depth,
stream subscriptions, follower applied state) alongside the always-exported
node-level aggregates. See monitoring.
tls.enabled = true serves TLS on the broker listener and the admin dashboard
from one certificate. Supply PEM files with tls.cert_path and tls.key_path,
or set tls.auto_self_signed = true to generate a per-deployment CA and server
certificate under <data_dir>/tls on first boot. The CA fingerprint is printed
at startup so clients can pin or trust it - self-signed material a client does
not verify defeats passive snooping only. Fibril never ships certificates.
tls.admin_enabled = false keeps the dashboard on plain HTTP for deployments
where a reverse proxy terminates TLS in front of it. Certificate material is
read once at startup, so replacing it requires a restart (live reload and
rotation are planned).
tls.inter_broker covers follower-to-owner replication and the coordination
raft channel. It follows tls.enabled, which assumes a homogeneous cluster
whose certificates chain to one CA - see
TLS across nodes for the shared-CA
lane and the rotation runbook. Set it to false when a service mesh or
tunnel already encrypts inter-broker traffic. tls.peer_ca_path names the CA
peers are verified against, falling back to the generated
<data_dir>/tls/ca.pem when present, then OS roots. The serving certificate
and key reload live via fibrilctl admin reload-tls or
POST /admin/api/tls/reload.
tls.client_auth turns client certificates into credentials: request
verifies a presented certificate (certless clients still connect and
password-auth, the migration lane), require rejects certless clients in the
handshake. A verified certificate whose identity (first DNS SAN, else CN)
names an existing user authenticates as that user with no password; the @
node namespace can never be claimed by certificate. tls.client_ca_path
names the CA client certificates chain to, falling back to the generated CA -
issue workload certificates from it with fibrilctl cert issue <identity>.
Brokers present their own certificate on inter-broker dials, so a cluster
converges at any client_auth setting.
The auth section governs broker authentication. The built-in default
credentials (fibril/fibril) are accepted from loopback connections only,
so local development works out of the box while remote access always requires
a real user. auth.seed_users creates users when the user store is empty
(first boot only, after which the persisted store owns the users), and the
FIBRIL_AUTH_USERNAME/FIBRIL_AUTH_PASSWORD pair seeds one user from the
environment. A real user named fibril replaces the built-in pair entirely.
Passwords are stored as argon2 hashes. With users configured and TLS off, the
server warns loudly: passwords travel in cleartext on non-loopback plaintext
connections.
The cluster shared secret authenticates node-to-node connections (replication)
as a node principal, never as a user account, and is required in ganglion
mode. Resolution order: FIBRIL_CLUSTER_SECRET (the value itself), then
coordination.secret_path, then <data_dir>/cluster.secret if present -
which is what fibrilctl secret generate writes. Every node holds the same
secret.
setup.mode = true arms first-boot setup: when the data dir holds no
setup_complete marker, the server serves only a setup page on
127.0.0.1:<admin port> and the broker listener stays down until the operator
chooses generated TLS material, supplies a certificate, or explicitly
continues without TLS. The choice is written to <data_dir>/config-overlay.toml,
which boot layers below explicit config, so anything set in the tls section
by file, environment, or CLI always wins over a setup choice. Completion
writes the marker and the broker starts in the same process. Deleting the
marker and booting with setup mode runs setup again. Deployments that
configure TLS explicitly never need setup mode, and with it armed they boot
straight through (the marker is written automatically).
storage.keratin.message_log.segment_max_bytes and storage.keratin.event_log.segment_max_bytes are rollover thresholds. A segment rolls after an append crosses the configured size, so an individual segment can be slightly larger than this value.
storage.keratin.writer_buffer_factor sets the startup capacity of each storage
writer input and notification channel to 64 × factor entries. Valid factors
are 1..=128: 1 gives 64 entries, 16 gives 1,024, and the default 128
gives 8,192. The channels allocate their slot arrays when a log opens, so smaller
values reduce the resident cost of materialized queues and streams and apply
backpressure earlier during bursts. A channel entry may own a batch; this
setting is a slot limit rather than a total memory limit. Measure throughput
and tail latency with representative payloads, storage and replication before
choosing a lower value. Changing it requires restart and applies to both the
message and event logs. Fsync pipeline depth remains controlled by
storage.keratin.max_inflight_fsyncs; caches, actor mailboxes and client buffers
have separate limits.
storage.keratin.adaptive_staging = true (the default) enables lazy staging allocations for
both message and event logs. Write buffers start empty with a 64 KiB reservation
floor; sparse-index buffers use a 4 KiB floor. They grow to fit batches, may halve
while empty every staging_decay_secs (default 10), and release fully after
staging_idle_release_secs without use (default 60). Decay preserves headroom
for recent batches and never discards staged records. For less frequent bursts,
longer retention delays reduce repeated allocation work.
Set adaptive_staging = false to retain eager 16 MiB write and 256 KiB index
reservations per log. These sizes and adaptive floors are initial reservations,
not maximum capacities. This startup-only setting changes allocation retention;
channel sizes, caches and durability policy remain independently configured.
Released capacity may remain resident in the allocator, so RSS can stay high or
even increase for some burst patterns. Evaluate memory, CPU and first-message
latency after idle when choosing retention delays or opting out.
coordination.ganglion.heartbeat_interval_ms controls how often a broker
renews its cluster liveness record. coordination.ganglion.liveness_ttl_ms
controls how long a broker can go without a fresh heartbeat before the cluster
considers it unavailable. The TTL must be at least twice the heartbeat interval.
For heavy replication benchmarks, a longer TTL can avoid false failover while
the node is under artificial load.
coordination.ganglion.repartition_adoption_timeout_ms bounds how long a live
repartition’s finalize (retiring shrunk-away partitions and clearing the
transition marker) waits for clients to adopt the new routing once the backlog
has drained. Adoption is observed from client topology acks. The timeout keeps a
silent or stuck client from stalling a cutover forever; publish version-fencing
is the correctness backstop regardless. See
live routing and cutover.
Linux Memory Policy
Section titled “Linux Memory Policy”The Linux broker uses mimalloc, which exposes a startup environment option for transparent huge pages (THP). Consider disabling THP when resident memory is a deployment constraint, especially with many materialized queues whose log buffers are lightly used. Huge pages can make a sparsely touched allocation consume substantially more resident RAM; disabling them can also reduce the footprint of an active broker.
MIMALLOC_ALLOW_THP=0 ./fibril-serverAppend the usual broker arguments. Set the same variable in a container’s
environment or a systemd service’s Environment=MIMALLOC_ALLOW_THP=0 directive,
then restart or recreate the broker. This allocator setting is read at process
startup and has no Fibril TOML field or admin runtime setting. It disables THP
for the broker process and its descendants without changing the host policy
for other services. Remove the variable and restart to restore the allocator’s
default behavior under the host’s policy.
The tradeoff depends on workload: huge pages can improve address-translation efficiency for large active working sets, while allocating and clearing larger pages can add latency and consume extra memory.
After restart, /proc/<broker-pid>/status should report THP_enabled: 0, and
AnonHugePages in /proc/<broker-pid>/smaps_rollup should be zero. These checks
confirm that the process policy took effect, including in restricted containers.
THP policy is independent of the writer channel factor and idle queue cleanup;
each addresses a different part of the materialized-resource footprint. See
the memory investigation notes
and the upstream mimalloc options
and Linux THP documentation.
Runtime Seeds
Section titled “Runtime Seeds”runtime_seed values initialize the persisted runtime settings document when no runtime settings exist yet.
After runtime settings exist, the persisted values own these settings. You can edit them through the admin settings page or the admin runtime settings API.
Delivery
Section titled “Delivery”| TOML field | Default | Meaning |
|---|---|---|
runtime_seed.delivery.inflight_ttl_ms | 30000 | How long a delivered message lease lasts before it can be retried. |
runtime_seed.delivery.expiry_poll_min_ms | 15000 | Minimum sleep between expiry checks when no earlier expiry is known. |
runtime_seed.delivery.expiry_batch_max | 8192 | Maximum expired messages to requeue in one expiry pass. Must be at least 1. |
runtime_seed.delivery.delivery_poll_max_ms | 5000 | Maximum idle poll delay for delivery loops. |
Idle Queue Cleanup
Section titled “Idle Queue Cleanup”| TOML field | Env/CLI compatibility | Default | Meaning |
|---|---|---|---|
runtime_seed.idle_queue_cleanup.enabled | enabled implicitly by FIBRIL_QUEUE_IDLE_EVICT_AFTER_MS or --queue-idle-evict-after-ms | false | Enables unloading idle queues from memory. |
runtime_seed.idle_queue_cleanup.evict_after_ms | FIBRIL_QUEUE_IDLE_EVICT_AFTER_MS, --queue-idle-evict-after-ms | 600000 | How long a queue must be idle before cleanup can unload it. |
runtime_seed.idle_queue_cleanup.sweep_interval_ms | FIBRIL_QUEUE_IDLE_SWEEP_INTERVAL_MS, --queue-idle-sweep-interval-ms | 60000 | How often the cleanup worker checks tracked queues. Must be at least 1. |
runtime_seed.idle_queue_cleanup.publisher_idle_timeout_ms | FIBRIL_PUBLISHER_CACHE_IDLE_TIMEOUT_MS, --publisher-idle-timeout-ms | unset | Lets unused publishers stop keeping queues active while a connection remains open. |
For sparse workloads, enable publisher idle expiry alongside queue cleanup. Without it, a long-lived connection that published to a queue can keep that queue active until the connection closes.
See many idle queues for the user-facing behavior.
Connections
Section titled “Connections”| TOML field | Env/CLI compatibility | Default | Meaning |
|---|---|---|---|
runtime_seed.connection.reconnect_grace_ms | FIBRIL_RECONNECT_GRACE_MS, --reconnect-grace-ms | 5000 | Keeps a disconnected resumable client alive for this long before cleaning up subscriptions and requeueing unsettled messages. On by default so a transient blip resumes transparently; set 0 to disable. |
runtime_seed.connection.resume_session_restart_ttl_ms | FIBRIL_RESUME_SESSION_RESTART_TTL_MS, --resume-session-restart-ttl-ms | 60000 | How stale a persisted session skeleton a restarted broker will honor. Within this window a resume after a broker restart succeeds and reconciles instead of being rejected; past it (or 0) it reports not-found. Independent of reconnect_grace_ms, which governs live-process dormancy. |
connection.drain_handoff_timeout_ms (runtime settings) | none | 30000 | Upper bound on how long a drain call waits for partition ownership to hand off to caught-up followers in coordinated mode. Only caps the wait: reactive failover stays the backstop either way. Adjustable live from the settings page. |
Reconnect grace is disabled when unset. It only helps clients that use the resume identity handshake.
Replication
Section titled “Replication”These settings apply to the experimental cluster replication path.
| TOML field | Default | Meaning |
|---|---|---|
runtime_seed.replication.confirm_timeout_ms | 5000 | How long a replica-durable publish confirm can wait for enough durable follower progress. |
runtime_seed.replication.agreed_checkpoint_interval_ms | 0 | Opt-in agreed queue checkpoints; 0 disables new attempts, otherwise 1000–86400000 ms between periodic attempts. Work thresholds can start earlier, at most once per second. Existing attempts finish and the last accepted base/suffix remain retained. All admitted replicas must agree. See checkpoint policy. |
runtime_seed.replication.agreed_checkpoint_max_events | 0 | Start an enabled checkpoint earlier after this many events beyond its agreed cut; 0 disables the event trigger. |
runtime_seed.replication.agreed_checkpoint_max_bytes | 0 | Start earlier after approximately this many local append content bytes since observing the last certificate; 0 disables. The counter resets on reopen and excludes record framing; the periodic interval remains the fallback. |
runtime_seed.replication.eager_failover | false | Opt-in explicit peer-connection failure detection on the active metadata controller. Recovery proof is still required; see eager failover. |
runtime_seed.replication.eager_failover_grace_ms | 1000 | Failed-reconnect grace, 0–60000 ms. Zero enables idle-disconnect monitoring and immediate reconnect verification on the active controller. Positive values require failed contact across the interval; heartbeat expiry handles silent failures. |
runtime_seed.replication.caught_up_poll_ms | 1000 | Caught-up pull interval and streaming owner long-poll timeout. Streaming reads wake when new work arrives, without waiting for the timeout. |
runtime_seed.replication.retry_poll_ms | 100 | Follower retry interval after a partial pull or transient replication error. |
runtime_seed.replication.checkpoint_retry_poll_ms | 5000 | Follower retry interval while it needs an owner checkpoint before it can continue. |
runtime_seed.replication.max_messages_per_read | 2048 | Maximum message records per owner read, for pull and streaming. |
runtime_seed.replication.max_events_per_read | 2048 | Maximum event records per owner read, for pull and streaming. |
runtime_seed.replication.max_bytes_per_read | 8388608 | Approximate byte budget per owner log scan and initial byte credit for a new follower stream. An oversized record and accompanying events can make the response exceed this budget. |
runtime_seed.replication.max_iterations_per_tick | 8 | Maximum pull/apply iterations a follower performs before yielding. |
runtime_seed.replication.min_in_sync_replicas | 1 | Minimum recently in-sync replicas required before accepting replica-durable publishes. 1 disables the floor. |
runtime_seed.replication.isr_timeout_ms | 10000 | How recently a follower must report durable progress to count as in sync. |
runtime_seed.replication.read_timeout_slack_ms | 10000 | Slack added to a follower read’s long-poll window before the read is abandoned and the connection dropped, so a dropped owner response cannot hang the follower. A read waits max_wait_ms + this. |
runtime_seed.replication.owner_connect_timeout_ms | 5000 | Upper bound on a follower establishing a connection to an owner (TCP connect plus the handshake) before it is abandoned and retried. |
runtime_seed.replication.stream_enabled | true | Use credit-based streaming replication on followers. When disabled, followers fall back to polling pulls. |
runtime_seed.replication.stream_apply_linger_us | 2000 | Microseconds a streaming follower gathers contiguous frames before one fsynced apply. Higher trades apply latency for fsync amortization. 0 is drain-only. |
runtime_seed.replication.stream_apply_max_merge_bytes | 16777216 | Byte threshold for a coalesced streaming apply. Whole frames can cross it. Trades peak memory and overlap against fsync amortization. |
runtime_seed.replication.stream_buffer_batches | 8 | Queued frame count for the streaming follower, separate from byte credit. Applied on the next stream. |
Streaming owner reads adopt updated record and byte limits at the next batch. The initial credit window changes when a new follower stream starts. Oversized batches consume that window and retain any excess as debt until the follower returns credit after durable application.
Read budgets trade batch size and memory use against throughput. Measure both throughput and latency under the intended workload before increasing them.
Partitioning and Consumer Groups
Section titled “Partitioning and Consumer Groups”| TOML field | Default | Meaning |
|---|---|---|
runtime_seed.partitioning.default_partition_count | 1 | Partition count for a new queue declared without an explicit count. |
runtime_seed.consumer_groups.default_target_per_consumer | unset | Optional soft target for exclusive consumer groups. When set, cohorts above the target can be reported as under-provisioned without reducing coverage. |
Runtime Locks
Section titled “Runtime Locks”Runtime locks let startup config own a runtime setting group and prevent admin edits.
[runtime_locks]idle_queue_cleanup = trueWhen idle_queue_cleanup is locked:
- startup config controls the effective idle queue cleanup settings
- the admin settings page shows the group as locked
- admin update attempts for that group are rejected
- updates to other runtime settings can still be saved
Use locks only when config management should intentionally own the setting. Do not use ordinary env vars as hidden runtime overrides.
Admin Runtime Settings
Section titled “Admin Runtime Settings”The admin UI exposes a settings page at:
/admin/settingsThe JSON API is:
GET /admin/api/runtime-settingsPUT /admin/api/runtime-settingsUpdate requests include an expected_version. If another operator changed settings first, the API returns 409 Conflict with the current settings instead of overwriting them.
Node-local Storage Settings
Section titled “Node-local Storage Settings”The dashboard’s Node-local Storage section edits the node serving that page. Use that node’s dashboard address; this endpoint does not forward edits to another node. These overrides remain local even in a coordinated cluster.
GET /admin/api/local-storage-settingsPUT /admin/api/local-storage-settingsA PUT contains the node identity and revision returned by GET:
{ "node_id": "broker-1", "expected_version": 0, "segment_preallocate_bytes": 1048576}segment_preallocate_bytes sets the allocation chunk for message, event and
internal metadata logs. Zero disables preallocation; null removes the override
and follows storage.keratin.segment_preallocate_bytes from startup configuration.
The override is durably saved before publication and survives restart. Startup
file/environment/CLI changes remain the fallback until the override is reset.
A stale revision returns 409 Conflict; a different node identity is rejected.
Existing segments keep their policy. Each log adopts the latest revision when creating or reopening an active segment, including recovery-created logs. Saving does not force rollover or resize current segments, so a quiet log can remain pending indefinitely. Future logs use the latest revision immediately. The shared configuration is sampled at these boundaries, without synchronization per record.
The response reports the accepted revision, requested and startup values, pending log count, and each open log’s adopted revision and effective allocation chunk. Filesystem allocation is best-effort: failure emits a warning, reports the error and effective zero, and retains normal extending writes. A later segment retries allocation. Closed logs are absent from the inventory. Applied status describes policy adoption, not a change to publish durability or an all-logs transaction.
Fsync cadence, batching and adaptive-buffer settings remain startup-only. This endpoint has the same administrator authentication as other settings endpoints.
Other Persisted Runtime Settings
Section titled “Other Persisted Runtime Settings”Some live settings are owned by storage-level state rather than the broker runtime settings document. The global dead-letter queue target is the current example.
The admin settings page also exposes:
GET /admin/api/global-dlqPUT /admin/api/global-dlqThis setting:
- applies live
- is persisted in Fibril’s storage state
- survives restart
- uses
expected_versionand returns409 Conflictif another update wins - is not seeded or overridden by TOML, environment variables, or CLI flags
See dead lettering for the setting shape and current limitations.
Validation
Section titled “Validation”Current validation rules:
server.data_dirmust not be empty- when
admin.auth.enabled = true,admin.auth.usernameandadmin.auth.passwordmust both be set - when
tls.enabled = true, exactly one certificate source must be configured: bothtls.cert_pathandtls.key_path, ortls.auto_self_signed = true tls.admin_enabled = truerequirestls.enabled = truestorage.keratin.fsync_interval_msmust be at least1storage.keratin.writer_buffer_factormust be in1..=128- Staging decay must be positive; idle release must be at least the decay interval and representable on the monotonic clock
storage.keratin.message_log.segment_max_bytesmust be at least1storage.keratin.event_log.segment_max_bytesmust be at least1coordination.ganglion.heartbeat_interval_msmust be at least1coordination.ganglion.liveness_ttl_msmust be at least twicecoordination.ganglion.heartbeat_interval_msruntime_seed.delivery.expiry_batch_maxmust be at least1runtime_seed.idle_queue_cleanup.sweep_interval_msmust be at least1runtime_seed.replicationpoll intervals and worker limits must be at least1runtime_seed.replication.min_in_sync_replicasmust be at least1runtime_seed.replication.isr_timeout_msmust be at least1runtime_seed.partitioning.default_partition_countmust be at least1
More validation will be added as more settings become user-facing.