Skip to content

Configuration

Fibril has two configuration layers:

  • Startup config decides how the server process starts.
  • Runtime settings decide live broker behavior and are persisted after first boot.

That split matters. Startup config is for things the process needs before it can run, such as bind addresses and the data directory. Runtime settings are for behavior that can be changed while the broker is running, such as delivery timing and idle queue cleanup.

The config file format is TOML. The repository includes fibril.example.toml:

[server]
data_dir = "server_data"
[broker.listener]
bind = "0.0.0.0:9876"
[admin.listener]
bind = "0.0.0.0:8081"
[admin.auth]
enabled = false
username = "fibril"
# password = "change-me"
[setup]
mode = false
[tls]
enabled = false
# Supply your own PEM files:
# cert_path = "/etc/fibril/tls/server.pem"
# key_path = "/etc/fibril/tls/server.key"
# Or generate per-deployment material under <data_dir>/tls on first boot:
# auto_self_signed = true
# The admin dashboard serves HTTPS from the same material when TLS is
# enabled. Opt it out when a reverse proxy terminates TLS for the dashboard:
# admin_enabled = false
# Follower-to-owner replication follows `enabled` too. Opt it out when a
# service mesh or tunnel already encrypts inter-broker traffic:
# inter_broker = false
# CA that peer certificates chain to. Unset falls back to the generated
# <data_dir>/tls/ca.pem, then OS roots:
# peer_ca_path = "/etc/fibril/tls/ca.pem"
[storage.keratin]
fsync_interval_ms = 5
# Floor between storage commits while the fsync worker is idle. 0 self-clocks
# group commit on fsync completions, the best setting for fast storage (NVMe,
# tmpfs). On slow-fsync storage such as SATA SSDs a floor around the fsync
# interval (5) gives the drive breathing room between write barriers.
min_fsync_interval_ms = 0
# Writer and notification channels each have 64 * factor slots per log.
# 128 preserves the default 8192 slots; smaller values reduce eager allocation.
writer_buffer_factor = 128
adaptive_staging = true
staging_decay_secs = 10
staging_idle_release_secs = 60
[storage.keratin.message_log]
segment_max_bytes = 268435456
[storage.keratin.event_log]
segment_max_bytes = 33554432
[runtime_seed.delivery]
inflight_ttl_ms = 30000
expiry_poll_min_ms = 15000
expiry_batch_max = 8192
delivery_poll_max_ms = 5000
[runtime_seed.idle_queue_cleanup]
enabled = false
evict_after_ms = 600000
sweep_interval_ms = 60000
# Set this for sparse workloads with long-lived publishing connections.
# publisher_idle_timeout_ms = 600000
[runtime_seed.connection]
# Reconnect grace is on by default (5000 ms). Set to 0 to disable it.
# reconnect_grace_ms = 5000
# How stale a persisted session skeleton a restarted broker will honor, so a
# fast restart lets clients resume and reconcile (60000 ms). Set to 0 to
# disable restart resume.
# resume_session_restart_ttl_ms = 60000
[runtime_seed.replication]
confirm_timeout_ms = 5000
caught_up_poll_ms = 1000
retry_poll_ms = 100
checkpoint_retry_poll_ms = 5000
max_messages_per_read = 2048
max_events_per_read = 2048
max_bytes_per_read = 8388608
max_iterations_per_tick = 8
min_in_sync_replicas = 1
isr_timeout_ms = 10000
read_timeout_slack_ms = 10000
owner_connect_timeout_ms = 5000
[runtime_seed.partitioning]
default_partition_count = 1
[runtime_seed.consumer_groups]
# Blank or omit to disable the under-provisioned signal.
# default_target_per_consumer = 4
[coordination.ganglion]
heartbeat_interval_ms = 3000
liveness_ttl_ms = 9000
[runtime_locks]
idle_queue_cleanup = false

Run with a config file:

Terminal window
cargo run --release --bin fibril-server -- --config fibril.toml

or:

Terminal window
FIBRIL_CONFIG=fibril.toml cargo run --release --bin fibril-server

Startup config is resolved in this order:

compiled defaults < TOML config file < environment variables < CLI arguments

This precedence applies to startup fields and first-boot runtime seeds. It does not mean environment variables keep overriding persisted runtime settings after runtime state exists.

These fields are read on process start.

TOML fieldEnv varCLI flagDefault
server.data_dirFIBRIL_DATA_DIR--data-dirserver_data
broker.listener.bindFIBRIL_BROKER_BIND--broker-bind0.0.0.0:9876
broker.listener.advertiseFIBRIL_BROKER_ADVERTISEnonederived (see below)
admin.listener.bindFIBRIL_ADMIN_BIND--admin-bind0.0.0.0:8081
admin.listener.advertiseFIBRIL_ADMIN_ADVERTISEnonederived (see below)
admin.auth.enabledFIBRIL_ADMIN_AUTH_ENABLED--admin-auth-enabledfalse
admin.auth.usernameFIBRIL_ADMIN_USERNAME--admin-usernamefibril
admin.auth.passwordFIBRIL_ADMIN_PASSWORD--admin-passwordunset
admin.metrics_per_channelFIBRIL_ADMIN_METRICS_PER_CHANNELnonetrue
tls.enabledFIBRIL_TLS_ENABLEDnonefalse
tls.cert_pathFIBRIL_TLS_CERT_PATHnoneunset
tls.key_pathFIBRIL_TLS_KEY_PATHnoneunset
tls.auto_self_signedFIBRIL_TLS_AUTO_SELF_SIGNEDnonefalse
tls.admin_enabledFIBRIL_TLS_ADMIN_ENABLEDnonefollows tls.enabled
tls.inter_brokerFIBRIL_TLS_INTER_BROKERnonefollows tls.enabled
tls.peer_ca_pathFIBRIL_TLS_PEER_CA_PATHnonegenerated CA, then OS roots
tls.client_authFIBRIL_TLS_CLIENT_AUTHnoneoff
tls.client_ca_pathFIBRIL_TLS_CLIENT_CA_PATHnonegenerated CA
setup.modeFIBRIL_SETUP_MODEnonefalse
auth.allow_default_loopbacknonenonetrue
auth.seed_usersFIBRIL_AUTH_USERNAME + FIBRIL_AUTH_PASSWORD (one entry)noneempty
coordination.secret_pathFIBRIL_CLUSTER_SECRET_PATH (or FIBRIL_CLUSTER_SECRET for the value)noneunset
storage.keratin.fsync_interval_msFIBRIL_KERATIN_FSYNC_INTERVAL_MS--keratin-fsync-interval-ms5
storage.keratin.min_fsync_interval_msFIBRIL_KERATIN_MIN_FSYNC_INTERVAL_MS--keratin-min-fsync-interval-ms0
storage.keratin.writer_buffer_factorFIBRIL_KERATIN_WRITER_BUFFER_FACTORnone128
storage.keratin.adaptive_stagingFIBRIL_KERATIN_ADAPTIVE_STAGINGnonetrue
storage.keratin.staging_decay_secsFIBRIL_KERATIN_STAGING_DECAY_SECSnone10
storage.keratin.staging_idle_release_secsFIBRIL_KERATIN_STAGING_IDLE_RELEASE_SECSnone60
storage.keratin.message_log.segment_max_bytesFIBRIL_KERATIN_MESSAGE_LOG_SEGMENT_MAX_BYTES--keratin-message-log-segment-max-bytes268435456
storage.keratin.event_log.segment_max_bytesFIBRIL_KERATIN_EVENT_LOG_SEGMENT_MAX_BYTES--keratin-event-log-segment-max-bytes33554432
coordination.modeFIBRIL_COORDINATION_MODEnonestatic
coordination.ganglion.heartbeat_interval_msFIBRIL_COORDINATION_HEARTBEAT_INTERVAL_MSnone3000
coordination.ganglion.liveness_ttl_msFIBRIL_COORDINATION_LIVENESS_TTL_MSnone9000
coordination.ganglion.target_followersnonenone1
coordination.ganglion.stream_replication_factornonenone1
coordination.ganglion.repartition_adoption_timeout_msnonenone30000
coordination.ganglion.assignment_durabilityFIBRIL_COORDINATION_ASSIGNMENT_DURABILITYnonelocal_durable
recovery.on_mismatchFIBRIL_RECOVERY_ON_MISMATCHnonequarantine

Changing these generally requires restarting the server.

broker.listener.advertise is the address (or addresses) the broker tells peers and clients to reach it on, separate from bind. This matters when bind is not itself dialable - the common case is binding 0.0.0.0 in a container, which a peer cannot connect back to. Give it a routable host:port (a service name is fine, it is resolved at connect time), or several comma-separated entries in FIBRIL_BROKER_ADVERTISE in priority order. When unset it is derived in ganglion mode from this node’s coordination peer host plus the broker port, and otherwise falls back to bind. Standalone single-broker deployments do not need it (clients connect to the broker directly). Only the first entry is dialed today; the rest are carried for forward compatibility.

admin.listener.advertise is the same idea for the dashboard: the address a node registers with the cluster so the Cluster page’s “open its admin” links and the broker switcher can point a browser at it. When unset and the admin bind host is unspecified (0.0.0.0), the node reuses the broker advertise host with the admin port; a raw unspecified address is never linked (the dashboard shows “no reachable admin address registered” instead). In the Docker cluster example each node advertises its host-mapped admin port.

coordination.mode is static for a standalone single-broker deployment (the default) or ganglion to run the embedded coordinator and form a cluster. The coordination.ganglion.* settings only apply in ganglion mode. See clustering and replication.

coordination.ganglion.target_followers is the desired follower count per queue partition. coordination.ganglion.stream_replication_factor is the equivalent for DURABLE Plexus stream partitions — tuned separately so stream and queue fault tolerance can differ; only the durable tier replicates, the express tiers stay owner-only. A value of one keeps a durable stream available across a single node loss; zero makes durable streams owner-only (durable on disk, not HA). See Plexus streams. coordination.ganglion.assignment_durability is the default durability policy for new assignments (local_durable, replica_accepted, replica_durable, or majority_durable).

recovery.on_mismatch controls what happens when recovery finds a damaged queue log: quarantine (default) isolates the partition, refuse reports not ready, and ignore truncates to the last valid record. See recovery quarantine.

admin.auth.enabled = true requires both admin.auth.username and admin.auth.password. The admin password is intentionally not shown in the dashboard startup summary.

admin.metrics_per_channel controls whether the Prometheus /metrics endpoint on the admin listener includes per-channel series (queue depth, stream subscriptions, follower applied state) alongside the always-exported node-level aggregates. See monitoring.

tls.enabled = true serves TLS on the broker listener and the admin dashboard from one certificate. Supply PEM files with tls.cert_path and tls.key_path, or set tls.auto_self_signed = true to generate a per-deployment CA and server certificate under <data_dir>/tls on first boot. The CA fingerprint is printed at startup so clients can pin or trust it - self-signed material a client does not verify defeats passive snooping only. Fibril never ships certificates. tls.admin_enabled = false keeps the dashboard on plain HTTP for deployments where a reverse proxy terminates TLS in front of it. Certificate material is read once at startup, so replacing it requires a restart (live reload and rotation are planned).

tls.inter_broker covers follower-to-owner replication and the coordination raft channel. It follows tls.enabled, which assumes a homogeneous cluster whose certificates chain to one CA - see TLS across nodes for the shared-CA lane and the rotation runbook. Set it to false when a service mesh or tunnel already encrypts inter-broker traffic. tls.peer_ca_path names the CA peers are verified against, falling back to the generated <data_dir>/tls/ca.pem when present, then OS roots. The serving certificate and key reload live via fibrilctl admin reload-tls or POST /admin/api/tls/reload.

tls.client_auth turns client certificates into credentials: request verifies a presented certificate (certless clients still connect and password-auth, the migration lane), require rejects certless clients in the handshake. A verified certificate whose identity (first DNS SAN, else CN) names an existing user authenticates as that user with no password; the @ node namespace can never be claimed by certificate. tls.client_ca_path names the CA client certificates chain to, falling back to the generated CA - issue workload certificates from it with fibrilctl cert issue <identity>. Brokers present their own certificate on inter-broker dials, so a cluster converges at any client_auth setting.

The auth section governs broker authentication. The built-in default credentials (fibril/fibril) are accepted from loopback connections only, so local development works out of the box while remote access always requires a real user. auth.seed_users creates users when the user store is empty (first boot only, after which the persisted store owns the users), and the FIBRIL_AUTH_USERNAME/FIBRIL_AUTH_PASSWORD pair seeds one user from the environment. A real user named fibril replaces the built-in pair entirely. Passwords are stored as argon2 hashes. With users configured and TLS off, the server warns loudly: passwords travel in cleartext on non-loopback plaintext connections.

The cluster shared secret authenticates node-to-node connections (replication) as a node principal, never as a user account, and is required in ganglion mode. Resolution order: FIBRIL_CLUSTER_SECRET (the value itself), then coordination.secret_path, then <data_dir>/cluster.secret if present - which is what fibrilctl secret generate writes. Every node holds the same secret.

setup.mode = true arms first-boot setup: when the data dir holds no setup_complete marker, the server serves only a setup page on 127.0.0.1:<admin port> and the broker listener stays down until the operator chooses generated TLS material, supplies a certificate, or explicitly continues without TLS. The choice is written to <data_dir>/config-overlay.toml, which boot layers below explicit config, so anything set in the tls section by file, environment, or CLI always wins over a setup choice. Completion writes the marker and the broker starts in the same process. Deleting the marker and booting with setup mode runs setup again. Deployments that configure TLS explicitly never need setup mode, and with it armed they boot straight through (the marker is written automatically).

storage.keratin.message_log.segment_max_bytes and storage.keratin.event_log.segment_max_bytes are rollover thresholds. A segment rolls after an append crosses the configured size, so an individual segment can be slightly larger than this value.

storage.keratin.writer_buffer_factor sets the startup capacity of each storage writer input and notification channel to 64 × factor entries. Valid factors are 1..=128: 1 gives 64 entries, 16 gives 1,024, and the default 128 gives 8,192. The channels allocate their slot arrays when a log opens, so smaller values reduce the resident cost of materialized queues and streams and apply backpressure earlier during bursts. A channel entry may own a batch; this setting is a slot limit rather than a total memory limit. Measure throughput and tail latency with representative payloads, storage and replication before choosing a lower value. Changing it requires restart and applies to both the message and event logs. Fsync pipeline depth remains controlled by storage.keratin.max_inflight_fsyncs; caches, actor mailboxes and client buffers have separate limits.

storage.keratin.adaptive_staging = true (the default) enables lazy staging allocations for both message and event logs. Write buffers start empty with a 64 KiB reservation floor; sparse-index buffers use a 4 KiB floor. They grow to fit batches, may halve while empty every staging_decay_secs (default 10), and release fully after staging_idle_release_secs without use (default 60). Decay preserves headroom for recent batches and never discards staged records. For less frequent bursts, longer retention delays reduce repeated allocation work.

Set adaptive_staging = false to retain eager 16 MiB write and 256 KiB index reservations per log. These sizes and adaptive floors are initial reservations, not maximum capacities. This startup-only setting changes allocation retention; channel sizes, caches and durability policy remain independently configured. Released capacity may remain resident in the allocator, so RSS can stay high or even increase for some burst patterns. Evaluate memory, CPU and first-message latency after idle when choosing retention delays or opting out.

coordination.ganglion.heartbeat_interval_ms controls how often a broker renews its cluster liveness record. coordination.ganglion.liveness_ttl_ms controls how long a broker can go without a fresh heartbeat before the cluster considers it unavailable. The TTL must be at least twice the heartbeat interval. For heavy replication benchmarks, a longer TTL can avoid false failover while the node is under artificial load.

coordination.ganglion.repartition_adoption_timeout_ms bounds how long a live repartition’s finalize (retiring shrunk-away partitions and clearing the transition marker) waits for clients to adopt the new routing once the backlog has drained. Adoption is observed from client topology acks. The timeout keeps a silent or stuck client from stalling a cutover forever; publish version-fencing is the correctness backstop regardless. See live routing and cutover.

The Linux broker uses mimalloc, which exposes a startup environment option for transparent huge pages (THP). Consider disabling THP when resident memory is a deployment constraint, especially with many materialized queues whose log buffers are lightly used. Huge pages can make a sparsely touched allocation consume substantially more resident RAM; disabling them can also reduce the footprint of an active broker.

Terminal window
MIMALLOC_ALLOW_THP=0 ./fibril-server

Append the usual broker arguments. Set the same variable in a container’s environment or a systemd service’s Environment=MIMALLOC_ALLOW_THP=0 directive, then restart or recreate the broker. This allocator setting is read at process startup and has no Fibril TOML field or admin runtime setting. It disables THP for the broker process and its descendants without changing the host policy for other services. Remove the variable and restart to restore the allocator’s default behavior under the host’s policy.

The tradeoff depends on workload: huge pages can improve address-translation efficiency for large active working sets, while allocating and clearing larger pages can add latency and consume extra memory.

After restart, /proc/<broker-pid>/status should report THP_enabled: 0, and AnonHugePages in /proc/<broker-pid>/smaps_rollup should be zero. These checks confirm that the process policy took effect, including in restricted containers. THP policy is independent of the writer channel factor and idle queue cleanup; each addresses a different part of the materialized-resource footprint. See the memory investigation notes and the upstream mimalloc options and Linux THP documentation.

runtime_seed values initialize the persisted runtime settings document when no runtime settings exist yet.

After runtime settings exist, the persisted values own these settings. You can edit them through the admin settings page or the admin runtime settings API.

TOML fieldDefaultMeaning
runtime_seed.delivery.inflight_ttl_ms30000How long a delivered message lease lasts before it can be retried.
runtime_seed.delivery.expiry_poll_min_ms15000Minimum sleep between expiry checks when no earlier expiry is known.
runtime_seed.delivery.expiry_batch_max8192Maximum expired messages to requeue in one expiry pass. Must be at least 1.
runtime_seed.delivery.delivery_poll_max_ms5000Maximum idle poll delay for delivery loops.
TOML fieldEnv/CLI compatibilityDefaultMeaning
runtime_seed.idle_queue_cleanup.enabledenabled implicitly by FIBRIL_QUEUE_IDLE_EVICT_AFTER_MS or --queue-idle-evict-after-msfalseEnables unloading idle queues from memory.
runtime_seed.idle_queue_cleanup.evict_after_msFIBRIL_QUEUE_IDLE_EVICT_AFTER_MS, --queue-idle-evict-after-ms600000How long a queue must be idle before cleanup can unload it.
runtime_seed.idle_queue_cleanup.sweep_interval_msFIBRIL_QUEUE_IDLE_SWEEP_INTERVAL_MS, --queue-idle-sweep-interval-ms60000How often the cleanup worker checks tracked queues. Must be at least 1.
runtime_seed.idle_queue_cleanup.publisher_idle_timeout_msFIBRIL_PUBLISHER_CACHE_IDLE_TIMEOUT_MS, --publisher-idle-timeout-msunsetLets unused publishers stop keeping queues active while a connection remains open.

For sparse workloads, enable publisher idle expiry alongside queue cleanup. Without it, a long-lived connection that published to a queue can keep that queue active until the connection closes.

See many idle queues for the user-facing behavior.

TOML fieldEnv/CLI compatibilityDefaultMeaning
runtime_seed.connection.reconnect_grace_msFIBRIL_RECONNECT_GRACE_MS, --reconnect-grace-ms5000Keeps a disconnected resumable client alive for this long before cleaning up subscriptions and requeueing unsettled messages. On by default so a transient blip resumes transparently; set 0 to disable.
runtime_seed.connection.resume_session_restart_ttl_msFIBRIL_RESUME_SESSION_RESTART_TTL_MS, --resume-session-restart-ttl-ms60000How stale a persisted session skeleton a restarted broker will honor. Within this window a resume after a broker restart succeeds and reconciles instead of being rejected; past it (or 0) it reports not-found. Independent of reconnect_grace_ms, which governs live-process dormancy.
connection.drain_handoff_timeout_ms (runtime settings)none30000Upper bound on how long a drain call waits for partition ownership to hand off to caught-up followers in coordinated mode. Only caps the wait: reactive failover stays the backstop either way. Adjustable live from the settings page.

Reconnect grace is disabled when unset. It only helps clients that use the resume identity handshake.

These settings apply to the experimental cluster replication path.

TOML fieldDefaultMeaning
runtime_seed.replication.confirm_timeout_ms5000How long a replica-durable publish confirm can wait for enough durable follower progress.
runtime_seed.replication.agreed_checkpoint_interval_ms0Opt-in agreed queue checkpoints; 0 disables new attempts, otherwise 1000–86400000 ms between periodic attempts. Work thresholds can start earlier, at most once per second. Existing attempts finish and the last accepted base/suffix remain retained. All admitted replicas must agree. See checkpoint policy.
runtime_seed.replication.agreed_checkpoint_max_events0Start an enabled checkpoint earlier after this many events beyond its agreed cut; 0 disables the event trigger.
runtime_seed.replication.agreed_checkpoint_max_bytes0Start earlier after approximately this many local append content bytes since observing the last certificate; 0 disables. The counter resets on reopen and excludes record framing; the periodic interval remains the fallback.
runtime_seed.replication.eager_failoverfalseOpt-in explicit peer-connection failure detection on the active metadata controller. Recovery proof is still required; see eager failover.
runtime_seed.replication.eager_failover_grace_ms1000Failed-reconnect grace, 0–60000 ms. Zero enables idle-disconnect monitoring and immediate reconnect verification on the active controller. Positive values require failed contact across the interval; heartbeat expiry handles silent failures.
runtime_seed.replication.caught_up_poll_ms1000Caught-up pull interval and streaming owner long-poll timeout. Streaming reads wake when new work arrives, without waiting for the timeout.
runtime_seed.replication.retry_poll_ms100Follower retry interval after a partial pull or transient replication error.
runtime_seed.replication.checkpoint_retry_poll_ms5000Follower retry interval while it needs an owner checkpoint before it can continue.
runtime_seed.replication.max_messages_per_read2048Maximum message records per owner read, for pull and streaming.
runtime_seed.replication.max_events_per_read2048Maximum event records per owner read, for pull and streaming.
runtime_seed.replication.max_bytes_per_read8388608Approximate byte budget per owner log scan and initial byte credit for a new follower stream. An oversized record and accompanying events can make the response exceed this budget.
runtime_seed.replication.max_iterations_per_tick8Maximum pull/apply iterations a follower performs before yielding.
runtime_seed.replication.min_in_sync_replicas1Minimum recently in-sync replicas required before accepting replica-durable publishes. 1 disables the floor.
runtime_seed.replication.isr_timeout_ms10000How recently a follower must report durable progress to count as in sync.
runtime_seed.replication.read_timeout_slack_ms10000Slack added to a follower read’s long-poll window before the read is abandoned and the connection dropped, so a dropped owner response cannot hang the follower. A read waits max_wait_ms + this.
runtime_seed.replication.owner_connect_timeout_ms5000Upper bound on a follower establishing a connection to an owner (TCP connect plus the handshake) before it is abandoned and retried.
runtime_seed.replication.stream_enabledtrueUse credit-based streaming replication on followers. When disabled, followers fall back to polling pulls.
runtime_seed.replication.stream_apply_linger_us2000Microseconds a streaming follower gathers contiguous frames before one fsynced apply. Higher trades apply latency for fsync amortization. 0 is drain-only.
runtime_seed.replication.stream_apply_max_merge_bytes16777216Byte threshold for a coalesced streaming apply. Whole frames can cross it. Trades peak memory and overlap against fsync amortization.
runtime_seed.replication.stream_buffer_batches8Queued frame count for the streaming follower, separate from byte credit. Applied on the next stream.

Streaming owner reads adopt updated record and byte limits at the next batch. The initial credit window changes when a new follower stream starts. Oversized batches consume that window and retain any excess as debt until the follower returns credit after durable application.

Read budgets trade batch size and memory use against throughput. Measure both throughput and latency under the intended workload before increasing them.

TOML fieldDefaultMeaning
runtime_seed.partitioning.default_partition_count1Partition count for a new queue declared without an explicit count.
runtime_seed.consumer_groups.default_target_per_consumerunsetOptional soft target for exclusive consumer groups. When set, cohorts above the target can be reported as under-provisioned without reducing coverage.

Runtime locks let startup config own a runtime setting group and prevent admin edits.

[runtime_locks]
idle_queue_cleanup = true

When idle_queue_cleanup is locked:

  • startup config controls the effective idle queue cleanup settings
  • the admin settings page shows the group as locked
  • admin update attempts for that group are rejected
  • updates to other runtime settings can still be saved

Use locks only when config management should intentionally own the setting. Do not use ordinary env vars as hidden runtime overrides.

The admin UI exposes a settings page at:

/admin/settings

The JSON API is:

GET /admin/api/runtime-settings
PUT /admin/api/runtime-settings

Update requests include an expected_version. If another operator changed settings first, the API returns 409 Conflict with the current settings instead of overwriting them.

The dashboard’s Node-local Storage section edits the node serving that page. Use that node’s dashboard address; this endpoint does not forward edits to another node. These overrides remain local even in a coordinated cluster.

GET /admin/api/local-storage-settings
PUT /admin/api/local-storage-settings

A PUT contains the node identity and revision returned by GET:

{
"node_id": "broker-1",
"expected_version": 0,
"segment_preallocate_bytes": 1048576
}

segment_preallocate_bytes sets the allocation chunk for message, event and internal metadata logs. Zero disables preallocation; null removes the override and follows storage.keratin.segment_preallocate_bytes from startup configuration. The override is durably saved before publication and survives restart. Startup file/environment/CLI changes remain the fallback until the override is reset. A stale revision returns 409 Conflict; a different node identity is rejected.

Existing segments keep their policy. Each log adopts the latest revision when creating or reopening an active segment, including recovery-created logs. Saving does not force rollover or resize current segments, so a quiet log can remain pending indefinitely. Future logs use the latest revision immediately. The shared configuration is sampled at these boundaries, without synchronization per record.

The response reports the accepted revision, requested and startup values, pending log count, and each open log’s adopted revision and effective allocation chunk. Filesystem allocation is best-effort: failure emits a warning, reports the error and effective zero, and retains normal extending writes. A later segment retries allocation. Closed logs are absent from the inventory. Applied status describes policy adoption, not a change to publish durability or an all-logs transaction.

Fsync cadence, batching and adaptive-buffer settings remain startup-only. This endpoint has the same administrator authentication as other settings endpoints.

Some live settings are owned by storage-level state rather than the broker runtime settings document. The global dead-letter queue target is the current example.

The admin settings page also exposes:

GET /admin/api/global-dlq
PUT /admin/api/global-dlq

This setting:

  • applies live
  • is persisted in Fibril’s storage state
  • survives restart
  • uses expected_version and returns 409 Conflict if another update wins
  • is not seeded or overridden by TOML, environment variables, or CLI flags

See dead lettering for the setting shape and current limitations.

Current validation rules:

  • server.data_dir must not be empty
  • when admin.auth.enabled = true, admin.auth.username and admin.auth.password must both be set
  • when tls.enabled = true, exactly one certificate source must be configured: both tls.cert_path and tls.key_path, or tls.auto_self_signed = true
  • tls.admin_enabled = true requires tls.enabled = true
  • storage.keratin.fsync_interval_ms must be at least 1
  • storage.keratin.writer_buffer_factor must be in 1..=128
  • Staging decay must be positive; idle release must be at least the decay interval and representable on the monotonic clock
  • storage.keratin.message_log.segment_max_bytes must be at least 1
  • storage.keratin.event_log.segment_max_bytes must be at least 1
  • coordination.ganglion.heartbeat_interval_ms must be at least 1
  • coordination.ganglion.liveness_ttl_ms must be at least twice coordination.ganglion.heartbeat_interval_ms
  • runtime_seed.delivery.expiry_batch_max must be at least 1
  • runtime_seed.idle_queue_cleanup.sweep_interval_ms must be at least 1
  • runtime_seed.replication poll intervals and worker limits must be at least 1
  • runtime_seed.replication.min_in_sync_replicas must be at least 1
  • runtime_seed.replication.isr_timeout_ms must be at least 1
  • runtime_seed.partitioning.default_partition_count must be at least 1

More validation will be added as more settings become user-facing.