Redis Memory Configuration
Why this exists
Every Redis instance ran with maxmemory 0 (unlimited) and maxmemory-policy noeviction,
with both AOF and RDB persistence active. Observed 2026-08-28:
| prod | dev | staging | |
|---|---|---|---|
| used / peak | 174 MB / 295 MB | 841 MB / 11.36 GB | 510 MB / 639 MB |
| host RAM / swap | 32 GB / 8 GB | 16 GB / 0 | 3.8 GB / 4 GB (1.4 GB used) |
All figures measured 2026-08-28 morning. Staging's swap use is climbing over the day -- it reached 1.9 GB by midday, with roughly 75% of the redis dataset paged out. Re-measure rather than trusting these numbers if you are reading this later. | container restarts | 0 | 144 | 0 | | pubsub client disconnects | 4 | 12 | 618 | | fragmentation ratio | 1.09 | 1.38 | 0.40 (swapped) |
Dev's 11.36 GB peak on a 16 GB box with no swap, plus 144 restarts, is unbounded growth being resolved by the OOM killer. Staging took a cgroup OOM kill at 06:17 on 2026-08-28.
Settings and rationale
| Setting | Why |
|---|---|
--maxmemory | Hard ceiling so Redis cannot take the host down. Per-environment values and their derivation are in the table below. |
--maxmemory-policy noeviction | Redis deletes nothing early; on reaching the ceiling, writes fail loudly with OOM command not allowed while reads keep working. See "Why noeviction" below — volatile-lru was considered and rejected. |
--save "" | Disables RDB. AOF already provides durability; running both meant a fork roughly every 60s (dev: 17,896 forks, staging: 20,746). Each fork briefly duplicates written pages and adds a second process to the cgroup. |
--auto-aof-rewrite-min-size 256mb | Default 64mb is below prod's aof_base_size of 96 MB, causing needless rewrite forks. |
--lazyfree-lazy-* | Frees memory on a background thread. Dev expires 838k keys; each synchronous free blocked the single-threaded event loop. |
--activedefrag + thresholds | Dev wastes ~337 MB to fragmentation (ratio 1.38). Thresholds keep it dormant on healthy instances like prod. |
--client-output-buffer-limit "pubsub 64mb 16mb 120" | Staging disconnected pubsub clients 618 times, silently dropping messages across 22 channels. Trade-off: one stuck consumer may now buffer up to 64 MB. |
--timeout 300 / --tcp-keepalive 60 | Reap abandoned connections. Pubsub and blocked clients are exempt from timeout. |
mem_limit (compose) | Backstop outside Redis. Sized from fragmentation, not fork overhead -- see "Sizing mem_limit against maxmemory" below. |
CONFIG stays removed
CONFIG is renamed to "" (i.e. removed) alongside FLUSHALL, FLUSHDB, SLAVEOF,
REPLICAOF, DEBUG and MODULE.
An earlier revision of this change renamed CONFIG to a per-environment alias instead, to
make settings inspectable without recreating the container. That was reverted: the alias is
committed in plaintext, so it grants every holder of REDIS_PASSWORD -- which includes the
application itself -- the ability to run CONFIG SET maxmemory 0 and silently undo the
limits this document describes. Binding to 127.0.0.1 does not help, because the threat is
a compromised backend rather than a remote attacker.
The justification for the alias was weaker than it appeared. The three settings that matter
most are readable from INFO with no CONFIG at all:
| setting | where to read it |
|---|---|
maxmemory | INFO memory -> maxmemory: |
maxmemory-policy | INFO memory -> maxmemory_policy: |
| AOF on/off | INFO persistence -> aof_enabled: |
| RDB on/off | absence of Background saving started in docker logs |
The secondary tuning flags (appendfsync, timeout, lazyfree-*, activedefrag, the
pubsub buffer) are not exposed by INFO. Their values are the compose file, which is the
source of truth; confirm a change took effect by restarting and reading the flags back from
docker inspect.
If live CONFIG SET is ever genuinely needed, the correct mechanism is a Redis ACL split --
an application user with -@admin and a separate operator credential -- not a shared
password with a renamed command.
The local docker-compose.yml
docker-compose.yml (the local stack, container hannibal-redis) also gets maxmemory 768mb, noeviction, --save "" and the lazyfree flags. It deliberately does NOT get the
activedefrag, pubsub-buffer, timeout or tcp-keepalive tuning — those address production
symptoms that do not arise locally.
Note this file has no --requirepass, no --protected-mode and no --rename-command
directives at all, and it publishes 6379:6379 on all interfaces. That predates this
change and is unchanged by it, but it means the "renamed CONFIG" section below does not
apply to the local stack: CONFIG, FLUSHALL and the rest are all available there
without a password.
Applying
AOF persists the dataset, so a restart loses no data. Expect a short AOF replay on
boot (prod's aof_current_size was 186 MB when last measured, roughly 2-5s).
ENV=dev # dev | staging | prod
docker compose -f "docker-compose.${ENV}.yml" up -d redis
Roll out dev -> staging -> prod, one environment at a time.
Verify before moving to the next environment
CONTAINER=strykr-redis # strykr-staging-redis on staging
R="docker exec $CONTAINER redis-cli -a $REDIS_PASSWORD --no-auth-warning"
docker inspect "$CONTAINER" --format 'restarts={{.RestartCount}} oom={{.State.OOMKilled}}'
docker logs "$CONTAINER" --tail 30 # expect no errors, no restart loop
# CONFIG is removed, so read the applied settings back from INFO and the container spec.
$R INFO memory | grep -E '^maxmemory:|^maxmemory_policy:' # ceiling + noeviction
$R INFO persistence | grep -E '^aof_enabled:|^aof_last_bgrewrite_status:'
$R INFO memory | grep -E 'used_memory_human|used_memory_rss_human|mem_fragmentation_ratio'
$R INFO stats | grep -E '^evicted_keys:|^expired_keys:'
# RDB off: no save points are configured, so this must print nothing.
docker logs "$CONTAINER" 2>&1 | grep -c 'Background saving started'
# The flags redis does not expose (appendfsync, timeout, lazyfree-*, activedefrag,
# pubsub buffer) are verified from the container spec instead.
docker inspect "$CONTAINER" --format '{{json .Config.Cmd}}' | tr ',' '\n' | grep -E 'save|lazyfree|activedefrag|timeout|buffer'
# Expect, in order: --save "" --lazyfree-lazy-eviction yes --lazyfree-lazy-expire yes
# --lazyfree-lazy-server-del yes --activedefrag yes --active-defrag-ignore-bytes 200mb
# --active-defrag-threshold-lower 20 --client-output-buffer-limit "pubsub 64mb 16mb 120"
# --timeout 300
# A missing line means the compose change did not roll out, not that redis rejected it.
Expect maxmemory to match the compose value in bytes, maxmemory_policy:noeviction,
aof_enabled:1, aof_last_bgrewrite_status:ok, a 0 count for Background saving started
(RDB is off), evicted_keys:0 (nothing evicts under noeviction) and expired_keys still
climbing — that last one confirms TTL expiry is reclaiming memory normally.
Watch for OOM command not allowed in the backend logs over the following day. Under
noeviction that is the signal the ceiling was reached. Loki query:
{container="strykr-backend"} |= "OOM command not allowed"
Reference-key retention rollout
TheSports team and competition records are cache data. New writes carry a sliding
seven-day TTL and use one atomic value+expiry command. Match writes also refresh
existing referenced keys because the provider's best-effort reference bundle may
omit them on a later diary response. If an environment predates that writer, its
legacy thesports:team:* and thesports:competition:* keys remain persistent until
an operator drains them.
Deploy and verify the writer first. A legacy plain SET clears an existing TTL, so
backfilling before deployment only creates temporary progress. After deployment,
confirm recent team and competition keys have a TTL near seven days.
Legacy-key mutation is deferred and blocked until a dedicated, version-controlled
tool has been reviewed and tested. Do not execute a hand-written SCAN/EXPIRE pass
from this document, even with separate authorization. That tool must default to
dry-run, fix the retention TTL at 604800 seconds, enforce scan/pipeline bounds and
rate limits, support safe resume/re-run, and account for seen, eligible, stamped,
declined, vanished, and errored results. Its apply path must use EXPIRE ... NX
only for the two reference prefixes and must never delete keys. Writer verification
and separate human authorization remain prerequisites once that tool exists.
Until then, dormant legacy keys retain their existing no-TTL state; this writer change alone does not reclaim them. Measure their remaining footprint read-only. Any seven-day legacy-drain observation window starts only after the reviewed tool is separately authorized and applied.
This is separate from backend/scripts/expire-market-meta-backlog.ts. That script
protects order-referenced market:* records and drains only the pre-ADR-0031 market
backlog; it does not cover TheSports reference keys.
Alerting
infra/monitoring/prometheus/alerts/redis.yml carries the rules. Three points:
RedisMemoryCriticalandRedisMemoryHighwere firing permanently on prod before this change. Their expression divides byredis_memory_max_bytes, which is0when maxmemory is unset; Prometheus evaluates that to+Inf, and+Inf > 95is always true. Four always-on redis alerts is why unbounded growth went unnoticed. Settingmaxmemoryrepairs them, and aredis_memory_max_bytes > 0guard now prevents a recurrence.RedisMaxmemoryUnsetreports a missing ceiling directly, instead of it surfacing as a nonsense percentage.RedisContainerOOMKilled/RedisContainerNearMemLimitwatch the cgroup limit via cadvisor. Redis refusing writes is the intended loud failure; a container kill is the failure mode that replaces it if RSS crossesmem_limitfirst, and it needs its own signal.RedisRDBBackupFailedis now gated onredis_aof_enabled == 0. Its expression istime() - redis_rdb_last_save_timestamp_seconds > 86400; once--save ""ships, no RDB save ever happens, the timestamp freezes, and the rule would fire permanently 24h later -- recreating the very problem this change removes.RedisAOFWriteFailedandRedisAOFRewriteFailedreplace that coverage for AOF-only instances.
In every guarded rule the percentage expression is the LEFT operand. PromQL and returns
the left vector's samples, so putting the guard first would make {{ $value }} render the
guard's raw byte count as a percentage. Verified on live data: ratio > 0 and metric > 0
returns the ratio, whereas metric > 0 and ratio > 0 returns the metric.
These rules DO ship, but through a separate path from the application containers.
Prometheus is defined in infra/docker/docker-compose.monitoring.yml, a stack that
docker-compose.<env>.yml does not include — so the alert rules are not carried by the
image build or the compose up. They are reconciled instead: deploy-prod.yml,
deploy-dev.yml and deploy-staging.yml each call scripts/reconcile-monitoring.sh,
which mounts infra/monitoring/prometheus/alerts into a config check, recreates
Prometheus, and then POSTs /-/reload. So a rule change lands on the next deploy of
those three environments and needs no manual step.
Two caveats. deploy-dev1.yml and deploy-dev2.yml do NOT reconcile monitoring — they
share the dev host's stack, so a rule change reaches them only via a dev deploy. And a
hand-run docker compose -f docker-compose.<env>.yml up bypasses reconciliation
entirely; reload Prometheus explicitly after one.
Sizing mem_limit against maxmemory
maxmemory bounds used_memory -- Redis's accounting of the dataset. mem_limit bounds
the container's RSS, which is what the kernel sees. The gap between them is what the
headroom must cover, and it is dominated by fragmentation, not by fork copy-on-write.
Measured 2026-08-28:
| used_memory | RSS | frag ratio | AOF rewrite CoW | |
|---|---|---|---|---|
| prod | 189 MB | 203 MB | 1.07 | 9.2 MB |
| dev | 889 MB | 1221 MB | 1.38 | 3.9 MB |
| staging | 536 MB | 227 MB | 0.42 (swapped out) | 6.5 MB |
Copy-on-write is 1-5% of the dataset, not the ~2x worst case often quoted: only pages written during the rewrite are duplicated, and these write rates are low relative to rewrite duration. Fragmentation is the real consumer -- dev holds 37% more RSS than data.
So the headroom rule is mem_limit >= ~1.5x maxmemory, driven by observed
fragmentation of up to 1.38 plus a margin, with fork CoW a rounding error on top.
Configured values
| env | maxmemory | mem_limit | memswap_limit | host | capped total |
|---|---|---|---|---|---|
| prod | 4gb | 6g | 6g | 32 GB | backend 12g + redis 6g = 18g |
| dev | 2gb | 3g | 3g | 16 GB, no swap | backend 3g + dev1 2g + redis 3g = 8g |
| staging | 700mb | 1g | 1g | 4 GB | backend 2g + redis 1g = 3g |
| local | 768mb | 1g | (unset) | n/a | n/a |
prod's 4gb is ~13x its 295 MB peak -- an alarm threshold, not a working constraint. staging's 700mb is a real working ceiling against a 639 MB peak; it is the low-resource forcing function and is meant to hit limits first. dev's 2gb is ~2.3x current usage and caps the 11.36 GB runaway.
Staging note: the host reports 4 GB (MemTotal 4009864 kB) and these values are sized for
that. It is currently ~1.9 GB into swap with roughly 75% of the redis dataset paged out
(mem_fragmentation_ratio 0.25, which indicates swap rather than fragmentation). Setting
memswap_limit == mem_limit stops that silent degradation by making the ceiling hard.
This matters because the AOF-rewrite child shares the container's cgroup, so parent and
child RSS are counted together. If RSS crosses mem_limit, the kernel OOM-kills the
container rather than Redis returning OOM command not allowed -- which would defeat the
loud-failure behaviour noeviction was chosen for. --activedefrag (enabled here) works
against this by reclaiming fragmented pages in place.
Reproducing the measurements
CONTAINER=strykr-redis # strykr-staging-redis on staging
R="docker exec $CONTAINER redis-cli -a $REDIS_PASSWORD --no-auth-warning"
$R INFO memory # used_memory, peak, fragmentation
$R INFO stats # evicted_keys, expired_keys, pubsub disconnects
$R INFO keyspace # key counts and how many carry TTLs
$R INFO commandstats # per-command call counts and CPU time
$R SLOWLOG GET 10
Known follow-ups (not fixed here)
vm.overcommit_memory=0on all three hosts. Redis warns this can cause background saves to fail, and failures even without memory pressure. Should be1(sysctl vm.overcommit_memory=1+/etc/sysctl.conf). Host-level change.- dev has zero swap on a 16 GB box — no cushion for fork spikes.
- Keyspace SCAN sweeps. Prod has spent 1,954 CPU-seconds across 16.4M
SCANcalls — more than all other commands combined, at only 32 ops/s. Application code is walking the keyspace withMATCHinstead of maintaining index sets. No Redis setting fixes this.