Application Observability
Hannibal phase-1 operational telemetry uses the existing self-hosted
Prometheus, Grafana, Alertmanager, Loki, and infrastructure exporters. The
backend exposes application and Node.js runtime metrics at GET /metrics on
the dedicated Docker-internal port 9464. The product HTTP listener on port
3001 does not serve metrics.
This phase does not include OpenTelemetry traces, continuous profiling, frontend browser telemetry, or AI-service metrics.
Backend metrics
HTTP RED metrics
| Metric | Type | Purpose |
|---|---|---|
http_requests_total | Counter | Completed requests by method, route template, and status |
http_request_duration_seconds | Histogram | End-to-end Express request duration |
http_requests_in_progress | Gauge | Requests currently executing |
http_request_aborted_total | Counter | Connections closed before completion |
http_response_size_bytes | Histogram | Response content length when known |
page_api_stage_duration_seconds | Histogram | Internal stages for selected page and agent read APIs |
HTTP metrics use these bounded labels:
service="hannibal-backend"environment— deployment identity fromENVIRONMENT(dev,dev1,dev2, orproduction), falling back toNODE_ENVonly when unsetmethod—GET,POST,PUT,PATCH,DELETE,OPTIONS,HEAD, orOTHER; unsupported methods are coalesced to keep cardinality boundedroute— a literal router mount plus the Express route templatestatus— the final HTTP status code emitted by Express
Raw URLs are never labels. Unmatched requests use route="unmatched", and
array/regular-expression routes use route="complex-route". User IDs,
usernames, order IDs, fixture IDs, provider account identifiers, request
bodies, credentials, cookies, and financial values must never be metric labels.
HTTP recording starts before CORS and body parsing, so malformed, oversized,
and otherwise early-rejected API requests remain observable. The recorder only
accepts /api paths, while scrapes use a separate HTTP server, so operational
traffic and static assets do not inflate product request-rate or latency
metrics.
Status instrumentation is independent of the response body shape: 201,
429, 503, and every other final status are recorded as sent. An endpoint
that embeds an application error in an HTTP 200 response will still be counted
as successful HTTP traffic. Such endpoints should be corrected to use the
appropriate HTTP status; the metrics layer deliberately does not inspect
response bodies because doing so would add overhead and risk exposing private
data.
Home and Sports API stages
page_api_stage_duration_seconds breaks the initial and recurring backend work for /home and
/sports into bounded internal stages. The existing HTTP histogram remains the route-total source
of truth; stage durations are diagnostic components and may overlap where a fan-out wall timer
contains its individual provider calls.
handler_total starts when the route-owned observer runs. Router-wide middleware registered before
that observer (for example optional/JWT authentication shared by an entire router) remains included
in http_request_duration_seconds, not a separate stage. Route-specific access resolution inside an
observed handler uses the authorization stage. The difference between the global HTTP duration and
handler_total is therefore the only bounded view of earlier shared middleware; this avoids adding
page-specific branches to the shared authentication or startup/HTTP-metrics modules.
The metric labels are restricted to code-owned allowlists:
route— the normalized Express template for a Home/Sports API, orother;stage— fixed processing stages such asprovider_fanout,catalogue_overlay,visibility_filter,exposure_query, andserialization;provider_class—none,betfair,bifrost,pinnacle,thesports,multi, orother;outcome—ok,error,degraded,cache_hit,cache_miss,skipped, orother.
Fixture, market, tournament, watchlist, user, agent, and game identifiers are never labels. Query
strings and raw URLs are never labels. A tolerated per-sport provider failure records the failed
child stage and marks both provider_fanout and the otherwise-successful handler degraded, so an
HTTP 200 cannot conceal a partial feed failure at either level.
Instrumented initial/recurring routes are:
- Home:
/api/fixtures/carousel/grouped,/api/odds/sports/all,/api/ai/watchlists,/api/casino/access, and/api/casino/lobbywhen access is enabled; - Sports:
/api/fixtures/live,/api/fixtures/upcoming,/api/fixtures/catalogue-version,/api/odds/tournaments/:sportId,/api/odds/sports/all,/api/ai/watchlists,/api/ai/watchlist/items,/api/orders/preview/exposure/batch/v2, and the shared-shell/api/casino/accessgate.
Click-only casino launch and watchlist mutations continue to use the global HTTP metrics. They are
not part of page-load or polling latency. Cross-application session/profile bootstrap (for example
/api/users/me) also remains under the global HTTP metrics; it is not part of the Home/Sports
fixture and fixture-adjacent data pipeline instrumented here.
Metrics endpoint network boundary
The internal GET :9464/metrics listener intentionally has no
application-level authentication. Access is enforced by topology: port 9464
is exposed only to the private Docker network and has no host port mapping.
Prometheus scrapes the unique dev container strykr-backend:9464 (staging/production use the
primary-only alias hannibal-primary-backend:9464), while the product listener on port
3001 returns 404 for /metrics. A reverse-proxy catch-all on the product port
therefore cannot publish runtime telemetry.
Prometheus itself is host-bound to loopback on port 9090, remains reachable
to Grafana and exporters over the monitoring Docker network, and does not enable
the destructive Prometheus admin API. Deployment reloads are performed through
the loopback-only lifecycle endpoint after promtool validates the mounted
configuration.
Deployment workflows require backend health and the private metrics listener
to return 200, then require the public /metrics path to return 404. A separate
blackbox probe continuously alerts if either the dev or production public path
returns anything other than the required 404, covering proxy changes made
outside the application deployment workflow.
The deploy check intentionally requires exactly 404. This proves the metrics route is absent from the public application listener. The continuous blackbox module disables redirects so it evaluates the same direct response. A 401 or 403 would mean the operational endpoint is publicly reachable behind application authentication, while a redirect or 5xx can hide an edge or deployment error; none satisfy the private-network contract.
Continuous Prometheus, blackbox, and Grafana monitoring covers the long-lived
dev and production environments. dev1 and dev2 are ephemeral isolated
validation stacks and are not continuous monitoring targets by default. Dev2 has an explicit,
temporary diagnostic opt-in described below; dev1 remains excluded.
Their deployment workflows still verify backend health, the private metrics
listener, and the exact public 404 privacy contract on every deployment.
If a future deployment cannot preserve this network boundary, add access control at the edge and update the Prometheus scrape configuration in the same change before making the endpoint reachable. Do not expose the endpoint first and rely on a later security change.
Node.js runtime metrics
prom-client default metrics provide:
process_cpu_seconds_totalprocess_resident_memory_bytesnodejs_heap_size_used_bytesandnodejs_heap_size_total_bytesnodejs_eventloop_lag_p50_seconds,p90, andp99nodejs_gc_duration_secondsnodejs_active_handles_totalnodejs_active_requests_total- process start time, open file descriptors, heap spaces, and Node version
GC histogram bucket series appear after the process observes its first garbage collection event; the metric descriptor is present from startup.
CPU is reported as CPU-seconds. rate(process_cpu_seconds_total[5m]) is the
average number of CPU cores consumed by the Node process during five minutes;
it is not a whole-host percentage.
Useful PromQL queries
Request rate:
sum(rate(http_requests_total{service="hannibal-backend"}[5m])) by (route)
5xx ratio:
sum(rate(http_requests_total{service="hannibal-backend",status=~"5.."}[5m]))
/
clamp_min(sum(rate(http_requests_total{service="hannibal-backend"}[5m])), 0.001)
p95 latency by route:
histogram_quantile(
0.95,
sum(rate(http_request_duration_seconds_bucket{service="hannibal-backend"}[5m])) by (le, route)
)
Home/Sports p95 stage latency:
histogram_quantile(
0.95,
sum(rate(page_api_stage_duration_seconds_bucket{
service="hannibal-backend",
environment="dev"
}[5m])) by (le, route, stage, provider_class, outcome)
)
Replace 0.95 with 0.50 and 0.99 for p50 and p99. Do not combine outcomes when
diagnosing a partial feed: compare ok, degraded, and error separately.
Provider calls per completed route request over 24 hours:
sum by (route, provider_class) (
increase(page_api_stage_duration_seconds_count{
service="hannibal-backend",
environment="dev",
stage="provider_fixture_fetch"
}[24h])
)
/
on (route) group_left
sum by (route) (
increase(http_request_duration_seconds_count{
service="hannibal-backend",
environment="dev"
}[24h])
)
Use the same expression with page_api_stage_duration_seconds_sum for cumulative provider seconds
per request. Compare it with the provider_fanout percentile to distinguish parallel fan-out wall
time from accumulated provider work.
Dev measurement window
Instrumentation precedes any new performance benchmark. After the instrumentation is deployed to
Dev, record the deployment timestamp, exclude the first 15 minutes for process/cache warm-up, then
measure the following 24 hours. Evaluate the queries at exactly deployment time + 24h15m so [24h]
means the post-warm-up window. Record sample count per route and treat a percentile with fewer than
100 completed requests as directional, not a baseline.
Route p50/p95/p99 for that window:
histogram_quantile(
0.95,
sum(increase(http_request_duration_seconds_bucket{
service="hannibal-backend",
environment="dev",
route=~"/api/(fixtures/(carousel/grouped|live|upcoming|catalogue-version)|odds/(sports/all|tournaments/:sportId)|ai/(watchlists|watchlist/items)|casino/(access|lobby)|orders/preview/exposure/batch/v2)"
}[24h])) by (le, route)
)
Run it with 0.50, 0.95, and 0.99. For a like-for-like route-total comparison with the 24 hours
immediately before deployment, run the same query with offset 24h15m. Stage metrics have no valid
pre-deployment baseline because the series did not exist; their first 24-hour window establishes it.
Stage p50/p95/p99 for the same window:
histogram_quantile(
0.95,
sum(increase(page_api_stage_duration_seconds_bucket{
service="hannibal-backend",
environment="dev"
}[24h])) by (le, route, stage, provider_class, outcome)
)
Completed samples per route:
sum(increase(http_request_duration_seconds_count{
service="hannibal-backend",
environment="dev"
}[24h])) by (route)
For each route, capture all of the following from the same window:
- route p50/p95/p99 and completed-request count;
- stage p50/p95/p99 split by outcome and provider class;
provider_fixture_fetchcount and cumulative seconds divided by completed requests;provider_fanoutwall-time p50/p95/p99;- cache hit/miss counts and degraded/error counts.
The comparison separates parallel wall time (provider_fanout) from total downstream cost (sum of
child provider_fixture_fetch observations). It does not prove a query or provider implementation is
the root cause by itself; use the stage evidence to choose the next bounded investigation.
Node process CPU cores:
rate(process_cpu_seconds_total{service="hannibal-backend"}[5m])
Resident and heap memory:
process_resident_memory_bytes{service="hannibal-backend"}
nodejs_heap_size_used_bytes{service="hannibal-backend"}
Event-loop delay:
nodejs_eventloop_lag_p99_seconds{service="hannibal-backend"}
Prometheus keeps the configured retention window, so the Grafana time picker or Prometheus HTTP API can evaluate these expressions at a specified historical time. Container CPU attribution remains limited by the current cAdvisor labels; the backend's own process metrics are the authoritative phase-1 Node.js view.
Dashboard and alerts
Grafana provisions Application Health from
infra/monitoring/grafana/dashboards/application-health.json. It shows target
health, request rate, 5xx ratio, latency percentiles, route-level performance,
process CPU, RSS/heap, event-loop delay, garbage collection, and active work.
Initial alerts cover:
- any configured Prometheus target down for two minutes, or the expected backend target disappearing from service discovery entirely;
- repeated 5xx responses and a sustained high 5xx ratio;
- p95/p99 API latency;
- sustained Node CPU-core saturation;
- backend RSS above 1 GiB;
- event-loop p99 above 100 ms and 1 second;
- any public dev or production
/metricspath returning a status other than the required 404; - absence of the continuous public-metrics privacy probe.
The thresholds are initial operational guardrails based on the current single Node.js process and observed production RSS baseline. Review them after 30 days of history rather than treating them as permanent capacity limits.
Viewing dev performance
After the observability changes are merged and deployed to dev, Grafana
automatically discovers the provisioned dashboard within about 30 seconds.
Sign in with the existing dev Grafana credentials and open Dashboards →
Application Health, select dev, and choose the time range you want to
inspect.
Grafana is bound to the server's loopback interface and is not a public admin
surface. Loki is bound the same way on port 3100; it runs with
auth_enabled: false, so loopback binding is the only thing keeping the log
corpus and its unauthenticated push API off the public internet. Open an SSH
tunnel from your workstation:
ssh -L 3030:127.0.0.1:3030 strykrdev
Then open http://localhost:3030. The former public
http://dev.strykr.io:3030 endpoint is intentionally unavailable. Do not
commit or share Grafana credentials.
Deployments explicitly load the server-managed
/root/strykr/infra/docker/.env for the monitoring stack and fail before
recreating Grafana if its required admin or Timescale datasource variables are
missing. The file remains untracked and its values must never be printed in
deployment logs.
The same reconciliation starts and verifies Prometheus, Grafana, Blackbox
Exporter, and Alertmanager together. A deploy fails if the private probes are
not active or if the loopback-only operator services are not ready. Prometheus,
Blackbox Exporter, and Alertmanager candidate configurations are validated in
pinned tool images before mutation; the config-owning containers are then
force-recreated so bind mounts cannot retain stale pre-pull file inodes.
Both dev and production call scripts/reconcile-monitoring.sh for this shared
contract. The script discovers each operator service's actual loopback host
port from Docker before readiness, reload, or query requests, so .env port
overrides cannot make deployment verification inspect the wrong listener.
Production materializes the helper from the commit that supplied the running
deployment workflow, rather than from the requested rollback ref. This keeps
the workflow's current verification contract in force while older application
revisions are deployed.
The dashboard environment selector is sourced from
process_start_time_seconds, an application metric emitted immediately at
process startup. It therefore works before the first API request while retaining
the backend registry's environment label; Prometheus's synthetic up metric
does not inherit labels emitted by the scraped application.
Use Request rate by route to see which APIs are busiest, p95 latency by route to find slow APIs, and the 5xx, CPU, memory, event-loop, garbage collection, and active-work panels to correlate slow or failing APIs with runtime pressure. For a status-by-route breakdown, run this in Grafana Explore:
sum(rate(http_requests_total{service="hannibal-backend",environment="dev"}[5m])) by (route, status)
Validation
Validation belongs in the Hannibal dev workflow, not on a developer laptop:
-
Run focused backend unit tests and typecheck in the selected dev workflow.
-
Validate Prometheus configuration and rules with
promtool check configandpromtool check rulesin the Prometheus image/environment. -
Parse all provisioned dashboard JSON.
-
Deploy to dev and confirm
hannibal-backendisupin Prometheus. -
Confirm port
9464has no host mapping, Prometheus can scrapestrykr-backend:9464, and the product listener returns 404 for/metrics. -
Verify the public negative case directly:
test "$(curl -sS -o /dev/null -w '%{http_code}' https://dev.strykr.io/metrics)" = "404" -
Exercise representative 2xx, 4xx, and 5xx API paths.
-
Confirm route labels contain templates and never real identifiers.
-
Confirm Application Health panels populate and the target-down alert is not firing for intentionally uninstrumented frontend or AI services.
Diagnosing intermittent latency
Select one deployment and the incident time range in Application Health. Start with route latency and request rate, then select Diagnostic route. Stage p95 and Stage sample count localize repeated work. Percentiles interpolate between histogram buckets: one 6s observation in the 5–10s bucket can produce a p95 close to 10s. Existing bucket boundaries are intentionally unchanged so old and new instances can be compared during rollout.
Exact slow requests and stage summaries retains every API response or aborted connection
lasting at least 1000ms from logger attachment as an info event, including HTTP 200 at LOG_LEVEL=info. durationMs
is measured using a monotonic clock from logger attachment through finish/abort. For regular
routes the logger attaches after body parsing; failures before it attaches have HTTP metrics only. diagnosticId is generated by the server and identifies this
single diagnostic record; it is not the incoming requestId and does not join earlier application
logs or trace downstream services. Search that ID in Loki to retrieve the record. The JSON message
survives the current Promtail message-only output; the panel also accepts future full JSON envelopes.
The diagnostic body includes deploymentEnvironment because NODE_ENV is production on dev too.
Only method, route template, HTTP status, completion kind, deployment identity, diagnosticId and bounded stage summaries are included. No URL, query string, user/account/order/fixture identity, SQL/parameters, financial values or request/response body is copied. Existing fast HTTP logging and error handling are unchanged. Slow diagnostic logs temporarily replace ambient request context so the logger cannot reintroduce user IDs or client-supplied request IDs into this event.
Each summary groups a fixed stage/provider-class/outcome combination with count, totalMs and
maxMs. At most 40 groups are retained; droppedStageObservations explicitly counts observations
which did not fit. Repeated observations aggregate in constant space. Nested timers and concurrent
Promise.all children overlap: never sum all stages and call that total request time. On abort,
only completed stages are present; work still pending after the connection closes is omitted.
handler_total starts at the route observer. Shared authentication before that observer is visible only in the difference from full HTTP
duration. The HTTP metric starts earlier than the logger and includes body parsing; authorization measures
route-owned authorization, not all middleware. Empty summaries mean the route is not instrumented
or no measured stage completed; they do not prove the request performed no work.
New route coverage:
/api/agent/downlineand/api/agent/dashboard: authorization, hierarchy/database reads, take, exposure/P&L, commission and allocation aggregation. Monetary logic is unchanged./api/exposure-limits/players/:playerId/effective: read authorization, configured limit defaults and the player's configured limit-set lookup. Enforcement/write paths are unchanged.- Catalogue lists: source resolution, display configuration, market metadata, dictionary,
synchronous transform/filter/sort/merge and TheSports score enrichment.
page_api_stage_itemsrecords input fixture and unique market-reference counts as numbers. - Betfair discovery: upcoming and live work, individual upcoming catalogue calls, actual cache
reads, refill lock acquisition, cache writes and lock release.
cache_bypassmeans no discovery cache policy, a disabled cache or explicit refresh; cache-enabledcache_missmeans a refill was attempted. Existing cache.get folds Redis errors into null, so cache_miss can also mean unavailable Redis. Existing cache.set swallows write errors: cache_write ok only means the helper returned, not verified persistence. A final cache hit can follow lock waiting, shown separately bycache_lock_wait. The existing outerprovider_fixture_fetchcached boolean is coarse; use the inner discovery outcome to distinguish bypass. Each catalogue-call timer includes the client's retries/session handling; it is not a pure network round-trip measurement.
database_query_duration_seconds observes Prisma's engine-reported query durations and its _count
provides query volume. These are process-wide, including jobs and all requests. Prisma query
events have no guaranteed request AsyncLocalStorage association, so no per-request database claim
or invented pool-wait metric is made. Compare query duration/volume with event-loop delay, GC, CPU
and in-progress work panels to distinguish likely database activity from process-wide pressure.
A stage identifies where elapsed time accrued; it does not itself prove CPU, network, provider
server, database-lock or connection-pool causation. Those may still require a targeted trace or
profile once the slow stage is known. Telemetry is best effort: sink/metric failures must never
change request results; missing events still require checking ingestion health.
Deterministic primary backend scrape
Only the primary backend in dev, staging and production declares hannibal-primary-backend.
The isolated dev1/dev2 Compose services must never declare it. Their shared default backend
Docker alias is ambiguous and must not be used for continuous scraping. Environment-specific
file-SD targets add expected_environment; the mismatch alert fires if an up endpoint does not
contain matching runtime identity. Existing down-target alerts cover unreachable endpoints.
reconcile-monitoring.sh derives BACKEND_METRICS_ENVIRONMENT from the selected primary container's
explicit ENVIRONMENT, validates dev/staging/production, and mounts the corresponding tracked target
file. For manual monitoring Compose operations, set BACKEND_METRICS_ENVIRONMENT=dev, staging,
or production explicitly in infra/docker/.env or the command environment. The new Compose
configuration refuses an unset value. Production rollback to an older checkout retains that
checkout's legacy target/Compose contract and reports that deterministic identity is unavailable.
Dev uses the existing unique strykr-backend container name so correcting its shared-network
alias collision does not require restarting the primary application. Staging and production retain
the primary alias contract. The Backend Target panel uses expected_environment to follow the
selected environment; an old target file without that label shows DOWN/unknown until targets and
dashboard are rolled out together. A full rollback to old files retains the old dashboard contract.
Promtail needs no pipeline change: it retains the bounded diagnostic JSON in the message.
Temporary dev2 diagnostics
The dev2 backend log label is an explicit per-deployment opt-in, defaulting to false. On the shared dev host, deploy the approved diagnostics branch through the supported isolated deployment:
DEV2_DIAGNOSTICS_COLLECT_LOGS=true /root/deploy-branch-to-dev-env.sh dev2 codex/latency-diagnostics
Then use the helper from that verified checkout (or a protected, versioned copy of it):
bash scripts/configure-dev2-monitoring.sh enable-dev2 /root/strykr/infra/docker --check
bash scripts/configure-dev2-monitoring.sh enable-dev2 /root/strykr/infra/docker
The helper merges onto the existing monitoring base, project and env file. It resolves candidate image tags to the same live image IDs and compares the complete effective
environment, command and entrypoint against those exact image defaults, including removed/null
overrides. It checks loopback ports, named data volumes and bind sources, refuses unknown existing overlays,
and validates the exact candidate Prometheus mount hierarchy before changing anything. It recreates
only Prometheus and Grafana with --no-deps. Three read-only child mounts select the dev+dev2 targets,
identity alerts and Application Health dashboard; primary checkout files and application containers
are untouched. Keep the versioned config source available while these mounts are in use. This focused
operation is separate from full monitoring reconciliation, which intentionally returns to primary-only
scraping. No credentials are written to tracked configuration or printed by the helper.
Dev2 must already be running with ENVIRONMENT=dev2 and logging.collect=true before enable passes.
Promtail discovers that label without a restart; no user/request/diagnostic identifier becomes a Loki
label. Targets carry expected_environment=dev or dev2, preserving the actual runtime environment
metric label for identity checks. Select exactly dev2 in Grafana for isolated results; selecting All
intentionally combines selected environments in aggregate panels. Confirm both targets are up with
matching runtime identities, the new panels are provisioned, and an exact emitted diagnosticId
can be found in Loki with deploymentEnvironment=dev2. A dashboard alone does not prove ingestion.
To stop the temporary scrape while keeping the deterministic dev target and diagnostic panels:
bash scripts/configure-dev2-monitoring.sh disable-dev2 /root/strykr/infra/docker --check
bash scripts/configure-dev2-monitoring.sh disable-dev2 /root/strykr/infra/docker
Then redeploy dev2 through its supported workflow without DEV2_DIAGNOSTICS_COLLECT_LOGS=true to stop
collection of new logs. Previously stored Loki logs remain subject to ordinary retention. The flag
is not persisted into the primary .env: any later ordinary dev2 deployment also resets it to false.
Re-enable it explicitly for every deployment while temporary diagnostics are needed, and re-check
both metrics and log ingestion afterwards. Full rollback of monitoring images/configuration uses the
operator's pre-change deployment snapshot; the helper's disable action only removes dev2 scraping.