Skip to main content

Application Observability

Hannibal phase-1 operational telemetry uses the existing self-hosted Prometheus, Grafana, Alertmanager, Loki, and infrastructure exporters. The backend exposes application and Node.js runtime metrics at GET /metrics on the dedicated Docker-internal port 9464. The product HTTP listener on port 3001 does not serve metrics.

This phase does not include OpenTelemetry traces, continuous profiling, frontend browser telemetry, or AI-service metrics.

Backend metrics​

HTTP RED metrics​

MetricTypePurpose
http_requests_totalCounterCompleted requests by method, route template, and status
http_request_duration_secondsHistogramEnd-to-end Express request duration
http_requests_in_progressGaugeRequests currently executing
http_request_aborted_totalCounterConnections closed before completion
http_response_size_bytesHistogramResponse content length when known
page_api_stage_duration_secondsHistogramInternal stages for selected page and agent read APIs

HTTP metrics use these bounded labels:

  • service="hannibal-backend"
  • environment — deployment identity from ENVIRONMENT (dev, dev1, dev2, or production), falling back to NODE_ENV only when unset
  • method — GET, POST, PUT, PATCH, DELETE, OPTIONS, HEAD, or OTHER; unsupported methods are coalesced to keep cardinality bounded
  • route — a literal router mount plus the Express route template
  • status — the final HTTP status code emitted by Express

Raw URLs are never labels. Unmatched requests use route="unmatched", and array/regular-expression routes use route="complex-route". User IDs, usernames, order IDs, fixture IDs, provider account identifiers, request bodies, credentials, cookies, and financial values must never be metric labels.

HTTP recording starts before CORS and body parsing, so malformed, oversized, and otherwise early-rejected API requests remain observable. The recorder only accepts /api paths, while scrapes use a separate HTTP server, so operational traffic and static assets do not inflate product request-rate or latency metrics.

Status instrumentation is independent of the response body shape: 201, 429, 503, and every other final status are recorded as sent. An endpoint that embeds an application error in an HTTP 200 response will still be counted as successful HTTP traffic. Such endpoints should be corrected to use the appropriate HTTP status; the metrics layer deliberately does not inspect response bodies because doing so would add overhead and risk exposing private data.

Home and Sports API stages​

page_api_stage_duration_seconds breaks the initial and recurring backend work for /home and /sports into bounded internal stages. The existing HTTP histogram remains the route-total source of truth; stage durations are diagnostic components and may overlap where a fan-out wall timer contains its individual provider calls.

handler_total starts when the route-owned observer runs. Router-wide middleware registered before that observer (for example optional/JWT authentication shared by an entire router) remains included in http_request_duration_seconds, not a separate stage. Route-specific access resolution inside an observed handler uses the authorization stage. The difference between the global HTTP duration and handler_total is therefore the only bounded view of earlier shared middleware; this avoids adding page-specific branches to the shared authentication or startup/HTTP-metrics modules.

The metric labels are restricted to code-owned allowlists:

  • route — the normalized Express template for a Home/Sports API, or other;
  • stage — fixed processing stages such as provider_fanout, catalogue_overlay, visibility_filter, exposure_query, and serialization;
  • provider_class — none, betfair, bifrost, pinnacle, thesports, multi, or other;
  • outcome — ok, error, degraded, cache_hit, cache_miss, skipped, or other.

Fixture, market, tournament, watchlist, user, agent, and game identifiers are never labels. Query strings and raw URLs are never labels. A tolerated per-sport provider failure records the failed child stage and marks both provider_fanout and the otherwise-successful handler degraded, so an HTTP 200 cannot conceal a partial feed failure at either level.

Instrumented initial/recurring routes are:

  • Home: /api/fixtures/carousel/grouped, /api/odds/sports/all, /api/ai/watchlists, /api/casino/access, and /api/casino/lobby when access is enabled;
  • Sports: /api/fixtures/live, /api/fixtures/upcoming, /api/fixtures/catalogue-version, /api/odds/tournaments/:sportId, /api/odds/sports/all, /api/ai/watchlists, /api/ai/watchlist/items, /api/orders/preview/exposure/batch/v2, and the shared-shell /api/casino/access gate.

Click-only casino launch and watchlist mutations continue to use the global HTTP metrics. They are not part of page-load or polling latency. Cross-application session/profile bootstrap (for example /api/users/me) also remains under the global HTTP metrics; it is not part of the Home/Sports fixture and fixture-adjacent data pipeline instrumented here.

Metrics endpoint network boundary​

The internal GET :9464/metrics listener intentionally has no application-level authentication. Access is enforced by topology: port 9464 is exposed only to the private Docker network and has no host port mapping. Prometheus scrapes the unique dev container strykr-backend:9464 (staging/production use the primary-only alias hannibal-primary-backend:9464), while the product listener on port 3001 returns 404 for /metrics. A reverse-proxy catch-all on the product port therefore cannot publish runtime telemetry.

Prometheus itself is host-bound to loopback on port 9090, remains reachable to Grafana and exporters over the monitoring Docker network, and does not enable the destructive Prometheus admin API. Deployment reloads are performed through the loopback-only lifecycle endpoint after promtool validates the mounted configuration.

Deployment workflows require backend health and the private metrics listener to return 200, then require the public /metrics path to return 404. A separate blackbox probe continuously alerts if either the dev or production public path returns anything other than the required 404, covering proxy changes made outside the application deployment workflow.

The deploy check intentionally requires exactly 404. This proves the metrics route is absent from the public application listener. The continuous blackbox module disables redirects so it evaluates the same direct response. A 401 or 403 would mean the operational endpoint is publicly reachable behind application authentication, while a redirect or 5xx can hide an edge or deployment error; none satisfy the private-network contract.

Continuous Prometheus, blackbox, and Grafana monitoring covers the long-lived dev and production environments. dev1 and dev2 are ephemeral isolated validation stacks and are not continuous monitoring targets by default. Dev2 has an explicit, temporary diagnostic opt-in described below; dev1 remains excluded. Their deployment workflows still verify backend health, the private metrics listener, and the exact public 404 privacy contract on every deployment.

If a future deployment cannot preserve this network boundary, add access control at the edge and update the Prometheus scrape configuration in the same change before making the endpoint reachable. Do not expose the endpoint first and rely on a later security change.

Node.js runtime metrics​

prom-client default metrics provide:

  • process_cpu_seconds_total
  • process_resident_memory_bytes
  • nodejs_heap_size_used_bytes and nodejs_heap_size_total_bytes
  • nodejs_eventloop_lag_p50_seconds, p90, and p99
  • nodejs_gc_duration_seconds
  • nodejs_active_handles_total
  • nodejs_active_requests_total
  • process start time, open file descriptors, heap spaces, and Node version

GC histogram bucket series appear after the process observes its first garbage collection event; the metric descriptor is present from startup.

CPU is reported as CPU-seconds. rate(process_cpu_seconds_total[5m]) is the average number of CPU cores consumed by the Node process during five minutes; it is not a whole-host percentage.

Useful PromQL queries​

Request rate:

sum(rate(http_requests_total{service="hannibal-backend"}[5m])) by (route)

5xx ratio:

sum(rate(http_requests_total{service="hannibal-backend",status=~"5.."}[5m]))
/
clamp_min(sum(rate(http_requests_total{service="hannibal-backend"}[5m])), 0.001)

p95 latency by route:

histogram_quantile(
0.95,
sum(rate(http_request_duration_seconds_bucket{service="hannibal-backend"}[5m])) by (le, route)
)

Home/Sports p95 stage latency:

histogram_quantile(
0.95,
sum(rate(page_api_stage_duration_seconds_bucket{
service="hannibal-backend",
environment="dev"
}[5m])) by (le, route, stage, provider_class, outcome)
)

Replace 0.95 with 0.50 and 0.99 for p50 and p99. Do not combine outcomes when diagnosing a partial feed: compare ok, degraded, and error separately.

Provider calls per completed route request over 24 hours:

sum by (route, provider_class) (
increase(page_api_stage_duration_seconds_count{
service="hannibal-backend",
environment="dev",
stage="provider_fixture_fetch"
}[24h])
)
/
on (route) group_left
sum by (route) (
increase(http_request_duration_seconds_count{
service="hannibal-backend",
environment="dev"
}[24h])
)

Use the same expression with page_api_stage_duration_seconds_sum for cumulative provider seconds per request. Compare it with the provider_fanout percentile to distinguish parallel fan-out wall time from accumulated provider work.

Dev measurement window​

Instrumentation precedes any new performance benchmark. After the instrumentation is deployed to Dev, record the deployment timestamp, exclude the first 15 minutes for process/cache warm-up, then measure the following 24 hours. Evaluate the queries at exactly deployment time + 24h15m so [24h] means the post-warm-up window. Record sample count per route and treat a percentile with fewer than 100 completed requests as directional, not a baseline.

Route p50/p95/p99 for that window:

histogram_quantile(
0.95,
sum(increase(http_request_duration_seconds_bucket{
service="hannibal-backend",
environment="dev",
route=~"/api/(fixtures/(carousel/grouped|live|upcoming|catalogue-version)|odds/(sports/all|tournaments/:sportId)|ai/(watchlists|watchlist/items)|casino/(access|lobby)|orders/preview/exposure/batch/v2)"
}[24h])) by (le, route)
)

Run it with 0.50, 0.95, and 0.99. For a like-for-like route-total comparison with the 24 hours immediately before deployment, run the same query with offset 24h15m. Stage metrics have no valid pre-deployment baseline because the series did not exist; their first 24-hour window establishes it.

Stage p50/p95/p99 for the same window:

histogram_quantile(
0.95,
sum(increase(page_api_stage_duration_seconds_bucket{
service="hannibal-backend",
environment="dev"
}[24h])) by (le, route, stage, provider_class, outcome)
)

Completed samples per route:

sum(increase(http_request_duration_seconds_count{
service="hannibal-backend",
environment="dev"
}[24h])) by (route)

For each route, capture all of the following from the same window:

  • route p50/p95/p99 and completed-request count;
  • stage p50/p95/p99 split by outcome and provider class;
  • provider_fixture_fetch count and cumulative seconds divided by completed requests;
  • provider_fanout wall-time p50/p95/p99;
  • cache hit/miss counts and degraded/error counts.

The comparison separates parallel wall time (provider_fanout) from total downstream cost (sum of child provider_fixture_fetch observations). It does not prove a query or provider implementation is the root cause by itself; use the stage evidence to choose the next bounded investigation.

Node process CPU cores:

rate(process_cpu_seconds_total{service="hannibal-backend"}[5m])

Resident and heap memory:

process_resident_memory_bytes{service="hannibal-backend"}
nodejs_heap_size_used_bytes{service="hannibal-backend"}

Event-loop delay:

nodejs_eventloop_lag_p99_seconds{service="hannibal-backend"}

Prometheus keeps the configured retention window, so the Grafana time picker or Prometheus HTTP API can evaluate these expressions at a specified historical time. Container CPU attribution remains limited by the current cAdvisor labels; the backend's own process metrics are the authoritative phase-1 Node.js view.

Dashboard and alerts​

Grafana provisions Application Health from infra/monitoring/grafana/dashboards/application-health.json. It shows target health, request rate, 5xx ratio, latency percentiles, route-level performance, process CPU, RSS/heap, event-loop delay, garbage collection, and active work.

Initial alerts cover:

  • any configured Prometheus target down for two minutes, or the expected backend target disappearing from service discovery entirely;
  • repeated 5xx responses and a sustained high 5xx ratio;
  • p95/p99 API latency;
  • sustained Node CPU-core saturation;
  • backend RSS above 1 GiB;
  • event-loop p99 above 100 ms and 1 second;
  • any public dev or production /metrics path returning a status other than the required 404;
  • absence of the continuous public-metrics privacy probe.

The thresholds are initial operational guardrails based on the current single Node.js process and observed production RSS baseline. Review them after 30 days of history rather than treating them as permanent capacity limits.

Viewing dev performance​

After the observability changes are merged and deployed to dev, Grafana automatically discovers the provisioned dashboard within about 30 seconds. Sign in with the existing dev Grafana credentials and open Dashboards → Application Health, select dev, and choose the time range you want to inspect.

Grafana is bound to the server's loopback interface and is not a public admin surface. Loki is bound the same way on port 3100; it runs with auth_enabled: false, so loopback binding is the only thing keeping the log corpus and its unauthenticated push API off the public internet. Open an SSH tunnel from your workstation:

ssh -L 3030:127.0.0.1:3030 strykrdev

Then open http://localhost:3030. The former public http://dev.strykr.io:3030 endpoint is intentionally unavailable. Do not commit or share Grafana credentials.

Deployments explicitly load the server-managed /root/strykr/infra/docker/.env for the monitoring stack and fail before recreating Grafana if its required admin or Timescale datasource variables are missing. The file remains untracked and its values must never be printed in deployment logs.

The same reconciliation starts and verifies Prometheus, Grafana, Blackbox Exporter, and Alertmanager together. A deploy fails if the private probes are not active or if the loopback-only operator services are not ready. Prometheus, Blackbox Exporter, and Alertmanager candidate configurations are validated in pinned tool images before mutation; the config-owning containers are then force-recreated so bind mounts cannot retain stale pre-pull file inodes. Both dev and production call scripts/reconcile-monitoring.sh for this shared contract. The script discovers each operator service's actual loopback host port from Docker before readiness, reload, or query requests, so .env port overrides cannot make deployment verification inspect the wrong listener. Production materializes the helper from the commit that supplied the running deployment workflow, rather than from the requested rollback ref. This keeps the workflow's current verification contract in force while older application revisions are deployed.

The dashboard environment selector is sourced from process_start_time_seconds, an application metric emitted immediately at process startup. It therefore works before the first API request while retaining the backend registry's environment label; Prometheus's synthetic up metric does not inherit labels emitted by the scraped application.

Use Request rate by route to see which APIs are busiest, p95 latency by route to find slow APIs, and the 5xx, CPU, memory, event-loop, garbage collection, and active-work panels to correlate slow or failing APIs with runtime pressure. For a status-by-route breakdown, run this in Grafana Explore:

sum(rate(http_requests_total{service="hannibal-backend",environment="dev"}[5m])) by (route, status)

Validation​

Validation belongs in the Hannibal dev workflow, not on a developer laptop:

  1. Run focused backend unit tests and typecheck in the selected dev workflow.

  2. Validate Prometheus configuration and rules with promtool check config and promtool check rules in the Prometheus image/environment.

  3. Parse all provisioned dashboard JSON.

  4. Deploy to dev and confirm hannibal-backend is up in Prometheus.

  5. Confirm port 9464 has no host mapping, Prometheus can scrape strykr-backend:9464, and the product listener returns 404 for /metrics.

  6. Verify the public negative case directly:

    test "$(curl -sS -o /dev/null -w '%{http_code}' https://dev.strykr.io/metrics)" = "404"
  7. Exercise representative 2xx, 4xx, and 5xx API paths.

  8. Confirm route labels contain templates and never real identifiers.

  9. Confirm Application Health panels populate and the target-down alert is not firing for intentionally uninstrumented frontend or AI services.

Diagnosing intermittent latency​

Select one deployment and the incident time range in Application Health. Start with route latency and request rate, then select Diagnostic route. Stage p95 and Stage sample count localize repeated work. Percentiles interpolate between histogram buckets: one 6s observation in the 5–10s bucket can produce a p95 close to 10s. Existing bucket boundaries are intentionally unchanged so old and new instances can be compared during rollout.

Exact slow requests and stage summaries retains every API response or aborted connection lasting at least 1000ms from logger attachment as an info event, including HTTP 200 at LOG_LEVEL=info. durationMs is measured using a monotonic clock from logger attachment through finish/abort. For regular routes the logger attaches after body parsing; failures before it attaches have HTTP metrics only. diagnosticId is generated by the server and identifies this single diagnostic record; it is not the incoming requestId and does not join earlier application logs or trace downstream services. Search that ID in Loki to retrieve the record. The JSON message survives the current Promtail message-only output; the panel also accepts future full JSON envelopes. The diagnostic body includes deploymentEnvironment because NODE_ENV is production on dev too.

Only method, route template, HTTP status, completion kind, deployment identity, diagnosticId and bounded stage summaries are included. No URL, query string, user/account/order/fixture identity, SQL/parameters, financial values or request/response body is copied. Existing fast HTTP logging and error handling are unchanged. Slow diagnostic logs temporarily replace ambient request context so the logger cannot reintroduce user IDs or client-supplied request IDs into this event.

Each summary groups a fixed stage/provider-class/outcome combination with count, totalMs and maxMs. At most 40 groups are retained; droppedStageObservations explicitly counts observations which did not fit. Repeated observations aggregate in constant space. Nested timers and concurrent Promise.all children overlap: never sum all stages and call that total request time. On abort, only completed stages are present; work still pending after the connection closes is omitted. handler_total starts at the route observer. Shared authentication before that observer is visible only in the difference from full HTTP duration. The HTTP metric starts earlier than the logger and includes body parsing; authorization measures route-owned authorization, not all middleware. Empty summaries mean the route is not instrumented or no measured stage completed; they do not prove the request performed no work.

New route coverage:

  • /api/agent/downline and /api/agent/dashboard: authorization, hierarchy/database reads, take, exposure/P&L, commission and allocation aggregation. Monetary logic is unchanged.
  • /api/exposure-limits/players/:playerId/effective: read authorization, configured limit defaults and the player's configured limit-set lookup. Enforcement/write paths are unchanged.
  • Catalogue lists: source resolution, display configuration, market metadata, dictionary, synchronous transform/filter/sort/merge and TheSports score enrichment. page_api_stage_items records input fixture and unique market-reference counts as numbers.
  • Betfair discovery: upcoming and live work, individual upcoming catalogue calls, actual cache reads, refill lock acquisition, cache writes and lock release. cache_bypass means no discovery cache policy, a disabled cache or explicit refresh; cache-enabled cache_miss means a refill was attempted. Existing cache.get folds Redis errors into null, so cache_miss can also mean unavailable Redis. Existing cache.set swallows write errors: cache_write ok only means the helper returned, not verified persistence. A final cache hit can follow lock waiting, shown separately by cache_lock_wait. The existing outer provider_fixture_fetch cached boolean is coarse; use the inner discovery outcome to distinguish bypass. Each catalogue-call timer includes the client's retries/session handling; it is not a pure network round-trip measurement.

database_query_duration_seconds observes Prisma's engine-reported query durations and its _count provides query volume. These are process-wide, including jobs and all requests. Prisma query events have no guaranteed request AsyncLocalStorage association, so no per-request database claim or invented pool-wait metric is made. Compare query duration/volume with event-loop delay, GC, CPU and in-progress work panels to distinguish likely database activity from process-wide pressure. A stage identifies where elapsed time accrued; it does not itself prove CPU, network, provider server, database-lock or connection-pool causation. Those may still require a targeted trace or profile once the slow stage is known. Telemetry is best effort: sink/metric failures must never change request results; missing events still require checking ingestion health.

Deterministic primary backend scrape​

Only the primary backend in dev, staging and production declares hannibal-primary-backend. The isolated dev1/dev2 Compose services must never declare it. Their shared default backend Docker alias is ambiguous and must not be used for continuous scraping. Environment-specific file-SD targets add expected_environment; the mismatch alert fires if an up endpoint does not contain matching runtime identity. Existing down-target alerts cover unreachable endpoints.

reconcile-monitoring.sh derives BACKEND_METRICS_ENVIRONMENT from the selected primary container's explicit ENVIRONMENT, validates dev/staging/production, and mounts the corresponding tracked target file. For manual monitoring Compose operations, set BACKEND_METRICS_ENVIRONMENT=dev, staging, or production explicitly in infra/docker/.env or the command environment. The new Compose configuration refuses an unset value. Production rollback to an older checkout retains that checkout's legacy target/Compose contract and reports that deterministic identity is unavailable. Dev uses the existing unique strykr-backend container name so correcting its shared-network alias collision does not require restarting the primary application. Staging and production retain the primary alias contract. The Backend Target panel uses expected_environment to follow the selected environment; an old target file without that label shows DOWN/unknown until targets and dashboard are rolled out together. A full rollback to old files retains the old dashboard contract. Promtail needs no pipeline change: it retains the bounded diagnostic JSON in the message.

Temporary dev2 diagnostics​

The dev2 backend log label is an explicit per-deployment opt-in, defaulting to false. On the shared dev host, deploy the approved diagnostics branch through the supported isolated deployment:

DEV2_DIAGNOSTICS_COLLECT_LOGS=true /root/deploy-branch-to-dev-env.sh dev2 codex/latency-diagnostics

Then use the helper from that verified checkout (or a protected, versioned copy of it):

bash scripts/configure-dev2-monitoring.sh enable-dev2 /root/strykr/infra/docker --check
bash scripts/configure-dev2-monitoring.sh enable-dev2 /root/strykr/infra/docker

The helper merges onto the existing monitoring base, project and env file. It resolves candidate image tags to the same live image IDs and compares the complete effective environment, command and entrypoint against those exact image defaults, including removed/null overrides. It checks loopback ports, named data volumes and bind sources, refuses unknown existing overlays, and validates the exact candidate Prometheus mount hierarchy before changing anything. It recreates only Prometheus and Grafana with --no-deps. Three read-only child mounts select the dev+dev2 targets, identity alerts and Application Health dashboard; primary checkout files and application containers are untouched. Keep the versioned config source available while these mounts are in use. This focused operation is separate from full monitoring reconciliation, which intentionally returns to primary-only scraping. No credentials are written to tracked configuration or printed by the helper.

Dev2 must already be running with ENVIRONMENT=dev2 and logging.collect=true before enable passes. Promtail discovers that label without a restart; no user/request/diagnostic identifier becomes a Loki label. Targets carry expected_environment=dev or dev2, preserving the actual runtime environment metric label for identity checks. Select exactly dev2 in Grafana for isolated results; selecting All intentionally combines selected environments in aggregate panels. Confirm both targets are up with matching runtime identities, the new panels are provisioned, and an exact emitted diagnosticId can be found in Loki with deploymentEnvironment=dev2. A dashboard alone does not prove ingestion.

To stop the temporary scrape while keeping the deterministic dev target and diagnostic panels:

bash scripts/configure-dev2-monitoring.sh disable-dev2 /root/strykr/infra/docker --check
bash scripts/configure-dev2-monitoring.sh disable-dev2 /root/strykr/infra/docker

Then redeploy dev2 through its supported workflow without DEV2_DIAGNOSTICS_COLLECT_LOGS=true to stop collection of new logs. Previously stored Loki logs remain subject to ordinary retention. The flag is not persisted into the primary .env: any later ordinary dev2 deployment also resets it to false. Re-enable it explicitly for every deployment while temporary diagnostics are needed, and re-check both metrics and log ingestion afterwards. Full rollback of monitoring images/configuration uses the operator's pre-change deployment snapshot; the helper's disable action only removes dev2 scraping.