Skip to main content

ADR-0038 — Live scores are delivered by push; REST is the recovery path

  • Status: Accepted
  • Date: 2026-08-27
  • Related: ADR-0035, docs/architecture/data-plane-integration.md §4b, §4c

Context​

The MQTT push path exists, is running, and is measured. Four topics (thesports/{cricket,tennis,football,basketball}/match/v1) subscribe with Granted QoS 0, carrying score, timer, stats, timeline, players and text-live deltas at several frames per second. Measured on dev, 2026-08-07, 203s aligned window, 15 live cricket matches: MQTT delivered 8 score deltas across 5 matches while a 5s REST ground-truth poll observed 6 transitions across 4 — and the set of matches that changed in REST but were silent on MQTT was empty. Push is a strict superset of what the poll could see. The unconditional 15s poller was removed on that evidence (#1187).

The delivery chain is complete on the backend:

MQTT delta → DeltaMergeBuffer → Redis thesports:live:* → liveScorePublisher
→ resolve OUR fixture id (approved link only) → Redis PUB → clientWs
→ room fixture:<id> → browser

And the last hop is not connected. useTheSportsFixtureScores — the hook that joins those rooms, reference-counts subscriptions and re-subscribes on reconnect — has zero mounted consumers in strykr-fe (verified 2026-08-27). Nothing subscribes, so nothing is delivered, so in practice a score reaches a screen only on the 10 s REST poll of the fixture list (useFixtures.ts:46; upcoming lists poll 60 s), while the fixture DETAIL page polls 30 s when the socket is connected and 5 s when it is not (fixtureDetailReconciliation.ts:3-4, selected at useFixtures.ts:332), through buildFixtureScoreOverlay.

Corrected 2026-08-29. This paragraph previously said "20-second REST poll". There is no 20-second timer anywhere in the frontend — the figure matched nothing in the code. Re-measured at deployed frontend ref de620007d. The conclusion is unchanged and the numbers are now the ones the code actually uses.

The cost is stated in the channel module's own header and is unchanged: the provider poll plus a browser poll made a score ~35 seconds stale before it reached a screen. MQTT puts the same change in Redis in hundreds of milliseconds. Everything between those two numbers is built, tested and idle.

Decision​

Push is the delivery mechanism for a live score. The REST fixture-list poll is the recovery and reconciliation path, not the delivery path.

  1. The browser subscribes. List surfaces call useTheSportsFixtureScores once with everything they render; a detail page calls the singular form. This is the whole remaining work on the delivery path.
  2. One transport, not two. Everything rides the existing Socket.IO server in clientWs.ts and the existing Redis pub/sub fan-out. There is no second websocket server and no second browser connection.
  3. Rooms are keyed by OUR fixture id, not by TheSports' id and not by sport. A match with no approved link is dropped at publish, not broadcast under a provider id — there is no room it could honestly go to. The previous shape was one room per sport, which sent every viewer every match in that sport (soccer alone had 1,771 on an observed day) to render the handful on their screen.
  4. Identity is resolved on the publish side, before the PUBLISH, so clientWs stays a pure fan-out with no database dependency in the path that also carries balance, settlement and odds.
  5. Publish strictly after the store write, and only on success. Publish-first would let a failed write put 2–1 on screen against a store holding 2–0, and the next re-read would roll the score backwards — the one direction a score must never move.
  6. A publish failure is latency, never correctness. It is logged and dropped; nothing retries. The store already holds the state and the client re-syncs on its next snapshot. A retry queue would add a second delivery path with its own ordering, to protect data the client can already re-read.
  7. A delta replaces the live projection as a unit — phase, scores, detail together. Exactly one record is authoritative per match and it supplies both the phase and the score, because a live phase beside a schedule score is how a stale number rides along with a fresh status. An absent score therefore clears a known one rather than leaving it: a blank score is honest, a stale one is not.
  8. REST keeps three jobs and loses one. It keeps the baseline (a partial delta has nothing to merge into), the lifecycle (only the provider's silence means a match ended), and the terminal status and final score (the diary is the only endpoint that reports them). It does not deliver in-play changes.

Consequences​

  • Wiring the hook is the single highest-value-per-unit-of-work item in the integration: roughly one useEffect per list surface, against ~35 seconds of avoidable staleness on every score in the product.
  • Until it lands, the push path's real load is approximately zero, and any capacity measurement taken today against socket traffic measures nothing. Sizing work must model the intended load, not the observed one.
  • The REST poll stays as reconciliation; it is not removed when push is connected. But it is not currently doing that job. Measured on dev 2026-08-29: a live fixture-detail page with the socket connected went 648 seconds with zero refetches, and the 30 s reconciliation was never observed to fire once — with visibilityState: visible, hasFocus(): true, background-timer throttling disabled, and the response body reading status: "live" on both sides of the gap. What actually refreshes the score today is odds:update invalidation as a side effect (useWebSocket.ts:585), so the refetch cadence tracks odds traffic rather than the clock: a busy market refetched at p50 0.8 s, a quiet one 7 times in 12 minutes. This ADR asserts a fallback that does not presently run — see the open item in the PR that mounted the consumer.
  • A perverse consequence, now resolved by the consumer landing. The detail page polls the SLOWER 30 s branch when the socket is connected, because the code assumed a connected socket was delivering scores. It was not. Being connected therefore put a user on the slower path. With useTheSportsFixtureScore mounted that assumption is true for the first time, which is the cleanest one-line justification for the change.
  • A score that arrives by push but whose fixture is not classified (ADR-0032) or not linked (ADR-0036) is dropped before it costs a PUBLISH. The three-gate chain applies to the push path exactly as it does to the poll path.

Amendment — 2026-09-03: the switches were the defect, and the fallback now runs​

Two things this ADR recorded as open are closed, and one of its premises was wrong.

The push path was dark in every environment, not just under-consumed. This ADR said the remaining work was mounting the hook, and PR #1484 mounted it. Measured 2026-09-03 by reading the container environments on strykrdev: dev had THESPORTS_SCORES_WS_ENABLED=false, so the backend published nothing at all; dev1 had it true; and neither frontend container carried NEXT_PUBLIC_THESPORTS_SCORES_WS_ENABLED, so the browser flag was false in both bundles and no client subscribed anywhere. Environments with the whole chain on: 0 of 2. Five switches (THESPORTS_ENABLED, INGEST, STREAM, SCORES_WS, and the NEXT_PUBLIC twin) had to agree, and every wrong combination failed as silence — indistinguishable from a fixture with no score, which is the one thing this path must never claim falsely.

Four of the five are deleted. THESPORTS_ENABLED is the only kill switch; under it REST ingestion and MQTT always start, deltas are always published, and every score-bearing surface always subscribes.

"This ADR asserts a fallback that does not presently run" is resolved, and the mechanism was not what the open item assumed. The 30 s reconciliation never fired because TanStack Query's QueryObserver.onQueryUpdate() calls #updateTimers() on every write to the query's cache, and #updateRefetchInterval() clears the interval and re-creates it with no same-value guard on that path (setOptions has one). Every odds:update writes the fixture's cache, so the countdown restarted before it could elapse — which is exactly why the observed cadence tracked odds traffic (648 s with zero refetches on a quiet market, p50 0.8 s on a busy one) rather than the clock. The timer is now owned outside the query observer and cannot be reset by a cache write. Read at query-core 5.101.2; strykr-fe pins ^5.90.18.

The final score is now pushed too. match/detail_live cannot report an ending (ADR-0035), so the push path went silent on a match's last in-play frame and the concluded scoreline waited for a REST poll. The diary — the only endpoint that reports a terminal status and the final score — now publishes through this same channel the moment it observes the transition, once per match. Completion remains authoritative from the 5-minute diary cycle and per-sport status_id mappings only; it is never inferred from time, score, silence, or a match disappearing from the live feed.

Rule 8 is unchanged and still binding. REST keeps the baseline, the lifecycle, and the terminal status and final score. What changed is that it no longer delivers in-play changes by accident, and its recovery cadence runs on a clock of its own rather than on odds traffic.

The cadence is 30 s connected / 5 s disconnected — the disconnected half deliberately UNCHANGED. An earlier revision of this work raised it to 10 s on the reasoning that push is unconditional now, so HTTP is only the backstop. That argument is inverted for the data this interval actually governs: getFixtureDetailRefetchInterval drives the whole fixture-detail query — markets, prices, market membership, fixture status — not just TheSports scores. When the socket is down there is no odds:update either, so HTTP is the only delivery path for a price, and doubling the interval doubles the worst-case staleness of a price a punter is looking at. That is a betting-surface freshness parameter and not this ADR's to move. Ruled 2026-09-04: 5 s stands; 30 s connected is the part this change sets.