Skip to main content

Sports Upcoming Feed Architecture — Before and After

GET /api/fixtures/upcoming — the endpoint behind the Sports page tournament list.

Companion to CAROUSEL_FEED_ARCHITECTURE.md, which covers /carousel/grouped (the Home rails). That work made discovery cheap and moved pricing out of discovery. This one applies the same rank-then-price idea to the Sports page, and then fixes the thing that was left dominating the request once pricing was isolated: the market-book fetch was serialised.

Numbers here were measured on dev1. Read §7 before quoting any of them — dev1 runs BIFROST_ENABLED=false, so one significant code path never executed in these measurements.


1. What this endpoint returns​

Upcoming fixtures for one sport (?sportId=) or across sports, grouped by tournament in the UI. Each fixture carries its match-odds market, the derived display fields (bestOdds, liquidityScore, totalVolume, marketCount, hasOdds), and viewer-currency stake limits.

The Sports page renders these as collapsible tournament panels, at most 5 matches per panel (FIXTURES_PER_TOURNAMENT), expanding client-side from data already in memory.


2. The problem, as measured​

PhaseWorkCost
Fixture discovery7 daily listMarketCatalogue per sport, uncachedlarge
Market bookslistMarketBook for every discovered fixture, in batches of 40, each batch awaited then followed by a 100 ms sleepdominant
Gatingcatalogue overlay + player visibility ran after pricing, discarding fixtures already paid forpure waste

Two independent problems, both invisible in a flame graph that only reports totals:

  1. Pricing was unconditional. Everything discovered inside a 7-day horizon got a market book, then the 36 h window, the admin overlay and the visibility filter threw much of it away.
  2. The book fetch was serial. N batches cost N×RTT + (N-1)×100 ms. For a ~350-market soccer request that is 9 sequential round-trips plus 800 ms of pure setTimeout.

3. Before → After​

The shape of the change: every box that can decide without a price moved above the pricing box, and the pricing box itself stopped waiting on itself.

Two measurement rounds, taken hours apart. The fixture counts differ between them (soccer inside a 36h window varies by time of day), so the rounds are not like-for-like with each other — each row is only comparable within its own round.

Round 1 — before the pacer was made process-wide, 274 fixtures:

BeforeAfter
Sequential median (soccer)3.768 s1.114 s
5 concurrent, each—~2.08 s
10 concurrent—all 10 in 3.61 s
provider_fanout (pricing stage)1.079 s0.789 s
provider_fixture_fetch1.386 s0.655 s
catalogue_overlay0.211 s0.152 s

Round 2 — current head, pacing process-wide, 149 fixtures:

Value
Sequential median0.842 s (best 0.584, worst 1.126)
5 concurrent0.770 – 2.456 s
10 concurrent1.61 – 4.67 s, 4.72 s wall clock

Round 3 — after merging dev (74 commits) into the branch, 57 fixtures. Both arms measured minutes apart on the same build, which makes this the only round with no data drift between the compared numbers:

Pre-PR path (?deferMarkets=false)This PR
median1.470 s0.285 s
runs1.466 / 1.470 / 1.7430.309 / 0.285 / 0.261
10 concurrent—1.83 – 3.99 s, 4.08 s wall

5.2× on identical data and build. Treat this as the headline rather than any cross-round comparison: soccer fixtures inside a 36h window ranged from 352 to 57 across these rounds depending on the hour, so absolute times are only meaningful within a round.

Round 4 — dispatch interval halved from 100ms to 50ms (ceiling 10 → 20 req/s), 60 fixtures. The only change from round 3 is that constant, and the fixture counts are close enough (57 vs 60) to compare directly:

100 ms (round 3)50 ms (round 4)
solo median0.285 s0.282 s
10 concurrent, wall clock4.08 s1.53 s
A/B on the same build1.470 s → 0.285 s1.024 s → 0.320 s

2.7× better under load, and solo is unchanged — which is the clearest confirmation that the ceiling only ever bound concurrent traffic, never a single request. Zero TOO_MANY_REQUESTS, TEMPORARY_BAN, TOO_MUCH_DATA or failed batches after the change.

Round 5 — pacing moved INSIDE the concurrency gate so it bounds actual dispatches rather than reservations (see §4.2). Same 60 fixtures as round 4, so directly comparable:

pacing outside the gatepacing inside the gate
solo median0.282 s0.343 s
10 concurrent, wall clock1.53 s1.95 s
rate ceiling actually heldno — 0 ms gaps observedyes

Correctness costs about 28% under load, and that is the right trade. Round 4's faster number was partly an artefact of a ceiling that did not really hold: reserving the slot before acquiring a permit paced the reservation, so waiters whose slots had elapsed fired together the moment permits freed. Zero rate-limit errors in both rounds — which is exactly why this could not have been found by watching for errors, only by measuring the dispatch timestamps.

The concurrency figure regressed on purpose: 3.61 s → 4.72 s wall clock for ten requests. That is the global rate ceiling doing its job — round 1 was fast partly because the ceiling was per-call and the process rate was unbounded. Note the regression is understated here, because round 2 had fewer fixtures (149 vs 274) and was still slower under load, which is what confirms the pacer is now the binding constraint rather than the book count.

Do not quote the sub-second sequential median as the win from this work: round 2 simply had a smaller fixture set. The defensible sequential claim is round 1's like-for-like 3.768 s → 1.114 s, measured against ?deferMarkets=false on the same backend minutes apart.

Output is unchanged, and that was verified rather than assumed: against the same backend with ?deferMarkets=false, both arms returned the same 168 fixture ids, the same markets, the same 504 outcomes, all five derived fields present on 168/168, and zero fixtures without an actionable price.


4. The three architectural changes​

4.1 The price-independent gates moved above pricing​

The catalogue overlay does two separable jobs. Its tournament-level half — catalogue visibility, the cricket fail-closed path, display-name overrides, and the tournamentTier stamp — needs no prices. Its market-level half — the dictionary bettability gate and per-market hide/rename — needs markets to act on.

So it runs twice, exactly as /carousel/grouped already does:

Pass 1 (pre-pricing)Pass 2 (post-pricing)
Inputevery windowed fixtureonly the priced survivors
Doestournament gate + tier stampmarket gate + hide/rename
trustExistingTierno — it is what stamps the tieryes

trustExistingTier is not an optimisation. buildTournamentGroupAggregates builds its DisjointSet over the tournaments present in its input and unions them by shared match fingerprint, so its result is set-dependent. Pass 2 sees a narrowed set, so a tournament whose tier came only from a union through a fixture that is no longer present would re-form as its own component, lose its tier, and be dropped after having already survived every gate — silently. Pass 2 therefore reuses pass 1's stamped values instead of re-deriving them.

The visibility filter also moved above pricing, and runs only once: filterFixturesForPlayer reads f.sportId and resolveTierForFixture(f.id) and never touches markets, so its verdict cannot change once prices land.

Why this is safe on unpriced fixtures. transformFixtureBasic assigns markets only when the canonical fixture has them, so a deferred fixture carries markets: undefined. The overlay's drop-if-nothing-visible gate is fixture.markets && visibleMarkets.length === 0, which is falsy on undefined. Were it ever markets: [] — truthy — pass 1 would drop every fixture and the route would return 200 with an empty list. That is not hypothetical: an unguarded filterFixturesWithMarkets emptied the Home rails exactly this way during the carousel work.

4.2 Market-book round-trips overlap instead of serialising​

fetchMarketBooksBatched awaited each listMarketBook and then slept 100 ms. Batches are now dispatched through a bounded-concurrency gate with the pacing applied to dispatches rather than completions.

BeforeAfter
Batch size4040 (unchanged — see below)
Concurrency1≤ 3, per process
Delay applies tocompletions (serialising RTT)dispatches
Interval100 ms50 ms (ceiling 20 req/s)
Cost of N batchesN×RTT + (N-1)×100ms≈ max(N×interval, RTT + …)

This is not rate-neutral, and the code says so. The old loop reached only about 1/(RTT + 100 ms) requests per second because the RTT sat inside the serial chain; this version actually reaches the 1-per-interval ceiling the constant always named — now 50 ms, i.e. 20 req/s.

The interval is the ONLY knob that raises throughput: BOOK_FETCH_CONCURRENCY (3) binds only when RTT exceeds concurrency × interval, so raising it alone changes nothing. 20 req/s was chosen to sit inside the headroom the OLD shape implies: that code had no global bound at all, so its peak rate scaled with concurrency and would have reached roughly 50 req/s at ten concurrent requests.

Correction to how that was evidenced. Earlier revisions of this document, and the commit that halved the interval, claimed "zero rate-limit errors across 7 days of dev logs". That claim was not obtainable. docker logs --since cannot see past the current container's lifetime, and the container had been up under an hour — so a 168h window returned the same lines a 24h window did, and the real sample was ~48 minutes on a low-traffic box. Verified again post-merge: a --since 168h query against a 35-minute-old container returns logs beginning at container start.

The choice of 20 req/s still stands, but on the MECHANISM — the old code genuinely had no global bound, so its peak genuinely scaled with users — not on an observed-clean week. What nobody in this work established is Betfair's actual documented data-request limit, so treat 20 req/s as conservative-pending-verification rather than validated.

Both bounds are process-wide, and the slot is reserved while HOLDING a permit, immediately before the provider call. Reserving it earlier — outside the gate, so a waiting task would not occupy a permit — paces the RESERVATION instead of the dispatch: a task whose slot elapsed while it queued fires as soon as a permit frees, so several long unevenly timed round-trips completing together dispatch in the same event-loop turn with no spacing, while the in-flight count still reads as 3. Reproduced at slots [0,50,100,150,200,250] against RTTs [1000,950,900], which dispatched at [0,50,100,1000,1000,1000]. Reserving inside the gate makes reservation order equal dispatch order, so spacing holds by construction. An earlier revision of this change paced per call, so ten simultaneous Sports requests each started their own pacer and the process rate was effectively 3/RTT while the in-flight bound still read as 3 — the ceiling existed only within one call. Caught in review; the pacing state is now a static slot reservation. Consequence worth naming: a real global ceiling means concurrent requests queue on it, so throughput under load is bounded by design rather than by luck, and the earlier "10 concurrent in 3.61 s" figure was measured without it.

Reads have no TOO_MANY_REQUESTS backoff — that retry path is placement-only — and a breach counts toward ban detection, so the concurrency bound is deliberately small. Over 20 minutes of dev1 traffic: TOO_MANY_REQUESTS 0, TEMPORARY_BAN 0, TOO_MUCH_DATA 0, failed batches 0.

The gate reuses BetfairClient's existing AsyncSemaphore (the same primitive behind LIST_CLEARED_ORDERS_GATE) and is per process, because Betfair rate-limits the account — a per-instance gate would multiply real concurrency by the number of adapters. It is built lazily rather than as a static field initialiser; see §7.

BOOK_BATCH_SIZE = 40 is a hard ceiling, not caution. Betfair budgets each request by weight, max 200, and EX_BEST_OFFERS costs 5 per market: 200 / 5 = 40. The 200 used elsewhere in the adapter is a MARKET_DESCRIPTION catalogue fetch at weight 1. The old comment here called 40 "the long-standing conservative batch size", which invited raising it — that would cost 1000 weight and be rejected.

4.3 The derived display fields come from one place​

The deferred path hand-rolled market attachment and got a subset of the fields transformFixtureBasic normally derives. Two stayed undefined on every fixture the route returned:

  • liquidityScore — sortFixtures('liquidity') reads liquidityScore ?? 999, so with it undefined everywhere that sort silently collapsed to a no-op.
  • bestOdds — MatchCard gates PmatchDisplay, the inline matched-exposure preview, on bestOdds.<side>.marketId, so Sports-page cards silently lost it.

Odds themselves were never affected: useMatchOddsMarket reads fixture.markets.

The fix is to call marginService.transformFixtureWithOdds, which is the definition of "a frontend fixture that has prices", instead of re-deriving a subset of it. Its output is grafted onto the existing fixture rather than returned wholesale, because that fixture carries overlay pass 1's stamping and pass 2 runs with trustExistingTier — a fresh object rebuilt from the canonical would hand pass 2 an unstamped tier.

This was the second bug from re-deriving that method's work; an earlier inlined totalVolume sum had already degraded sort=popular to start-time order.


5. The Sports list is primary-exchange only​

/fixtures/upcoming passes supplementFixtures: 'primary-only'. This is a product decision — the Sports fixture list shows primary-exchange fixtures — and it was taken knowingly:

  • Supplement-only events stop appearing in this list. Measured on dev: 4 of 500 upcoming soccer fixtures were Bifrost-sourced. They are gone from the Sports page. Accepted deliberately; it was previously rejected when proposed purely as a latency optimisation, because deleting real content is not a performance win.
  • It does NOT remove Bifrost markets from a match page. Supplement markets reach a fixture through getOdds / mergeSupplementMarkets on the detail route, which never calls getFixtures. A match page is unchanged.
  • The AI-chat welcome shares this route. StrykrWelcome's suggestion chips lose supplement-only matches too. Cosmetic — the chat can still analyse any fixture by id.

It also removes a real hazard rather than papering over one, which is why it interacts with §4.1. With the merge on, the catalogue overlay collapses same-match provider rows by matchFingerprint and keeps ...base — whichever row it happened to see first. Under this route's pre-pricing overlay pass those rows are still unpriced, and the pricing step resolves only the surviving base id, so a Bifrost-first row would carry Bifrost identity and drop Betfair's markets. That is non-deterministic and in the wrong direction. With no supplement fixtures there is no cross-provider merge on this route at all.

/fixtures/live, the flat /carousel, and the fixture detail route keep the default additive merge. Only this list opts out.


6. What was tried and rejected​

Recorded so nobody spends the afternoon again. Each was measured, not reasoned away.

IdeaVerdictEvidence
Serve prices from the streamed BetfairCache, as the match page doesRejected hereThe stream subscribes only in-play or viewer-retained markets and never firehoses. An upcoming list is 352/352 not_started → ~0 % hit rate. Valuable for the Home live rail, not here.
Raise BOOK_BATCH_SIZE 40 → 200Rejected40 already is Betfair's 200-weight ceiling for EX_BEST_OFFERS. 200 would be 1000 weight.
Cap priced fixtures at 5 per tournament, matching the UIRejected65 tournaments with a long tail: only 92 of 352 (26 %) are never rendered. ~26 % fewer books for ~1.16 s, and it breaks instant client-side expand.
Cache prices in RedisRejected by designPrices must stay live on a betting surface. Redis holds discovery catalogues and MarketMeta identity only — no prices anywhere.

7. Deliberate trade-offs​

  • Upcoming fixture membership can lag by the discovery TTL. Prices and live-status never lag — the in-play scan and listMarketBook are uncached.
  • The 36h window cannot be switched off by accident. ?windowHours= takes a positive number (clamped to the 168h discovery horizon) or the explicit string all for the full sweep. Anything else — abc, 0, -5, a repeated parameter — falls back to the 36h default and is logged. The original parse skipped the filter entirely for all of those, which turned a typo in a public query string into the full-horizon price-everything workload this route exists to avoid.
  • ?deferMarkets=false preserves the old shape exactly: markets arrive during discovery, one full-recompute overlay pass, visibility after pricing. It is the escape hatch and the A/B arm.
  • The concurrency change is wider than this route. getFixtures shares fetchMarketBooksBatched, so /fixtures/live, the flat /carousel and /api/odds/* all got faster too (that arm went 3.768 s → 2.428 s). Wider blast radius than the route-scoped changes.
  • /fixtures/live was left alone. It carries a byte-identical overlay block; it is not in this change's scope.

8. Known limits of these measurements​

State these before quoting the numbers.

  • BIFROST_ENABLED=false on dev1. The supplement-merge path — which appends ~9,900 Bifrost soccer events for the overlay to resolve and then fail-close — never ran in any measurement here. Dev and prod have it on. This is the largest untested area.
  • API-level only. No browser verification of the rendered Sports page; dev1's frontend container was a different branch throughout.
  • Soccer only since the concurrency change. Cricket measured 0.41 s, but before it.
  • Small samples. Stage histograms are n≈9–17 requests; latency figures are 5 runs.
  • Rate-limit observation windows are short, and shorter than once stated. docker logs --since is bounded by container lifetime, so every "N days" figure in earlier revisions was really minutes-to-an-hour. All rate-limit counts quoted here (TOO_MANY_REQUESTS, TEMPORARY_BAN, TOO_MUCH_DATA, failed batches) are zero, but over tens of minutes, not days. This is the second log-derived overstatement in this work; the first was counting warns with a bracketed [warn] grep against JSON logs, which could only ever return zero. TOO_MANY_REQUESTS, TEMPORARY_BAN, TOO_MUCH_DATA and failed batches were all zero across both rounds.
  • An earlier revision of this document reported "0 warns/errors" on dev1. That figure was wrong — it came from a grep for a bracketed [warn] token, while dev1 logs JSON ("level":"warn"), so it could only ever return zero. Counted properly, dev1 shows ~137 warns and 8 errors per 25 minutes, none attributable to this change: the bulk are populateMarketMetaCache: unsupported dictionary outcome space (62) and no match_odds outcome bound to away team (54), both pre-existing dictionary issues, plus line-market FINANCIAL_INTEGRITY errors and a missing TIMESCALE_URL on the fraud worker. The substring greps for the Betfair rate-limit codes were unaffected and remain valid.
  • Re-verified after merging dev's 74 commits, 18 of which touched files this change depends on (BetfairAdapter.ts 9, tournamentMarketViewService.ts 6, marginService.ts 3). dev changed the overlay's gate — isFixtureVisible now returns a string and isValidResolution takes an archived flag — but both remain price-independent, and the load-bearing empty-markets gate is textually unchanged (fixture.markets && visibleMarkets.length === 0), so the pre-pricing pass is still safe on markets: undefined. Checked deliberately: a semantic change there would merge cleanly and return 200 with an empty list.
  • primary-only is still unverified in effect. dev1 has BIFROST_ENABLED=false, so there are no supplement fixtures for it to exclude there. Its measured basis is the dev observation of 4 Bifrost fixtures in 500.

One process note worth keeping: the book-fetch gate is built lazily, not as a static field initialiser, because this adapter sits on the import path of route modules that carry a latent test-only registration race (a request landing before the router mounts, surfacing as expected 404 to be 400). With an eager initialiser nettingExplorer.market.test.ts failed 2 of 6 runs against 0 of 6 on baseline; lazily, 0 of 8. That race is pre-existing and is not fixed here — settlements.reportParams.test.ts shows the same signature independently of this branch. Lazy construction only keeps this change out of its timing.