Sports Upcoming Feed Architecture — Before and After
GET /api/fixtures/upcoming — the endpoint behind the Sports page tournament list.
Companion to CAROUSEL_FEED_ARCHITECTURE.md, which covers
/carousel/grouped (the Home rails). That work made discovery cheap and moved pricing out of
discovery. This one applies the same rank-then-price idea to the Sports page, and then fixes the
thing that was left dominating the request once pricing was isolated: the market-book fetch was
serialised.
Numbers here were measured on dev1. Read §7 before quoting any of them — dev1 runs
BIFROST_ENABLED=false, so one significant code path never executed in these measurements.
1. What this endpoint returns
Upcoming fixtures for one sport (?sportId=) or across sports, grouped by tournament in the UI.
Each fixture carries its match-odds market, the derived display fields (bestOdds,
liquidityScore, totalVolume, marketCount, hasOdds), and viewer-currency stake limits.
The Sports page renders these as collapsible tournament panels, at most 5 matches per panel
(FIXTURES_PER_TOURNAMENT), expanding client-side from data already in memory.
2. The problem, as measured
| Phase | Work | Cost |
|---|---|---|
| Fixture discovery | 7 daily listMarketCatalogue per sport, uncached | large |
| Market books | listMarketBook for every discovered fixture, in batches of 40, each batch awaited then followed by a 100 ms sleep | dominant |
| Gating | catalogue overlay + player visibility ran after pricing, discarding fixtures already paid for | pure waste |
Two independent problems, both invisible in a flame graph that only reports totals:
- Pricing was unconditional. Everything discovered inside a 7-day horizon got a market book, then the 36 h window, the admin overlay and the visibility filter threw much of it away.
- The book fetch was serial.
Nbatches costN×RTT + (N-1)×100 ms. For a ~350-market soccer request that is 9 sequential round-trips plus 800 ms of puresetTimeout.
3. Before → After
The shape of the change: every box that can decide without a price moved above the pricing box, and the pricing box itself stopped waiting on itself.
Two measurement rounds, taken hours apart. The fixture counts differ between them (soccer inside a 36h window varies by time of day), so the rounds are not like-for-like with each other — each row is only comparable within its own round.
Round 1 — before the pacer was made process-wide, 274 fixtures:
| Before | After | |
|---|---|---|
| Sequential median (soccer) | 3.768 s | 1.114 s |
| 5 concurrent, each | — | ~2.08 s |
| 10 concurrent | — | all 10 in 3.61 s |
provider_fanout (pricing stage) | 1.079 s | 0.789 s |
provider_fixture_fetch | 1.386 s | 0.655 s |
catalogue_overlay | 0.211 s | 0.152 s |
Round 2 — current head, pacing process-wide, 149 fixtures:
| Value | |
|---|---|
| Sequential median | 0.842 s (best 0.584, worst 1.126) |
| 5 concurrent | 0.770 – 2.456 s |
| 10 concurrent | 1.61 – 4.67 s, 4.72 s wall clock |
Round 3 — after merging dev (74 commits) into the branch, 57 fixtures. Both arms measured
minutes apart on the same build, which makes this the only round with no data drift between
the compared numbers:
Pre-PR path (?deferMarkets=false) | This PR | |
|---|---|---|
| median | 1.470 s | 0.285 s |
| runs | 1.466 / 1.470 / 1.743 | 0.309 / 0.285 / 0.261 |
| 10 concurrent | — | 1.83 – 3.99 s, 4.08 s wall |
5.2× on identical data and build. Treat this as the headline rather than any cross-round comparison: soccer fixtures inside a 36h window ranged from 352 to 57 across these rounds depending on the hour, so absolute times are only meaningful within a round.
Round 4 — dispatch interval halved from 100ms to 50ms (ceiling 10 → 20 req/s), 60 fixtures. The only change from round 3 is that constant, and the fixture counts are close enough (57 vs 60) to compare directly:
| 100 ms (round 3) | 50 ms (round 4) | |
|---|---|---|
| solo median | 0.285 s | 0.282 s |
| 10 concurrent, wall clock | 4.08 s | 1.53 s |
| A/B on the same build | 1.470 s → 0.285 s | 1.024 s → 0.320 s |
2.7× better under load, and solo is unchanged — which is the clearest confirmation that the
ceiling only ever bound concurrent traffic, never a single request. Zero TOO_MANY_REQUESTS,
TEMPORARY_BAN, TOO_MUCH_DATA or failed batches after the change.
Round 5 — pacing moved INSIDE the concurrency gate so it bounds actual dispatches rather than reservations (see §4.2). Same 60 fixtures as round 4, so directly comparable:
| pacing outside the gate | pacing inside the gate | |
|---|---|---|
| solo median | 0.282 s | 0.343 s |
| 10 concurrent, wall clock | 1.53 s | 1.95 s |
| rate ceiling actually held | no — 0 ms gaps observed | yes |
Correctness costs about 28% under load, and that is the right trade. Round 4's faster number was partly an artefact of a ceiling that did not really hold: reserving the slot before acquiring a permit paced the reservation, so waiters whose slots had elapsed fired together the moment permits freed. Zero rate-limit errors in both rounds — which is exactly why this could not have been found by watching for errors, only by measuring the dispatch timestamps.
The concurrency figure regressed on purpose: 3.61 s → 4.72 s wall clock for ten requests. That is the global rate ceiling doing its job — round 1 was fast partly because the ceiling was per-call and the process rate was unbounded. Note the regression is understated here, because round 2 had fewer fixtures (149 vs 274) and was still slower under load, which is what confirms the pacer is now the binding constraint rather than the book count.
Do not quote the sub-second sequential median as the win from this work: round 2 simply had a
smaller fixture set. The defensible sequential claim is round 1's like-for-like 3.768 s → 1.114 s,
measured against ?deferMarkets=false on the same backend minutes apart.
Output is unchanged, and that was verified rather than assumed: against the same backend with
?deferMarkets=false, both arms returned the same 168 fixture ids, the same markets, the same
504 outcomes, all five derived fields present on 168/168, and zero fixtures without an
actionable price.
4. The three architectural changes
4.1 The price-independent gates moved above pricing
The catalogue overlay does two separable jobs. Its tournament-level half — catalogue
visibility, the cricket fail-closed path, display-name overrides, and the tournamentTier stamp —
needs no prices. Its market-level half — the dictionary bettability gate and per-market
hide/rename — needs markets to act on.
So it runs twice, exactly as /carousel/grouped already does:
| Pass 1 (pre-pricing) | Pass 2 (post-pricing) | |
|---|---|---|
| Input | every windowed fixture | only the priced survivors |
| Does | tournament gate + tier stamp | market gate + hide/rename |
trustExistingTier | no — it is what stamps the tier | yes |
trustExistingTier is not an optimisation. buildTournamentGroupAggregates builds its DisjointSet
over the tournaments present in its input and unions them by shared match fingerprint, so its
result is set-dependent. Pass 2 sees a narrowed set, so a tournament whose tier came only from a
union through a fixture that is no longer present would re-form as its own component, lose its
tier, and be dropped after having already survived every gate — silently. Pass 2 therefore reuses
pass 1's stamped values instead of re-deriving them.
The visibility filter also moved above pricing, and runs only once: filterFixturesForPlayer
reads f.sportId and resolveTierForFixture(f.id) and never touches markets, so its verdict
cannot change once prices land.
Why this is safe on unpriced fixtures. transformFixtureBasic assigns markets only when the
canonical fixture has them, so a deferred fixture carries markets: undefined. The overlay's
drop-if-nothing-visible gate is fixture.markets && visibleMarkets.length === 0, which is falsy on
undefined. Were it ever markets: [] — truthy — pass 1 would drop every fixture and the
route would return 200 with an empty list. That is not hypothetical: an unguarded
filterFixturesWithMarkets emptied the Home rails exactly this way during the carousel work.
4.2 Market-book round-trips overlap instead of serialising
fetchMarketBooksBatched awaited each listMarketBook and then slept 100 ms. Batches are now
dispatched through a bounded-concurrency gate with the pacing applied to dispatches rather than
completions.
| Before | After | |
|---|---|---|
| Batch size | 40 | 40 (unchanged — see below) |
| Concurrency | 1 | ≤ 3, per process |
| Delay applies to | completions (serialising RTT) | dispatches |
| Interval | 100 ms | 50 ms (ceiling 20 req/s) |
| Cost of N batches | N×RTT + (N-1)×100ms | ≈ max(N×interval, RTT + …) |
This is not rate-neutral, and the code says so. The old loop reached only about
1/(RTT + 100 ms) requests per second because the RTT sat inside the serial chain; this version
actually reaches the 1-per-interval ceiling the constant always named — now 50 ms, i.e. 20 req/s.
The interval is the ONLY knob that raises throughput: BOOK_FETCH_CONCURRENCY (3) binds only
when RTT exceeds concurrency × interval, so raising it alone changes nothing. 20 req/s was
chosen to sit inside the headroom the OLD shape implies: that code had no global bound at all, so
its peak rate scaled with concurrency and would have reached roughly 50 req/s at ten concurrent
requests.
Correction to how that was evidenced. Earlier revisions of this document, and the commit that
halved the interval, claimed "zero rate-limit errors across 7 days of dev logs". That claim was not
obtainable. docker logs --since cannot see past the current container's lifetime, and the
container had been up under an hour — so a 168h window returned the same lines a 24h window did,
and the real sample was ~48 minutes on a low-traffic box. Verified again post-merge: a
--since 168h query against a 35-minute-old container returns logs beginning at container start.
The choice of 20 req/s still stands, but on the MECHANISM — the old code genuinely had no global bound, so its peak genuinely scaled with users — not on an observed-clean week. What nobody in this work established is Betfair's actual documented data-request limit, so treat 20 req/s as conservative-pending-verification rather than validated.
Both bounds are process-wide, and the slot is reserved while HOLDING a permit, immediately
before the provider call. Reserving it earlier — outside the gate, so a waiting task would not
occupy a permit — paces the RESERVATION instead of the dispatch: a task whose slot elapsed while
it queued fires as soon as a permit frees, so several long unevenly timed round-trips completing
together dispatch in the same event-loop turn with no spacing, while the in-flight count still
reads as 3. Reproduced at slots [0,50,100,150,200,250] against RTTs [1000,950,900], which
dispatched at [0,50,100,1000,1000,1000]. Reserving inside the gate makes reservation order equal
dispatch order, so spacing holds by construction.
An earlier revision of this change paced per call, so ten
simultaneous Sports requests each started their own pacer and the process rate was effectively
3/RTT while the in-flight bound still read as 3 — the ceiling existed only within one call.
Caught in review; the pacing state is now a static slot reservation. Consequence worth naming: a
real global ceiling means concurrent requests queue on it, so throughput under load is bounded
by design rather than by luck, and the earlier "10 concurrent in 3.61 s" figure was measured
without it.
Reads have no
TOO_MANY_REQUESTS backoff — that retry path is placement-only — and a breach counts toward ban
detection, so the concurrency bound is deliberately small. Over 20 minutes of dev1 traffic:
TOO_MANY_REQUESTS 0, TEMPORARY_BAN 0, TOO_MUCH_DATA 0, failed batches 0.
The gate reuses BetfairClient's existing AsyncSemaphore (the same primitive behind
LIST_CLEARED_ORDERS_GATE) and is per process, because Betfair rate-limits the account — a
per-instance gate would multiply real concurrency by the number of adapters. It is built lazily
rather than as a static field initialiser; see §7.
BOOK_BATCH_SIZE = 40 is a hard ceiling, not caution. Betfair budgets each request by weight,
max 200, and EX_BEST_OFFERS costs 5 per market: 200 / 5 = 40. The 200 used elsewhere in the
adapter is a MARKET_DESCRIPTION catalogue fetch at weight 1. The old comment here called 40 "the
long-standing conservative batch size", which invited raising it — that would cost 1000 weight and
be rejected.
4.3 The derived display fields come from one place
The deferred path hand-rolled market attachment and got a subset of the fields
transformFixtureBasic normally derives. Two stayed undefined on every fixture the route
returned:
liquidityScore—sortFixtures('liquidity')readsliquidityScore ?? 999, so with it undefined everywhere that sort silently collapsed to a no-op.bestOdds—MatchCardgatesPmatchDisplay, the inline matched-exposure preview, onbestOdds.<side>.marketId, so Sports-page cards silently lost it.
Odds themselves were never affected: useMatchOddsMarket reads fixture.markets.
The fix is to call marginService.transformFixtureWithOdds, which is the definition of "a
frontend fixture that has prices", instead of re-deriving a subset of it. Its output is grafted
onto the existing fixture rather than returned wholesale, because that fixture carries overlay
pass 1's stamping and pass 2 runs with trustExistingTier — a fresh object rebuilt from the
canonical would hand pass 2 an unstamped tier.
This was the second bug from re-deriving that method's work; an earlier inlined totalVolume sum
had already degraded sort=popular to start-time order.
5. The Sports list is primary-exchange only
/fixtures/upcoming passes supplementFixtures: 'primary-only'. This is a product decision —
the Sports fixture list shows primary-exchange fixtures — and it was taken knowingly:
- Supplement-only events stop appearing in this list. Measured on dev: 4 of 500 upcoming soccer fixtures were Bifrost-sourced. They are gone from the Sports page. Accepted deliberately; it was previously rejected when proposed purely as a latency optimisation, because deleting real content is not a performance win.
- It does NOT remove Bifrost markets from a match page. Supplement markets reach a fixture
through
getOdds/mergeSupplementMarketson the detail route, which never callsgetFixtures. A match page is unchanged. - The AI-chat welcome shares this route.
StrykrWelcome's suggestion chips lose supplement-only matches too. Cosmetic — the chat can still analyse any fixture by id.
It also removes a real hazard rather than papering over one, which is why it interacts with §4.1.
With the merge on, the catalogue overlay collapses same-match provider rows by matchFingerprint
and keeps ...base — whichever row it happened to see first. Under this route's pre-pricing
overlay pass those rows are still unpriced, and the pricing step resolves only the surviving base
id, so a Bifrost-first row would carry Bifrost identity and drop Betfair's markets. That is
non-deterministic and in the wrong direction. With no supplement fixtures there is no cross-provider
merge on this route at all.
/fixtures/live, the flat /carousel, and the fixture detail route keep the default additive
merge. Only this list opts out.
6. What was tried and rejected
Recorded so nobody spends the afternoon again. Each was measured, not reasoned away.
| Idea | Verdict | Evidence |
|---|---|---|
Serve prices from the streamed BetfairCache, as the match page does | Rejected here | The stream subscribes only in-play or viewer-retained markets and never firehoses. An upcoming list is 352/352 not_started → ~0 % hit rate. Valuable for the Home live rail, not here. |
Raise BOOK_BATCH_SIZE 40 → 200 | Rejected | 40 already is Betfair's 200-weight ceiling for EX_BEST_OFFERS. 200 would be 1000 weight. |
| Cap priced fixtures at 5 per tournament, matching the UI | Rejected | 65 tournaments with a long tail: only 92 of 352 (26 %) are never rendered. ~26 % fewer books for ~1.16 s, and it breaks instant client-side expand. |
| Cache prices in Redis | Rejected by design | Prices must stay live on a betting surface. Redis holds discovery catalogues and MarketMeta identity only — no prices anywhere. |
7. Deliberate trade-offs
- Upcoming fixture membership can lag by the discovery TTL. Prices and live-status never lag —
the in-play scan and
listMarketBookare uncached. - The 36h window cannot be switched off by accident.
?windowHours=takes a positive number (clamped to the 168h discovery horizon) or the explicit stringallfor the full sweep. Anything else —abc,0,-5, a repeated parameter — falls back to the 36h default and is logged. The original parse skipped the filter entirely for all of those, which turned a typo in a public query string into the full-horizon price-everything workload this route exists to avoid. ?deferMarkets=falsepreserves the old shape exactly: markets arrive during discovery, one full-recompute overlay pass, visibility after pricing. It is the escape hatch and the A/B arm.- The concurrency change is wider than this route.
getFixturessharesfetchMarketBooksBatched, so/fixtures/live, the flat/carouseland/api/odds/*all got faster too (that arm went 3.768 s → 2.428 s). Wider blast radius than the route-scoped changes. /fixtures/livewas left alone. It carries a byte-identical overlay block; it is not in this change's scope.
8. Known limits of these measurements
State these before quoting the numbers.
BIFROST_ENABLED=falseon dev1. The supplement-merge path — which appends ~9,900 Bifrost soccer events for the overlay to resolve and then fail-close — never ran in any measurement here. Dev and prod have it on. This is the largest untested area.- API-level only. No browser verification of the rendered Sports page; dev1's frontend container was a different branch throughout.
- Soccer only since the concurrency change. Cricket measured 0.41 s, but before it.
- Small samples. Stage histograms are n≈9–17 requests; latency figures are 5 runs.
- Rate-limit observation windows are short, and shorter than once stated.
docker logs --sinceis bounded by container lifetime, so every "N days" figure in earlier revisions was really minutes-to-an-hour. All rate-limit counts quoted here (TOO_MANY_REQUESTS,TEMPORARY_BAN,TOO_MUCH_DATA, failed batches) are zero, but over tens of minutes, not days. This is the second log-derived overstatement in this work; the first was counting warns with a bracketed[warn]grep against JSON logs, which could only ever return zero.TOO_MANY_REQUESTS,TEMPORARY_BAN,TOO_MUCH_DATAand failed batches were all zero across both rounds. - An earlier revision of this document reported "0 warns/errors" on dev1. That figure was
wrong — it came from a grep for a bracketed
[warn]token, while dev1 logs JSON ("level":"warn"), so it could only ever return zero. Counted properly, dev1 shows ~137 warns and 8 errors per 25 minutes, none attributable to this change: the bulk arepopulateMarketMetaCache: unsupported dictionary outcome space(62) andno match_odds outcome bound to away team(54), both pre-existing dictionary issues, plus line-marketFINANCIAL_INTEGRITYerrors and a missingTIMESCALE_URLon the fraud worker. The substring greps for the Betfair rate-limit codes were unaffected and remain valid. - Re-verified after merging dev's 74 commits, 18 of which touched files this change depends
on (
BetfairAdapter.ts9,tournamentMarketViewService.ts6,marginService.ts3). dev changed the overlay's gate —isFixtureVisiblenow returns a string andisValidResolutiontakes an archived flag — but both remain price-independent, and the load-bearing empty-markets gate is textually unchanged (fixture.markets && visibleMarkets.length === 0), so the pre-pricing pass is still safe onmarkets: undefined. Checked deliberately: a semantic change there would merge cleanly and return 200 with an empty list. primary-onlyis still unverified in effect. dev1 hasBIFROST_ENABLED=false, so there are no supplement fixtures for it to exclude there. Its measured basis is the dev observation of 4 Bifrost fixtures in 500.
One process note worth keeping: the book-fetch gate is built lazily, not as a static field
initialiser, because this adapter sits on the import path of route modules that carry a latent
test-only registration race (a request landing before the router mounts, surfacing as
expected 404 to be 400). With an eager initialiser
nettingExplorer.market.test.ts failed 2 of 6 runs against 0 of 6 on baseline; lazily, 0 of 8.
That race is pre-existing and is not fixed here — settlements.reportParams.test.ts shows the
same signature independently of this branch. Lazy construction only keeps this change out of its
timing.