CI/CD Pipeline
How code gets from a feature branch to a running environment, what it costs, and where the gaps are.
Last verified against the repo and the live hosts on 2026-09-05. Timings and run IDs come from real runs. Where a number is an estimate it says so.
Branching model
feature/* ──PR──▶ dev ──PR──▶ staging ──PR──▶ main ──automatic──▶ production
Two workflows enforce this:
restrict-staging-base.yml— PRs intostagingmust come fromdev, or from ahotfix/*branch (allowed with a notice telling you to cherry-pick back todev). The head repo must be this repo.restrict-main-base.yml— PRs intomainmust come fromstaging.
Merging to main deploys production automatically. workflow_dispatch is retained as the
rollback path — dispatch with ref set to an earlier commit or tag.
Note that restrict-main-base.yml gates pull requests into main only. A direct push to
main bypasses it and deploys, and no branch protection currently prevents that.
Workflow inventory
| Workflow | Trigger | What it does |
|---|---|---|
sanity-tests.yml | push + PR to dev, paths backend/** | Financial sanity integration tests against a real Postgres, plus a backend build + unit test job |
typecheck-frontend.yml | PR to dev / main | Design-system check, tsc --noEmit, frontend unit specs |
build-images.yml | PR to dev / staging | Builds the service Dockerfiles without pushing. Holds no packages: permission |
publish-images.yml | workflow_call | Builds and pushes one environment's images to GHCR. The only holder of packages: write |
restrict-staging-base.yml | PR to staging | Head must be dev or hotfix/* |
restrict-main-base.yml | PR to main | Head must be staging |
deploy-dev.yml | push to dev | Publishes dev-<sha> images, then deploys dev.strykr.io |
deploy-staging.yml | push to staging | Publishes staging-<sha> images, then deploys staging.strykr.io |
deploy-dev1.yml | push to feat/dev1 | Deploys dev1.strykr.io — still builds on the server |
deploy-dev2.yml | push to feat/dev2 | Deploys dev2.strykr.io — still builds on the server |
deploy-prod.yml | push to main, plus workflow_dispatch | Deploys strykr.io — still builds on the server |
Why typecheck-frontend has no path filter
It runs on every PR to dev and main, then decides internally whether to do the
expensive work by diffing strykr-fe/. Filtering at the on: level would report the
check as skipped on backend-only PRs, and a required-check rule on main would then
block those PRs forever.
Environments
Every port binds to 127.0.0.1; host nginx terminates TLS and proxies in.
| Env | Branch | Trigger | Host dir | Compose file | FE | BE | Images from |
|---|---|---|---|---|---|---|---|
| dev | dev | automatic | /root/strykr | docker-compose.dev.yml | 3000 | 3001 | GHCR (dev-<sha>) |
| dev1 | feat/dev1 | automatic | /root/strykr-dev1 | docker-compose.dev1.yml | 3002 | 3003 | built on host |
| dev2 | feat/dev2 | automatic | /root/strykr-dev2 | docker-compose.dev2.yml | 3004 | 3005 | built on host |
| staging | staging | automatic | /root/strykr-staging | docker-compose.staging.yml | 3000 | 3001 | GHCR (staging-<sha>) |
| prod | any ref, default main | manual | /root/strykr | docker-compose.prod.yml | 3000 | 3001 | built on host |
dev, dev1 and dev2 share one VPS and one Docker daemon. dev1 and dev2 run no data
layer of their own — their deploy first brings up postgres, redis and timescaledb
from /root/strykr, then joins that network with a separate database. staging and prod
are separate hosts; the /root/strykr path collision between dev and prod is
coincidental.
Images
Each environment has its own image reference. Nothing is shared:
ghcr.io/forsyt-io/hannibal-<service>:<env>-<sha> what deploys pin
ghcr.io/forsyt-io/hannibal-<service>:<env> moving, convenience only
The <env>- prefix is the safeguard — no tag one environment resolves is reachable by
another, so prod cannot pull a dev image even by mistake.
<env>-<sha> is commit-scoped, not immutable: re-running a workflow overwrites the
tag, and floating base images mean the same commit can rebuild to different bytes.
Deploying by digest (image@sha256:…) is the only real guarantee and is not yet done.
Backend and ai-chat only. Both Dockerfiles declare zero build args and read all
config from compose env_file at runtime. The frontend bakes 14 NEXT_PUBLIC_* values
at build time and those differ per environment, so it is still built on each host.
Publishing it requires GitHub Environments holding those values — none exist today.
Registry auth uses the job-scoped GITHUB_TOKEN passed over SSH and logged out after
use. No long-lived PAT lives on any host.
Deploy anatomy
dev and staging
publish-imagesbuildsbackendandai-chaton a runner and pushes<env>-<sha>- The deploy job runs with
if: ${{ !cancelled() }}— a superseded run must not deploy git fetch+git checkout --detach <sha>— the exact commit the images were built from, not the branch tipdocker login ghcr.io→docker compose pull backend ai-chat→docker logout- Build the frontend. If the pull failed, build backend and ai-chat too
prisma migrate deploy→docker compose up -dverify-deployment-observability.sh, thenreconcile-monitoring.sh
Step 3 matters. Pinning images to a commit while the host took the branch tip would let one commit's backend run against another commit's migrations — staging queues rather than cancels, so that was a live race, not a theoretical one.
Step 4/5 is best-effort by design: if login fails, the registry is unreachable, or publishing failed, the deploy falls back to building on the host. A registry problem degrades to the old behaviour instead of blocking a release.
dev1, dev2, prod
Unchanged — still docker compose build on the host.
Prod additionally refuses to run on a dirty tree, loads reconcile-monitoring.sh from
the workflow's commit rather than the deploy ref, greps the target ref for six
observability-baseline markers, restarts host nginx after container rotation, and
asserts https://strykr.io/metrics returns 404. Its user inputs are passed via env:
and consumed as shell variables, never interpolated into run:.
What the registry actually cost
Honest accounting, because the expectation was a speed-up and that is not what happened.
Wall clock for deploy-dev, same workflow, consecutive days:
| Run | Date | Total | Notes |
|---|---|---|---|
| 33617145511 | 09-02 | 4m 56s | before the Next cache mount |
| 33874971699 | 09-04 | 5m 13s | cache mount merged, still building on host |
| 33887197360 | 09-04 | 6m 42s | same |
| 33902137585 | 09-04 | 6m 25s | same |
| 33932455282 | 09-05 | 8m 07s | first run pulling from GHCR |
Breakdown of the GHCR run: publish-images 3m 43s (cold GHA layer cache), then the
deploy job 4m 18s — of which the frontend build is 2m 44s and the image pull 30s.
So moving image builds into CI made dev's wall clock worse by roughly two minutes. The image build went from a near-free layer-cache hit on the host to a cold build on a runner that runs serially before the deploy. The GHA cache should warm on subsequent runs; that has not been measured yet.
Two further cautions on these numbers: the spread across identical pre-GHCR runs is 4m 56s to 6m 42s, so single-run comparisons are weak; and the Next cache mount shows no clear wall-clock win either, though the frontend build alone did drop from roughly 3m 35s to 2m 44s.
The registry was not, in the end, a dev performance win. What it bought is real but different: one artefact per environment that CI actually produced, per-environment isolation, rollback by retag, and — the one that matters most — image builds off the staging droplet, which has 2 vCPUs and needed a 40-minute timeout because building three Node images locally thrashed it.
Known gaps
Each with the evidence, so it can be re-checked rather than trusted.
Production deploys have never worked
All five runs of deploy-prod.yml failed, each in 41–54s, ending at
dial tcp …: i/o timeout. The runner cannot open SSH to the prod host, so the script
never executes. Production is released by hand, mirroring those steps manually.
The cleanest fix is a self-hosted runner on the prod box — it dials out, so no inbound SSH from GitHub's ranges is needed.
Branch protection is not visible
gh api …/branches/{main,dev}/protection returns 404 and /rulesets returns [].
Either nothing is configured, or the querying token lacks admin scope. If the former,
every check here is advisory and a red PR can be merged.
Git LFS assets are broken on every host
.gitattributes puts backend/public/assets/cricket/team-logos/* in LFS, and
git-lfs is not installed on the dev or prod servers. git pull therefore leaves
130-byte pointer files, which the backend serves happily:
GET /assets/cricket/team-logos/<file>
→ 200, Content-Type: image/png, 131 bytes
→ version https://git-lfs.github.com/spec/v1
TeamLogo falls back to initials on error, so it degrades silently. ipl-2026-logos
renders correctly only because it was never LFS-tracked.
CI-built images do not have this problem — both image workflows check out with
lfs: true, so GHCR images carry real bytes. That makes the registry path strictly
better than a host build here, and it means the fix for dev/staging is already in
place; dev1, dev2 and prod still serve pointers.
Sanity tests do not run on the release PRs
sanity-tests.yml is scoped to branches: [dev]. The dev → staging and
staging → main PRs get frontend typecheck and nothing else.
No CI for services/ai-chat-assistant
No tests, no typecheck, no lint — yet it is built and deployed.
No frontend build check in CI
CI runs tsc --noEmit, never next build. A build-only failure surfaces on the server.
Deploy summaries do not distinguish pull from fallback
Both paths end as Status: Deployed. Reading the job log for
Pulled from registry: yes|no is currently the only way to tell.
Tags are overwritable
See "Images" above. Deploying by digest is the fix.
dev1 and dev2 are stale snapshots
feat/dev1 is a strict ancestor of dev — 575 commits behind, 0 ahead, last deployed
2026-08-06. feat/dev2 last deployed 2026-06-03. Bringing either current is a clean
fast-forward, but would deploy hundreds of commits and weeks of migrations at once.
docs/DEPLOYMENT.md is obsolete
It documents hannibal.forsyt.io, /root/Hannibal/frontend and raw docker run
commands — the topology from before the Strykr rebrand and before GitHub Actions.
docs/infra/DEV-ENVIRONMENTS.md is stale in the same way.
What is next
- Fix production deploys — self-hosted runner. Until then prod releases stay manual and the workflow's careful gating is unexercised.
- Turn on required status checks.
- Deploy by digest rather than by tag.
- Publish the frontend image — needs GitHub Environments holding the 14
NEXT_PUBLIC_*values per environment. This is also what would let a singledocker compose pullreplace the last host build. - Materialise LFS on the remaining hosts, or move those assets to object storage.
- Smaller:
npm cicache mounts,docker builder pruneafter deploys, report pull-vs-fallback in the deploy summary, pin third-party actions to commit SHAs.
Re-measuring
Phase timings for a deploy run:
gh run view <run-id> --repo ForSyt-io/Hannibal --log \
| grep -E "out: (Checking out|Pulled from registry|Building frontend|Running database|Deployment Complete)"
Whether a deploy pulled or fell back:
gh run view <run-id> --repo ForSyt-io/Hannibal --log | grep "Pulled from registry"
Deploy success rates:
gh run list --repo ForSyt-io/Hannibal --workflow deploy-prod.yml --limit 100 \
--json conclusion --jq 'group_by(.conclusion)|map({c:.[0].conclusion,n:length})'