Skip to main content

CI/CD Pipeline

How code gets from a feature branch to a running environment, what it costs, and where the gaps are.

Last verified against the repo and the live hosts on 2026-09-05. Timings and run IDs come from real runs. Where a number is an estimate it says so.

Branching model​

feature/* ──PR──▶ dev ──PR──▶ staging ──PR──▶ main ──automatic──▶ production

Two workflows enforce this:

  • restrict-staging-base.yml — PRs into staging must come from dev, or from a hotfix/* branch (allowed with a notice telling you to cherry-pick back to dev). The head repo must be this repo.
  • restrict-main-base.yml — PRs into main must come from staging.

Merging to main deploys production automatically. workflow_dispatch is retained as the rollback path — dispatch with ref set to an earlier commit or tag.

Note that restrict-main-base.yml gates pull requests into main only. A direct push to main bypasses it and deploys, and no branch protection currently prevents that.

Workflow inventory​

WorkflowTriggerWhat it does
sanity-tests.ymlpush + PR to dev, paths backend/**Financial sanity integration tests against a real Postgres, plus a backend build + unit test job
typecheck-frontend.ymlPR to dev / mainDesign-system check, tsc --noEmit, frontend unit specs
build-images.ymlPR to dev / stagingBuilds the service Dockerfiles without pushing. Holds no packages: permission
publish-images.ymlworkflow_callBuilds and pushes one environment's images to GHCR. The only holder of packages: write
restrict-staging-base.ymlPR to stagingHead must be dev or hotfix/*
restrict-main-base.ymlPR to mainHead must be staging
deploy-dev.ymlpush to devPublishes dev-<sha> images, then deploys dev.strykr.io
deploy-staging.ymlpush to stagingPublishes staging-<sha> images, then deploys staging.strykr.io
deploy-dev1.ymlpush to feat/dev1Deploys dev1.strykr.io — still builds on the server
deploy-dev2.ymlpush to feat/dev2Deploys dev2.strykr.io — still builds on the server
deploy-prod.ymlpush to main, plus workflow_dispatchDeploys strykr.io — still builds on the server

Why typecheck-frontend has no path filter​

It runs on every PR to dev and main, then decides internally whether to do the expensive work by diffing strykr-fe/. Filtering at the on: level would report the check as skipped on backend-only PRs, and a required-check rule on main would then block those PRs forever.

Environments​

Every port binds to 127.0.0.1; host nginx terminates TLS and proxies in.

EnvBranchTriggerHost dirCompose fileFEBEImages from
devdevautomatic/root/strykrdocker-compose.dev.yml30003001GHCR (dev-<sha>)
dev1feat/dev1automatic/root/strykr-dev1docker-compose.dev1.yml30023003built on host
dev2feat/dev2automatic/root/strykr-dev2docker-compose.dev2.yml30043005built on host
stagingstagingautomatic/root/strykr-stagingdocker-compose.staging.yml30003001GHCR (staging-<sha>)
prodany ref, default mainmanual/root/strykrdocker-compose.prod.yml30003001built on host

dev, dev1 and dev2 share one VPS and one Docker daemon. dev1 and dev2 run no data layer of their own — their deploy first brings up postgres, redis and timescaledb from /root/strykr, then joins that network with a separate database. staging and prod are separate hosts; the /root/strykr path collision between dev and prod is coincidental.

Images​

Each environment has its own image reference. Nothing is shared:

ghcr.io/forsyt-io/hannibal-<service>:<env>-<sha>   what deploys pin
ghcr.io/forsyt-io/hannibal-<service>:<env> moving, convenience only

The <env>- prefix is the safeguard — no tag one environment resolves is reachable by another, so prod cannot pull a dev image even by mistake.

<env>-<sha> is commit-scoped, not immutable: re-running a workflow overwrites the tag, and floating base images mean the same commit can rebuild to different bytes. Deploying by digest (image@sha256:…) is the only real guarantee and is not yet done.

Backend and ai-chat only. Both Dockerfiles declare zero build args and read all config from compose env_file at runtime. The frontend bakes 14 NEXT_PUBLIC_* values at build time and those differ per environment, so it is still built on each host. Publishing it requires GitHub Environments holding those values — none exist today.

Registry auth uses the job-scoped GITHUB_TOKEN passed over SSH and logged out after use. No long-lived PAT lives on any host.

Deploy anatomy​

dev and staging​

  1. publish-images builds backend and ai-chat on a runner and pushes <env>-<sha>
  2. The deploy job runs with if: ${{ !cancelled() }} — a superseded run must not deploy
  3. git fetch + git checkout --detach <sha> — the exact commit the images were built from, not the branch tip
  4. docker login ghcr.io → docker compose pull backend ai-chat → docker logout
  5. Build the frontend. If the pull failed, build backend and ai-chat too
  6. prisma migrate deploy → docker compose up -d
  7. verify-deployment-observability.sh, then reconcile-monitoring.sh

Step 3 matters. Pinning images to a commit while the host took the branch tip would let one commit's backend run against another commit's migrations — staging queues rather than cancels, so that was a live race, not a theoretical one.

Step 4/5 is best-effort by design: if login fails, the registry is unreachable, or publishing failed, the deploy falls back to building on the host. A registry problem degrades to the old behaviour instead of blocking a release.

dev1, dev2, prod​

Unchanged — still docker compose build on the host.

Prod additionally refuses to run on a dirty tree, loads reconcile-monitoring.sh from the workflow's commit rather than the deploy ref, greps the target ref for six observability-baseline markers, restarts host nginx after container rotation, and asserts https://strykr.io/metrics returns 404. Its user inputs are passed via env: and consumed as shell variables, never interpolated into run:.

What the registry actually cost​

Honest accounting, because the expectation was a speed-up and that is not what happened.

Wall clock for deploy-dev, same workflow, consecutive days:

RunDateTotalNotes
3361714551109-024m 56sbefore the Next cache mount
3387497169909-045m 13scache mount merged, still building on host
3388719736009-046m 42ssame
3390213758509-046m 25ssame
3393245528209-058m 07sfirst run pulling from GHCR

Breakdown of the GHCR run: publish-images 3m 43s (cold GHA layer cache), then the deploy job 4m 18s — of which the frontend build is 2m 44s and the image pull 30s.

So moving image builds into CI made dev's wall clock worse by roughly two minutes. The image build went from a near-free layer-cache hit on the host to a cold build on a runner that runs serially before the deploy. The GHA cache should warm on subsequent runs; that has not been measured yet.

Two further cautions on these numbers: the spread across identical pre-GHCR runs is 4m 56s to 6m 42s, so single-run comparisons are weak; and the Next cache mount shows no clear wall-clock win either, though the frontend build alone did drop from roughly 3m 35s to 2m 44s.

The registry was not, in the end, a dev performance win. What it bought is real but different: one artefact per environment that CI actually produced, per-environment isolation, rollback by retag, and — the one that matters most — image builds off the staging droplet, which has 2 vCPUs and needed a 40-minute timeout because building three Node images locally thrashed it.

Known gaps​

Each with the evidence, so it can be re-checked rather than trusted.

Production deploys have never worked​

All five runs of deploy-prod.yml failed, each in 41–54s, ending at dial tcp …: i/o timeout. The runner cannot open SSH to the prod host, so the script never executes. Production is released by hand, mirroring those steps manually.

The cleanest fix is a self-hosted runner on the prod box — it dials out, so no inbound SSH from GitHub's ranges is needed.

Branch protection is not visible​

gh api …/branches/{main,dev}/protection returns 404 and /rulesets returns []. Either nothing is configured, or the querying token lacks admin scope. If the former, every check here is advisory and a red PR can be merged.

Git LFS assets are broken on every host​

.gitattributes puts backend/public/assets/cricket/team-logos/* in LFS, and git-lfs is not installed on the dev or prod servers. git pull therefore leaves 130-byte pointer files, which the backend serves happily:

GET /assets/cricket/team-logos/<file>
→ 200, Content-Type: image/png, 131 bytes
→ version https://git-lfs.github.com/spec/v1

TeamLogo falls back to initials on error, so it degrades silently. ipl-2026-logos renders correctly only because it was never LFS-tracked.

CI-built images do not have this problem — both image workflows check out with lfs: true, so GHCR images carry real bytes. That makes the registry path strictly better than a host build here, and it means the fix for dev/staging is already in place; dev1, dev2 and prod still serve pointers.

Sanity tests do not run on the release PRs​

sanity-tests.yml is scoped to branches: [dev]. The dev → staging and staging → main PRs get frontend typecheck and nothing else.

No CI for services/ai-chat-assistant​

No tests, no typecheck, no lint — yet it is built and deployed.

No frontend build check in CI​

CI runs tsc --noEmit, never next build. A build-only failure surfaces on the server.

Deploy summaries do not distinguish pull from fallback​

Both paths end as Status: Deployed. Reading the job log for Pulled from registry: yes|no is currently the only way to tell.

Tags are overwritable​

See "Images" above. Deploying by digest is the fix.

dev1 and dev2 are stale snapshots​

feat/dev1 is a strict ancestor of dev — 575 commits behind, 0 ahead, last deployed 2026-08-06. feat/dev2 last deployed 2026-06-03. Bringing either current is a clean fast-forward, but would deploy hundreds of commits and weeks of migrations at once.

docs/DEPLOYMENT.md is obsolete​

It documents hannibal.forsyt.io, /root/Hannibal/frontend and raw docker run commands — the topology from before the Strykr rebrand and before GitHub Actions. docs/infra/DEV-ENVIRONMENTS.md is stale in the same way.

What is next​

  1. Fix production deploys — self-hosted runner. Until then prod releases stay manual and the workflow's careful gating is unexercised.
  2. Turn on required status checks.
  3. Deploy by digest rather than by tag.
  4. Publish the frontend image — needs GitHub Environments holding the 14 NEXT_PUBLIC_* values per environment. This is also what would let a single docker compose pull replace the last host build.
  5. Materialise LFS on the remaining hosts, or move those assets to object storage.
  6. Smaller: npm ci cache mounts, docker builder prune after deploys, report pull-vs-fallback in the deploy summary, pin third-party actions to commit SHAs.

Re-measuring​

Phase timings for a deploy run:

gh run view <run-id> --repo ForSyt-io/Hannibal --log \
| grep -E "out: (Checking out|Pulled from registry|Building frontend|Running database|Deployment Complete)"

Whether a deploy pulled or fell back:

gh run view <run-id> --repo ForSyt-io/Hannibal --log | grep "Pulled from registry"

Deploy success rates:

gh run list --repo ForSyt-io/Hannibal --workflow deploy-prod.yml --limit 100 \
--json conclusion --jq 'group_by(.conclusion)|map({c:.[0].conclusion,n:length})'