ci: deploy the image the commit built, not whatever the tag points at
Every service tracked the mutable `:prod` tag, so a deploy applied whatever that tag happened to name at the time rather than the commit it was deploying. A rollback had no way to state what it was rolling back to, and two deploys of one commit could land different images. CI now publishes an immutable `sha-<commit12>` tag beside `:prod` on main, and re-tags it for every image a push did not rebuild. That re-tag copies the manifest list, so no layer moves. The deploy resolves the immutable tag to a digest and pins the workload to it, and only falls back to the moving tag when the immutable one cannot be resolved -- which it says out loud, because that fallback is the deploy quietly ceasing to be reproducible from its own commit. The image list comes out of the tree with git grep rather than being written out a second time, so adding a service no longer means keeping two lists in step. build also gains the three jobs it was skipping -- scan-deps, test-backend, test-frontend -- so a change that breaks them cannot be tagged at all. The two run blocks where a mid-loop failure was survivable now run under set -euo pipefail: the build loop and the service detector both carried on past an error and could report a green build having produced nothing. The registry password moves from run: substitution into an env: block. A quote, a backtick or a $(...) in the password is parsed as shell before the command ever runs, and a login that failed that way looked exactly like a build that failed. The apply and verify timeouts stay at 45 and 30 minutes. The comments now record the arithmetic that says so rather than leaving the numbers to be raised on the next scare: three no-op helm upgrades run 3-5 minutes, one broken release is a single 10 minute rollback because the loop aborts on the first failure, and the apply loop itself is about a minute. That is roughly 15 minutes of work against a 45 minute budget. verify is 32 workloads at 8 wide -- four waves of 300 seconds, 20 minutes -- which leaves room for two serial rollbacks, and only becomes derivable at 45 once rollback_workloads is parallelised.
This commit is contained in:
1 parent
24dd82e801
commit
c70d2db3a1
3 files changed
+194
-56
No files matched your search
@@ -73,8 +73,34 @@ jobs:
|
||||
needs: [validate]
|
||||
runs-on: [self-hosted, linux, arch, homelab, prod]
|
||||
# Apply only, no verification, so this is just the work itself: snapshot,
|
||||
# then up to three sequential `helm upgrade --atomic --timeout 10m`, then the
|
||||
# apply loop. Verification has its own job and its own budget.
|
||||
# then sequential `helm upgrade --atomic --timeout 10m`, then the apply loop.
|
||||
# Verification has its own job and its own budget.
|
||||
#
|
||||
# 45 is roughly four times the measured cost of the stage, which is
|
||||
# deliberately not raised on a theory:
|
||||
#
|
||||
# helm, healthy 3 no-op upgrades ~3-5 min
|
||||
# helm, one release bad --atomic spends its 10m, ~10-15 min
|
||||
# then rolls that one back
|
||||
# apply loop ~40 manifests, 4 of which ~1 min
|
||||
# resolve an image digest
|
||||
# restart_stale_images 7.6s to find 8 workloads, ~0.5 min
|
||||
# 9.8s to resolve their digests
|
||||
#
|
||||
# The helm figure is one release, not three: `set -e` aborts
|
||||
# upgrade_helm_releases on the first failure, so a broken release costs
|
||||
# 10m and the other two are never attempted. Multiplying 10m by three
|
||||
# overstates the worst case by 20 minutes.
|
||||
#
|
||||
# The 45 minutes this was last raised to 45 were still not enough, and the
|
||||
# job logs for those runs no longer exist, so what actually consumed the
|
||||
# budget is not known - the two measurable candidates above account for
|
||||
# ~15 of it. The one unbounded thing left in this stage is
|
||||
# `docker manifest inspect` at deploy-lib.sh:236, which has no timeout
|
||||
# against a registry with a known hang mode. Bound it, and make the stage
|
||||
# announce what it is working on, before spending any of that on a larger
|
||||
# ceiling: a stage that is killed with a diagnosable last line is a bug
|
||||
# report, one that vanishes is not.
|
||||
timeout-minutes: 45
|
||||
steps:
|
||||
- name: Checkout repository
|
||||
@@ -112,7 +138,22 @@ jobs:
|
||||
needs.apply-k8s.result != 'skipped' &&
|
||||
needs.apply-compose.result != 'skipped'
|
||||
runs-on: [self-hosted, linux, arch, homelab, prod]
|
||||
# ceil(changed_workloads / 8) waves of ROLLOUT_TIMEOUT each, plus rollback.
|
||||
# Not raised, because the arithmetic does not close.
|
||||
#
|
||||
# 32 workloads are under management and the wave width is 8, so the verify
|
||||
# itself is 4 waves of ROLLOUT_TIMEOUT (300s) = 20 minutes worst case, when
|
||||
# every rollout times out rather than converging. That is already 20 of 30.
|
||||
#
|
||||
# The other 10 would have to absorb rollback, and rollback_workloads is a
|
||||
# serial `while read` loop at 300s per failed workload. 10 minutes buys two.
|
||||
# Any larger number is buying a bigger multiple of an unbounded term rather
|
||||
# than covering a known cost: 60 minutes buys eight, and 60 minutes is
|
||||
# therefore not a bound, it is a guess with two digits.
|
||||
#
|
||||
# The number becomes derivable the moment rollback uses the same wave width
|
||||
# as the verify: 32 failures then cost 4 waves = 20 minutes instead of 160,
|
||||
# and 45 covers verify plus rollback at full width. That change is to the
|
||||
# recovery path and is not folded into a timeout edit.
|
||||
timeout-minutes: 30
|
||||
steps:
|
||||
- name: Checkout repository
|
||||
|
||||
Reference in new issue
Block a user