f49d91b63d249ef0aeb59e5bf6858446121e3ba2
8
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
f49d91b63d |
Update deploy-lib.sh
ci / lint-compose (push) Successful in 3s
ci / lint-actionlint (push) Successful in 2s
ci / lint-shellcheck (push) Successful in 2s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 3s
ci / lint-dockerfiles (push) Successful in 2s
ci / scan-deps (push) Successful in 17s
ci / test-backend (push) Successful in 7s
ci / test-frontend (push) Successful in 10s
ci / validate (push) Successful in 2s
renovate-ci / validate-renovate (push) Successful in 13s
ci / build (push) Successful in 0s
|
||
|
|
a5d384a4d8 |
feat(deploy): probe every active service after a deploy, rollouts included
ci / lint-compose (push) Successful in 4s
ci / lint-actionlint (push) Successful in 1s
ci / lint-shellcheck (push) Successful in 2s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 2s
ci / scan-deps (push) Failing after 13s
ci / test-backend (push) Failing after 10s
ci / test-frontend (push) Failing after 1s
ci / validate (push) Successful in 2s
renovate-ci / validate-renovate (push) Successful in 59s
ci / build (push) Successful in 1m50s
verify-k8s watches rollouts, which reports that pods converged. It cannot tell
a converged pod from a serving one. A Service selector pointing at a port
nothing listens on, a 500 from the app itself, a Traefik route that stopped
matching, a pod that OOMKilled early enough to still count as Available for the
duration of the check -- all of those are green at the rollout level and broken
for whoever opens the URL.
So ask what users ask. A smoke stage probes the public route of every active
service and fails on a transport error, a 5xx, or a 000, which curl reports
when it exits cleanly and nothing replied. Everything else passes, including 4xx:
a 404 from a path the service does not serve and a 302 to a login both prove
Traefik matched the host, the Service resolved to a pod and the pod answered,
which is the whole claim being tested.
An empty host list is an error, not a pass. Zero names means the extraction
broke, and reporting a clean deploy off a broken grep is the failure mode this
job exists to catch.
It runs on always() and after verify-k8s rather than before it, because a
rollback is when a route most needs re-checking. It only skips when verify-k8s
did, which is when nothing was deployed at all.
Two things worth writing down, because both were wrong on the first pass:
Stripping comments before reading the routes is not optional. naio and xui are
still in the tree commented out, and a plain grep picks both up and then reports
two services as unreachable when nobody ever deployed them. The apex
forust.xyz also needs a filter that admits it, so /\.forust\.xyz$/ quietly
dropped the site root.
And the 5xx test was written as ${code%%[0-9]*} != 5, which is empty for every
three-digit code, so a 500 was reported as ok. A case glob on the leading digit
is what actually works.
23 routes answer today, in 1.6s. The 5xx branch is the one part no live service
here exercises, so it was checked by running the block over 200 through 599 and
000 rather than against a real response.
|
||
|
|
0691536f28 |
fix(deploy): roll out our images by digest instead of a moving tag
`kubectl rollout undo` restores the previous ReplicaSet's pod template
verbatim. While that template names a tag, the rollback does not roll back
the image: the tag has already moved, so the reverted pod pulls the very
build that just failed and the cluster stays broken. The safety net added
in
|
||
|
|
30d2b83efe |
chore(reloader): manage the reloader release from the repository
Reloader has been running since 23 September and is what makes the reloader.stakater.com/auto annotation on a pod template do anything, but the repository only held a namespace. It was a release someone installed by hand, so it was invisible to review, invisible to Renovate, and one reinstall away from being silently dropped. Declaring it in HELM_RELEASES pins the chart version somewhere Renovate can update it, and the active marker means the namespace is applied before the upgrade instead of only existing as a side effect of the original install. The values file sets nothing the running release does not already do, apart from resource requests and limits, which the chart leaves empty. |
||
|
|
db7bccfd89 |
fix(deploy): restart workloads whose image tag moved past what they run
Our manifests pin images to `:latest`, so a rebuild leaves the pod template byte-identical. kubectl apply sees no change, creates no ReplicaSet and pulls nothing, and the cluster keeps serving the previous build. imagePullPolicy: Always does not help, because it only decides whether a pod that *is* starting pulls, and no pod ever starts. All eight workloads that consume an image from our own registry were affected. Three of them had been running code from 23 September, and the single hardcoded `rollout restart deployment/userbot-panel` covered one of the eight. Restarting everything unconditionally was not the answer either: that bounces healthy services on every deploy, error-pages included, and the brief window where nothing answers is exactly what error-pages exists to prevent. So compare what each workload actually runs against what the tag resolves to now, and restart only the ones that differ. When the tag still points at the running digest nothing happens, so a redeploy that changed no image is a no-op. Scope is the repository, deliberately. Five more workloads run our images but have no manifest here, and they are applied out of band. Walking the manifests rather than the cluster means this can never reach them. The digest is resolved for the node architecture. A multi-arch tag also carries `unknown/unknown` entries for the build attestation, and a pod's imageID is always the per-platform digest, so comparing the wrong entry would mark everything stale forever. Once a restart happens it bumps the generation, which is what makes the change visible to changed_workloads and therefore watchable and revertible by the verify stage. |
||
|
|
1505b638ce |
fix(deploy): verify and roll back in a separate job
verify_workloads ended on `[ -s "$failed_file" ]`, which is the opposite of what its own contract says. A non-empty file means something failed, so the function returned success exactly when a workload never came up, and failure when everything was fine. Every rollback was therefore skipped, and every deploy that changed anything ended red with an empty failure list and a bogus "Rolled back successfully". Worse, the check only ever ran at the end of stage_apply_k8s, inside the same process as the apply. A job killed by timeout-minutes, cancelled by a new push, or cut off by a dropped SSH connection never reached it, which is precisely when a rollback matters. The three helm upgrades alone can consume the whole 30-minute job budget, so that path was reachable. Verification now lives in its own job, gated on always(), so it runs whatever happened to the apply. The apply stage publishes its pre-apply snapshot through DEPLOY_SNAPSHOT_DIR/current before touching anything, and the verify stage picks it up from there. A snapshot whose recorded commit does not match the deploy is refused rather than trusted, so a stale pointer from an earlier run cannot make the rollback revert the wrong workloads. An unwritable snapshot directory now fails the deploy up front instead of silently continuing without a way back. cancel-in-progress becomes false for the same reason: cancelling a run kills the apply job and takes the verify job with it, which is the failure this change exists to prevent. Both applies are idempotent, so queueing costs little. The SSH key moves to a per-run directory removed on exit, and the deploy is pinned to the exact commit CI validated. |
||
|
|
a2ff9515a3 |
fix(deploy): validate compose without workstation secrets
renovate-ci / validate-renovate (push) Skipped
deploy / validate (push) Skipped
ci / lint-prettier (push) Successful in 2s
ci / validate (push) Successful in 2s
ci / build (push) Skipped
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 1s
docker compose config required real values for gitignored .env files and secrets, so validate always failed on stacks with :? guards (netbird, netbox). Validate structure only via --no-interpolate, --no-env-resolution and --no-path-resolution, keeping normalization and consistency checks. |
||
|
|
648b354951 |
refactor(deploy): marker-driven selection (k8s/active, root active); enable headscale/nextcloud hybrid, disable dockmon/kener/downtify/n8n
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / validate (push) Successful in 2s
deploy / preflight (push) Successful in 1s
renovate-ci / validate-renovate (push) Successful in 8s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 1s
ci / build (push) Successful in 1s
deploy / validate (push) Successful in 1m40s
deploy / apply-k8s (push) Successful in 1m41s
deploy / apply-compose (push) Successful in 13s
|