These five stacks still asked for :latest, but the build stopped pushing
it - on main it only pushes main and prod, on dev only dev. Every one of
these images therefore resolved only because the registry still had a
stale :latest from before that change, and the next time one of them was
actually built the reference would have dangled.
Named for dtek-notif: of the seven, only that one still resolved at
:prod, because it is the only image not rebuilt since the build dropped
:latest - and it is the only one of these five stacks the deploy does not
manage (no `active` marker, and its file is docker-compose.yaml, which
the COMPOSE_STACKS glob does not even match). Touching
dtek_notif/docker-compose.yaml matches dtek_notif/* in the build's
changed-service detector, so pushing this builds it and publishes
:prod for it too.
Verified the other six resolve at :prod in the registry.
`kubectl rollout undo` restores the previous ReplicaSet's pod template
verbatim. While that template names a tag, the rollback does not roll back
the image: the tag has already moved, so the reverted pod pulls the very
build that just failed and the cluster stays broken. The safety net added
in 1505b63 therefore could not recover from a bad image.
Pin the digest at apply time. A digest is not knowable when a manifest is
written, so render_pinned resolves it on the way into the cluster and the
digest is never committed. Git keeps a readable `:prod`, Renovate keeps
seeing exactly the manifests it saw before, and the previous revision of
each workload now holds the digest that was actually serving, so undo
restores those exact bytes.
imagePullPolicy is dropped from the manifests rather than set to
IfNotPresent: a reference that is not `:latest` already defaults to it, and
that is what the Kubernetes docs ask for alongside a digest.
An unresolvable image is fatal instead of a warning, because carrying on
would quietly apply a mutable tag again.
restart_stale_images keeps its comparison but is no longer how a rebuild
reaches the cluster -- the pinned template rolls out on its own now. What
is left is a drift check for hand-run `kubectl set image`, so it matches
the container by repository: a pod's status now reports `repo@sha256:...`
while the manifest still says `:prod`.
The build job stops pushing `:latest` altogether, which removes the tag
that a dev branch could otherwise move under a prod deploy.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Six of eight workloads had no readinessProbe, so a pod turned Ready the
moment its process started. The verify job relies on `rollout status`, so
it passed for images that crash-looped or served errors, which left the
rollback safety net inert.
Each probe targets the path the service is actually reached on:
- homepages: / (verified 200)
- error-pages: /404.html, the path Traefik's errorPages middleware
requests. / returns 403 by design and would never pass.
- webinar-checker: /health (verified 200). /metrics also answers, but it
is a Prometheus endpoint, not a readiness signal.
The two userbot deployments stay without probes: they expose no port and
no session file, and the panel reaches Telegram through its own client. A
truthful signal there needs a health endpoint in the app itself.
Our manifests pin images to `:latest`, so a rebuild leaves the pod template
byte-identical. kubectl apply sees no change, creates no ReplicaSet and pulls
nothing, and the cluster keeps serving the previous build. imagePullPolicy:
Always does not help, because it only decides whether a pod that *is* starting
pulls, and no pod ever starts.
All eight workloads that consume an image from our own registry were affected.
Three of them had been running code from 23 September, and the single hardcoded
`rollout restart deployment/userbot-panel` covered one of the eight.
Restarting everything unconditionally was not the answer either: that bounces
healthy services on every deploy, error-pages included, and the brief window
where nothing answers is exactly what error-pages exists to prevent. So compare
what each workload actually runs against what the tag resolves to now, and
restart only the ones that differ. When the tag still points at the running
digest nothing happens, so a redeploy that changed no image is a no-op.
Scope is the repository, deliberately. Five more workloads run our images but
have no manifest here, and they are applied out of band. Walking the manifests
rather than the cluster means this can never reach them.
The digest is resolved for the node architecture. A multi-arch tag also carries
`unknown/unknown` entries for the build attestation, and a pod's imageID is
always the per-platform digest, so comparing the wrong entry would mark
everything stale forever.
Once a restart happens it bumps the generation, which is what makes the change
visible to changed_workloads and therefore watchable and revertible by the
verify stage.