11e92fdf4e449f07a81dd328ec383a6de80daffe
621
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
11e92fdf4e |
fix(deploy): bound ssh hangs and retry the stage on transport loss
A connection that died silently used to hang until the job timeout, and the stage was never re-run. One flaky TCP session cost a whole 45-minute apply, and the symptom - a job that stops mid-output with no error - is what made the last few deploy failures expensive to read. ServerAliveInterval/CountMax cap how long a dead peer goes unnoticed at ~60s, ConnectTimeout caps setup. Only exit 255 - ssh's own transport failures - is retried, up to three attempts with a growing gap. A stage that fails on its own merits exits with the remote's status, so a real failure surfaces its own log immediately instead of being repeated three times over 45 minutes. The stages are declarative applies, so re-running one that had already committed is harmless. The stage environment now goes through `env` as separate argv entries rather than one interpolated string, so nothing in REPO, DEPLOY_SHA or DEPLOY_SNAPSHOT_DIR is re-split by the remote shell. Verified against a stubbed ssh: clean run attempts once, a single transport failure recovers on attempt 2 and exits 0, three failures give up preserving 255, and a stage failing with 1 or 7 attempts once and passes the code through unchanged. Also records why USERBOT_IMAGE stays on the prod tag: render_pinned rewrites only plain `image:` lines, and this ref is what the panel injects into the per-instance Deployments it creates, so those instances track the tag rather than the panel's own resolved digest. The two panel-created instances currently in the cluster are digest-pinned, so the panel does accept one either way; the tag is the choice, not a limitation. |
||
|
|
b1f98fc148 |
Revert "fix(compose): stop pointing at the tag the build dropped" for dtek_notif
dtek-notif is being picked up again, so leave its compose alone. It is also the one image not rebuilt since the build dropped :latest, so :prod does not exist for it yet - unlike the other six, which resolve. Keeping this file out of the change also keeps dtek_notif/* out of the build's changed-service detector, so pushing does not build an image nobody asked for. |
||
|
|
9b91b5847e |
fix(compose): stop pointing at the tag the build dropped
These five stacks still asked for :latest, but the build stopped pushing it - on main it only pushes main and prod, on dev only dev. Every one of these images therefore resolved only because the registry still had a stale :latest from before that change, and the next time one of them was actually built the reference would have dangled. Named for dtek-notif: of the seven, only that one still resolved at :prod, because it is the only image not rebuilt since the build dropped :latest - and it is the only one of these five stacks the deploy does not manage (no `active` marker, and its file is docker-compose.yaml, which the COMPOSE_STACKS glob does not even match). Touching dtek_notif/docker-compose.yaml matches dtek_notif/* in the build's changed-service detector, so pushing this builds it and publishes :prod for it too. Verified the other six resolve at :prod in the registry. |
||
|
|
892790822d |
fix(crowdsec): stop the 403 loop that banned our own VPS and runner
Chasing why apply-k8s kept dying mid-run turned up a self-inflicted
ban loop. 585 of 586 LePresidente/http-generic-403-bf alerts in the LAPI
came from 193.181.211.79 - our own VPS - POSTing
/management.ManagementService/GetServerKey, i.e. the NetBird client's
own management call. netbird-server had never seen a single one of them,
so the 403 was not NetBird's: the bouncer was rejecting the request
before it got there. A banned peer keeps retrying, each retry is another
403, and the scenario turns five 403s in ten seconds into a 4h ban, so
the loop kept re-arming the ban it was serving. The hourly janitor step
that deleted those decisions hourly was masking all of it.
* netbird/k8s/ingress.yaml - drop the bouncer from the mesh API routes
(gRPC-gateway management, signal, relay, /api, /oauth2). A ban there
locks a peer out of the network it needs to reach anything else, and
those endpoints authenticate by NetBird token, not by a login form.
The dashboard keeps the bouncer; it is a real login surface.
netbird-local was already exempt, so this makes prod match.
* crowdsec-middleware.yaml - CrowdsecMode stream instead of live. live
blocked on GET /v1/decisions per request, so a burst saturated the
LAPI and the plugin 403'd IPs that were never banned. v1.3.3 ignores
UpdateMaxFailure in live, so fail-open is only reachable in stream;
-1 now means an unreachable LAPI degrades to "no protection" rather
than "every site 403". 15s poll instead of the 60s default, because
the runner shares one public IP with the house.
Also corrects the key name: HTTPTimeoutSeconds, not
CrowdsecLapiTimeout, which never existed and was being silently
dropped, leaving the 10s default. Back at 10, not the 2s
|
||
|
|
f7cd75d65e |
fix(crowdsec): stop the bouncer from 403-ing our own deploys
ci / lint-compose (push) Successful in 4s
ci / lint-actionlint (push) Successful in 1s
ci / lint-shellcheck (push) Successful in 2s
ci / lint-prettier (push) Successful in 3s
ci / validate (push) Successful in 3s
renovate-ci / validate-renovate (push) Successful in 32s
ci / lint-ruff (push) Successful in 0s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 2s
ci / scan-deps (push) Successful in 17s
ci / test-backend (push) Successful in 7s
ci / test-frontend (push) Successful in 10s
ci / build (push) Successful in 1s
Every bouncer-protected request blocks on a synchronous GET /v1/decisions against the LAPI, and the plugin fails CLOSED with 403 when that lookup exceeds its timeout. The LAPI was capped at 400m/500Mi on a single replica, so idle lookups measured 1.3-7.4s and a deploy burst pushed them past the fork's implicit 10s default: gitea answered 403 for ten seconds straight, and containerd turned those 403s on gcr.forust.xyz into ErrImagePull/ImagePullBackOff on freshly rolled pods. Three changes, plus the 403 feedback loop that made it sticky: * gitea/k8s/ingress.yaml - drop the bouncer from the /v2 registry route. A deploy fires hundreds of parallel authenticated OCI requests (runner Action API, manifest inspect per own image, containerd pulls, smoke probes) and scanners gain nothing from a registry that already does its own token auth. The web UI route keeps the bouncer. * crowdsec-values.yaml - LAPI to 1500m/1Gi. Replicas stay at 1 on purpose: LAPI is stateful (BoltDB plus credentials on two RWO PVCs), so a second replica would corrupt the decision store. * crowdsec-middleware.yaml - CrowdsecLapiTimeout: 2s instead of the implicit 10s, so a slow LAPI costs a fast 403 rather than a 10s hang. LePresidente/http-generic-403-bf then banned us for our own 403s: five POSTs answered 403 within 10s earn a 4h ban, and the hairpin-NAT address 192.168.88.1 that the Gitea Actions runner presents to Traefik is not covered by the home-dynamic-IP whitelist. That scenario cannot be dropped per-scenario - it is baked into the hub item crowdsecurity/http-generic-bf v0.9, and disabling the whole base-http-scenarios collection would cost ~40 useful detections. So the janitor now deletes its decisions hourly and a new forust/lan whitelist postoverflow covers 192.168.88.0/24. First janitor run removed 165 decisions; none of the remaining ones are local. Measured after: gitea 200 in 25-148ms (was 403 at 10001ms), a 60-way parallel burst all 200 with a 422ms max, gcr /v2 back to its 401 auth challenge, LAPI at 60m CPU with no throttling. |
||
|
|
6a9a460769 |
fix(deploy): let a pinning failure explain itself
ci / lint-compose (push) Successful in 3s
ci / lint-actionlint (push) Successful in 1s
ci / lint-shellcheck (push) Successful in 2s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 3s
ci / lint-dockerfiles (push) Successful in 2s
ci / scan-deps (push) Successful in 15s
ci / test-backend (push) Successful in 6s
ci / test-frontend (push) Successful in 10s
ci / validate (push) Successful in 2s
renovate-ci / validate-renovate (push) Successful in 26s
ci / build (push) Successful in 1s
registry_digest was written to return an empty string for a ref the registry does not have, so render_pinned could print "cannot resolve <ref>" and stop. It could not do that. Every caller runs under set -euo pipefail, pipefail reports the rightmost non-zero stage, and the failed docker manifest inspect made the assignment itself fail, which set -e turns into an immediate exit. render_pinned therefore died silently on the first unresolvable ref: nothing on stderr, nothing on stdout, exit 1. The apply loop piped that empty stream into kubectl, so the whole deploy stopped with "error: no objects passed to apply" - kubectl guessing at a cause, with the actual reason nowhere in the log. The missing message is the reason the |
||
|
|
ac0f845da6 |
fix(deploy): stop the panel from pointing at a tag the build dropped
ci / lint-compose (push) Successful in 3s
ci / lint-actionlint (push) Successful in 1s
ci / lint-shellcheck (push) Successful in 2s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 1s
ci / scan-deps (push) Successful in 15s
ci / test-backend (push) Successful in 7s
ci / test-frontend (push) Successful in 11s
renovate-ci / validate-renovate (push) Successful in 12s
ci / validate (push) Successful in 3s
ci / build (push) Successful in 37s
USERBOT_IMAGE named userbot:latest, but the build stopped pushing latest when images became :prod. render_pinned resolves every gcr.forust.xyz reference it finds in a manifest, not just the ones it rewrites, so that env var made the whole apply abort with "cannot resolve userbot:latest; applying nothing". Naming it :prod also makes the build job rebuild userbot, which is what restores forust/userbot:prod - the tag is listed in the registry but its manifest 404s, so the image: field in userbots.yaml had nothing to resolve either. Found by resolving every digest apply-k8s will need before spending a run on it: 6 of 8 resolved, and both failures were userbot. |
||
|
|
f54589a05c |
fix(deploy): put the snapshot where the deploy user can write it
ci / lint-compose (push) Successful in 4s
ci / lint-actionlint (push) Successful in 1s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 2s
ci / lint-shellcheck (push) Successful in 2s
ci / lint-prettier (push) Successful in 3s
ci / test-backend (push) Successful in 7s
ci / lint-dockerfiles (push) Successful in 3s
ci / scan-deps (push) Successful in 15s
ci / test-frontend (push) Successful in 11s
ci / validate (push) Successful in 7s
renovate-ci / validate-renovate (push) Successful in 29s
ci / build (push) Successful in 1s
The first deploy to actually run died on its very first action, and the error the other job reported was only the consequence. DEPLOY_SNAPSHOT_DIR defaulted to /var/backups/homelab-deploy. The deploy is unprivileged, and this Arch host has no /var/backups at all, so snapshot_dir's mkdir -p had to create it under root-owned /var and got Permission denied. It refused to go on, which is exactly what the guard is for, so no workload was touched - but the verify job then found no pointer and could only say to go look by hand. Defaulting to the deploy user's own XDG state directory fixes it with no root and no setup step, and keeps the guard: an unwritable snapshot dir still stops the deploy before the first apply. ssh-run.sh now forwards DEPLOY_SNAPSHOT_DIR too, so the path is overridable without editing the library. Verified on the workstation as the unprivileged user: pointer published, commit recorded, 71 workload generations and three helm releases captured, and the stale-pointer refusal still works. |
||
|
|
f49d91b63d |
Update deploy-lib.sh
ci / lint-compose (push) Successful in 3s
ci / lint-actionlint (push) Successful in 2s
ci / lint-shellcheck (push) Successful in 2s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 3s
ci / lint-dockerfiles (push) Successful in 2s
ci / scan-deps (push) Successful in 17s
ci / test-backend (push) Successful in 7s
ci / test-frontend (push) Successful in 10s
ci / validate (push) Successful in 2s
renovate-ci / validate-renovate (push) Successful in 13s
ci / build (push) Successful in 0s
|
||
|
|
aafa74b70a |
chore: give local verification tooling a gitignored home
ci / lint-compose (push) Successful in 3s
ci / lint-actionlint (push) Successful in 1s
ci / lint-shellcheck (push) Successful in 2s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 2s
ci / scan-deps (push) Successful in 17s
ci / test-backend (push) Successful in 10s
ci / test-frontend (push) Successful in 11s
ci / validate (push) Successful in 3s
renovate-ci / validate-renovate (push) Successful in 25s
ci / build (push) Successful in 1s
The pinned CI tools and the pre-push verification script were living in a temp directory that did not survive. tmp/ is now ignored, so ruff (which respects .gitignore by default) and the git ls-files globs the lint jobs use both skip it without needing to be told. |
||
|
|
4f74fe1778 |
ci: stop inheriting the runner's python and node
The runner executes jobs on the host rather than in a container, so a
workflow that says 'python3' or 'npm' is really saying 'whatever this
machine happens to have today'. Both of the test jobs added in
|
||
|
|
a5d384a4d8 |
feat(deploy): probe every active service after a deploy, rollouts included
ci / lint-compose (push) Successful in 4s
ci / lint-actionlint (push) Successful in 1s
ci / lint-shellcheck (push) Successful in 2s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 2s
ci / scan-deps (push) Failing after 13s
ci / test-backend (push) Failing after 10s
ci / test-frontend (push) Failing after 1s
ci / validate (push) Successful in 2s
renovate-ci / validate-renovate (push) Successful in 59s
ci / build (push) Successful in 1m50s
verify-k8s watches rollouts, which reports that pods converged. It cannot tell
a converged pod from a serving one. A Service selector pointing at a port
nothing listens on, a 500 from the app itself, a Traefik route that stopped
matching, a pod that OOMKilled early enough to still count as Available for the
duration of the check -- all of those are green at the rollout level and broken
for whoever opens the URL.
So ask what users ask. A smoke stage probes the public route of every active
service and fails on a transport error, a 5xx, or a 000, which curl reports
when it exits cleanly and nothing replied. Everything else passes, including 4xx:
a 404 from a path the service does not serve and a 302 to a login both prove
Traefik matched the host, the Service resolved to a pod and the pod answered,
which is the whole claim being tested.
An empty host list is an error, not a pass. Zero names means the extraction
broke, and reporting a clean deploy off a broken grep is the failure mode this
job exists to catch.
It runs on always() and after verify-k8s rather than before it, because a
rollback is when a route most needs re-checking. It only skips when verify-k8s
did, which is when nothing was deployed at all.
Two things worth writing down, because both were wrong on the first pass:
Stripping comments before reading the routes is not optional. naio and xui are
still in the tree commented out, and a plain grep picks both up and then reports
two services as unreachable when nobody ever deployed them. The apex
forust.xyz also needs a filter that admits it, so /\.forust\.xyz$/ quietly
dropped the site root.
And the 5xx test was written as ${code%%[0-9]*} != 5, which is empty for every
three-digit code, so a 500 was reported as ok. A case glob on the leading digit
is what actually works.
23 routes answer today, in 1.6s. The 5xx branch is the one part no live service
here exercises, so it was checked by running the block over 200 through 599 and
000 rather than against a real response.
|
||
|
|
af9a22fea9 |
fix(converters): spell workstation right in the compose router label
The local bentopdf router matched Host(`pdf.wokstation.internal`), a hostname that resolves to nothing. converters/k8s next to it has always had the correct spelling, so the service is reachable in the path that actually deploys and this one is dormant -- converters/ has no active marker, so select_manifests skips the stack. It would only bite whoever switches the service back to Compose and then wonders why one of the three internal routes 404s. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
baedea504d |
ci: scan dependencies for new advisories and stop handing out a write token
Two things, both about not finding out late. No workflow declared `permissions`, so all eighteen jobs across the four workflows ran on a token with the default full repository scope. Every one of them only checks out code, and deploy reaches the cluster over SSH with the deploy key, and Renovate writes through its own bot PAT rather than the Actions token. So `contents: read` is all any of them needed. The panel image ships 15 known advisories and nothing was looking. Add a scan-deps job that fails on anything new, and record the eight current ones by ID in the workflow. It is a list rather than a baseline count so that the diff that accepts an advisory says so in words, and it lives in our workflow instead of the package manifest so a subtree sync from forust/userbot cannot quietly widen the exemption. Both halves were checked to fail on a regression, not just to pass today: removing one --ignore-vuln turns the Python step red, and dropping --audit-level to moderate turns the npm one red on the devalue advisory. npm audits production dependencies only. All seven findings in the full tree are build- or test-time: the esbuild advisory needs a vite dev server exposed to the internet, and nanoid's infinite loop needs a custom generator called with size 0, which postcss does not do. None are in the 91 kB bundle the panel serves, so failing on them would be noise that trains people to ignore the job. The starlette entries are the reason the job is not "fail on everything": fastapi 0.115.12 pins starlette<0.47.0 and the last four fixes need 0.49.1 through 1.3.1, so clearing them is a jump to fastapi 0.141.x and is upstream's call, not a drive-by. Four of the seven are reachable in principle, which the comment on the job sets out. The panel answers only on userbot.workstation.internal with no public route, which is what keeps those four from being an internet-facing DoS. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
30995ee009 |
ci: pin the last four linters instead of trusting the runner
prettier, ruff, yamllint and hadolint were the only CI tools still called bare, straight off whatever the runner happened to have installed. Pin them in tool-versions.env like the other three and install them the same way, so the versions Renovate moves are the versions CI runs. Each pinned version equals what is already on the runner, so this changes what CI does not at all today. It changes what CI does on a rebuilt runner: the pinned one gets installed over the drift. The four need four different mechanisms, which is why this is not one pattern: hadolint a bare binary per platform, like actionlint ruff, yamllint PyPI wheels, unpacked by uv prettier an npm tarball, unpacked by tar prettier is the awkward one. Its entry point requires ../package.json relative to its own real path, so copying the single file out -- which is what every other installer here does -- yields a module-not-found at the first run. It keeps its package directory in a versioned one next to a relative symlink, and the tarball ships bin/ without the exec bit, so that needs chmod too. hadolint's release names one platform uname-style and the other Go-style (x86_64 but arm64), which 404s on the first architecture if you assume otherwise. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
3a05d86e3e |
ci: run svelte-check, which was already a dependency with no script
svelte-check sits in devDependencies at ^4.7.3 and nothing in the repository ever invoked it, so the type errors it reports had no path to a human. Point a script at it and run it in the frontend job, and it is clean: 0 errors, 0 warnings. It shares the one `npm ci` with the test step. A second install would have doubled the slowest part of the job to learn exactly the same thing. The lockfile is untouched, because scripts are not part of what it pins. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
c00a4724f5 |
ci: actually run the test suites that exist in the tree
The panel ships 25 pytest tests and 2 vitest tests. Nothing executed them:
there was no job, no local dev loop, and nothing that would have noticed when
one of them rotted. They pass, and they are 8 seconds of work, which is the
argument for having them.
Both jobs mirror how the image is built rather than how a developer would run
them by hand: `npm ci` because that is what the Dockerfile does, so the tree
under test is the tree that ships, and requirements-dev.txt through uv, which
is now pinned like the other CI tools.
The backend job runs `python -m pytest`, not bare `pytest`. The tests import
`app.*` relative to the backend directory, and only the `-m` form puts the
working directory on sys.path.
ruff format --check joins ruff check in the lint job. It needed
|
||
|
|
284e19ef88 |
ci: pin uv, the tool that builds the pytest venv
The panel backend has 25 pytest tests that no workflow has ever run. Making them run needs a throwaway virtualenv, and uv is what builds it in seconds against the pinned requirements-dev.txt. Installing it through install-ci-tools.sh rather than assuming it is on the runner keeps the version in one place, where the other three tools already live, and where the Renovate regex manager can move it. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
0ae0df7473 |
style: format the last 10 files that ruff format disagreed with
ruff.toml has declared `quote-style = "single"` and line-length 120 since the lint job landed, and 118 of 128 files follow it. The panel backend and the netbox configuration were written in black/prettier style instead, so a `ruff format --check` would have failed on them from the start. Bring them onto the style the repository already declares, which is what makes the check adoptable at all. Formatting only: apart from quote style the diff is multi-line expressions joined where they fit inside 120 columns. Both suites still pass afterwards (25 pytest, and ruff check is clean). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
fddd82704f |
ci: keep pull requests away from the production admission webhooks
`kubectl apply --dry-run=server` persists nothing, but it does execute the admission webhooks of the real API server. The validate job runs on pull_request with no branch guard, so anyone able to open a PR could run arbitrary manifest content through cert-manager and Traefik in production. Limit the step to pushes to main. A pull request loses nothing by it: only main is ever deployed, and this job has to complete successfully before the deploy workflow is allowed to start, so a bad CRD is still caught before anything reaches the cluster -- on the push instead of on the PR. The skip is announced rather than silent, so a missing server-side pass does not read as a pass. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
0691536f28 |
fix(deploy): roll out our images by digest instead of a moving tag
`kubectl rollout undo` restores the previous ReplicaSet's pod template
verbatim. While that template names a tag, the rollback does not roll back
the image: the tag has already moved, so the reverted pod pulls the very
build that just failed and the cluster stays broken. The safety net added
in
|
||
|
|
2b9e34ba4a |
fix(k8s): add readiness probes so a bad image cannot look healthy
Six of eight workloads had no readinessProbe, so a pod turned Ready the moment its process started. The verify job relies on `rollout status`, so it passed for images that crash-looped or served errors, which left the rollback safety net inert. Each probe targets the path the service is actually reached on: - homepages: / (verified 200) - error-pages: /404.html, the path Traefik's errorPages middleware requests. / returns 403 by design and would never pass. - webinar-checker: /health (verified 200). /metrics also answers, but it is a Prometheus endpoint, not a readiness signal. The two userbot deployments stay without probes: they expose no port and no session file, and the panel reaches Telegram through its own client. A truthful signal there needs a health endpoint in the app itself. |
||
|
|
30d2b83efe |
chore(reloader): manage the reloader release from the repository
Reloader has been running since 23 September and is what makes the reloader.stakater.com/auto annotation on a pod template do anything, but the repository only held a namespace. It was a release someone installed by hand, so it was invisible to review, invisible to Renovate, and one reinstall away from being silently dropped. Declaring it in HELM_RELEASES pins the chart version somewhere Renovate can update it, and the active marker means the namespace is applied before the upgrade instead of only existing as a side effect of the original install. The values file sets nothing the running release does not already do, apart from resource requests and limits, which the chart leaves empty. |
||
|
|
db7bccfd89 |
fix(deploy): restart workloads whose image tag moved past what they run
Our manifests pin images to `:latest`, so a rebuild leaves the pod template byte-identical. kubectl apply sees no change, creates no ReplicaSet and pulls nothing, and the cluster keeps serving the previous build. imagePullPolicy: Always does not help, because it only decides whether a pod that *is* starting pulls, and no pod ever starts. All eight workloads that consume an image from our own registry were affected. Three of them had been running code from 23 September, and the single hardcoded `rollout restart deployment/userbot-panel` covered one of the eight. Restarting everything unconditionally was not the answer either: that bounces healthy services on every deploy, error-pages included, and the brief window where nothing answers is exactly what error-pages exists to prevent. So compare what each workload actually runs against what the tag resolves to now, and restart only the ones that differ. When the tag still points at the running digest nothing happens, so a redeploy that changed no image is a no-op. Scope is the repository, deliberately. Five more workloads run our images but have no manifest here, and they are applied out of band. Walking the manifests rather than the cluster means this can never reach them. The digest is resolved for the node architecture. A multi-arch tag also carries `unknown/unknown` entries for the build attestation, and a pod's imageID is always the per-platform digest, so comparing the wrong entry would mark everything stale forever. Once a restart happens it bumps the generation, which is what makes the change visible to changed_workloads and therefore watchable and revertible by the verify stage. |
||
|
|
1505b638ce |
fix(deploy): verify and roll back in a separate job
verify_workloads ended on `[ -s "$failed_file" ]`, which is the opposite of what its own contract says. A non-empty file means something failed, so the function returned success exactly when a workload never came up, and failure when everything was fine. Every rollback was therefore skipped, and every deploy that changed anything ended red with an empty failure list and a bogus "Rolled back successfully". Worse, the check only ever ran at the end of stage_apply_k8s, inside the same process as the apply. A job killed by timeout-minutes, cancelled by a new push, or cut off by a dropped SSH connection never reached it, which is precisely when a rollback matters. The three helm upgrades alone can consume the whole 30-minute job budget, so that path was reachable. Verification now lives in its own job, gated on always(), so it runs whatever happened to the apply. The apply stage publishes its pre-apply snapshot through DEPLOY_SNAPSHOT_DIR/current before touching anything, and the verify stage picks it up from there. A snapshot whose recorded commit does not match the deploy is refused rather than trusted, so a stale pointer from an earlier run cannot make the rollback revert the wrong workloads. An unwritable snapshot directory now fails the deploy up front instead of silently continuing without a way back. cancel-in-progress becomes false for the same reason: cancelling a run kills the apply job and takes the verify job with it, which is the failure this change exists to prevent. Both applies are idempotent, so queueing costs little. The SSH key moves to a per-run directory removed on exit, and the deploy is pinned to the exact commit CI validated. |
||
|
|
7ce727bc8a |
chore(renovate): move config under renovate/ and validate it in CI
The config lived in renovate.json at the repo root while everything else Renovate-related sat under renovate/, and renovate/config.js was a second, unused source of truth. Both are gone: renovate/renovate.json is now the only config file. Because the CronJob in the cluster cannot read the repository, its ConfigMap carries an inlined copy of the config. That copy is generated, and sync-renovate-configmap.sh --check now fails the build when it drifts from the source file. The workflows also stop carrying a copy of the renovate/renovate image tag. They read it from renovate/k8s/cronjob.yaml, so the version validated in CI is the version that actually runs in the cluster. ci.yaml validates the config with renovate-config-validator, checks the generated ConfigMap, and kubeconforms the CronJob's own manifests. |
||
|
|
f22793e32e |
ci: lint workflows and shell scripts, validate k8s against the API server
Adds three lint jobs (actionlint, shellcheck, compose) and a server-side dry-run of the active manifests. Previously the only k8s check was kubeconform, which has no schemas for CRDs, so every IngressRoute, Certificate, PrometheusRule and Middleware was silently skipped. The server-side pass needs the live API server because that is the only place the real CRD schemas and the cert-manager / Traefik admission webhooks exist. It is scoped to services carrying a k8s/active marker, since dry-run needs the target namespace to exist. userbot/ is excluded from shellcheck: it is a git subtree, and linting upstream's scripts would let a routine subtree pull turn the deploy gate red on code we do not own. kubeconform, shellcheck and actionlint are now installed from pinned versions in tool-versions.env rather than picked up from the runner's PATH. The Compose helper is shared with the deploy workflow so both check the same file set the same way. |
||
|
|
4a8d4feea0 |
fix(netbox): run probes with curl directly, not python -c
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 1s
ci / validate (push) Successful in 2s
renovate-ci / validate-renovate (push) Successful in 20s
ci / build (push) Successful in 1s
deploy / preflight (push) Successful in 2s
deploy / validate (push) Successful in 1m53s
deploy / apply-k8s (push) Successful in 1m47s
deploy / apply-compose (push) Successful in 11s
python -c 'exec /usr/bin/curl ...' is shell syntax, not Python, so every probe raised SyntaxError and the pod never became ready. Verified curl against /login/ returns HTTP 200. |
||
|
|
1d9a85bef9 |
fix(edu-master): stop 20722d false alert on zeroed last_success gauge
ci / validate (push) Successful in 2s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 1s
deploy / preflight (push) Successful in 2s
renovate-ci / validate-renovate (push) Successful in 14s
ci / build (push) Successful in 2s
deploy / validate (push) Successful in 1m52s
deploy / apply-k8s (push) Successful in 1m48s
deploy / apply-compose (push) Successful in 13s
checker.py initialises last_success to 0, so right after a pod restart `time() - last_success` equals the current epoch. The rule compared that against 300, went firing instantly, and humanizeDuration rendered the raw epoch delta as ~20722d. The last_run > 0 guard did not help because a run happens long before the first success. Guard the duration rule on last_success > 0, keeping the duration expression on the left of `and` so $value stays the real gap, and add a separate WebinarCheckerNeverSucceeded rule for the zeroed-gauge case so a checker that has never succeeded is still caught. |
||
|
|
b625d30568 |
ci(deploy): cancel superseded deploys on new push
deploy / preflight (push) Successful in 1s
deploy / validate (push) Successful in 1m53s
deploy / apply-k8s (push) Successful in 1m48s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 2s
ci / lint-yaml (push) Successful in 3s
ci / lint-dockerfiles (push) Successful in 6s
ci / validate (push) Successful in 1s
renovate-ci / validate-renovate (push) Successful in 13s
ci / build (push) Successful in 1s
deploy / apply-compose (push) Successful in 12s
Queued deploy runs were deploying origin/main at start anyway (preflight reset), so waiting runs duplicated the newest deploy instead of their own commit. Cancel them. |
||
|
|
d018a441af |
fix(deploy): move netbox from compose to k8s
ci / lint-dockerfiles (push) Successful in 1s
ci / validate (push) Successful in 2s
renovate-ci / validate-renovate (push) Successful in 17s
deploy / preflight (push) Successful in 2s
ci / build (push) Successful in 1s
deploy / apply-k8s (push) Canceled after 0s
deploy / apply-compose (push) Canceled after 0s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 3s
deploy / validate (push) Canceled after 4s
compose.yaml is a local-only stand on 127.0.0.1:8000 per README; the live service runs in-cluster. The root active marker made apply-compose fail on gitignored .env vars. |
||
|
|
ecb254017d |
fix(searxng): use existing image tag 2026.9.25-12f8b6515
ci / validate (push) Successful in 2s
ci / build (push) Successful in 1s
deploy / apply-k8s (push) Successful in 2m23s
deploy / apply-compose (push) Failing after 8s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 1s
deploy / preflight (push) Successful in 2s
renovate-ci / validate-renovate (push) Successful in 18s
deploy / validate (push) Successful in 1m46s
2026.09.13-d4ce87c23 was never published (upstream tags month without leading zero); rollout stuck in ImagePullBackOff. Verified replacement tag exists on Docker Hub. |
||
|
|
41f18ea993 |
fix(deploy): move netbird from compose to k8s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 3s
ci / lint-dockerfiles (push) Successful in 1s
ci / validate (push) Successful in 2s
deploy / preflight (push) Successful in 2s
renovate-ci / validate-renovate (push) Successful in 19s
ci / build (push) Successful in 2s
deploy / validate (push) Successful in 1m45s
deploy / apply-k8s (push) Successful in 1m52s
deploy / apply-compose (push) Failing after 24s
netbird runs in-cluster; the root active marker made apply-compose pick up netbird/compose.yaml and fail on gitignored .env vars. Drop the compose marker and enable k8s/active instead. |
||
|
|
91c344fe2c |
fix(monitoring): disable control-plane alerts and scrapes on k0s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 1s
ci / validate (push) Successful in 3s
deploy / preflight (push) Successful in 2s
renovate-ci / validate-renovate (push) Successful in 24s
deploy / validate (push) Successful in 1m42s
deploy / apply-k8s (push) Successful in 1m45s
ci / build (push) Successful in 2s
deploy / apply-compose (push) Failing after 9s
k0s runs kube-controller-manager, kube-scheduler and etcd inside its own process rather than as pods, so the chart's Services never get endpoints and the targets stay permanently absent. Drop the matching ServiceMonitors and their Down/HighCommitDurations rules; kube-proxy and kubelet do get endpoints on k0s and stay enabled. |
||
|
|
cab6ef2102 |
Merge branch 'feat/netbird'
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 1s
ci / validate (push) Successful in 2s
deploy / preflight (push) Successful in 2s
deploy / validate (push) Successful in 1m43s
renovate-ci / validate-renovate (push) Successful in 14s
ci / build (push) Successful in 1s
deploy / apply-k8s (push) Successful in 3m15s
deploy / apply-compose (push) Failing after 41s
|
||
|
|
a2ff9515a3 |
fix(deploy): validate compose without workstation secrets
renovate-ci / validate-renovate (push) Skipped
deploy / validate (push) Skipped
ci / lint-prettier (push) Successful in 2s
ci / validate (push) Successful in 2s
ci / build (push) Skipped
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 1s
docker compose config required real values for gitignored .env files and secrets, so validate always failed on stacks with :? guards (netbird, netbox). Validate structure only via --no-interpolate, --no-env-resolution and --no-path-resolution, keeping normalization and consistency checks. |
||
|
|
62d39ee4f1 |
Merge pull request 'chore(deps): update netbirdio/dashboard docker tag to v2.93.0' (#53) from renovate/netbirdio-dashboard-2.x into main
ci / lint-dockerfiles (push) Successful in 1s
deploy / preflight (push) Successful in 2s
renovate-ci / validate-renovate (push) Successful in 9s
deploy / validate (push) Failing after 2s
deploy / apply-k8s (push) Skipped
deploy / apply-compose (push) Skipped
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 3s
ci / validate (push) Successful in 1s
ci / build (push) Successful in 1s
Reviewed-on: #53 |
||
|
|
f1f7dd4a0a |
chore(deps): update netbirdio/dashboard docker tag to v2.93.0
deploy / validate (push) Skipped
renovate-ci / validate-renovate (push) Skipped
ci / lint-prettier (pull_request) Successful in 4s
ci / lint-ruff (pull_request) Successful in 2s
ci / lint-yaml (pull_request) Successful in 2s
ci / lint-dockerfiles (pull_request) Successful in 1s
ci / validate (pull_request) Successful in 2s
ci / build (pull_request) Skipped
renovate-ci / validate-renovate (pull_request) Successful in 9s
ci / lint-prettier (push) Successful in 2s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 1s
ci / validate (push) Successful in 1s
ci / build (push) Skipped
|
||
|
|
f4df6d4437 |
Merge pull request 'chore(deps): update container patch updates' (#38) from renovate/container-patch-updates into main
renovate-ci / validate-renovate (push) Skipped
deploy / preflight (push) Successful in 3s
deploy / validate (push) Failing after 2s
deploy / apply-k8s (push) Skipped
deploy / apply-compose (push) Skipped
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 1s
ci / validate (push) Successful in 2s
ci / lint-prettier (pull_request) Successful in 3s
ci / lint-ruff (pull_request) Successful in 1s
ci / lint-yaml (pull_request) Successful in 2s
ci / validate (pull_request) Successful in 1s
ci / build (pull_request) Skipped
ci / build (push) Skipped
ci / lint-dockerfiles (pull_request) Successful in 1s
renovate-ci / validate-renovate (pull_request) Successful in 21s
Reviewed-on: #38 |
||
|
|
987a89f022 |
chore(deps): update container patch updates
deploy / validate (push) Skipped
renovate-ci / validate-renovate (push) Skipped
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 1s
ci / validate (push) Successful in 1s
ci / build (push) Skipped
ci / lint-prettier (pull_request) Canceled after 0s
ci / lint-yaml (pull_request) Canceled after 0s
ci / lint-dockerfiles (pull_request) Canceled after 0s
ci / validate (pull_request) Canceled after 0s
ci / lint-ruff (pull_request) Canceled after 0s
ci / build (pull_request) Canceled after 0s
renovate-ci / validate-renovate (pull_request) Successful in 18s
|
||
|
|
b423119632 |
Merge pull request 'chore(deps): update renovate/renovate docker tag to v44.115.9' (#47) from renovate/renovate-renovate-44.x into main
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 1s
ci / validate (push) Successful in 3s
ci / build (push) Canceled after 0s
deploy / preflight (push) Successful in 3s
deploy / validate (push) Canceled after 0s
deploy / apply-k8s (push) Canceled after 0s
deploy / apply-compose (push) Canceled after 0s
renovate-ci / validate-renovate (push) Successful in 1m30s
Reviewed-on: #47 |
||
|
|
ddbc0cc7e6 | chore(deps): update renovate/renovate docker tag to v44.115.9 | ||
|
|
6ad33d75e6 |
Merge pull request 'chore(deps): update prom/prometheus docker tag to v3.15.0' (#48) from renovate/prom-prometheus-3.x into main
ci / lint-prettier (push) Successful in 3s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 1s
ci / validate (push) Successful in 2s
ci / lint-ruff (push) Successful in 1s
deploy / preflight (push) Successful in 2s
ci / build (push) Canceled after 0s
deploy / validate (push) Canceled after 0s
deploy / apply-k8s (push) Canceled after 0s
deploy / apply-compose (push) Canceled after 0s
renovate-ci / validate-renovate (push) Successful in 24s
Reviewed-on: #48 |
||
|
|
11bfb426c0 | chore(deps): update prom/prometheus docker tag to v3.15.0 | ||
|
|
b7a1835adb |
Merge pull request 'feat(netbird): add tailscale-alternative' (#50) from feat/netbird into main
ci / lint-dockerfiles (push) Successful in 1s
ci / validate (push) Successful in 1s
deploy / preflight (push) Successful in 2s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 3s
ci / build (push) Canceled after 0s
deploy / validate (push) Canceled after 0s
deploy / apply-k8s (push) Canceled after 0s
deploy / apply-compose (push) Canceled after 0s
renovate-ci / validate-renovate (push) Successful in 21s
Reviewed-on: #50 |
||
|
|
ad4bb8750d |
chore(deploy): enable netbird
deploy / validate (push) Skipped
renovate-ci / validate-renovate (push) Skipped
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 1s
ci / validate (push) Successful in 2s
ci / build (push) Skipped
ci / lint-prettier (pull_request) Successful in 3s
ci / lint-ruff (pull_request) Successful in 1s
ci / lint-yaml (pull_request) Successful in 2s
ci / lint-dockerfiles (pull_request) Successful in 1s
ci / validate (pull_request) Successful in 1s
ci / build (pull_request) Skipped
renovate-ci / validate-renovate (pull_request) Successful in 8s
|
||
|
|
f1e00b946f |
Merge pull request 'Feat/documenting services' (#51) from feat/documenting-services into main
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 0s
ci / lint-yaml (push) Successful in 3s
ci / lint-dockerfiles (push) Successful in 1s
ci / validate (push) Successful in 1s
deploy / preflight (push) Successful in 2s
renovate-ci / validate-renovate (push) Successful in 9s
ci / build (push) Successful in 1s
deploy / validate (push) Failing after 1s
deploy / apply-k8s (push) Skipped
deploy / apply-compose (push) Skipped
Reviewed-on: #51 |
||
|
|
7ee7d0c961 |
Update README.md
deploy / validate (push) Skipped
renovate-ci / validate-renovate (push) Skipped
ci / lint-dockerfiles (push) Successful in 1s
ci / validate (push) Successful in 1s
ci / build (push) Skipped
ci / lint-prettier (pull_request) Successful in 3s
ci / lint-ruff (pull_request) Successful in 0s
ci / lint-yaml (pull_request) Successful in 2s
ci / validate (pull_request) Successful in 1s
ci / build (pull_request) Skipped
renovate-ci / validate-renovate (pull_request) Successful in 8s
ci / lint-prettier (push) Successful in 2s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (pull_request) Successful in 1s
|
||
|
|
27ab6b859e |
chore(deploy): enable netbird
deploy / validate (push) Skipped
renovate-ci / validate-renovate (push) Skipped
ci / lint-prettier (pull_request) Successful in 3s
ci / lint-ruff (pull_request) Successful in 1s
ci / validate (pull_request) Successful in 1s
ci / build (pull_request) Skipped
renovate-ci / validate-renovate (pull_request) Successful in 13s
ci / lint-yaml (pull_request) Successful in 2s
ci / lint-dockerfiles (pull_request) Successful in 0s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 0s
ci / validate (push) Successful in 2s
ci / build (push) Skipped
|
||
|
|
71e769c002 |
chore(deploy): disable checkmk, enable netbox and rackpeek
deploy / validate (push) Skipped
renovate-ci / validate-renovate (push) Skipped
ci / lint-prettier (push) Failing after 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 1s
ci / validate (push) Successful in 2s
ci / build (push) Skipped
ci / lint-prettier (pull_request) Failing after 3s
ci / lint-ruff (pull_request) Successful in 1s
ci / lint-yaml (pull_request) Successful in 3s
ci / lint-dockerfiles (pull_request) Successful in 1s
ci / validate (pull_request) Successful in 1s
ci / build (pull_request) Skipped
renovate-ci / validate-renovate (pull_request) Successful in 21s
|