436fd1ecae8be688fd5065e637fab6f6ea4af4eb
660
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
436fd1ecae |
chore: extract userbot subtree to its own repo
ci / lint-compose (push) Successful in 3s
ci / lint-actionlint (push) Successful in 1s
ci / lint-shellcheck (push) Successful in 2s
ci / lint-prettier (push) Successful in 2s
ci / lint-ruff (push) Successful in 2s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 1s
ci / validate (push) Successful in 2s
renovate-ci / validate-renovate (push) Successful in 17s
ci / build (push) Successful in 8s
userbot/ (bot plus panel) now lives at /home/forust/userbot as a clone of forust/userbot instead of a subtree in homelab. Cleans up the pipeline references that only existed for it: scan-deps, test-backend and test-frontend jobs, the userbot build matrix entries, the deploy panel hook, and the userbot-only pyrightconfig. Running cluster workloads are untouched; homelab just stops building and testing upstream's code. |
||
|
|
a564dd8b67 |
fix(alerts): exclude netbird streams from traefik latency alert
ci / test-frontend (push) Successful in 10s
ci / validate (push) Successful in 3s
renovate-ci / validate-renovate (push) Successful in 12s
ci / lint-compose (push) Successful in 3s
ci / lint-actionlint (push) Successful in 2s
ci / lint-shellcheck (push) Successful in 2s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 2s
ci / lint-yaml (push) Successful in 3s
ci / lint-dockerfiles (push) Successful in 2s
ci / scan-deps (push) Successful in 17s
ci / test-backend (push) Successful in 8s
ci / build (push) Successful in 10s
SignalExchange ConnectStream holds 60s gRPC streams by design; at night they exceed 5% of samples and pin P95 to the 5.0s bucket ceiling, flapping the alert. Also fixes the stale xui exclusion pattern, which matched no real service label. |
||
|
|
c2c90c490d |
Merge pull request 'chore(deps): update ghcr.io/autobrr/netronome docker tag to v0.15.0' (#62) from renovate/ghcr.io-autobrr-netronome-0.x into main
ci / lint-compose (push) Successful in 2s
ci / lint-actionlint (push) Successful in 1s
ci / lint-shellcheck (push) Successful in 2s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 2s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 2s
ci / scan-deps (push) Successful in 15s
ci / test-backend (push) Successful in 7s
ci / test-frontend (push) Successful in 11s
ci / validate (push) Successful in 2s
renovate-ci / validate-renovate (push) Successful in 13s
ci / build (push) Successful in 16s
Reviewed-on: #62 |
||
|
|
da0f9e84c3 | chore(deps): update ghcr.io/autobrr/netronome docker tag to v0.15.0 | ||
|
|
3ecc12300a |
Merge pull request 'chore(deps): update ghcr.io/alexta69/metube docker tag to v2026.09.28' (#61) from renovate/container-patch-updates into main
renovate-ci / validate-renovate (push) Successful in 13s
ci / lint-compose (push) Canceled after 0s
ci / lint-actionlint (push) Canceled after 0s
ci / lint-shellcheck (push) Canceled after 0s
ci / lint-prettier (push) Canceled after 0s
ci / lint-ruff (push) Canceled after 0s
ci / lint-yaml (push) Canceled after 0s
ci / lint-dockerfiles (push) Canceled after 0s
ci / scan-deps (push) Canceled after 0s
ci / test-backend (push) Canceled after 0s
ci / test-frontend (push) Canceled after 0s
ci / validate (push) Canceled after 0s
ci / build (push) Canceled after 0s
Reviewed-on: #61 |
||
|
|
19029fa012 | chore(deps): update ghcr.io/alexta69/metube docker tag to v2026.09.28 | ||
|
|
2a5d690e34 |
Merge pull request 'chore(deps): update netbirdio/dashboard docker tag to v2.94.0' (#63) from renovate/netbirdio-dashboard-2.x into main
ci / lint-compose (push) Canceled after 0s
ci / lint-actionlint (push) Canceled after 0s
ci / lint-shellcheck (push) Canceled after 0s
ci / lint-prettier (push) Canceled after 0s
ci / lint-ruff (push) Canceled after 0s
ci / lint-yaml (push) Canceled after 0s
ci / lint-dockerfiles (push) Canceled after 0s
ci / scan-deps (push) Canceled after 0s
ci / test-backend (push) Canceled after 0s
ci / test-frontend (push) Canceled after 0s
ci / validate (push) Canceled after 0s
ci / build (push) Canceled after 0s
renovate-ci / validate-renovate (push) Successful in 19s
Reviewed-on: #63 |
||
|
|
b53ce36d89 | chore(deps): update netbirdio/dashboard docker tag to v2.94.0 | ||
|
|
a8c4e9bcbe |
Merge pull request 'chore(deps): update renovate/renovate docker tag to v44.117.0' (#64) from renovate/renovate-self-update into main
renovate-ci / validate-renovate (push) Successful in 16s
ci / lint-compose (push) Successful in 3s
ci / lint-actionlint (push) Successful in 1s
ci / lint-shellcheck (push) Successful in 2s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 2s
ci / lint-yaml (push) Successful in 3s
ci / lint-dockerfiles (push) Successful in 2s
ci / scan-deps (push) Successful in 15s
ci / test-backend (push) Successful in 7s
ci / test-frontend (push) Successful in 10s
ci / validate (push) Successful in 2s
ci / build (push) Successful in 9s
Reviewed-on: #64 |
||
|
|
761f97ef8e | chore(deps): update renovate/renovate docker tag to v44.117.0 | ||
|
|
0c9743e241 |
Merge pull request 'chore(deps): update ghcr.io/immich-app/postgres docker tag to v16' (#65) from renovate/ghcr.io-immich-app-postgres-16.x into main
ci / lint-compose (push) Successful in 4s
ci / lint-actionlint (push) Successful in 1s
ci / lint-shellcheck (push) Successful in 2s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 2s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 3s
ci / scan-deps (push) Successful in 14s
ci / test-backend (push) Successful in 8s
ci / test-frontend (push) Successful in 10s
ci / validate (push) Successful in 3s
renovate-ci / validate-renovate (push) Successful in 27s
ci / build (push) Successful in 8s
Reviewed-on: #65 |
||
|
|
2b3c28a46b | chore(deps): update ghcr.io/immich-app/postgres docker tag to v16 | ||
|
|
2cb06debc5 |
fix(deploy): retry registry lookups with a timeout
ci / lint-compose (push) Successful in 3s
ci / lint-actionlint (push) Successful in 1s
ci / lint-shellcheck (push) Successful in 2s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 1s
ci / scan-deps (push) Successful in 19s
ci / test-backend (push) Successful in 8s
ci / test-frontend (push) Successful in 10s
ci / validate (push) Successful in 2s
renovate-ci / validate-renovate (push) Successful in 14s
ci / build (push) Successful in 7s
A single blink of the registry failed render_pinned for the whole file and redded the apply stage. registry_digest now retries 3 times under a 25s timeout with a warning per attempt; empty still means unresolvable and callers report it by name as before. |
||
|
|
a6af69dca0 |
fix(k8s): Recreate singletons and trim requests for scheduler headroom
ci / lint-actionlint (push) Successful in 1s
ci / lint-shellcheck (push) Successful in 2s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 2s
ci / lint-compose (push) Successful in 3s
ci / test-backend (push) Successful in 8s
ci / test-frontend (push) Successful in 11s
ci / validate (push) Successful in 3s
renovate-ci / validate-renovate (push) Successful in 13s
ci / lint-dockerfiles (push) Successful in 2s
ci / scan-deps (push) Successful in 18s
ci / build (push) Successful in 1m30s
RollingUpdate with default maxSurge needs a spare pod the single node does not have (99% CPU requested), so multi-workload restarts end Pending and verify times out. Recreate on all replicas:1 Deployments (immich-server and bentopdf keep RollingUpdate at replicas 2). Also trims CPU/memory requests toward measured use (adguard, authentik, gitea, netbox, uptime-kuma, netbird-server) and gives traefik requests/limits so it is no longer BestEffort. |
||
|
|
76f39da90c |
fix(ci): skip heavy jobs on renovate branches, automerge digest and patch
ci / lint-compose (push) Canceled after 0s
ci / lint-actionlint (push) Canceled after 0s
ci / lint-shellcheck (push) Canceled after 0s
ci / lint-prettier (push) Canceled after 0s
ci / lint-ruff (push) Canceled after 0s
ci / lint-yaml (push) Canceled after 0s
ci / lint-dockerfiles (push) Canceled after 0s
ci / scan-deps (push) Canceled after 0s
ci / test-backend (push) Canceled after 0s
ci / test-frontend (push) Canceled after 0s
ci / validate (push) Canceled after 0s
ci / build (push) Canceled after 0s
renovate-ci / validate-renovate (push) Canceled after 0s
Renovate branches only carry version/digest bumps, so scan-deps, test-backend, test-frontend and build just burn runner time on the box that also serves prod. Static checks and validate still run. Digest and patch updates automerge (playwright, helm and major rules below still override to no-automerge). Also replaces deprecated helm --atomic with --wait --rollback-on-failure. |
||
|
|
86730ff0c4 |
Merge pull request 'chore(deps): update ghcr.io/c4illin/convertx docker tag to v0.19.0' (#55) from renovate/ghcr.io-c4illin-convertx-0.x into main
renovate-ci / validate-renovate (push) Successful in 6s
ci / lint-compose (push) Successful in 3s
ci / lint-actionlint (push) Successful in 1s
ci / lint-shellcheck (push) Successful in 2s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 2s
ci / lint-yaml (push) Successful in 3s
ci / lint-dockerfiles (push) Successful in 2s
ci / scan-deps (push) Successful in 14s
ci / test-backend (push) Successful in 7s
ci / test-frontend (push) Successful in 10s
ci / validate (push) Successful in 3s
ci / build (push) Successful in 10s
Reviewed-on: #55 |
||
|
|
af66d3e7fd | chore(deps): update ghcr.io/c4illin/convertx docker tag to v0.19.0 | ||
|
|
deeaefa695 |
Merge pull request 'chore(deps): update ghcr.io/henriquesebastiao/downtify docker tag to v3.2.0' (#60) from renovate/ghcr.io-henriquesebastiao-downtify-3.x into main
ci / lint-compose (push) Canceled after 0s
ci / lint-actionlint (push) Canceled after 0s
ci / lint-shellcheck (push) Canceled after 0s
ci / lint-prettier (push) Canceled after 0s
ci / lint-ruff (push) Canceled after 0s
ci / lint-yaml (push) Canceled after 0s
ci / lint-dockerfiles (push) Canceled after 0s
ci / scan-deps (push) Canceled after 0s
ci / test-backend (push) Canceled after 0s
ci / test-frontend (push) Canceled after 0s
ci / validate (push) Canceled after 0s
ci / build (push) Canceled after 0s
renovate-ci / validate-renovate (push) Successful in 18s
Reviewed-on: #60 |
||
|
|
6999dd2728 | chore(deps): update ghcr.io/henriquesebastiao/downtify docker tag to v3.2.0 | ||
|
|
742d78944b |
Merge pull request 'chore(deps): update ghcr.io/alexta69/metube docker tag to v2026.09.27' (#54) from renovate/container-patch-updates into main
ci / lint-compose (push) Canceled after 0s
ci / lint-actionlint (push) Canceled after 0s
ci / lint-shellcheck (push) Canceled after 0s
ci / lint-prettier (push) Canceled after 0s
ci / lint-ruff (push) Canceled after 0s
ci / lint-yaml (push) Canceled after 0s
ci / lint-dockerfiles (push) Canceled after 0s
ci / scan-deps (push) Canceled after 0s
ci / test-backend (push) Canceled after 0s
ci / test-frontend (push) Canceled after 0s
ci / validate (push) Canceled after 0s
ci / build (push) Canceled after 0s
renovate-ci / validate-renovate (push) Successful in 23s
Reviewed-on: #54 |
||
|
|
e451c97dfc |
chore(deps): update container patch updates
renovate-ci / validate-renovate (push) Skipped
ci / lint-actionlint (push) Successful in 1s
ci / lint-shellcheck (push) Successful in 2s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 2s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 1s
ci / scan-deps (push) Successful in 15s
ci / lint-actionlint (pull_request) Successful in 1s
ci / lint-compose (push) Successful in 3s
ci / test-backend (push) Successful in 7s
ci / test-frontend (push) Successful in 11s
ci / validate (push) Successful in 2s
ci / build (push) Skipped
ci / lint-compose (pull_request) Successful in 3s
ci / lint-shellcheck (pull_request) Successful in 2s
ci / lint-prettier (pull_request) Successful in 2s
ci / lint-ruff (pull_request) Successful in 1s
ci / lint-yaml (pull_request) Successful in 3s
ci / lint-dockerfiles (pull_request) Successful in 2s
ci / scan-deps (pull_request) Successful in 14s
ci / test-backend (pull_request) Successful in 7s
ci / test-frontend (pull_request) Successful in 11s
ci / validate (pull_request) Successful in 2s
ci / build (pull_request) Skipped
renovate-ci / validate-renovate (pull_request) Successful in 20s
|
||
|
|
33c54ac830 |
Merge pull request 'chore(deps): update docker.io/valkey/valkey:9 docker digest to 418652c' (#59) from renovate/docker.io-valkey-valkey-9 into main
ci / lint-compose (push) Successful in 4s
ci / lint-actionlint (push) Successful in 1s
ci / lint-shellcheck (push) Successful in 2s
ci / lint-prettier (push) Successful in 3s
ci / lint-yaml (push) Successful in 2s
ci / lint-ruff (push) Successful in 3s
ci / lint-dockerfiles (push) Successful in 1s
ci / scan-deps (push) Successful in 19s
ci / test-backend (push) Successful in 8s
ci / test-frontend (push) Successful in 10s
ci / validate (push) Successful in 3s
renovate-ci / validate-renovate (push) Successful in 14s
ci / build (push) Successful in 11s
Reviewed-on: #59 |
||
|
|
b09d718310 |
chore(deps): update docker.io/valkey/valkey:9 docker digest to 418652c
renovate-ci / validate-renovate (push) Skipped
ci / lint-compose (pull_request) Successful in 6s
ci / lint-actionlint (pull_request) Successful in 2s
ci / lint-shellcheck (pull_request) Successful in 2s
ci / lint-prettier (pull_request) Successful in 3s
ci / lint-ruff (pull_request) Successful in 2s
ci / lint-yaml (pull_request) Successful in 2s
ci / lint-dockerfiles (pull_request) Successful in 3s
ci / scan-deps (pull_request) Successful in 55s
ci / test-backend (pull_request) Successful in 9s
ci / test-frontend (pull_request) Successful in 11s
ci / validate (pull_request) Successful in 3s
ci / build (pull_request) Skipped
ci / lint-compose (push) Successful in 3s
ci / lint-actionlint (push) Successful in 1s
ci / lint-shellcheck (push) Successful in 2s
ci / lint-prettier (push) Successful in 2s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 2s
ci / scan-deps (push) Successful in 14s
ci / test-backend (push) Successful in 8s
ci / test-frontend (push) Successful in 11s
ci / validate (push) Successful in 1s
ci / build (push) Skipped
renovate-ci / validate-renovate (pull_request) Successful in 53s
|
||
|
|
b4f76373bb |
fix(gitea): serve issue search from postgres instead of reindexing on boot
renovate-ci / validate-renovate (push) Successful in 5s
ci / lint-compose (push) Successful in 4s
ci / lint-actionlint (push) Successful in 1s
ci / lint-shellcheck (push) Successful in 3s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 2s
ci / lint-yaml (push) Successful in 3s
ci / lint-dockerfiles (push) Successful in 2s
ci / scan-deps (push) Successful in 17s
ci / test-backend (push) Successful in 9s
ci / test-frontend (push) Successful in 13s
ci / validate (push) Successful in 2s
ci / build (push) Successful in 7s
cron.rebuild_issue_indexer runs at start, so every gitea pod restart reindexed the whole issue index and read ~7MB/s off the rotational disk for an hour. |
||
|
|
74adf38d63 |
feat(alerts): cover OOM kills, restart loops and evictions
The OOMKilled container behind the immich crash loop was invisible: PodCrashLooping only fires once kubelet has already given up and started the backoff. |
||
|
|
f5b2f89f38 |
style(immich): match repo prettier quoting in compose file
ci / lint-compose (push) Successful in 4s
ci / lint-actionlint (push) Successful in 2s
ci / lint-shellcheck (push) Successful in 2s
renovate-ci / validate-renovate (push) Successful in 9s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 2s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 2s
ci / scan-deps (push) Successful in 16s
ci / test-backend (push) Successful in 9s
ci / test-frontend (push) Successful in 13s
ci / validate (push) Successful in 3s
ci / build (push) Successful in 9s
|
||
|
|
8e63284240 |
feat(immich): add self-hosted photo backup with dedicated postgres
ci / lint-compose (push) Successful in 4s
ci / lint-actionlint (push) Successful in 1s
ci / lint-shellcheck (push) Successful in 3s
ci / lint-prettier (push) Failing after 4s
ci / lint-ruff (push) Successful in 3s
ci / lint-yaml (push) Successful in 3s
ci / lint-dockerfiles (push) Successful in 2s
ci / scan-deps (push) Successful in 20s
ci / test-backend (push) Successful in 8s
ci / validate (push) Successful in 3s
ci / build (push) Skipped
renovate-ci / validate-renovate (push) Successful in 11s
ci / test-frontend (push) Successful in 14s
Server x2, machine learning, valkey and VectorChord postgres on local storage, Traefik routes for external and internal access. |
||
|
|
1d81410cd8 |
fix(traefik): persist plugin storage on a PVC
ci / lint-compose (push) Successful in 4s
ci / lint-actionlint (push) Successful in 3s
ci / lint-shellcheck (push) Successful in 4s
ci / lint-prettier (push) Successful in 5s
ci / lint-ruff (push) Successful in 3s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 3s
ci / scan-deps (push) Successful in 18s
ci / test-backend (push) Successful in 8s
ci / test-frontend (push) Successful in 11s
ci / validate (push) Successful in 4s
renovate-ci / validate-renovate (push) Successful in 1m28s
ci / build (push) Successful in 37s
Mount traefik-plugins PVC at /plugins-storage instead of the chart default emptyDir, so the crowdsec-bouncer download survives node reboots. Without this Traefik starts before the network is ready, the download from plugins.traefik.io times out, plugins get disabled and every route behind the middleware returns 404/503 until a manual restart. |
||
|
|
2f891a5d31 |
fix(postgres): give probes room on an I/O-bound single node
renovate-ci / validate-renovate (push) Canceled after 21s
ci / lint-compose (push) Successful in 8s
ci / lint-actionlint (push) Successful in 3s
ci / lint-shellcheck (push) Successful in 4s
ci / lint-prettier (push) Successful in 4s
ci / lint-ruff (push) Successful in 10s
ci / lint-yaml (push) Successful in 4s
ci / lint-dockerfiles (push) Successful in 4s
ci / scan-deps (push) Successful in 19s
ci / test-backend (push) Successful in 9s
ci / test-frontend (push) Successful in 21s
ci / validate (push) Successful in 5s
ci / build (push) Successful in 22s
pg_isready with the 1s default times out under I/O stall and kubelet kills a healthy postgres mid-recovery; each kill restarts a multi-minute fsync from zero and loops forever. readiness/liveness timeout 5s, liveness threshold 5, startup budget 15min. |
||
|
|
cda0022d81 |
fix(renovate): reap finished job pods with ttlSecondsAfterFinished
ci / lint-compose (push) Successful in 5s
ci / lint-actionlint (push) Successful in 2s
ci / lint-shellcheck (push) Successful in 2s
ci / lint-prettier (push) Successful in 4s
ci / lint-ruff (push) Successful in 2s
ci / lint-yaml (push) Successful in 3s
ci / lint-dockerfiles (push) Successful in 3s
ci / scan-deps (push) Successful in 17s
ci / test-backend (push) Successful in 7s
ci / test-frontend (push) Successful in 12s
ci / validate (push) Successful in 5s
renovate-ci / validate-renovate (push) Successful in 9m17s
ci / build (push) Failing after 28m20s
History limits never delete manual 'create job --from' runs, so Failed pods accumulated for a week. Keep a day for debugging, reap the rest. |
||
|
|
dde6b1c743 |
fix(deploy): recover helm releases from pending-* and skip helm-owned rollbacks
An --atomic upgrade whose own rollback never finishes leaves the release in pending-*, blocking every future run until a human rolls back (loki rev 18/21). Recover automatically before and after each upgrade, and fail loud when recovery does not land on deployed. Also skip helm-managed workloads in rollback_workloads: rollout undo there would step back to the revision --atomic just escaped. |
||
|
|
7c4843c88c |
fix(loki): unblock gateway rollout on a single node
Chart default is required podAntiAffinity on hostname plus RollingUpdate 25%/25%, which is maxUnavailable=0 at replicas=1: the new pod stays Unschedulable while the old one lives, and the old one never leaves while the new one is not Ready. Null the affinity (an empty map deep-merges with the default and keeps the rule) and set maxSurge/maxUnavailable to 1. |
||
|
|
f9e4623ade |
fix(k8s): size the remaining workloads against measured use
Finishes the sizing pass over every workload the deploy actually manages. Each request is at or above the container's p95 over the last seven days, so nothing is sized below what it is known to use, and each limit is between 1.6x and 5x the observed max, which is the figure that decides whether a burst gets an OOMKill. Some of these go up, and that is the point. adguard was holding 975M against a 500Mi request and netbox 962M against 512Mi, so both sat permanently above their own request and were standing eviction candidates on a node that has about 300M of headroom. Raising a request costs scheduler room; leaving it low costs the pod its place in the queue when the node gets tight. Others come down. loki ran with a 2Gi limit on 224M, gitea 1.5Gi on 305M, the authentik worker 1Gi on 305M, and a tail of single-purpose pods -- redis, glance, the two homepages, cfddns, session-keeper, the netbird dashboard, the loki gateway -- each reserved 4x to 16x more than they have ever touched. prometheus gets the opposite treatment: 768Mi/2560Mi, above its p95, because it compacts its TSDB in place and that is a burst worth budgeting for rather than throttling. Two of these limits are close enough to the observed max to be worth watching rather than trusting: adguard at 1.5x, and its DNS cache grows monotonically, so the ceiling is a date, not a margin. That was true before this change too; the pod sizing does not fix it and the cache needs bounding. CPU limits are untouched throughout. Leaving postgres alone as well: it sits in an uncommitted file that belongs to other work in progress. Verified: every request is at or above p95 and every limit above the observed max across all 74 containers, and 16/16 local gates pass. |
||
|
|
2ad4fa1b82 |
chore(portainer): stop deploying a container manager nothing routes to
Portainer had been running for 111 days with a 512Mi request and a 2Gi limit against 50M of measured use, on a node that is short of memory. It is a UI over the Docker socket; nothing in the repo or the cluster depends on it. The marker goes, not the manifests. `portainer/k8s/active` is what puts these files in the deploy's manifest set, so without it the next push leaves the namespace alone and the manifests stay on disk for a one-command return. This also matters for the smoke stage: that host list is built from the active directories, so `portainer.forust.xyz` leaves it and the new router check does not go looking for a route to a service we just retired. In the cluster the Deployment, the Service and both IngressRoutes are deleted. The routes go first: leaving an IngressRoute behind a deleted Service keeps a Traefik router pointing at nothing, which answers 502 while looking perfectly healthy to the stage that just started checking for routers. Deliberately kept, so this is reversible rather than destructive: the namespace, the 2Gi `portainer-data-pvc` and both Certificates stay. Re-enabling is `git checkout HEAD~1 -- portainer/k8s/active` plus an apply, and no Let's Encrypt quota is spent reissuing the production certificate. `glance` still links to `portainer.forust.xyz` and that tile will now be a 404. Left alone on purpose. The Cloudflare record is manual and cfddns only ever creates records, so `portainer.forust.xyz` keeps resolving until it is removed in the dashboard, same as `dockmon.forust.xyz`. Verified: no Traefik router matches portainer any more, all 19 hosts left in the smoke list still have a router, and 16/16 local gates pass. |
||
|
|
2a4f215546 |
fix(k8s): bring the over-reserved memory limits down to measured use
Six pods reserved far more memory than they have ever touched. uptime-kuma held a 3Gi limit against 469M of measured p95, metube 2Gi against 72M, convertx 1.5Gi against 85M, netbird-server 1Gi against 97M, searxng 700Mi against 134M and bentopdf 700Mi against 4M. Every one of them is a ceiling the scheduler counts against the node while the memory sits unused. Requests move down with the limits but never below the measured p95, so none of these becomes an eviction candidate as a side effect of being right-sized. The limits keep between 2.2x and 11.6x over the observed max, which is the figure that decides whether a pod gets OOM-killed during a burst. Net effect across the six: requests -557M, limits -4.6Gi, all of it ceiling that was never in use. This is the first change that actually gives memory back. CPU limits are left exactly as they were. They were not part of the sizing pass, they are not being hit on a node sitting at 5% CPU, and removing them is a separate decision from moving memory. Verified: each limit is above the container's own observed max and each request is above its p95, and 16/16 local gates pass. |
||
|
|
16aaeb60c1 |
fix(k8s): set requests and limits on the pods that shipped with neither
Eighteen containers had no memory limit at all, so nothing on the node could bound them. Three of the values files even claimed to set resources: Helm does not complain about a key it does not recognise, so the block sat there looking like a limit while the pod ran unbounded. alloy is the one that mattered. The chart reads `alloy.resources`; the file had `controller.resources`, so the DaemonSet that tails every pod log on the node shipped with nothing at all. `kubeStateMetrics` is the same trap in a different shape -- that is the condition key, the values live under `kube-state-metrics` -- and `configReloader` in the alloy chart sits at the top level rather than under `alloy`. Each one is verified by rendering the chart and reading the resources back off the containers, because a values key that is ignored looks exactly like one that works. reloader turned out to be set and still wrong: 64Mi request against a measured p95 of 73M, so the pod ran permanently above its own request and stayed a standing eviction candidate. That is the pod that restarts every other pod, so it is the last one that should be evicted. Raised to 96Mi. Requests are set at p95 throughout, grafana, playwright and alloy included. Left at the values first proposed they would have sat below their own p95 and queued for eviction ahead of everything smaller. CPU limits are deliberately absent: the node is I/O bound at 5% CPU, and CFS throttling would turn disk wait into runnable-throttled, which is the failure mode that took the node down. The prometheus and alertmanager configReloader sidecars are left open: chart 86.2.3 does not template the key, so reaching those two containers needs a postRenderer. Verified: all four charts render with the resources landing on the intended containers, and 16/16 local gates pass. |
||
|
|
a5409edbf2 |
fix(deploy): fail the smoke stage when Traefik has no route for a host
ci / lint-compose (push) Successful in 5s
ci / lint-actionlint (push) Successful in 2s
ci / lint-shellcheck (push) Successful in 3s
ci / lint-prettier (push) Successful in 2s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 3s
ci / lint-dockerfiles (push) Successful in 3s
ci / scan-deps (push) Successful in 15s
renovate-ci / validate-renovate (push) Successful in 1m15s
ci / test-frontend (push) Successful in 14s
ci / validate (push) Successful in 4s
ci / test-backend (push) Failing after 13m28s
ci / build (push) Skipped
The smoke stage treats any HTTP response as proof the service is serving, which is right -- a 302 to a login or a 404 from a path the app does not serve still means the chain is intact. But a 404 is not evidence of that on its own: a router Traefik refused to build answers with exactly the same 404 and nothing behind it. That is not hypothetical. The crowdsec bouncer is a plugin, and when Traefik cannot fetch it at startup it disables the plugin without failing, then drops every router whose chain referenced it. Sixteen routes answered 404 and the stage printed `ok` for all sixteen, because a dropped router and an unserved path are indistinguishable from outside. The Kubernetes objects cannot tell us either: the IngressRoute is still sitting there looking healthy, the router Traefik built from it is simply not there. So ask Traefik. api.insecure is already on for the internal entrypoint and the router list says which hosts it matches right now. Every probed host has to appear in that list. HTTP routers only -- the TCP ones match on a HostSNI wildcard and the UDP ones carry no rule at all, both selected by entrypoint and port, so neither can answer the question. A router mid-rollout is legitimately absent for a moment, so the list is re-read twice over 20s; a plugin that failed to load stays absent and waiting cannot rescue it. An unreadable router list fails the stage rather than skipping the check, since a check that cannot run is not a passing check. Verified against the live cluster: all 23 public routes have a router and the stage passes. With gitea, grafana and uptime removed from that list the probes still answer and the stage fails on exactly those three. |
||
|
|
c70d2db3a1 |
ci: deploy the image the commit built, not whatever the tag points at
Every service tracked the mutable `:prod` tag, so a deploy applied whatever that tag happened to name at the time rather than the commit it was deploying. A rollback had no way to state what it was rolling back to, and two deploys of one commit could land different images. CI now publishes an immutable `sha-<commit12>` tag beside `:prod` on main, and re-tags it for every image a push did not rebuild. That re-tag copies the manifest list, so no layer moves. The deploy resolves the immutable tag to a digest and pins the workload to it, and only falls back to the moving tag when the immutable one cannot be resolved -- which it says out loud, because that fallback is the deploy quietly ceasing to be reproducible from its own commit. The image list comes out of the tree with git grep rather than being written out a second time, so adding a service no longer means keeping two lists in step. build also gains the three jobs it was skipping -- scan-deps, test-backend, test-frontend -- so a change that breaks them cannot be tagged at all. The two run blocks where a mid-loop failure was survivable now run under set -euo pipefail: the build loop and the service detector both carried on past an error and could report a green build having produced nothing. The registry password moves from run: substitution into an env: block. A quote, a backtick or a $(...) in the password is parsed as shell before the command ever runs, and a login that failed that way looked exactly like a build that failed. The apply and verify timeouts stay at 45 and 30 minutes. The comments now record the arithmetic that says so rather than leaving the numbers to be raised on the next scare: three no-op helm upgrades run 3-5 minutes, one broken release is a single 10 minute rollback because the loop aborts on the first failure, and the apply loop itself is about a minute. That is roughly 15 minutes of work against a 45 minute budget. verify is 32 workloads at 8 wide -- four waves of 300 seconds, 20 minutes -- which leaves room for two serial rollbacks, and only becomes derivable at 45 once rollback_workloads is parallelised. |
||
|
|
24dd82e801 |
fix(crowdsec): make the bouncer trust the mobile range independently
The parser-stage whitelist already covers 84.245.64.0/18, so an event from the phone is dropped before it reaches a bucket and no decision is ever created for it - confirmed against 72h of traefik access logs, where the phone shows up as 84.245.120.147, inside that /18. But that left the bouncer's own ClientTrustedIPs without the range, so the guarantee rested on a single config. If the parser whitelist ever stops matching, a ban would be created and then served against the phone, which is the one thing that must not happen: the address belongs to a carrier, so it comes back to us by rotation and a 4h ban is not survivable from the device. ClientTrustedIPs bypasses the bouncer and the decision cache entirely, so repeating the range there holds even if a decision exists for any reason. All nine parser-stage ranges are now mirrored in the bouncer, and the bouncer has no range the parser stage does not know about. The file header now records that this middleware must be applied together with a traefik restart. Applying it alone wedges the plugin: in stream mode handleStreamTicker runs over package-level globals that no reconfiguration stops, so every route referencing the middleware answers 404 with 'invalid middleware crowdsec-crowdsec-bouncer@kubernetescrd' until the pod is replaced. Re-applying the prior config does not recover it and the config is not the cause - NewChecker is a plain net.ParseCIDR and cannot fail on a valid range. That cost 21 routes down before the restart requirement was found; recovery is a pod replace, ~35s. Verified live: middleware applied, traefik restarted, 17 of 20 hosts serving (the three exceptions are unchanged and unrelated - searxng returns its own 429, checkmk is down with 503, and one host is local-only), zero invalid-middleware errors, and the LAPI still shows /v1/decisions/stream polls at the 15s interval. |
||
|
|
11e92fdf4e |
fix(deploy): bound ssh hangs and retry the stage on transport loss
A connection that died silently used to hang until the job timeout, and the stage was never re-run. One flaky TCP session cost a whole 45-minute apply, and the symptom - a job that stops mid-output with no error - is what made the last few deploy failures expensive to read. ServerAliveInterval/CountMax cap how long a dead peer goes unnoticed at ~60s, ConnectTimeout caps setup. Only exit 255 - ssh's own transport failures - is retried, up to three attempts with a growing gap. A stage that fails on its own merits exits with the remote's status, so a real failure surfaces its own log immediately instead of being repeated three times over 45 minutes. The stages are declarative applies, so re-running one that had already committed is harmless. The stage environment now goes through `env` as separate argv entries rather than one interpolated string, so nothing in REPO, DEPLOY_SHA or DEPLOY_SNAPSHOT_DIR is re-split by the remote shell. Verified against a stubbed ssh: clean run attempts once, a single transport failure recovers on attempt 2 and exits 0, three failures give up preserving 255, and a stage failing with 1 or 7 attempts once and passes the code through unchanged. Also records why USERBOT_IMAGE stays on the prod tag: render_pinned rewrites only plain `image:` lines, and this ref is what the panel injects into the per-instance Deployments it creates, so those instances track the tag rather than the panel's own resolved digest. The two panel-created instances currently in the cluster are digest-pinned, so the panel does accept one either way; the tag is the choice, not a limitation. |
||
|
|
b1f98fc148 |
Revert "fix(compose): stop pointing at the tag the build dropped" for dtek_notif
dtek-notif is being picked up again, so leave its compose alone. It is also the one image not rebuilt since the build dropped :latest, so :prod does not exist for it yet - unlike the other six, which resolve. Keeping this file out of the change also keeps dtek_notif/* out of the build's changed-service detector, so pushing does not build an image nobody asked for. |
||
|
|
9b91b5847e |
fix(compose): stop pointing at the tag the build dropped
These five stacks still asked for :latest, but the build stopped pushing it - on main it only pushes main and prod, on dev only dev. Every one of these images therefore resolved only because the registry still had a stale :latest from before that change, and the next time one of them was actually built the reference would have dangled. Named for dtek-notif: of the seven, only that one still resolved at :prod, because it is the only image not rebuilt since the build dropped :latest - and it is the only one of these five stacks the deploy does not manage (no `active` marker, and its file is docker-compose.yaml, which the COMPOSE_STACKS glob does not even match). Touching dtek_notif/docker-compose.yaml matches dtek_notif/* in the build's changed-service detector, so pushing this builds it and publishes :prod for it too. Verified the other six resolve at :prod in the registry. |
||
|
|
892790822d |
fix(crowdsec): stop the 403 loop that banned our own VPS and runner
Chasing why apply-k8s kept dying mid-run turned up a self-inflicted
ban loop. 585 of 586 LePresidente/http-generic-403-bf alerts in the LAPI
came from 193.181.211.79 - our own VPS - POSTing
/management.ManagementService/GetServerKey, i.e. the NetBird client's
own management call. netbird-server had never seen a single one of them,
so the 403 was not NetBird's: the bouncer was rejecting the request
before it got there. A banned peer keeps retrying, each retry is another
403, and the scenario turns five 403s in ten seconds into a 4h ban, so
the loop kept re-arming the ban it was serving. The hourly janitor step
that deleted those decisions hourly was masking all of it.
* netbird/k8s/ingress.yaml - drop the bouncer from the mesh API routes
(gRPC-gateway management, signal, relay, /api, /oauth2). A ban there
locks a peer out of the network it needs to reach anything else, and
those endpoints authenticate by NetBird token, not by a login form.
The dashboard keeps the bouncer; it is a real login surface.
netbird-local was already exempt, so this makes prod match.
* crowdsec-middleware.yaml - CrowdsecMode stream instead of live. live
blocked on GET /v1/decisions per request, so a burst saturated the
LAPI and the plugin 403'd IPs that were never banned. v1.3.3 ignores
UpdateMaxFailure in live, so fail-open is only reachable in stream;
-1 now means an unreachable LAPI degrades to "no protection" rather
than "every site 403". 15s poll instead of the 60s default, because
the runner shares one public IP with the house.
Also corrects the key name: HTTPTimeoutSeconds, not
CrowdsecLapiTimeout, which never existed and was being silently
dropped, leaving the 10s default. Back at 10, not the 2s
|
||
|
|
f7cd75d65e |
fix(crowdsec): stop the bouncer from 403-ing our own deploys
ci / lint-compose (push) Successful in 4s
ci / lint-actionlint (push) Successful in 1s
ci / lint-shellcheck (push) Successful in 2s
ci / lint-prettier (push) Successful in 3s
ci / validate (push) Successful in 3s
renovate-ci / validate-renovate (push) Successful in 32s
ci / lint-ruff (push) Successful in 0s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 2s
ci / scan-deps (push) Successful in 17s
ci / test-backend (push) Successful in 7s
ci / test-frontend (push) Successful in 10s
ci / build (push) Successful in 1s
Every bouncer-protected request blocks on a synchronous GET /v1/decisions against the LAPI, and the plugin fails CLOSED with 403 when that lookup exceeds its timeout. The LAPI was capped at 400m/500Mi on a single replica, so idle lookups measured 1.3-7.4s and a deploy burst pushed them past the fork's implicit 10s default: gitea answered 403 for ten seconds straight, and containerd turned those 403s on gcr.forust.xyz into ErrImagePull/ImagePullBackOff on freshly rolled pods. Three changes, plus the 403 feedback loop that made it sticky: * gitea/k8s/ingress.yaml - drop the bouncer from the /v2 registry route. A deploy fires hundreds of parallel authenticated OCI requests (runner Action API, manifest inspect per own image, containerd pulls, smoke probes) and scanners gain nothing from a registry that already does its own token auth. The web UI route keeps the bouncer. * crowdsec-values.yaml - LAPI to 1500m/1Gi. Replicas stay at 1 on purpose: LAPI is stateful (BoltDB plus credentials on two RWO PVCs), so a second replica would corrupt the decision store. * crowdsec-middleware.yaml - CrowdsecLapiTimeout: 2s instead of the implicit 10s, so a slow LAPI costs a fast 403 rather than a 10s hang. LePresidente/http-generic-403-bf then banned us for our own 403s: five POSTs answered 403 within 10s earn a 4h ban, and the hairpin-NAT address 192.168.88.1 that the Gitea Actions runner presents to Traefik is not covered by the home-dynamic-IP whitelist. That scenario cannot be dropped per-scenario - it is baked into the hub item crowdsecurity/http-generic-bf v0.9, and disabling the whole base-http-scenarios collection would cost ~40 useful detections. So the janitor now deletes its decisions hourly and a new forust/lan whitelist postoverflow covers 192.168.88.0/24. First janitor run removed 165 decisions; none of the remaining ones are local. Measured after: gitea 200 in 25-148ms (was 403 at 10001ms), a 60-way parallel burst all 200 with a 422ms max, gcr /v2 back to its 401 auth challenge, LAPI at 60m CPU with no throttling. |
||
|
|
6a9a460769 |
fix(deploy): let a pinning failure explain itself
ci / lint-compose (push) Successful in 3s
ci / lint-actionlint (push) Successful in 1s
ci / lint-shellcheck (push) Successful in 2s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 3s
ci / lint-dockerfiles (push) Successful in 2s
ci / scan-deps (push) Successful in 15s
ci / test-backend (push) Successful in 6s
ci / test-frontend (push) Successful in 10s
ci / validate (push) Successful in 2s
renovate-ci / validate-renovate (push) Successful in 26s
ci / build (push) Successful in 1s
registry_digest was written to return an empty string for a ref the registry does not have, so render_pinned could print "cannot resolve <ref>" and stop. It could not do that. Every caller runs under set -euo pipefail, pipefail reports the rightmost non-zero stage, and the failed docker manifest inspect made the assignment itself fail, which set -e turns into an immediate exit. render_pinned therefore died silently on the first unresolvable ref: nothing on stderr, nothing on stdout, exit 1. The apply loop piped that empty stream into kubectl, so the whole deploy stopped with "error: no objects passed to apply" - kubectl guessing at a cause, with the actual reason nowhere in the log. The missing message is the reason the |
||
|
|
ac0f845da6 |
fix(deploy): stop the panel from pointing at a tag the build dropped
ci / lint-compose (push) Successful in 3s
ci / lint-actionlint (push) Successful in 1s
ci / lint-shellcheck (push) Successful in 2s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 1s
ci / scan-deps (push) Successful in 15s
ci / test-backend (push) Successful in 7s
ci / test-frontend (push) Successful in 11s
renovate-ci / validate-renovate (push) Successful in 12s
ci / validate (push) Successful in 3s
ci / build (push) Successful in 37s
USERBOT_IMAGE named userbot:latest, but the build stopped pushing latest when images became :prod. render_pinned resolves every gcr.forust.xyz reference it finds in a manifest, not just the ones it rewrites, so that env var made the whole apply abort with "cannot resolve userbot:latest; applying nothing". Naming it :prod also makes the build job rebuild userbot, which is what restores forust/userbot:prod - the tag is listed in the registry but its manifest 404s, so the image: field in userbots.yaml had nothing to resolve either. Found by resolving every digest apply-k8s will need before spending a run on it: 6 of 8 resolved, and both failures were userbot. |
||
|
|
f54589a05c |
fix(deploy): put the snapshot where the deploy user can write it
ci / lint-compose (push) Successful in 4s
ci / lint-actionlint (push) Successful in 1s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 2s
ci / lint-shellcheck (push) Successful in 2s
ci / lint-prettier (push) Successful in 3s
ci / test-backend (push) Successful in 7s
ci / lint-dockerfiles (push) Successful in 3s
ci / scan-deps (push) Successful in 15s
ci / test-frontend (push) Successful in 11s
ci / validate (push) Successful in 7s
renovate-ci / validate-renovate (push) Successful in 29s
ci / build (push) Successful in 1s
The first deploy to actually run died on its very first action, and the error the other job reported was only the consequence. DEPLOY_SNAPSHOT_DIR defaulted to /var/backups/homelab-deploy. The deploy is unprivileged, and this Arch host has no /var/backups at all, so snapshot_dir's mkdir -p had to create it under root-owned /var and got Permission denied. It refused to go on, which is exactly what the guard is for, so no workload was touched - but the verify job then found no pointer and could only say to go look by hand. Defaulting to the deploy user's own XDG state directory fixes it with no root and no setup step, and keeps the guard: an unwritable snapshot dir still stops the deploy before the first apply. ssh-run.sh now forwards DEPLOY_SNAPSHOT_DIR too, so the path is overridable without editing the library. Verified on the workstation as the unprivileged user: pointer published, commit recorded, 71 workload generations and three helm releases captured, and the stale-pointer refusal still works. |
||
|
|
f49d91b63d |
Update deploy-lib.sh
ci / lint-compose (push) Successful in 3s
ci / lint-actionlint (push) Successful in 2s
ci / lint-shellcheck (push) Successful in 2s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 3s
ci / lint-dockerfiles (push) Successful in 2s
ci / scan-deps (push) Successful in 17s
ci / test-backend (push) Successful in 7s
ci / test-frontend (push) Successful in 10s
ci / validate (push) Successful in 2s
renovate-ci / validate-renovate (push) Successful in 13s
ci / build (push) Successful in 0s
|
||
|
|
aafa74b70a |
chore: give local verification tooling a gitignored home
ci / lint-compose (push) Successful in 3s
ci / lint-actionlint (push) Successful in 1s
ci / lint-shellcheck (push) Successful in 2s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 2s
ci / scan-deps (push) Successful in 17s
ci / test-backend (push) Successful in 10s
ci / test-frontend (push) Successful in 11s
ci / validate (push) Successful in 3s
renovate-ci / validate-renovate (push) Successful in 25s
ci / build (push) Successful in 1s
The pinned CI tools and the pre-push verification script were living in a temp directory that did not survive. tmp/ is now ignored, so ruff (which respects .gitignore by default) and the git ls-files globs the lint jobs use both skip it without needing to be told. |
||
|
|
4f74fe1778 |
ci: stop inheriting the runner's python and node
The runner executes jobs on the host rather than in a container, so a
workflow that says 'python3' or 'npm' is really saying 'whatever this
machine happens to have today'. Both of the test jobs added in
|