Compare commits

..
Author SHA1 Message Date
renovate-bot e451c97dfc chore(deps): update container patch updates
renovate-ci / validate-renovate (push) Skipped
ci / lint-actionlint (push) Successful in 1s
ci / lint-shellcheck (push) Successful in 2s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 2s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 1s
ci / scan-deps (push) Successful in 15s
ci / lint-actionlint (pull_request) Successful in 1s
ci / lint-compose (push) Successful in 3s
ci / test-backend (push) Successful in 7s
ci / test-frontend (push) Successful in 11s
ci / validate (push) Successful in 2s
ci / build (push) Skipped
ci / lint-compose (pull_request) Successful in 3s
ci / lint-shellcheck (pull_request) Successful in 2s
ci / lint-prettier (pull_request) Successful in 2s
ci / lint-ruff (pull_request) Successful in 1s
ci / lint-yaml (pull_request) Successful in 3s
ci / lint-dockerfiles (pull_request) Successful in 2s
ci / scan-deps (pull_request) Successful in 14s
ci / test-backend (pull_request) Successful in 7s
ci / test-frontend (pull_request) Successful in 11s
ci / validate (pull_request) Successful in 2s
ci / build (pull_request) Skipped
renovate-ci / validate-renovate (pull_request) Successful in 20s
2026-09-28 19:56:45 +00:00
forust 33c54ac830 Merge pull request 'chore(deps): update docker.io/valkey/valkey:9 docker digest to 418652c' (#59) from renovate/docker.io-valkey-valkey-9 into main
ci / lint-compose (push) Successful in 4s
ci / lint-actionlint (push) Successful in 1s
ci / lint-shellcheck (push) Successful in 2s
ci / lint-prettier (push) Successful in 3s
ci / lint-yaml (push) Successful in 2s
ci / lint-ruff (push) Successful in 3s
ci / lint-dockerfiles (push) Successful in 1s
ci / scan-deps (push) Successful in 19s
ci / test-backend (push) Successful in 8s
ci / test-frontend (push) Successful in 10s
ci / validate (push) Successful in 3s
renovate-ci / validate-renovate (push) Successful in 14s
ci / build (push) Successful in 11s
Reviewed-on: #59
2026-09-28 19:56:27 +00:00
renovate-bot b09d718310 chore(deps): update docker.io/valkey/valkey:9 docker digest to 418652c
renovate-ci / validate-renovate (push) Skipped
ci / lint-compose (pull_request) Successful in 6s
ci / lint-actionlint (pull_request) Successful in 2s
ci / lint-shellcheck (pull_request) Successful in 2s
ci / lint-prettier (pull_request) Successful in 3s
ci / lint-ruff (pull_request) Successful in 2s
ci / lint-yaml (pull_request) Successful in 2s
ci / lint-dockerfiles (pull_request) Successful in 3s
ci / scan-deps (pull_request) Successful in 55s
ci / test-backend (pull_request) Successful in 9s
ci / test-frontend (pull_request) Successful in 11s
ci / validate (pull_request) Successful in 3s
ci / build (pull_request) Skipped
ci / lint-compose (push) Successful in 3s
ci / lint-actionlint (push) Successful in 1s
ci / lint-shellcheck (push) Successful in 2s
ci / lint-prettier (push) Successful in 2s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 2s
ci / scan-deps (push) Successful in 14s
ci / test-backend (push) Successful in 8s
ci / test-frontend (push) Successful in 11s
ci / validate (push) Successful in 1s
ci / build (push) Skipped
renovate-ci / validate-renovate (pull_request) Successful in 53s
2026-09-28 16:20:21 +00:00
forust b4f76373bb fix(gitea): serve issue search from postgres instead of reindexing on boot
renovate-ci / validate-renovate (push) Successful in 5s
ci / lint-compose (push) Successful in 4s
ci / lint-actionlint (push) Successful in 1s
ci / lint-shellcheck (push) Successful in 3s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 2s
ci / lint-yaml (push) Successful in 3s
ci / lint-dockerfiles (push) Successful in 2s
ci / scan-deps (push) Successful in 17s
ci / test-backend (push) Successful in 9s
ci / test-frontend (push) Successful in 13s
ci / validate (push) Successful in 2s
ci / build (push) Successful in 7s
cron.rebuild_issue_indexer runs at start, so every gitea pod restart reindexed the whole issue index and read ~7MB/s off the rotational disk for an hour.
2026-09-28 17:01:36 +02:00
forust 74adf38d63 feat(alerts): cover OOM kills, restart loops and evictions
The OOMKilled container behind the immich crash loop was invisible: PodCrashLooping only fires once kubelet has already given up and started the backoff.
2026-09-28 17:01:36 +02:00
forust f5b2f89f38 style(immich): match repo prettier quoting in compose file
ci / lint-compose (push) Successful in 4s
ci / lint-actionlint (push) Successful in 2s
ci / lint-shellcheck (push) Successful in 2s
renovate-ci / validate-renovate (push) Successful in 9s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 2s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 2s
ci / scan-deps (push) Successful in 16s
ci / test-backend (push) Successful in 9s
ci / test-frontend (push) Successful in 13s
ci / validate (push) Successful in 3s
ci / build (push) Successful in 9s
2026-09-28 17:00:41 +02:00
forust 8e63284240 feat(immich): add self-hosted photo backup with dedicated postgres
ci / lint-compose (push) Successful in 4s
ci / lint-actionlint (push) Successful in 1s
ci / lint-shellcheck (push) Successful in 3s
ci / lint-prettier (push) Failing after 4s
ci / lint-ruff (push) Successful in 3s
ci / lint-yaml (push) Successful in 3s
ci / lint-dockerfiles (push) Successful in 2s
ci / scan-deps (push) Successful in 20s
ci / test-backend (push) Successful in 8s
ci / validate (push) Successful in 3s
ci / build (push) Skipped
renovate-ci / validate-renovate (push) Successful in 11s
ci / test-frontend (push) Successful in 14s
Server x2, machine learning, valkey and VectorChord postgres on local storage, Traefik routes for external and internal access.
2026-09-28 16:42:05 +02:00
forust 1d81410cd8 fix(traefik): persist plugin storage on a PVC
ci / lint-compose (push) Successful in 4s
ci / lint-actionlint (push) Successful in 3s
ci / lint-shellcheck (push) Successful in 4s
ci / lint-prettier (push) Successful in 5s
ci / lint-ruff (push) Successful in 3s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 3s
ci / scan-deps (push) Successful in 18s
ci / test-backend (push) Successful in 8s
ci / test-frontend (push) Successful in 11s
ci / validate (push) Successful in 4s
renovate-ci / validate-renovate (push) Successful in 1m28s
ci / build (push) Successful in 37s
Mount traefik-plugins PVC at /plugins-storage instead of the chart default emptyDir, so the crowdsec-bouncer download survives node reboots. Without this Traefik starts before the network is ready, the download from plugins.traefik.io times out, plugins get disabled and every route behind the middleware returns 404/503 until a manual restart.
2026-09-28 13:36:46 +02:00
forust 2f891a5d31 fix(postgres): give probes room on an I/O-bound single node
renovate-ci / validate-renovate (push) Canceled after 21s
ci / lint-compose (push) Successful in 8s
ci / lint-actionlint (push) Successful in 3s
ci / lint-shellcheck (push) Successful in 4s
ci / lint-prettier (push) Successful in 4s
ci / lint-ruff (push) Successful in 10s
ci / lint-yaml (push) Successful in 4s
ci / lint-dockerfiles (push) Successful in 4s
ci / scan-deps (push) Successful in 19s
ci / test-backend (push) Successful in 9s
ci / test-frontend (push) Successful in 21s
ci / validate (push) Successful in 5s
ci / build (push) Successful in 22s
pg_isready with the 1s default times out under I/O stall and kubelet kills a healthy postgres mid-recovery; each kill restarts a multi-minute fsync from zero and loops forever. readiness/liveness timeout 5s, liveness threshold 5, startup budget 15min.
2026-09-28 12:35:16 +02:00
forust cda0022d81 fix(renovate): reap finished job pods with ttlSecondsAfterFinished
ci / lint-compose (push) Successful in 5s
ci / lint-actionlint (push) Successful in 2s
ci / lint-shellcheck (push) Successful in 2s
ci / lint-prettier (push) Successful in 4s
ci / lint-ruff (push) Successful in 2s
ci / lint-yaml (push) Successful in 3s
ci / lint-dockerfiles (push) Successful in 3s
ci / scan-deps (push) Successful in 17s
ci / test-backend (push) Successful in 7s
ci / test-frontend (push) Successful in 12s
ci / validate (push) Successful in 5s
renovate-ci / validate-renovate (push) Successful in 9m17s
ci / build (push) Failing after 28m20s
History limits never delete manual 'create job --from' runs, so Failed pods accumulated for a week. Keep a day for debugging, reap the rest.
2026-09-28 12:03:26 +02:00
forust dde6b1c743 fix(deploy): recover helm releases from pending-* and skip helm-owned rollbacks
An --atomic upgrade whose own rollback never finishes leaves the release in pending-*, blocking every future run until a human rolls back (loki rev 18/21). Recover automatically before and after each upgrade, and fail loud when recovery does not land on deployed. Also skip helm-managed workloads in rollback_workloads: rollout undo there would step back to the revision --atomic just escaped.
2026-09-28 12:03:26 +02:00
forust 7c4843c88c fix(loki): unblock gateway rollout on a single node
Chart default is required podAntiAffinity on hostname plus RollingUpdate 25%/25%, which is maxUnavailable=0 at replicas=1: the new pod stays Unschedulable while the old one lives, and the old one never leaves while the new one is not Ready. Null the affinity (an empty map deep-merges with the default and keeps the rule) and set maxSurge/maxUnavailable to 1.
2026-09-28 11:33:38 +02:00
forust f9e4623ade fix(k8s): size the remaining workloads against measured use
Finishes the sizing pass over every workload the deploy actually manages. Each
request is at or above the container's p95 over the last seven days, so nothing
is sized below what it is known to use, and each limit is between 1.6x and 5x
the observed max, which is the figure that decides whether a burst gets an
OOMKill.

Some of these go up, and that is the point. adguard was holding 975M against a
500Mi request and netbox 962M against 512Mi, so both sat permanently above
their own request and were standing eviction candidates on a node that has
about 300M of headroom. Raising a request costs scheduler room; leaving it low
costs the pod its place in the queue when the node gets tight.

Others come down. loki ran with a 2Gi limit on 224M, gitea 1.5Gi on 305M, the
authentik worker 1Gi on 305M, and a tail of single-purpose pods -- redis,
glance, the two homepages, cfddns, session-keeper, the netbird dashboard, the
loki gateway -- each reserved 4x to 16x more than they have ever touched.

prometheus gets the opposite treatment: 768Mi/2560Mi, above its p95, because it
compacts its TSDB in place and that is a burst worth budgeting for rather than
throttling.

Two of these limits are close enough to the observed max to be worth watching
rather than trusting: adguard at 1.5x, and its DNS cache grows monotonically, so
the ceiling is a date, not a margin. That was true before this change too; the
pod sizing does not fix it and the cache needs bounding.

CPU limits are untouched throughout. Leaving postgres alone as well: it sits in
an uncommitted file that belongs to other work in progress.

Verified: every request is at or above p95 and every limit above the observed
max across all 74 containers, and 16/16 local gates pass.
2026-09-28 10:38:19 +02:00
forust 2ad4fa1b82 chore(portainer): stop deploying a container manager nothing routes to
Portainer had been running for 111 days with a 512Mi request and a 2Gi limit
against 50M of measured use, on a node that is short of memory. It is a UI over
the Docker socket; nothing in the repo or the cluster depends on it.

The marker goes, not the manifests. `portainer/k8s/active` is what puts these
files in the deploy's manifest set, so without it the next push leaves the
namespace alone and the manifests stay on disk for a one-command return. This
also matters for the smoke stage: that host list is built from the active
directories, so `portainer.forust.xyz` leaves it and the new router check does
not go looking for a route to a service we just retired.

In the cluster the Deployment, the Service and both IngressRoutes are deleted.
The routes go first: leaving an IngressRoute behind a deleted Service keeps a
Traefik router pointing at nothing, which answers 502 while looking perfectly
healthy to the stage that just started checking for routers.

Deliberately kept, so this is reversible rather than destructive: the namespace,
the 2Gi `portainer-data-pvc` and both Certificates stay. Re-enabling is
`git checkout HEAD~1 -- portainer/k8s/active` plus an apply, and no Let's Encrypt
quota is spent reissuing the production certificate.

`glance` still links to `portainer.forust.xyz` and that tile will now be a 404.
Left alone on purpose. The Cloudflare record is manual and cfddns only ever
creates records, so `portainer.forust.xyz` keeps resolving until it is removed
in the dashboard, same as `dockmon.forust.xyz`.

Verified: no Traefik router matches portainer any more, all 19 hosts left in the
smoke list still have a router, and 16/16 local gates pass.
2026-09-28 10:22:51 +02:00
forust 2a4f215546 fix(k8s): bring the over-reserved memory limits down to measured use
Six pods reserved far more memory than they have ever touched. uptime-kuma held
a 3Gi limit against 469M of measured p95, metube 2Gi against 72M, convertx
1.5Gi against 85M, netbird-server 1Gi against 97M, searxng 700Mi against 134M
and bentopdf 700Mi against 4M. Every one of them is a ceiling the scheduler
counts against the node while the memory sits unused.

Requests move down with the limits but never below the measured p95, so none of
these becomes an eviction candidate as a side effect of being right-sized. The
limits keep between 2.2x and 11.6x over the observed max, which is the figure
that decides whether a pod gets OOM-killed during a burst.

Net effect across the six: requests -557M, limits -4.6Gi, all of it ceiling that
was never in use. This is the first change that actually gives memory back.

CPU limits are left exactly as they were. They were not part of the sizing pass,
they are not being hit on a node sitting at 5% CPU, and removing them is a
separate decision from moving memory.

Verified: each limit is above the container's own observed max and each request
is above its p95, and 16/16 local gates pass.
2026-09-28 10:20:24 +02:00
forust 16aaeb60c1 fix(k8s): set requests and limits on the pods that shipped with neither
Eighteen containers had no memory limit at all, so nothing on the node could
bound them. Three of the values files even claimed to set resources: Helm does
not complain about a key it does not recognise, so the block sat there looking
like a limit while the pod ran unbounded.

alloy is the one that mattered. The chart reads `alloy.resources`; the file had
`controller.resources`, so the DaemonSet that tails every pod log on the node
shipped with nothing at all. `kubeStateMetrics` is the same trap in a different
shape -- that is the condition key, the values live under `kube-state-metrics` --
and `configReloader` in the alloy chart sits at the top level rather than under
`alloy`. Each one is verified by rendering the chart and reading the resources
back off the containers, because a values key that is ignored looks exactly
like one that works.

reloader turned out to be set and still wrong: 64Mi request against a measured
p95 of 73M, so the pod ran permanently above its own request and stayed a
standing eviction candidate. That is the pod that restarts every other pod, so
it is the last one that should be evicted. Raised to 96Mi.

Requests are set at p95 throughout, grafana, playwright and alloy included.
Left at the values first proposed they would have sat below their own p95 and
queued for eviction ahead of everything smaller. CPU limits are deliberately
absent: the node is I/O bound at 5% CPU, and CFS throttling would turn disk
wait into runnable-throttled, which is the failure mode that took the node down.

The prometheus and alertmanager configReloader sidecars are left open: chart
86.2.3 does not template the key, so reaching those two containers needs a
postRenderer.

Verified: all four charts render with the resources landing on the intended
containers, and 16/16 local gates pass.
2026-09-28 10:14:52 +02:00
forust a5409edbf2 fix(deploy): fail the smoke stage when Traefik has no route for a host
ci / lint-compose (push) Successful in 5s
ci / lint-actionlint (push) Successful in 2s
ci / lint-shellcheck (push) Successful in 3s
ci / lint-prettier (push) Successful in 2s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 3s
ci / lint-dockerfiles (push) Successful in 3s
ci / scan-deps (push) Successful in 15s
renovate-ci / validate-renovate (push) Successful in 1m15s
ci / test-frontend (push) Successful in 14s
ci / validate (push) Successful in 4s
ci / test-backend (push) Failing after 13m28s
ci / build (push) Skipped
The smoke stage treats any HTTP response as proof the service is serving,
which is right -- a 302 to a login or a 404 from a path the app does not
serve still means the chain is intact. But a 404 is not evidence of that on
its own: a router Traefik refused to build answers with exactly the same
404 and nothing behind it.

That is not hypothetical. The crowdsec bouncer is a plugin, and when Traefik
cannot fetch it at startup it disables the plugin without failing, then drops
every router whose chain referenced it. Sixteen routes answered 404 and the
stage printed `ok` for all sixteen, because a dropped router and an unserved
path are indistinguishable from outside.

The Kubernetes objects cannot tell us either: the IngressRoute is still
sitting there looking healthy, the router Traefik built from it is simply not
there. So ask Traefik. api.insecure is already on for the internal entrypoint
and the router list says which hosts it matches right now.

Every probed host has to appear in that list. HTTP routers only -- the TCP
ones match on a HostSNI wildcard and the UDP ones carry no rule at all, both
selected by entrypoint and port, so neither can answer the question. A router
mid-rollout is legitimately absent for a moment, so the list is re-read twice
over 20s; a plugin that failed to load stays absent and waiting cannot rescue
it. An unreadable router list fails the stage rather than skipping the check,
since a check that cannot run is not a passing check.

Verified against the live cluster: all 23 public routes have a router and the
stage passes. With gitea, grafana and uptime removed from that list the
probes still answer and the stage fails on exactly those three.
2026-09-28 09:13:24 +02:00
forust c70d2db3a1 ci: deploy the image the commit built, not whatever the tag points at
Every service tracked the mutable `:prod` tag, so a deploy applied whatever
that tag happened to name at the time rather than the commit it was
deploying. A rollback had no way to state what it was rolling back to, and
two deploys of one commit could land different images.

CI now publishes an immutable `sha-<commit12>` tag beside `:prod` on main,
and re-tags it for every image a push did not rebuild. That re-tag copies
the manifest list, so no layer moves. The deploy resolves the immutable tag
to a digest and pins the workload to it, and only falls back to the moving
tag when the immutable one cannot be resolved -- which it says out loud,
because that fallback is the deploy quietly ceasing to be reproducible from
its own commit.

The image list comes out of the tree with git grep rather than being written
out a second time, so adding a service no longer means keeping two lists in
step.

build also gains the three jobs it was skipping -- scan-deps, test-backend,
test-frontend -- so a change that breaks them cannot be tagged at all. The
two run blocks where a mid-loop failure was survivable now run under
set -euo pipefail: the build loop and the service detector both carried on
past an error and could report a green build having produced nothing.

The registry password moves from run: substitution into an env: block. A
quote, a backtick or a $(...) in the password is parsed as shell before the
command ever runs, and a login that failed that way looked exactly like a
build that failed.

The apply and verify timeouts stay at 45 and 30 minutes. The comments now
record the arithmetic that says so rather than leaving the numbers to be
raised on the next scare: three no-op helm upgrades run 3-5 minutes, one
broken release is a single 10 minute rollback because the loop aborts on
the first failure, and the apply loop itself is about a minute. That is
roughly 15 minutes of work against a 45 minute budget. verify is 32
workloads at 8 wide -- four waves of 300 seconds, 20 minutes -- which
leaves room for two serial rollbacks, and only becomes derivable at 45 once
rollback_workloads is parallelised.
2026-09-28 09:12:01 +02:00
forust 24dd82e801 fix(crowdsec): make the bouncer trust the mobile range independently
The parser-stage whitelist already covers 84.245.64.0/18, so an event
from the phone is dropped before it reaches a bucket and no decision is
ever created for it - confirmed against 72h of traefik access logs, where
the phone shows up as 84.245.120.147, inside that /18. But that left the
bouncer's own ClientTrustedIPs without the range, so the guarantee rested
on a single config. If the parser whitelist ever stops matching, a ban
would be created and then served against the phone, which is the one
thing that must not happen: the address belongs to a carrier, so it comes
back to us by rotation and a 4h ban is not survivable from the device.

ClientTrustedIPs bypasses the bouncer and the decision cache entirely, so
repeating the range there holds even if a decision exists for any reason.
All nine parser-stage ranges are now mirrored in the bouncer, and the
bouncer has no range the parser stage does not know about.

The file header now records that this middleware must be applied together
with a traefik restart. Applying it alone wedges the plugin: in stream
mode handleStreamTicker runs over package-level globals that no
reconfiguration stops, so every route referencing the middleware answers
404 with 'invalid middleware crowdsec-crowdsec-bouncer@kubernetescrd'
until the pod is replaced. Re-applying the prior config does not recover
it and the config is not the cause - NewChecker is a plain net.ParseCIDR
and cannot fail on a valid range. That cost 21 routes down before the
restart requirement was found; recovery is a pod replace, ~35s.

Verified live: middleware applied, traefik restarted, 17 of 20 hosts
serving (the three exceptions are unchanged and unrelated - searxng
returns its own 429, checkmk is down with 503, and one host is
local-only), zero invalid-middleware errors, and the LAPI still shows
/v1/decisions/stream polls at the 15s interval.
2026-09-27 18:52:27 +02:00
forust 11e92fdf4e fix(deploy): bound ssh hangs and retry the stage on transport loss
A connection that died silently used to hang until the job timeout, and
the stage was never re-run. One flaky TCP session cost a whole
45-minute apply, and the symptom - a job that stops mid-output with no
error - is what made the last few deploy failures expensive to read.

ServerAliveInterval/CountMax cap how long a dead peer goes unnoticed at
~60s, ConnectTimeout caps setup. Only exit 255 - ssh's own transport
failures - is retried, up to three attempts with a growing gap. A stage
that fails on its own merits exits with the remote's status, so a real
failure surfaces its own log immediately instead of being repeated three
times over 45 minutes. The stages are declarative applies, so re-running
one that had already committed is harmless.

The stage environment now goes through `env` as separate argv entries
rather than one interpolated string, so nothing in REPO, DEPLOY_SHA or
DEPLOY_SNAPSHOT_DIR is re-split by the remote shell.

Verified against a stubbed ssh: clean run attempts once, a single
transport failure recovers on attempt 2 and exits 0, three failures give
up preserving 255, and a stage failing with 1 or 7 attempts once and
passes the code through unchanged.

Also records why USERBOT_IMAGE stays on the prod tag: render_pinned
rewrites only plain `image:` lines, and this ref is what the panel
injects into the per-instance Deployments it creates, so those instances
track the tag rather than the panel's own resolved digest. The two
panel-created instances currently in the cluster are digest-pinned, so
the panel does accept one either way; the tag is the choice, not a
limitation.
2026-09-27 18:26:25 +02:00
forust b1f98fc148 Revert "fix(compose): stop pointing at the tag the build dropped" for dtek_notif
dtek-notif is being picked up again, so leave its compose alone. It is
also the one image not rebuilt since the build dropped :latest, so
:prod does not exist for it yet - unlike the other six, which resolve.
Keeping this file out of the change also keeps dtek_notif/* out of the
build's changed-service detector, so pushing does not build an image
nobody asked for.
2026-09-27 18:17:32 +02:00
forust 9b91b5847e fix(compose): stop pointing at the tag the build dropped
These five stacks still asked for :latest, but the build stopped pushing
it - on main it only pushes main and prod, on dev only dev. Every one of
these images therefore resolved only because the registry still had a
stale :latest from before that change, and the next time one of them was
actually built the reference would have dangled.

Named for dtek-notif: of the seven, only that one still resolved at
:prod, because it is the only image not rebuilt since the build dropped
:latest - and it is the only one of these five stacks the deploy does not
manage (no `active` marker, and its file is docker-compose.yaml, which
the COMPOSE_STACKS glob does not even match). Touching
dtek_notif/docker-compose.yaml matches dtek_notif/* in the build's
changed-service detector, so pushing this builds it and publishes
:prod for it too.

Verified the other six resolve at :prod in the registry.
2026-09-27 18:11:45 +02:00
forust 892790822d fix(crowdsec): stop the 403 loop that banned our own VPS and runner
Chasing why apply-k8s kept dying mid-run turned up a self-inflicted
ban loop. 585 of 586 LePresidente/http-generic-403-bf alerts in the LAPI
came from 193.181.211.79 - our own VPS - POSTing
/management.ManagementService/GetServerKey, i.e. the NetBird client's
own management call. netbird-server had never seen a single one of them,
so the 403 was not NetBird's: the bouncer was rejecting the request
before it got there. A banned peer keeps retrying, each retry is another
403, and the scenario turns five 403s in ten seconds into a 4h ban, so
the loop kept re-arming the ban it was serving. The hourly janitor step
that deleted those decisions hourly was masking all of it.

* netbird/k8s/ingress.yaml - drop the bouncer from the mesh API routes
  (gRPC-gateway management, signal, relay, /api, /oauth2). A ban there
  locks a peer out of the network it needs to reach anything else, and
  those endpoints authenticate by NetBird token, not by a login form.
  The dashboard keeps the bouncer; it is a real login surface.
  netbird-local was already exempt, so this makes prod match.

* crowdsec-middleware.yaml - CrowdsecMode stream instead of live. live
  blocked on GET /v1/decisions per request, so a burst saturated the
  LAPI and the plugin 403'd IPs that were never banned. v1.3.3 ignores
  UpdateMaxFailure in live, so fail-open is only reachable in stream;
  -1 now means an unreachable LAPI degrades to "no protection" rather
  than "every site 403". 15s poll instead of the 60s default, because
  the runner shares one public IP with the house.

  Also corrects the key name: HTTPTimeoutSeconds, not
  CrowdsecLapiTimeout, which never existed and was being silently
  dropped, leaving the 10s default. Back at 10, not the 2s f7cd75d
  guessed - a pull that times out leaves the ban cache frozen at its
  startup contents, so new bans would silently never apply.

* crowdsec-values.yaml - CIDR allowlisting moves to parsers/s02-enrich,
  where CrowdSec's docs put it: a parser whitelist drops the event before
  it reaches a bucket, so those addresses never become a decision at
  all. The old postoverflow LAN list was checked only after the ban
  existed, which is the window the deploy kept landing in. Added
  100.64.0.0/10, which the RFC 1918 blocks miss and where the
  workstation, the k0s node and the VPS actually live. The DDNS
  home-IP whitelist stays in postoverflows, because resolving a hostname
  is the expensive check the docs reserve that stage for.

  ClientTrustedIPs mirrors that list so the bouncer skips the LAPI
  round-trip entirely for those addresses.

* janitor-cronjob.yaml - drop step 5. The bouncer can no longer
  manufacture 403s, so the only remaining firings of that scenario are
  real scanners, and deleting their decisions hourly was undoing a
  working ban.

* Also lands the LAPI config.yaml.local (SQLite WAL, Central API off,
  bounded flush) that f7cd75d's comments referenced but never included:
  the LAPI was blocked in fsync on its rollback journal, and the CAPI
  resolver held a write transaction while timing out against a host
  this network cannot reach.

Verified in-cluster: no new 403-bf alerts in the 3.5min after applying,
the management endpoint answers 404 from the backend instead of 403 from
the bouncer in 0.17s, and the VPS client reports Management and Signal
connected with 2/2 relays.
2026-09-27 18:09:07 +02:00
forust f7cd75d65e fix(crowdsec): stop the bouncer from 403-ing our own deploys
ci / lint-compose (push) Successful in 4s
ci / lint-actionlint (push) Successful in 1s
ci / lint-shellcheck (push) Successful in 2s
ci / lint-prettier (push) Successful in 3s
ci / validate (push) Successful in 3s
renovate-ci / validate-renovate (push) Successful in 32s
ci / lint-ruff (push) Successful in 0s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 2s
ci / scan-deps (push) Successful in 17s
ci / test-backend (push) Successful in 7s
ci / test-frontend (push) Successful in 10s
ci / build (push) Successful in 1s
Every bouncer-protected request blocks on a synchronous
GET /v1/decisions against the LAPI, and the plugin fails CLOSED with 403
when that lookup exceeds its timeout. The LAPI was capped at 400m/500Mi
on a single replica, so idle lookups measured 1.3-7.4s and a deploy
burst pushed them past the fork's implicit 10s default: gitea answered
403 for ten seconds straight, and containerd turned those 403s on
gcr.forust.xyz into ErrImagePull/ImagePullBackOff on freshly rolled
pods.

Three changes, plus the 403 feedback loop that made it sticky:

* gitea/k8s/ingress.yaml - drop the bouncer from the /v2 registry route.
  A deploy fires hundreds of parallel authenticated OCI requests
  (runner Action API, manifest inspect per own image, containerd pulls,
  smoke probes) and scanners gain nothing from a registry that already
  does its own token auth. The web UI route keeps the bouncer.
* crowdsec-values.yaml - LAPI to 1500m/1Gi. Replicas stay at 1 on
  purpose: LAPI is stateful (BoltDB plus credentials on two RWO PVCs),
  so a second replica would corrupt the decision store.
* crowdsec-middleware.yaml - CrowdsecLapiTimeout: 2s instead of the
  implicit 10s, so a slow LAPI costs a fast 403 rather than a 10s hang.

LePresidente/http-generic-403-bf then banned us for our own 403s: five
POSTs answered 403 within 10s earn a 4h ban, and the hairpin-NAT
address 192.168.88.1 that the Gitea Actions runner presents to Traefik
is not covered by the home-dynamic-IP whitelist. That scenario cannot
be dropped per-scenario - it is baked into the hub item
crowdsecurity/http-generic-bf v0.9, and disabling the whole
base-http-scenarios collection would cost ~40 useful detections. So the
janitor now deletes its decisions hourly and a new forust/lan
whitelist postoverflow covers 192.168.88.0/24. First janitor run
removed 165 decisions; none of the remaining ones are local.

Measured after: gitea 200 in 25-148ms (was 403 at 10001ms), a 60-way
parallel burst all 200 with a 422ms max, gcr /v2 back to its 401 auth
challenge, LAPI at 60m CPU with no throttling.
2026-09-27 17:08:22 +02:00
forust 6a9a460769 fix(deploy): let a pinning failure explain itself
ci / lint-compose (push) Successful in 3s
ci / lint-actionlint (push) Successful in 1s
ci / lint-shellcheck (push) Successful in 2s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 3s
ci / lint-dockerfiles (push) Successful in 2s
ci / scan-deps (push) Successful in 15s
ci / test-backend (push) Successful in 6s
ci / test-frontend (push) Successful in 10s
ci / validate (push) Successful in 2s
renovate-ci / validate-renovate (push) Successful in 26s
ci / build (push) Successful in 1s
registry_digest was written to return an empty string for a ref the registry
does not have, so render_pinned could print "cannot resolve <ref>" and stop.
It could not do that. Every caller runs under set -euo pipefail, pipefail
reports the rightmost non-zero stage, and the failed docker manifest inspect
made the assignment itself fail, which set -e turns into an immediate exit.

render_pinned therefore died silently on the first unresolvable ref: nothing on
stderr, nothing on stdout, exit 1. The apply loop piped that empty stream into
kubectl, so the whole deploy stopped with "error: no objects passed to apply" -
kubectl guessing at a cause, with the actual reason nowhere in the log. The
missing message is the reason the f54589a run looked like a network death.

Reproduced against the old file with a docker stub that always fails: identical
to the ac0f845 log. The || true makes the empty string reachable, and the apply
loop now names the file that failed instead of letting kubectl speak.
2026-09-27 16:33:37 +02:00
forust ac0f845da6 fix(deploy): stop the panel from pointing at a tag the build dropped
ci / lint-compose (push) Successful in 3s
ci / lint-actionlint (push) Successful in 1s
ci / lint-shellcheck (push) Successful in 2s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 1s
ci / scan-deps (push) Successful in 15s
ci / test-backend (push) Successful in 7s
ci / test-frontend (push) Successful in 11s
renovate-ci / validate-renovate (push) Successful in 12s
ci / validate (push) Successful in 3s
ci / build (push) Successful in 37s
USERBOT_IMAGE named userbot:latest, but the build stopped pushing latest when
images became :prod. render_pinned resolves every gcr.forust.xyz reference it
finds in a manifest, not just the ones it rewrites, so that env var made the
whole apply abort with "cannot resolve userbot:latest; applying nothing".

Naming it :prod also makes the build job rebuild userbot, which is what
restores forust/userbot:prod - the tag is listed in the registry but its
manifest 404s, so the image: field in userbots.yaml had nothing to resolve
either.

Found by resolving every digest apply-k8s will need before spending a run on
it: 6 of 8 resolved, and both failures were userbot.
2026-09-27 16:14:43 +02:00
forust f54589a05c fix(deploy): put the snapshot where the deploy user can write it
ci / lint-compose (push) Successful in 4s
ci / lint-actionlint (push) Successful in 1s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 2s
ci / lint-shellcheck (push) Successful in 2s
ci / lint-prettier (push) Successful in 3s
ci / test-backend (push) Successful in 7s
ci / lint-dockerfiles (push) Successful in 3s
ci / scan-deps (push) Successful in 15s
ci / test-frontend (push) Successful in 11s
ci / validate (push) Successful in 7s
renovate-ci / validate-renovate (push) Successful in 29s
ci / build (push) Successful in 1s
The first deploy to actually run died on its very first action, and the
error the other job reported was only the consequence.

DEPLOY_SNAPSHOT_DIR defaulted to /var/backups/homelab-deploy. The deploy
is unprivileged, and this Arch host has no /var/backups at all, so
snapshot_dir's mkdir -p had to create it under root-owned /var and got
Permission denied. It refused to go on, which is exactly what the guard
is for, so no workload was touched - but the verify job then found no
pointer and could only say to go look by hand.

Defaulting to the deploy user's own XDG state directory fixes it with no
root and no setup step, and keeps the guard: an unwritable snapshot dir
still stops the deploy before the first apply.

ssh-run.sh now forwards DEPLOY_SNAPSHOT_DIR too, so the path is
overridable without editing the library. Verified on the workstation as
the unprivileged user: pointer published, commit recorded, 71 workload
generations and three helm releases captured, and the stale-pointer
refusal still works.
2026-09-27 15:40:32 +02:00
forust f49d91b63d Update deploy-lib.sh
ci / lint-compose (push) Successful in 3s
ci / lint-actionlint (push) Successful in 2s
ci / lint-shellcheck (push) Successful in 2s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 3s
ci / lint-dockerfiles (push) Successful in 2s
ci / scan-deps (push) Successful in 17s
ci / test-backend (push) Successful in 7s
ci / test-frontend (push) Successful in 10s
ci / validate (push) Successful in 2s
renovate-ci / validate-renovate (push) Successful in 13s
ci / build (push) Successful in 0s
2026-09-27 12:02:24 +02:00
forust aafa74b70a chore: give local verification tooling a gitignored home
ci / lint-compose (push) Successful in 3s
ci / lint-actionlint (push) Successful in 1s
ci / lint-shellcheck (push) Successful in 2s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 2s
ci / scan-deps (push) Successful in 17s
ci / test-backend (push) Successful in 10s
ci / test-frontend (push) Successful in 11s
ci / validate (push) Successful in 3s
renovate-ci / validate-renovate (push) Successful in 25s
ci / build (push) Successful in 1s
The pinned CI tools and the pre-push verification script were living in
a temp directory that did not survive. tmp/ is now ignored, so ruff
(which respects .gitignore by default) and the git ls-files globs the
lint jobs use both skip it without needing to be told.
2026-09-27 11:43:00 +02:00
forust 4f74fe1778 ci: stop inheriting the runner's python and node
The runner executes jobs on the host rather than in a container, so a
workflow that says 'python3' or 'npm' is really saying 'whatever this
machine happens to have today'. Both of the test jobs added in c00a472
were red on the first run for exactly that reason.

uv venv with no --python takes the first interpreter it finds, which is
the host's 3.14 here. pyrogram's sync.py calls the bare
asyncio.get_event_loop() that 3.14 no longer auto-creates, so three
tests died at collection. Pinned to 3.13, which is both what uv will
fetch when the host has none and what python:3.13-slim actually builds.

npm was missing outright, and turned up an hour later as npm 12 on node
26 - the same push, minutes apart. Neither is the panel image's
node:22-alpine, so node is now installed from the official tarball the
way the other tools are, at the image's major. npm is checked by running
it rather than by looking it up, so a name that resolves to something
broken reads as not installed.

Renovate keeps NODE_VERSION in step with the Dockerfile's node: tag, and
the two are one grouped dependency: CI that tests on a different major
than it builds on is a gate that can pass over a real break.
2026-09-27 11:42:56 +02:00
forust a5d384a4d8 feat(deploy): probe every active service after a deploy, rollouts included
ci / lint-compose (push) Successful in 4s
ci / lint-actionlint (push) Successful in 1s
ci / lint-shellcheck (push) Successful in 2s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 2s
ci / scan-deps (push) Failing after 13s
ci / test-backend (push) Failing after 10s
ci / test-frontend (push) Failing after 1s
ci / validate (push) Successful in 2s
renovate-ci / validate-renovate (push) Successful in 59s
ci / build (push) Successful in 1m50s
verify-k8s watches rollouts, which reports that pods converged. It cannot tell
a converged pod from a serving one. A Service selector pointing at a port
nothing listens on, a 500 from the app itself, a Traefik route that stopped
matching, a pod that OOMKilled early enough to still count as Available for the
duration of the check -- all of those are green at the rollout level and broken
for whoever opens the URL.

So ask what users ask. A smoke stage probes the public route of every active
service and fails on a transport error, a 5xx, or a 000, which curl reports
when it exits cleanly and nothing replied. Everything else passes, including 4xx:
a 404 from a path the service does not serve and a 302 to a login both prove
Traefik matched the host, the Service resolved to a pod and the pod answered,
which is the whole claim being tested.

An empty host list is an error, not a pass. Zero names means the extraction
broke, and reporting a clean deploy off a broken grep is the failure mode this
job exists to catch.

It runs on always() and after verify-k8s rather than before it, because a
rollback is when a route most needs re-checking. It only skips when verify-k8s
did, which is when nothing was deployed at all.

Two things worth writing down, because both were wrong on the first pass:

Stripping comments before reading the routes is not optional. naio and xui are
still in the tree commented out, and a plain grep picks both up and then reports
two services as unreachable when nobody ever deployed them. The apex
forust.xyz also needs a filter that admits it, so /\.forust\.xyz$/ quietly
dropped the site root.

And the 5xx test was written as ${code%%[0-9]*} != 5, which is empty for every
three-digit code, so a 500 was reported as ok. A case glob on the leading digit
is what actually works.

23 routes answer today, in 1.6s. The 5xx branch is the one part no live service
here exercises, so it was checked by running the block over 200 through 599 and
000 rather than against a real response.
2026-09-27 10:50:07 +02:00
forustandClaude Opus 4.8 af9a22fea9 fix(converters): spell workstation right in the compose router label
The local bentopdf router matched Host(`pdf.wokstation.internal`), a hostname
that resolves to nothing. converters/k8s next to it has always had the correct
spelling, so the service is reachable in the path that actually deploys and this
one is dormant -- converters/ has no active marker, so select_manifests skips
the stack. It would only bite whoever switches the service back to Compose and
then wonders why one of the three internal routes 404s.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-09-27 10:46:52 +02:00
forustandClaude Opus 4.8 baedea504d ci: scan dependencies for new advisories and stop handing out a write token
Two things, both about not finding out late.

No workflow declared `permissions`, so all eighteen jobs across the four
workflows ran on a token with the default full repository scope. Every one of
them only checks out code, and deploy reaches the cluster over SSH with the
deploy key, and Renovate writes through its own bot PAT rather than the
Actions token. So `contents: read` is all any of them needed.

The panel image ships 15 known advisories and nothing was looking. Add a
scan-deps job that fails on anything new, and record the eight current ones by
ID in the workflow. It is a list rather than a baseline count so that the diff
that accepts an advisory says so in words, and it lives in our workflow
instead of the package manifest so a subtree sync from forust/userbot cannot
quietly widen the exemption.

Both halves were checked to fail on a regression, not just to pass today:
removing one --ignore-vuln turns the Python step red, and dropping
--audit-level to moderate turns the npm one red on the devalue advisory.

npm audits production dependencies only. All seven findings in the full tree
are build- or test-time: the esbuild advisory needs a vite dev server exposed
to the internet, and nanoid's infinite loop needs a custom generator called
with size 0, which postcss does not do. None are in the 91 kB bundle the panel
serves, so failing on them would be noise that trains people to ignore the
job.

The starlette entries are the reason the job is not "fail on everything":
fastapi 0.115.12 pins starlette<0.47.0 and the last four fixes need 0.49.1
through 1.3.1, so clearing them is a jump to fastapi 0.141.x and is upstream's
call, not a drive-by. Four of the seven are reachable in principle, which the
comment on the job sets out. The panel answers only on
userbot.workstation.internal with no public route, which is what keeps those
four from being an internet-facing DoS.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-09-27 10:46:01 +02:00
forustandClaude Opus 4.8 30995ee009 ci: pin the last four linters instead of trusting the runner
prettier, ruff, yamllint and hadolint were the only CI tools still called bare,
straight off whatever the runner happened to have installed. Pin them in
tool-versions.env like the other three and install them the same way, so the
versions Renovate moves are the versions CI runs.

Each pinned version equals what is already on the runner, so this changes what
CI does not at all today. It changes what CI does on a rebuilt runner: the
pinned one gets installed over the drift.

The four need four different mechanisms, which is why this is not one pattern:

  hadolint  a bare binary per platform, like actionlint
  ruff,
  yamllint  PyPI wheels, unpacked by uv
  prettier  an npm tarball, unpacked by tar

prettier is the awkward one. Its entry point requires ../package.json relative
to its own real path, so copying the single file out -- which is what every
other installer here does -- yields a module-not-found at the first run. It
keeps its package directory in a versioned one next to a relative symlink, and
the tarball ships bin/ without the exec bit, so that needs chmod too.

hadolint's release names one platform uname-style and the other Go-style
(x86_64 but arm64), which 404s on the first architecture if you assume
otherwise.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-09-27 10:28:42 +02:00
forustandClaude Opus 4.8 3a05d86e3e ci: run svelte-check, which was already a dependency with no script
svelte-check sits in devDependencies at ^4.7.3 and nothing in the repository
ever invoked it, so the type errors it reports had no path to a human. Point a
script at it and run it in the frontend job, and it is clean: 0 errors, 0
warnings.

It shares the one `npm ci` with the test step. A second install would have
doubled the slowest part of the job to learn exactly the same thing.

The lockfile is untouched, because scripts are not part of what it pins.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-09-27 10:24:18 +02:00
forustandClaude Opus 4.8 c00a4724f5 ci: actually run the test suites that exist in the tree
The panel ships 25 pytest tests and 2 vitest tests. Nothing executed them:
there was no job, no local dev loop, and nothing that would have noticed when
one of them rotted. They pass, and they are 8 seconds of work, which is the
argument for having them.

Both jobs mirror how the image is built rather than how a developer would run
them by hand: `npm ci` because that is what the Dockerfile does, so the tree
under test is the tree that ships, and requirements-dev.txt through uv, which
is now pinned like the other CI tools.

The backend job runs `python -m pytest`, not bare `pytest`. The tests import
`app.*` relative to the backend directory, and only the `-m` form puts the
working directory on sys.path.

ruff format --check joins ruff check in the lint job. It needed 0ae0df7 to be
addable, since ten files disagreed with the style ruff.toml has always
declared.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-09-27 10:20:54 +02:00
forustandClaude Opus 4.8 284e19ef88 ci: pin uv, the tool that builds the pytest venv
The panel backend has 25 pytest tests that no workflow has ever run. Making
them run needs a throwaway virtualenv, and uv is what builds it in seconds
against the pinned requirements-dev.txt.

Installing it through install-ci-tools.sh rather than assuming it is on the
runner keeps the version in one place, where the other three tools already
live, and where the Renovate regex manager can move it.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-09-27 10:19:27 +02:00
forustandClaude Opus 4.8 0ae0df7473 style: format the last 10 files that ruff format disagreed with
ruff.toml has declared `quote-style = "single"` and line-length 120 since the
lint job landed, and 118 of 128 files follow it. The panel backend and the
netbox configuration were written in black/prettier style instead, so a
`ruff format --check` would have failed on them from the start.

Bring them onto the style the repository already declares, which is what makes
the check adoptable at all. Formatting only: apart from quote style the diff is
multi-line expressions joined where they fit inside 120 columns.

Both suites still pass afterwards (25 pytest, and ruff check is clean).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-09-27 10:19:23 +02:00
forustandClaude Opus 4.8 fddd82704f ci: keep pull requests away from the production admission webhooks
`kubectl apply --dry-run=server` persists nothing, but it does execute the
admission webhooks of the real API server. The validate job runs on
pull_request with no branch guard, so anyone able to open a PR could run
arbitrary manifest content through cert-manager and Traefik in production.

Limit the step to pushes to main. A pull request loses nothing by it: only
main is ever deployed, and this job has to complete successfully before the
deploy workflow is allowed to start, so a bad CRD is still caught before
anything reaches the cluster -- on the push instead of on the PR.

The skip is announced rather than silent, so a missing server-side pass does
not read as a pass.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-09-27 10:15:43 +02:00
forustandClaude Opus 4.8 0691536f28 fix(deploy): roll out our images by digest instead of a moving tag
`kubectl rollout undo` restores the previous ReplicaSet's pod template
verbatim. While that template names a tag, the rollback does not roll back
the image: the tag has already moved, so the reverted pod pulls the very
build that just failed and the cluster stays broken. The safety net added
in 1505b63 therefore could not recover from a bad image.

Pin the digest at apply time. A digest is not knowable when a manifest is
written, so render_pinned resolves it on the way into the cluster and the
digest is never committed. Git keeps a readable `:prod`, Renovate keeps
seeing exactly the manifests it saw before, and the previous revision of
each workload now holds the digest that was actually serving, so undo
restores those exact bytes.

imagePullPolicy is dropped from the manifests rather than set to
IfNotPresent: a reference that is not `:latest` already defaults to it, and
that is what the Kubernetes docs ask for alongside a digest.

An unresolvable image is fatal instead of a warning, because carrying on
would quietly apply a mutable tag again.

restart_stale_images keeps its comparison but is no longer how a rebuild
reaches the cluster -- the pinned template rolls out on its own now. What
is left is a drift check for hand-run `kubectl set image`, so it matches
the container by repository: a pod's status now reports `repo@sha256:...`
while the manifest still says `:prod`.

The build job stops pushing `:latest` altogether, which removes the tag
that a dev branch could otherwise move under a prod deploy.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-09-27 10:09:06 +02:00
forust 2b9e34ba4a fix(k8s): add readiness probes so a bad image cannot look healthy
Six of eight workloads had no readinessProbe, so a pod turned Ready the
moment its process started. The verify job relies on `rollout status`, so
it passed for images that crash-looped or served errors, which left the
rollback safety net inert.

Each probe targets the path the service is actually reached on:
- homepages: / (verified 200)
- error-pages: /404.html, the path Traefik's errorPages middleware
  requests. / returns 403 by design and would never pass.
- webinar-checker: /health (verified 200). /metrics also answers, but it
  is a Prometheus endpoint, not a readiness signal.

The two userbot deployments stay without probes: they expose no port and
no session file, and the panel reaches Telegram through its own client. A
truthful signal there needs a health endpoint in the app itself.
2026-09-27 10:01:09 +02:00
forust 30d2b83efe chore(reloader): manage the reloader release from the repository
Reloader has been running since 23 September and is what makes the
reloader.stakater.com/auto annotation on a pod template do anything, but the
repository only held a namespace. It was a release someone installed by hand,
so it was invisible to review, invisible to Renovate, and one reinstall away from
being silently dropped.

Declaring it in HELM_RELEASES pins the chart version somewhere Renovate can
update it, and the active marker means the namespace is applied before the
upgrade instead of only existing as a side effect of the original install.

The values file sets nothing the running release does not already do, apart from
resource requests and limits, which the chart leaves empty.
2026-09-27 09:48:32 +02:00
forust db7bccfd89 fix(deploy): restart workloads whose image tag moved past what they run
Our manifests pin images to `:latest`, so a rebuild leaves the pod template
byte-identical. kubectl apply sees no change, creates no ReplicaSet and pulls
nothing, and the cluster keeps serving the previous build. imagePullPolicy:
Always does not help, because it only decides whether a pod that *is* starting
pulls, and no pod ever starts.

All eight workloads that consume an image from our own registry were affected.
Three of them had been running code from 23 September, and the single hardcoded
`rollout restart deployment/userbot-panel` covered one of the eight.

Restarting everything unconditionally was not the answer either: that bounces
healthy services on every deploy, error-pages included, and the brief window
where nothing answers is exactly what error-pages exists to prevent. So compare
what each workload actually runs against what the tag resolves to now, and
restart only the ones that differ. When the tag still points at the running
digest nothing happens, so a redeploy that changed no image is a no-op.

Scope is the repository, deliberately. Five more workloads run our images but
have no manifest here, and they are applied out of band. Walking the manifests
rather than the cluster means this can never reach them.

The digest is resolved for the node architecture. A multi-arch tag also carries
`unknown/unknown` entries for the build attestation, and a pod's imageID is
always the per-platform digest, so comparing the wrong entry would mark
everything stale forever.

Once a restart happens it bumps the generation, which is what makes the change
visible to changed_workloads and therefore watchable and revertible by the
verify stage.
2026-09-27 09:48:04 +02:00
forust 1505b638ce fix(deploy): verify and roll back in a separate job
verify_workloads ended on `[ -s "$failed_file" ]`, which is the opposite
of what its own contract says. A non-empty file means something failed, so
the function returned success exactly when a workload never came up, and
failure when everything was fine. Every rollback was therefore skipped,
and every deploy that changed anything ended red with an empty failure
list and a bogus "Rolled back successfully".

Worse, the check only ever ran at the end of stage_apply_k8s, inside the
same process as the apply. A job killed by timeout-minutes, cancelled by
a new push, or cut off by a dropped SSH connection never reached it, which
is precisely when a rollback matters. The three helm upgrades alone can
consume the whole 30-minute job budget, so that path was reachable.

Verification now lives in its own job, gated on always(), so it runs
whatever happened to the apply. The apply stage publishes its pre-apply
snapshot through DEPLOY_SNAPSHOT_DIR/current before touching anything,
and the verify stage picks it up from there. A snapshot whose recorded
commit does not match the deploy is refused rather than trusted, so a
stale pointer from an earlier run cannot make the rollback revert the
wrong workloads. An unwritable snapshot directory now fails the deploy up
front instead of silently continuing without a way back.

cancel-in-progress becomes false for the same reason: cancelling a run
kills the apply job and takes the verify job with it, which is the failure
this change exists to prevent. Both applies are idempotent, so queueing
costs little. The SSH key moves to a per-run directory removed on exit,
and the deploy is pinned to the exact commit CI validated.
2026-09-26 20:09:20 +02:00
forust 7ce727bc8a chore(renovate): move config under renovate/ and validate it in CI
The config lived in renovate.json at the repo root while everything else
Renovate-related sat under renovate/, and renovate/config.js was a second,
unused source of truth. Both are gone: renovate/renovate.json is now the
only config file.

Because the CronJob in the cluster cannot read the repository, its
ConfigMap carries an inlined copy of the config. That copy is generated,
and sync-renovate-configmap.sh --check now fails the build when it drifts
from the source file.

The workflows also stop carrying a copy of the renovate/renovate image
tag. They read it from renovate/k8s/cronjob.yaml, so the version validated
in CI is the version that actually runs in the cluster.

ci.yaml validates the config with renovate-config-validator, checks the
generated ConfigMap, and kubeconforms the CronJob's own manifests.
2026-09-26 20:09:13 +02:00
forust f22793e32e ci: lint workflows and shell scripts, validate k8s against the API server
Adds three lint jobs (actionlint, shellcheck, compose) and a server-side
dry-run of the active manifests. Previously the only k8s check was
kubeconform, which has no schemas for CRDs, so every IngressRoute,
Certificate, PrometheusRule and Middleware was silently skipped.

The server-side pass needs the live API server because that is the only
place the real CRD schemas and the cert-manager / Traefik admission
webhooks exist. It is scoped to services carrying a k8s/active marker,
since dry-run needs the target namespace to exist. userbot/ is excluded
from shellcheck: it is a git subtree, and linting upstream's scripts would
let a routine subtree pull turn the deploy gate red on code we do not own.

kubeconform, shellcheck and actionlint are now installed from pinned
versions in tool-versions.env rather than picked up from the runner's
PATH. The Compose helper is shared with the deploy workflow so both
check the same file set the same way.
2026-09-26 20:09:09 +02:00
forust 4a8d4feea0 fix(netbox): run probes with curl directly, not python -c
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 1s
ci / validate (push) Successful in 2s
renovate-ci / validate-renovate (push) Successful in 20s
ci / build (push) Successful in 1s
deploy / preflight (push) Successful in 2s
deploy / validate (push) Successful in 1m53s
deploy / apply-k8s (push) Successful in 1m47s
deploy / apply-compose (push) Successful in 11s
python -c 'exec /usr/bin/curl ...' is shell syntax, not Python, so every probe raised SyntaxError and the pod never became ready. Verified curl against /login/ returns HTTP 200.
2026-09-26 19:07:47 +02:00
forust 1d9a85bef9 fix(edu-master): stop 20722d false alert on zeroed last_success gauge
ci / validate (push) Successful in 2s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 1s
deploy / preflight (push) Successful in 2s
renovate-ci / validate-renovate (push) Successful in 14s
ci / build (push) Successful in 2s
deploy / validate (push) Successful in 1m52s
deploy / apply-k8s (push) Successful in 1m48s
deploy / apply-compose (push) Successful in 13s
checker.py initialises last_success to 0, so right after a pod restart
`time() - last_success` equals the current epoch. The rule compared that
against 300, went firing instantly, and humanizeDuration rendered the raw
epoch delta as ~20722d. The last_run > 0 guard did not help because a run
happens long before the first success.

Guard the duration rule on last_success > 0, keeping the duration
expression on the left of `and` so $value stays the real gap, and add a
separate WebinarCheckerNeverSucceeded rule for the zeroed-gauge case so a
checker that has never succeeded is still caught.
2026-09-26 17:50:45 +02:00
forust b625d30568 ci(deploy): cancel superseded deploys on new push
deploy / preflight (push) Successful in 1s
deploy / validate (push) Successful in 1m53s
deploy / apply-k8s (push) Successful in 1m48s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 2s
ci / lint-yaml (push) Successful in 3s
ci / lint-dockerfiles (push) Successful in 6s
ci / validate (push) Successful in 1s
renovate-ci / validate-renovate (push) Successful in 13s
ci / build (push) Successful in 1s
deploy / apply-compose (push) Successful in 12s
Queued deploy runs were deploying origin/main at start anyway (preflight reset), so waiting runs duplicated the newest deploy instead of their own commit. Cancel them.
2026-09-26 17:42:36 +02:00
forust d018a441af fix(deploy): move netbox from compose to k8s
ci / lint-dockerfiles (push) Successful in 1s
ci / validate (push) Successful in 2s
renovate-ci / validate-renovate (push) Successful in 17s
deploy / preflight (push) Successful in 2s
ci / build (push) Successful in 1s
deploy / apply-k8s (push) Canceled after 0s
deploy / apply-compose (push) Canceled after 0s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 3s
deploy / validate (push) Canceled after 4s
compose.yaml is a local-only stand on 127.0.0.1:8000 per README; the live service runs in-cluster. The root active marker made apply-compose fail on gitignored .env vars.
2026-09-26 17:38:20 +02:00
forust ecb254017d fix(searxng): use existing image tag 2026.9.25-12f8b6515
ci / validate (push) Successful in 2s
ci / build (push) Successful in 1s
deploy / apply-k8s (push) Successful in 2m23s
deploy / apply-compose (push) Failing after 8s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 1s
deploy / preflight (push) Successful in 2s
renovate-ci / validate-renovate (push) Successful in 18s
deploy / validate (push) Successful in 1m46s
2026.09.13-d4ce87c23 was never published (upstream tags month without leading zero); rollout stuck in ImagePullBackOff. Verified replacement tag exists on Docker Hub.
2026-09-26 17:37:20 +02:00
forust 41f18ea993 fix(deploy): move netbird from compose to k8s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 3s
ci / lint-dockerfiles (push) Successful in 1s
ci / validate (push) Successful in 2s
deploy / preflight (push) Successful in 2s
renovate-ci / validate-renovate (push) Successful in 19s
ci / build (push) Successful in 2s
deploy / validate (push) Successful in 1m45s
deploy / apply-k8s (push) Successful in 1m52s
deploy / apply-compose (push) Failing after 24s
netbird runs in-cluster; the root active marker made apply-compose pick up netbird/compose.yaml and fail on gitignored .env vars. Drop the compose marker and enable k8s/active instead.
2026-09-26 17:24:03 +02:00
forust 91c344fe2c fix(monitoring): disable control-plane alerts and scrapes on k0s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 1s
ci / validate (push) Successful in 3s
deploy / preflight (push) Successful in 2s
renovate-ci / validate-renovate (push) Successful in 24s
deploy / validate (push) Successful in 1m42s
deploy / apply-k8s (push) Successful in 1m45s
ci / build (push) Successful in 2s
deploy / apply-compose (push) Failing after 9s
k0s runs kube-controller-manager, kube-scheduler and etcd inside its own
process rather than as pods, so the chart's Services never get endpoints
and the targets stay permanently absent. Drop the matching ServiceMonitors
and their Down/HighCommitDurations rules; kube-proxy and kubelet do get
endpoints on k0s and stay enabled.
2026-09-26 17:16:55 +02:00
forust cab6ef2102 Merge branch 'feat/netbird'
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 1s
ci / validate (push) Successful in 2s
deploy / preflight (push) Successful in 2s
deploy / validate (push) Successful in 1m43s
renovate-ci / validate-renovate (push) Successful in 14s
ci / build (push) Successful in 1s
deploy / apply-k8s (push) Successful in 3m15s
deploy / apply-compose (push) Failing after 41s
2026-09-26 17:05:55 +02:00
forust a2ff9515a3 fix(deploy): validate compose without workstation secrets
renovate-ci / validate-renovate (push) Skipped
deploy / validate (push) Skipped
ci / lint-prettier (push) Successful in 2s
ci / validate (push) Successful in 2s
ci / build (push) Skipped
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 1s
docker compose config required real values for gitignored .env files and secrets, so validate always failed on stacks with :? guards (netbird, netbox). Validate structure only via --no-interpolate, --no-env-resolution and --no-path-resolution, keeping normalization and consistency checks.
2026-09-26 17:00:43 +02:00
forust 62d39ee4f1 Merge pull request 'chore(deps): update netbirdio/dashboard docker tag to v2.93.0' (#53) from renovate/netbirdio-dashboard-2.x into main
ci / lint-dockerfiles (push) Successful in 1s
deploy / preflight (push) Successful in 2s
renovate-ci / validate-renovate (push) Successful in 9s
deploy / validate (push) Failing after 2s
deploy / apply-k8s (push) Skipped
deploy / apply-compose (push) Skipped
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 3s
ci / validate (push) Successful in 1s
ci / build (push) Successful in 1s
Reviewed-on: #53
2026-09-25 22:27:21 +00:00
renovate-bot f1f7dd4a0a chore(deps): update netbirdio/dashboard docker tag to v2.93.0
deploy / validate (push) Skipped
renovate-ci / validate-renovate (push) Skipped
ci / lint-prettier (pull_request) Successful in 4s
ci / lint-ruff (pull_request) Successful in 2s
ci / lint-yaml (pull_request) Successful in 2s
ci / lint-dockerfiles (pull_request) Successful in 1s
ci / validate (pull_request) Successful in 2s
ci / build (pull_request) Skipped
renovate-ci / validate-renovate (pull_request) Successful in 9s
ci / lint-prettier (push) Successful in 2s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 1s
ci / validate (push) Successful in 1s
ci / build (push) Skipped
2026-09-25 22:26:10 +00:00
forust f4df6d4437 Merge pull request 'chore(deps): update container patch updates' (#38) from renovate/container-patch-updates into main
renovate-ci / validate-renovate (push) Skipped
deploy / preflight (push) Successful in 3s
deploy / validate (push) Failing after 2s
deploy / apply-k8s (push) Skipped
deploy / apply-compose (push) Skipped
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 1s
ci / validate (push) Successful in 2s
ci / lint-prettier (pull_request) Successful in 3s
ci / lint-ruff (pull_request) Successful in 1s
ci / lint-yaml (pull_request) Successful in 2s
ci / validate (pull_request) Successful in 1s
ci / build (pull_request) Skipped
ci / build (push) Skipped
ci / lint-dockerfiles (pull_request) Successful in 1s
renovate-ci / validate-renovate (pull_request) Successful in 21s
Reviewed-on: #38
2026-09-25 22:18:21 +00:00
renovate-bot 987a89f022 chore(deps): update container patch updates
deploy / validate (push) Skipped
renovate-ci / validate-renovate (push) Skipped
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 1s
ci / validate (push) Successful in 1s
ci / build (push) Skipped
ci / lint-prettier (pull_request) Canceled after 0s
ci / lint-yaml (pull_request) Canceled after 0s
ci / lint-dockerfiles (pull_request) Canceled after 0s
ci / validate (pull_request) Canceled after 0s
ci / lint-ruff (pull_request) Canceled after 0s
ci / build (pull_request) Canceled after 0s
renovate-ci / validate-renovate (pull_request) Successful in 18s
2026-09-25 22:18:07 +00:00
forust b423119632 Merge pull request 'chore(deps): update renovate/renovate docker tag to v44.115.9' (#47) from renovate/renovate-renovate-44.x into main
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 1s
ci / validate (push) Successful in 3s
ci / build (push) Canceled after 0s
deploy / preflight (push) Successful in 3s
deploy / validate (push) Canceled after 0s
deploy / apply-k8s (push) Canceled after 0s
deploy / apply-compose (push) Canceled after 0s
renovate-ci / validate-renovate (push) Successful in 1m30s
Reviewed-on: #47
2026-09-25 22:17:46 +00:00
renovate-bot ddbc0cc7e6 chore(deps): update renovate/renovate docker tag to v44.115.9 2026-09-25 22:17:46 +00:00
forust 6ad33d75e6 Merge pull request 'chore(deps): update prom/prometheus docker tag to v3.15.0' (#48) from renovate/prom-prometheus-3.x into main
ci / lint-prettier (push) Successful in 3s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 1s
ci / validate (push) Successful in 2s
ci / lint-ruff (push) Successful in 1s
deploy / preflight (push) Successful in 2s
ci / build (push) Canceled after 0s
deploy / validate (push) Canceled after 0s
deploy / apply-k8s (push) Canceled after 0s
deploy / apply-compose (push) Canceled after 0s
renovate-ci / validate-renovate (push) Successful in 24s
Reviewed-on: #48
2026-09-25 22:17:15 +00:00
renovate-bot 11bfb426c0 chore(deps): update prom/prometheus docker tag to v3.15.0 2026-09-25 22:17:15 +00:00
forust b7a1835adb Merge pull request 'feat(netbird): add tailscale-alternative' (#50) from feat/netbird into main
ci / lint-dockerfiles (push) Successful in 1s
ci / validate (push) Successful in 1s
deploy / preflight (push) Successful in 2s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 3s
ci / build (push) Canceled after 0s
deploy / validate (push) Canceled after 0s
deploy / apply-k8s (push) Canceled after 0s
deploy / apply-compose (push) Canceled after 0s
renovate-ci / validate-renovate (push) Successful in 21s
Reviewed-on: #50
2026-09-25 22:16:51 +00:00
forust ad4bb8750d chore(deploy): enable netbird
deploy / validate (push) Skipped
renovate-ci / validate-renovate (push) Skipped
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 1s
ci / validate (push) Successful in 2s
ci / build (push) Skipped
ci / lint-prettier (pull_request) Successful in 3s
ci / lint-ruff (pull_request) Successful in 1s
ci / lint-yaml (pull_request) Successful in 2s
ci / lint-dockerfiles (pull_request) Successful in 1s
ci / validate (pull_request) Successful in 1s
ci / build (pull_request) Skipped
renovate-ci / validate-renovate (pull_request) Successful in 8s
2026-09-25 22:16:02 +00:00
forust f1e00b946f Merge pull request 'Feat/documenting services' (#51) from feat/documenting-services into main
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 0s
ci / lint-yaml (push) Successful in 3s
ci / lint-dockerfiles (push) Successful in 1s
ci / validate (push) Successful in 1s
deploy / preflight (push) Successful in 2s
renovate-ci / validate-renovate (push) Successful in 9s
ci / build (push) Successful in 1s
deploy / validate (push) Failing after 1s
deploy / apply-k8s (push) Skipped
deploy / apply-compose (push) Skipped
Reviewed-on: #51
2026-09-25 22:15:45 +00:00
forust 7ee7d0c961 Update README.md
deploy / validate (push) Skipped
renovate-ci / validate-renovate (push) Skipped
ci / lint-dockerfiles (push) Successful in 1s
ci / validate (push) Successful in 1s
ci / build (push) Skipped
ci / lint-prettier (pull_request) Successful in 3s
ci / lint-ruff (pull_request) Successful in 0s
ci / lint-yaml (pull_request) Successful in 2s
ci / validate (pull_request) Successful in 1s
ci / build (pull_request) Skipped
renovate-ci / validate-renovate (pull_request) Successful in 8s
ci / lint-prettier (push) Successful in 2s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (pull_request) Successful in 1s
2026-09-26 00:15:01 +02:00
forust 27ab6b859e chore(deploy): enable netbird
deploy / validate (push) Skipped
renovate-ci / validate-renovate (push) Skipped
ci / lint-prettier (pull_request) Successful in 3s
ci / lint-ruff (pull_request) Successful in 1s
ci / validate (pull_request) Successful in 1s
ci / build (pull_request) Skipped
renovate-ci / validate-renovate (pull_request) Successful in 13s
ci / lint-yaml (pull_request) Successful in 2s
ci / lint-dockerfiles (pull_request) Successful in 0s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 0s
ci / validate (push) Successful in 2s
ci / build (push) Skipped
2026-09-26 00:04:50 +02:00
forust 71e769c002 chore(deploy): disable checkmk, enable netbox and rackpeek
deploy / validate (push) Skipped
renovate-ci / validate-renovate (push) Skipped
ci / lint-prettier (push) Failing after 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 1s
ci / validate (push) Successful in 2s
ci / build (push) Skipped
ci / lint-prettier (pull_request) Failing after 3s
ci / lint-ruff (pull_request) Successful in 1s
ci / lint-yaml (pull_request) Successful in 3s
ci / lint-dockerfiles (pull_request) Successful in 1s
ci / validate (pull_request) Successful in 1s
ci / build (pull_request) Skipped
renovate-ci / validate-renovate (pull_request) Successful in 21s
2026-09-26 00:04:37 +02:00
forust 16dd67c2c0 feat(rackpeek): add internal-only rack visualization service
Compose and k8s manifests behind workstation/gigaforust internal hosts.
2026-09-26 00:00:49 +02:00
forust c536a16a2a feat(netbox): add enterprise-grade server documenting app w/ shared postgres 2026-09-25 23:24:21 +02:00
forust a4a4bb4cc5 feat(netbird): add tailscale-alternative
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 1s
ci / validate (push) Successful in 2s
ci / build (push) Skipped
deploy / validate (push) Skipped
renovate-ci / validate-renovate (push) Skipped
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-prettier (pull_request) Successful in 3s
ci / lint-ruff (pull_request) Successful in 1s
ci / lint-yaml (pull_request) Successful in 5s
ci / lint-dockerfiles (pull_request) Successful in 1s
ci / validate (pull_request) Successful in 1s
ci / build (pull_request) Skipped
renovate-ci / validate-renovate (pull_request) Successful in 10s
2026-09-25 20:26:11 +02:00
forust 648b354951 refactor(deploy): marker-driven selection (k8s/active, root active); enable headscale/nextcloud hybrid, disable dockmon/kener/downtify/n8n
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / validate (push) Successful in 2s
deploy / preflight (push) Successful in 1s
renovate-ci / validate-renovate (push) Successful in 8s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 1s
ci / build (push) Successful in 1s
deploy / validate (push) Successful in 1m40s
deploy / apply-k8s (push) Successful in 1m41s
deploy / apply-compose (push) Successful in 13s
2026-09-23 18:11:51 +02:00
forust b0a9b3476d fix(deploy): drop broken %q quoting that wrapped remote vars in literal quotes
ci / lint-prettier (push) Successful in 2s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 1s
ci / validate (push) Successful in 2s
renovate-ci / validate-renovate (push) Successful in 8s
ci / build (push) Successful in 1s
deploy / redeploy (push) Failing after 1m43s
2026-09-23 16:33:51 +02:00
forust ac3bf4a55a fix(deploy): strip quotes and CR from DEPLOY_PATH
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 1s
ci / validate (push) Successful in 2s
renovate-ci / validate-renovate (push) Successful in 8s
ci / build (push) Successful in 1s
deploy / redeploy (push) Failing after 1s
2026-09-23 16:31:06 +02:00
forust 34fb6f85ba Merge pull request 'chore(deps): update darthnorse/dockmon docker tag to v2.5.0' (#39) from renovate/darthnorse-dockmon-2.x into main
renovate-ci / validate-renovate (push) Successful in 6s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 2s
ci / build (push) Successful in 1s
ci / lint-dockerfiles (push) Successful in 1s
ci / validate (push) Successful in 1s
deploy / redeploy (push) Failing after 1s
Reviewed-on: #39
2026-09-23 14:16:21 +00:00
renovate-bot 66eacd86d1 chore(deps): update darthnorse/dockmon docker tag to v2.5.0 2026-09-23 14:16:21 +00:00
forust 54f43fc2c7 Merge pull request 'chore(deps): update ghcr.io/lukegus/termix docker tag to v2.8.0' (#40) from renovate/ghcr.io-lukegus-termix-2.x into main
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 1s
ci / lint-prettier (push) Successful in 2s
ci / lint-ruff (push) Successful in 1s
ci / validate (push) Successful in 2s
deploy / redeploy (push) Failing after 0s
renovate-ci / validate-renovate (push) Successful in 7s
ci / build (push) Successful in 1s
Reviewed-on: #40
2026-09-23 14:16:06 +00:00
renovate-bot 451d0dc6b7 chore(deps): update ghcr.io/lukegus/termix docker tag to v2.8.0 2026-09-23 14:16:06 +00:00
forust 88ae20a543 Merge pull request 'chore(deps): update ghcr.io/henriquesebastiao/downtify docker tag to v3' (#42) from renovate/ghcr.io-henriquesebastiao-downtify-3.x into main
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 0s
ci / validate (push) Successful in 1s
deploy / redeploy (push) Failing after 0s
ci / build (push) Canceled after 0s
renovate-ci / validate-renovate (push) Successful in 7s
Reviewed-on: #42
2026-09-23 14:15:51 +00:00
renovate-bot 6bd183ba73 chore(deps): update ghcr.io/henriquesebastiao/downtify docker tag to v3
renovate-ci / validate-renovate (push) Skipped
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 1s
ci / validate (push) Successful in 1s
ci / lint-ruff (pull_request) Successful in 1s
ci / lint-yaml (pull_request) Successful in 2s
ci / lint-dockerfiles (pull_request) Successful in 1s
ci / validate (pull_request) Successful in 2s
ci / build (pull_request) Skipped
renovate-ci / validate-renovate (pull_request) Successful in 10s
ci / build (push) Skipped
ci / lint-prettier (pull_request) Successful in 3s
2026-09-23 14:14:58 +00:00
forust aeefdd8560 Merge pull request 'chore(deps): update docker.n8n.io/n8nio/n8n docker tag to v2.41.0' (#43) from renovate/docker.n8n.io-n8nio-n8n-2.x into main
ci / lint-prettier (push) Successful in 2s
ci / lint-ruff (push) Successful in 0s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 1s
ci / validate (push) Successful in 1s
deploy / redeploy (push) Failing after 1s
renovate-ci / validate-renovate (push) Successful in 8s
ci / build (push) Successful in 1s
Reviewed-on: #43
2026-09-23 14:14:39 +00:00
renovate-bot c10d344fd7 chore(deps): update docker.n8n.io/n8nio/n8n docker tag to v2.41.0 2026-09-23 14:14:39 +00:00
forust 6e5f80611b Merge pull request 'Cicd/deploy rework' (#44) from cicd/deploy-rework into main
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 0s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 1s
ci / validate (push) Successful in 1s
deploy / redeploy (push) Failing after 0s
renovate-ci / validate-renovate (push) Successful in 6s
ci / build (push) Successful in 1m47s
Reviewed-on: #44
2026-09-23 14:11:42 +00:00
139 changed files with 5882 additions and 923 deletions

No files matched your search

+10
View File
@@ -0,0 +1,10 @@
# actionlint configuration. Passed explicitly from the ci workflow:
# actionlint -config-file .gitea/actionlint.yaml .gitea/workflows/*.yaml
#
# The self-hosted act_runner registers custom labels that actionlint cannot know
# about, so declare them here instead of silencing the whole runner-label check.
self-hosted-runner:
labels:
- arch
- homelab
- prod
+465 -57
View File
@@ -7,6 +7,12 @@ on:
pull_request:
workflow_dispatch:
# Every job here is checkout plus local tools. The token needs to read the tree
# and nothing else, and saying so keeps a future step that reaches for the API
# from quietly holding a token that can write to the repository.
permissions:
contents: read
concurrency:
group: ci-${{ github.ref }}
cancel-in-progress: ${{ github.ref != 'refs/heads/main' }}
@@ -15,8 +21,90 @@ env:
REGISTRY: gcr.forust.xyz
jobs:
lint-compose:
runs-on: [self-hosted, linux, arch, homelab]
timeout-minutes: 10
steps:
- name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
# Structure check for every committed Compose file, active or not.
# Interpolation, env-file and bind-mount resolution are all switched off,
# because inactive stacks have no .env here and would only fail on their
# ${VAR:?} guards. Active stacks get the full check with interpolation in
# the deploy workflow, where the real .env files live.
- name: Validate Compose files
shell: bash
run: |
set -euo pipefail
source .gitea/workflows/compose-lint.sh
mapfile -t safe_flags < <(compose_safe_flags)
echo "docker compose config ${safe_flags[*]-}"
mapfile -t files < <(compose_files)
if [ "${#files[@]}" -eq 0 ]; then
echo "No Compose files found."
exit 0
fi
failed=0
for f in "${files[@]}"; do
if ! out="$(validate_compose_file "$f" ${safe_flags[@]+"${safe_flags[@]}"} 2>&1)"; then
failed=1
echo "::error file=${f}::$(printf '%s' "$out" | head -1)"
fi
done
if [ "$failed" -ne 0 ]; then
echo "Compose validation failed."
exit 1
fi
echo "checked ${#files[@]} Compose file(s)"
lint-actionlint:
runs-on: [self-hosted, linux, arch, homelab]
timeout-minutes: 10
steps:
- name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
- name: Lint Gitea Actions workflows with actionlint
shell: bash
run: |
set -euo pipefail
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh actionlint)"
export PATH="$tools_dir:$PATH"
actionlint -config-file .gitea/actionlint.yaml -color .gitea/workflows/*.yaml
lint-shellcheck:
runs-on: [self-hosted, linux, arch, homelab]
timeout-minutes: 10
steps:
- name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
- name: Lint shell scripts with ShellCheck
shell: bash
run: |
set -euo pipefail
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh shellcheck)"
export PATH="$tools_dir:$PATH"
# userbot/ is a git subtree synced from forust/userbot, so its shell
# scripts are upstream's to maintain, not ours. Linting them would let a
# routine subtree pull turn the deploy gate red on code we do not own.
mapfile -t scripts < <(
git ls-files '*.sh' ':(glob)**/*.bash' ':!userbot/**'
)
if [ "${#scripts[@]}" -eq 0 ]; then
echo "No shell scripts found."
exit 0
fi
shellcheck --external-sources --source-path=SCRIPTDIR --severity=style "${scripts[@]}"
lint-prettier:
runs-on: [self-hosted, linux, arch, homelab]
timeout-minutes: 10
steps:
- name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
@@ -24,6 +112,10 @@ jobs:
- name: Check formatting with Prettier
shell: bash
run: |
set -euo pipefail
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh prettier)"
export PATH="$tools_dir:$PATH"
mapfile -t prettier_files < <(
git ls-files \
| grep -E '\.(md|json|ya?ml|html|css)$' \
@@ -39,17 +131,23 @@ jobs:
lint-ruff:
runs-on: [self-hosted, linux, arch, homelab]
timeout-minutes: 10
steps:
- name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
- name: Lint Python with Ruff
- name: Lint and format-check Python with Ruff
shell: bash
run: |
set -euo pipefail
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh ruff)"
export PATH="$tools_dir:$PATH"
ruff check .
ruff format --check .
lint-yaml:
runs-on: [self-hosted, linux, arch, homelab]
timeout-minutes: 10
steps:
- name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
@@ -57,6 +155,10 @@ jobs:
- name: Lint YAML syntax
shell: bash
run: |
set -euo pipefail
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh yamllint)"
export PATH="$tools_dir:$PATH"
mapfile -t yaml_files < <(
git ls-files '*.yaml' '*.yml' \
':!node_modules/**' \
@@ -72,6 +174,7 @@ jobs:
lint-dockerfiles:
runs-on: [self-hosted, linux, arch, homelab]
timeout-minutes: 10
steps:
- name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
@@ -79,6 +182,10 @@ jobs:
- name: Lint Dockerfiles
shell: bash
run: |
set -euo pipefail
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh hadolint)"
export PATH="$tools_dir:$PATH"
mapfile -t dockerfiles < <(
git ls-files ':(glob)**/Dockerfile' ':(glob)**/Dockerfile.*'
)
@@ -90,15 +197,169 @@ jobs:
hadolint -c .hadolint.yaml "${dockerfiles[@]}"
validate:
# Known, accepted, and recorded. Each line is a real advisory against a
# package we build into the panel image, kept in this workflow rather than in
# the package manifest so that a subtree sync from forust/userbot cannot
# silently widen the exemption.
#
# starlette is the reason this job is not simply "fail on everything":
# fastapi 0.115.12 pins `starlette<0.47.0`, and the fixes for the last four
# below need 0.49.1 through 1.3.1, so clearing them means a jump from fastapi
# 0.115.12 to 0.141.x. That is upstream's call, not a drive-by in a lint
# commit. Of the seven, four are reachable here in principle: 1942 is a
# crafted Range header hitting FileResponse, and the panel serves its built
# SPA through exactly that; 249 is request.form() ignoring max_fields for
# x-www-form-urlencoded, which is the login form; 1941 is a large multipart
# body blocking the event loop; 161 and 248 are unvalidated Host and request
# path reaching request.url. 2280 needs HTTPEndpoint, which the panel does
# not use, and 2281 is Windows-only, and this deploys on Linux.
#
# The panel answers on userbot.workstation.internal and has no public
# forust.xyz route, which is what keeps the four reachable ones from being
# an internet-facing DoS. It still manages Telegram credentials.
#
# Deleting an entry here is how you accept a new advisory, so the diff says
# so out loud.
scan-deps:
runs-on: [self-hosted, linux, arch, homelab]
timeout-minutes: 15
steps:
- name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
- name: Validate Kubernetes manifests
- name: Audit the Python dependencies that ship in the image
shell: bash
run: |
set -euo pipefail
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh pip-audit)"
export PATH="$tools_dir:$PATH"
# requirements.txt, not requirements-dev.txt: this is what the image
# installs, and the test tooling is not a shipped attack surface.
pip-audit -r userbot/panel/backend/requirements.txt --strict \
--ignore-vuln CVE-2025-67720 \
--ignore-vuln PYSEC-2026-161 \
--ignore-vuln PYSEC-2026-1941 \
--ignore-vuln PYSEC-2026-1942 \
--ignore-vuln PYSEC-2026-2280 \
--ignore-vuln PYSEC-2026-2281 \
--ignore-vuln PYSEC-2026-248 \
--ignore-vuln PYSEC-2026-249
# devDependencies are excluded on purpose. `npm audit` on the full tree
# reports 7 findings, and every one of them is a build- or test-time
# package: the esbuild CORS advisory needs a vite dev server serving to
# the internet, and nanoid's infinite loop needs a custom generator
# called with size 0, which postcss does not do. None of them are in the
# 91 kB bundle the panel serves. The one production finding, devalue
# via svelte, is moderate, which is where --audit-level draws the line;
# this fails on the next high or critical one.
- name: Audit the production npm dependencies
shell: bash
run: |
set -euo pipefail
# The pinned node, not whatever the runner has. Its system node is a
# rolling Arch package: during this very push its npm was missing
# entirely, and an hour later it was npm 12 on node 26. Both are the
# wrong major anyway — the panel image is node:22-alpine.
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh node)"
export PATH="$tools_dir:$PATH"
cd userbot/panel/frontend
npm ci
npm audit --omit=dev --audit-level=high
test-backend:
runs-on: [self-hosted, linux, arch, homelab]
timeout-minutes: 15
steps:
- name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
# 25 tests over the panel's pydantic models, its auth flow, the SPA
# fallback and the Kubernetes client it shells out with. They existed and
# had never been executed by anything.
#
# Note that userbot/ is a subtree synced from forust/userbot, so a routine
# sync can turn this red on upstream's code. Unlike the shellcheck job,
# which skips that tree because style disagreements there are ours to
# lose, a failing test here is a real defect in a service we deploy.
- name: Run the panel backend test suite
shell: bash
run: |
set -euo pipefail
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh uv)"
export PATH="$tools_dir:$PATH"
# A venv in a temp dir rather than a checked-out one: the runner is
# shared, and a leftover .venv would let a dependency the
# requirements no longer pin still satisfy an import.
#
# --python is not optional. uv otherwise takes whatever interpreter it
# finds first, and which one that is depends on the machine: this
# runner runs jobs on the host, where the only interpreter is 3.14,
# and pyrogram's sync.py calls the bare asyncio.get_event_loop() that
# 3.14 no longer auto-creates, so three tests fail at collection. The
# image is python:3.13-slim, so 3.13 is also the version worth
# testing: uv fetches a managed build of it when the host has none,
# which is what makes this job independent of the runner.
venv="$(mktemp -d)/venv"
uv venv --python 3.13 --quiet "$venv"
uv pip install --quiet --python "$venv/bin/python" \
-r userbot/panel/backend/requirements-dev.txt
# `python -m`, not bare `pytest`: the tests import `app.*` relative to
# the backend directory, which only works if the cwd is on sys.path,
# and only `python -m` puts it there.
cd userbot/panel/backend
"$venv/bin/python" -m pytest tests/ -q
test-frontend:
runs-on: [self-hosted, linux, arch, homelab]
timeout-minutes: 15
steps:
- name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
# One `npm ci` for both checks below: it is by far the slowest part of
# this job, and a second one would learn nothing the first did not.
#
# `npm ci`, not `npm install`, for the same reason the Dockerfile uses it:
# the lockfile is what makes the tree that gets checked the tree that
# gets shipped.
- name: Type-check and test the panel frontend
shell: bash
run: |
set -euo pipefail
# The pinned node, not whatever the runner has. Its system node is a
# rolling Arch package: during this very push its npm was missing
# entirely, and an hour later it was npm 12 on node 26. Both are the
# wrong major anyway — the panel image is node:22-alpine.
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh node)"
export PATH="$tools_dir:$PATH"
cd userbot/panel/frontend
npm ci
# svelte-check has been a devDependency all along with no script
# pointing at it, so the type errors it reports had nowhere to
# surface. It is clean today, which is the only reason it can be a
# gate: it stops at whatever upstream introduces rather than
# reporting a backlog we inherited.
npm run check
npm test
validate:
runs-on: [self-hosted, linux, arch, homelab]
timeout-minutes: 20
steps:
- name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
- name: Validate Kubernetes manifests against JSON schemas
shell: bash
run: |
set -euo pipefail
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh kubeconform)"
export PATH="$tools_dir:$PATH"
mapfile -t manifests < <(
git ls-files ':(glob)**/k8s/**/*.yaml' ':(glob)**/k8s/**/*.yml' \
| grep -Ev '(^|/)(kustomization\.ya?ml|.*\.example\.ya?ml|.*values\.ya?ml|patch-.*\.ya?ml)$'
@@ -115,10 +376,115 @@ jobs:
-summary \
"${manifests[@]}"
# kubeconform has no schemas for CRDs, so every IngressRoute, Certificate,
# PrometheusRule, Middleware, ServersTransport and ServiceMonitor is silently
# skipped above. The live API server knows the real CRD schemas (and runs the
# cert-manager / Traefik admission webhooks), so validate there too.
#
# Only services marked with a k8s/active marker are checked: server-side
# dry-run needs the target namespace to exist, and inactive services are not
# deployed. Services being enabled for the first time are still covered by
# the JSON-schema pass above.
#
# Main pushes only. `--dry-run=server` persists nothing, but it does execute
# the admission webhooks of the production API server, so anyone able to open
# a pull request would be able to run arbitrary manifest content through
# cert-manager and Traefik. A pull request has nothing to gain from it either:
# only main is ever deployed, and this job runs to completion before the
# deploy workflow is allowed to start, so a bad CRD is still caught before
# anything reaches the cluster -- just on the push rather than on the PR.
- name: Note the server-side check is not running here
if: github.event_name == 'pull_request' || github.ref != 'refs/heads/main'
shell: bash
run: |
echo "::notice::Skipping the server-side dry-run. It executes the cert-manager and" \
"Traefik admission webhooks against the production API server, so it is limited" \
"to pushes to main. CRDs are still schema-checked by kubeconform above, and the" \
"server-side pass still runs on main before the deploy."
- name: Validate active manifests against the live API server
if: github.event_name != 'pull_request' && github.ref == 'refs/heads/main'
shell: bash
run: |
set -euo pipefail
if ! kubectl get --raw='/readyz' --request-timeout=10s >/dev/null 2>&1; then
echo "::warning::Cluster unreachable — skipped server-side validation of CRDs (IngressRoute, Certificate, PrometheusRule). Review manifest changes manually."
exit 0
fi
mapfile -t k8s_dirs < <(
git ls-files '*.yaml' '*.yml' \
| grep -E '(^|/)k8s/' \
| sed -E 's#((^|.*/)k8s)/.*#\1#' \
| sort -u
)
manifests=()
kustomize_apps=()
for dir in "${k8s_dirs[@]}"; do
if [ ! -f "${dir}/active" ]; then
echo "skip (no k8s/active): ${dir}"
continue
fi
if [ -f "${dir}/overlays/prod/kustomization.yaml" ]; then
kustomize_apps+=("${dir}/overlays/prod")
elif [ -f "${dir}/base/kustomization.yaml" ]; then
kustomize_apps+=("${dir}/base")
else
while IFS= read -r f; do
[ -n "$f" ] && manifests+=("$f")
done < <(
git ls-files "${dir}/*.yaml" "${dir}/*.yml" \
| grep -Ev '(^|/)(kustomization\.ya?ml|.*\.example\.ya?ml|.*values\.ya?ml|patch-.*\.ya?ml)$'
)
fi
done
echo "server-side dry-run: ${#manifests[@]} manifests, ${#kustomize_apps[@]} kustomize apps"
failed=0
for m in ${manifests[@]+"${manifests[@]}"}; do
if ! out="$(kubectl apply --dry-run=server -f "$m" 2>&1)"; then
failed=1
echo "::error file=${m}::$(printf '%s' "$out" | head -1)"
fi
done
for k in ${kustomize_apps[@]+"${kustomize_apps[@]}"}; do
if ! out="$(kubectl apply -k "$k" --dry-run=server 2>&1)"; then
failed=1
echo "::error file=${k}::$(printf '%s' "$out" | head -1)"
fi
done
if [ "$failed" -ne 0 ]; then
echo "Server-side validation failed. The API server (or an admission webhook) rejected these manifests."
exit 1
fi
echo "server-side dry-run: all active manifests accepted by the API server"
build:
needs: [lint-prettier, lint-ruff, lint-yaml, lint-dockerfiles, validate]
needs:
# scan-deps and the two test jobs were missing here, so a commit with a
# known-vulnerable dependency or a failing test still moved the :prod tag.
# The deploy was blocked either way - it requires the whole workflow to
# have succeeded - but the tag had already moved, and the next deploy to
# run resolved it. Publishing and passing the checks are the same gate.
[
lint-actionlint,
lint-shellcheck,
lint-compose,
lint-prettier,
lint-ruff,
lint-yaml,
lint-dockerfiles,
scan-deps,
test-backend,
test-frontend,
validate,
]
if: github.event_name != 'pull_request' && (github.ref_name == 'main' || github.ref_name == 'dev')
runs-on: [self-hosted, linux, arch, homelab]
timeout-minutes: 60
outputs:
services: ${{ steps.services.outputs.services }}
steps:
@@ -131,12 +497,20 @@ jobs:
id: services
shell: bash
run: |
set -euo pipefail
base="${{ github.event.before }}"
if [ -z "$base" ] || [ "$base" = "0000000000000000000000000000000000000000" ]; then
base="$(git rev-list --max-parents=0 HEAD)"
fi
mapfile -t changed_files < <(git diff --name-only "$base" "${GITHUB_SHA}")
# A failed diff used to leave changed_files empty, which reads exactly
# like "nothing to build": the job went green having built nothing and
# the tag never moved. The status is checked, not assumed.
if ! changed="$(git diff --name-only "$base" "${GITHUB_SHA}")"; then
echo "::error::cannot diff ${base}..${GITHUB_SHA}"
exit 1
fi
mapfile -t changed_files <<<"$changed"
services=()
@@ -186,30 +560,57 @@ jobs:
- name: Log in to registry
if: steps.services.outputs.services != ''
shell: bash
# Through env, not by substitution into the script. A secret written
# into a run: block is pasted into the shell source before bash parses
# it, so a password containing a quote, a backtick or $(...) becomes
# code that runs. Masking the value in the log does not prevent that.
env:
REGISTRY_USERNAME: ${{ secrets.REGISTRY_USERNAME }}
REGISTRY_PASSWORD: ${{ secrets.REGISTRY_PASSWORD }}
run: |
echo "${{ secrets.REGISTRY_PASSWORD }}" | docker login "${REGISTRY}" \
-u "${{ secrets.REGISTRY_USERNAME }}" \
set -euo pipefail
printf '%s' "$REGISTRY_PASSWORD" | docker login "${REGISTRY}" \
-u "$REGISTRY_USERNAME" \
--password-stdin
- name: Build and push changed images
if: steps.services.outputs.services != ''
shell: bash
run: |
# This step was the one run: block in the workflow without it, and it
# is the one that cannot afford it: a docker push that failed partway
# through the loop used to be followed by more pushes, the loop's exit
# status came from the last one, and the job went green with half the
# images missing from the registry.
set -euo pipefail
IFS=, read -r -a services <<< "${{ steps.services.outputs.services }}"
# Tags for this push. The commit-pinned name is the point of this
# step: the deploy resolves it in preference to :prod, so a deploy
# that sat in the queue behind a later push still gets the build of
# the commit CI validated, instead of whatever :prod points at by the
# time it runs. See render_pinned in deploy-lib.sh.
commit_tag=""
if [ "${GITHUB_REF_NAME}" = "main" ]; then
commit_tag="sha-${GITHUB_SHA:0:12}"
fi
set_tags() {
tags=()
case "${GITHUB_REF_NAME}" in
main) tags+=("main" "prod") ;;
dev) tags+=("dev") ;;
esac
if [ -n "$commit_tag" ]; then
tags+=("$commit_tag")
fi
}
for service in "${services[@]}"; do
case "$service" in
dtek_notif)
image="${REGISTRY}/forust/dtek-notif"
tags=("latest")
case "${GITHUB_REF_NAME}" in
main)
tags+=("main" "prod")
;;
dev)
tags+=("dev")
;;
esac
set_tags
build_args=()
for tag in "${tags[@]}"; do
build_args+=(-t "${image}:${tag}")
@@ -224,15 +625,7 @@ jobs:
;;
errorpages)
image="${REGISTRY}/forust/error-pages"
tags=("latest")
case "${GITHUB_REF_NAME}" in
main)
tags+=("main" "prod")
;;
dev)
tags+=("dev")
;;
esac
set_tags
build_args=()
for tag in "${tags[@]}"; do
build_args+=(-t "${image}:${tag}")
@@ -246,15 +639,7 @@ jobs:
done
;;
userbot)
tags=("latest")
case "${GITHUB_REF_NAME}" in
main)
tags+=("main" "prod")
;;
dev)
tags+=("dev")
;;
esac
set_tags
for target in runtime panel; do
case "$target" in
runtime)
@@ -280,8 +665,8 @@ jobs:
done
;;
homepages)
for service in forust xdfnx; do
case "$service" in
for variant in forust xdfnx; do
case "$variant" in
forust)
image="${REGISTRY}/forust/forust-homepage"
;;
@@ -289,15 +674,7 @@ jobs:
image="${REGISTRY}/forust/xdfnx-homepage"
;;
esac
tags=("latest")
case "${GITHUB_REF_NAME}" in
main)
tags+=("main" "prod")
;;
dev)
tags+=("dev")
;;
esac
set_tags
build_args=()
for tag in "${tags[@]}"; do
build_args+=(-t "${image}:${tag}")
@@ -305,15 +682,15 @@ jobs:
docker build \
--cache-from "type=registry,ref=${image}:buildcache" \
--cache-to "type=registry,ref=${image}:buildcache,mode=max" \
"${build_args[@]}" -f "homepages/Dockerfile.${service}" homepages
"${build_args[@]}" -f "homepages/Dockerfile.${variant}" homepages
for tag in "${tags[@]}"; do
docker push "${image}:${tag}"
done
done
;;
edu_master)
for service in session-keeper webinar-checker; do
case "$service" in
for variant in session-keeper webinar-checker; do
case "$variant" in
session-keeper)
context="edu_master/phpsessid-bot"
image="${REGISTRY}/forust/session-keeper"
@@ -323,15 +700,7 @@ jobs:
image="${REGISTRY}/forust/webinar-checker"
;;
esac
tags=("latest")
case "${GITHUB_REF_NAME}" in
main)
tags+=("main" "prod")
;;
dev)
tags+=("dev")
;;
esac
set_tags
build_args=()
for tag in "${tags[@]}"; do
build_args+=(-t "${image}:${tag}")
@@ -347,3 +716,42 @@ jobs:
;;
esac
done
# Every image the tree names has to carry the commit-pinned name, not only
# the ones this push rebuilt. A push that touches nothing but manifests
# builds nothing, and its deploy would then find no commit-pinned tag to
# resolve and quietly fall back to the moving :prod - which is the whole
# failure the commit-pinned name exists to remove.
#
# Re-tagging copies the manifest list and transfers no layers, so pinning
# six images that already exist costs six registry writes.
#
# The list is derived from the tree rather than written out here, so an
# image added to a manifest is covered without a second place to update.
- name: Pin the commit name on the images this push did not rebuild
if: github.ref_name == 'main'
shell: bash
run: |
set -euo pipefail
commit_tag="sha-${GITHUB_SHA:0:12}"
mapfile -t repos < <(
git grep -hoE 'gcr\.forust\.xyz/forust/[A-Za-z0-9._-]+' -- '*.yaml' '*.yml' \
| sort -u
)
if [ "${#repos[@]}" -eq 0 ]; then
echo "No own images referenced by the tree."
exit 0
fi
echo "pinning ${#repos[@]} image(s) to $commit_tag"
for repo in "${repos[@]}"; do
if docker buildx imagetools inspect "$repo:$commit_tag" >/dev/null 2>&1; then
echo " already built by this push: ${repo##*/}"
continue
fi
if ! docker buildx imagetools inspect "$repo:prod" >/dev/null 2>&1; then
echo " WARNING: ${repo##*/} has no :prod to pin and no build produced it"
continue
fi
docker buildx imagetools create --tag "$repo:$commit_tag" "$repo:prod"
echo " pinned ${repo##*/}"
done
+46
View File
@@ -0,0 +1,46 @@
#!/usr/bin/env bash
# Shared helpers for validating Compose files. Sourced both by steps in
# .gitea/workflows/ci.yaml and by deploy-lib.sh on the workstation.
#
# Two levels of checking, matching how the repo is structured:
#
# general every committed Compose file, active or not. Pure structure check:
# no ${VAR} interpolation, no .env lookup, no bind-mount path
# resolution. Disabled stacks deliberately have no .env in the repo
# and no values on the CI runner, so a full `config` run would fail on
# their `${VAR:?}` guards for reasons that have nothing to do with the
# change under review.
#
# full active stacks only, with interpolation and env-file resolution, so
# required variables and referenced files are actually resolved. Needs
# the gitignored .env files, so this only runs in the deploy workflow
# on the workstation.
#
# This file is meant to be sourced, not executed.
# All committed Compose files, including the ones deploy never starts.
compose_files() {
git ls-files \
'*/compose.yaml' '*/compose.yml' 'compose.yaml' 'compose.yml' \
'*/docker-compose.yaml' '*/docker-compose.yml'
}
# Prints the flags that turn `docker compose config` into the general check.
# Probed rather than hardcoded so an older Compose without --no-env-resolution
# still gets the flags it does support.
compose_safe_flags() {
local help flag
help="$(docker compose config --help 2>/dev/null || true)"
for flag in --no-interpolate --no-env-resolution --no-path-resolution; do
if printf '%s' "$help" | grep -q -- "$flag"; then
printf '%s\n' "$flag"
fi
done
}
# validate_compose_file <file> [extra docker compose config flags...]
validate_compose_file() {
local file="$1"
shift
docker compose -f "$file" config --quiet "$@"
}
File diff suppressed because it is too large. Load diff
+166 -281
View File
@@ -1,307 +1,192 @@
name: deploy
on:
push:
branches:
- main
# Deploy only what CI already validated. workflow_run is used instead of
# workflow_dispatch so a red lint/validate run can never reach the cluster.
workflow_run:
workflows: [ci]
types: [completed]
workflow_dispatch:
# The deploy jobs read the tree, then reach the cluster over SSH with the
# deploy key. The Actions token itself is not part of that path, so it gets
# read-only contents and no more.
permissions:
contents: read
concurrency:
group: deploy-main
# Queue instead of cancelling. Cancelling a run kills the apply job mid-loop and
# takes the verify job down with it, so a superseded deploy would leave the
# cluster half-applied and unchecked — the exact failure the verify job exists
# to catch. kubectl apply and docker compose up are both idempotent, so letting
# the older run finish and then deploying the newer commit costs little.
cancel-in-progress: false
env:
DEPLOY_HOST: ${{ secrets.DEPLOY_HOST }}
DEPLOY_PORT: ${{ secrets.DEPLOY_PORT }}
DEPLOY_USER: ${{ secrets.DEPLOY_USER }}
DEPLOY_PATH: ${{ secrets.DEPLOY_PATH }}
DEPLOY_KEY: ${{ secrets.DEPLOY_SSH_KEY }}
APPLY_PRUNE: ${{ vars.APPLY_PRUNE }}
# workflow_run's own GITHUB_SHA points at the branch head, not at the commit the
# finished ci run checked. Pin the exact validated commit instead, so a push
# landing mid-deploy cannot make the workstation deploy something else. Also
# what the verify job checks the snapshot against. Empty for workflow_dispatch,
# which falls back to the current origin/main.
DEPLOY_SHA: ${{ github.event.workflow_run.head_sha }}
jobs:
redeploy:
preflight:
if: >-
github.event_name != 'workflow_run' ||
(github.event.workflow_run.conclusion == 'success' &&
github.event.workflow_run.head_branch == 'main')
runs-on: [self-hosted, linux, arch, homelab, prod]
timeout-minutes: 10
steps:
- name: Redeploy workstation
- name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
- name: Fetch and reset workstation
shell: bash
env:
DEPLOY_HOST: ${{ secrets.DEPLOY_HOST }}
DEPLOY_PORT: ${{ secrets.DEPLOY_PORT }}
DEPLOY_USER: ${{ secrets.DEPLOY_USER }}
DEPLOY_PATH: ${{ secrets.DEPLOY_PATH }}
DEPLOY_KEY: ${{ secrets.DEPLOY_SSH_KEY }}
APPLY_PRUNE: ${{ vars.APPLY_PRUNE }}
run: |
set -euo pipefail
./.gitea/workflows/ssh-run.sh preflight
: "${DEPLOY_HOST:?missing DEPLOY_HOST}"
: "${DEPLOY_USER:?missing DEPLOY_USER}"
: "${DEPLOY_KEY:?missing DEPLOY_SSH_KEY}"
validate:
needs: [preflight]
runs-on: [self-hosted, linux, arch, homelab, prod]
timeout-minutes: 20
steps:
- name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
deploy_port="${DEPLOY_PORT:-22}"
deploy_path="${DEPLOY_PATH:-/srv/homelab}"
ssh_key="$RUNNER_TEMP/deploy_key"
mkdir -p "$RUNNER_TEMP"
printf '%s\n' "$DEPLOY_KEY" > "$ssh_key"
chmod 600 "$ssh_key"
ssh_opts=(
-i "$ssh_key"
-p "$deploy_port"
-o BatchMode=yes
-o StrictHostKeyChecking=accept-new
)
ssh "${ssh_opts[@]}" "${DEPLOY_USER}@${DEPLOY_HOST}" \
"DEPLOY_PATH=$(printf '%q' \"$deploy_path\") APPLY_PRUNE=$(printf '%q' \"${APPLY_PRUNE:-false}\") bash -se" <<'EOF'
- name: Dry-run manifests and check Secrets
shell: bash
run: |
set -euo pipefail
./.gitea/workflows/ssh-run.sh validate
repo="${DEPLOY_PATH:-/srv/homelab}"
apply-k8s:
needs: [validate]
runs-on: [self-hosted, linux, arch, homelab, prod]
# Apply only, no verification, so this is just the work itself: snapshot,
# then sequential `helm upgrade --atomic --timeout 10m`, then the apply loop.
# Verification has its own job and its own budget.
#
# 45 is roughly four times the measured cost of the stage, which is
# deliberately not raised on a theory:
#
# helm, healthy 3 no-op upgrades ~3-5 min
# helm, one release bad --atomic spends its 10m, ~10-15 min
# then rolls that one back
# apply loop ~40 manifests, 4 of which ~1 min
# resolve an image digest
# restart_stale_images 7.6s to find 8 workloads, ~0.5 min
# 9.8s to resolve their digests
#
# The helm figure is one release, not three: `set -e` aborts
# upgrade_helm_releases on the first failure, so a broken release costs
# 10m and the other two are never attempted. Multiplying 10m by three
# overstates the worst case by 20 minutes.
#
# The 45 minutes this was last raised to 45 were still not enough, and the
# job logs for those runs no longer exist, so what actually consumed the
# budget is not known - the two measurable candidates above account for
# ~15 of it. The one unbounded thing left in this stage is
# `docker manifest inspect` at deploy-lib.sh:236, which has no timeout
# against a registry with a known hang mode. Bound it, and make the stage
# announce what it is working on, before spending any of that on a larger
# ceiling: a stage that is killed with a diagnosable last line is a bug
# report, one that vanishes is not.
timeout-minutes: 45
steps:
- name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
if [ ! -d "$repo/.git" ]; then
echo "Repository not found at $repo"
exit 1
fi
- name: Apply Kubernetes manifests
shell: bash
run: |
set -euo pipefail
./.gitea/workflows/ssh-run.sh apply-k8s
git -C "$repo" fetch origin main
apply-compose:
needs: [validate]
runs-on: [self-hosted, linux, arch, homelab, prod]
timeout-minutes: 30
steps:
- name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
echo "== Workstation state =="
echo " local: $(git -C "$repo" rev-parse --short HEAD)"
echo " remote: $(git -C "$repo" rev-parse --short origin/main)"
- name: Redeploy docker compose stacks
shell: bash
run: |
set -euo pipefail
./.gitea/workflows/ssh-run.sh apply-compose
if [ -n "$(git -C "$repo" status --porcelain --untracked-files=no)" ]; then
echo "ERROR: workstation has local tracked modifications, refusing reset:"
git -C "$repo" status --porcelain --untracked-files=no
git -C "$repo" diff --stat
exit 1
fi
# Watches the workloads this deploy changed and rolls back the ones that never
# became healthy. Runs even when the apply jobs failed, timed out or were
# cancelled — that is the whole point of splitting it out. `always()` is what
# lets it start after a failed dependency; the needs on apply-compose are a
# barrier, so verification begins only once both applies are done.
verify-k8s:
needs: [apply-k8s, apply-compose]
if: >-
always() &&
needs.apply-k8s.result != 'skipped' &&
needs.apply-compose.result != 'skipped'
runs-on: [self-hosted, linux, arch, homelab, prod]
# Not raised, because the arithmetic does not close.
#
# 32 workloads are under management and the wave width is 8, so the verify
# itself is 4 waves of ROLLOUT_TIMEOUT (300s) = 20 minutes worst case, when
# every rollout times out rather than converging. That is already 20 of 30.
#
# The other 10 would have to absorb rollback, and rollback_workloads is a
# serial `while read` loop at 300s per failed workload. 10 minutes buys two.
# Any larger number is buying a bigger multiple of an unbounded term rather
# than covering a known cost: 60 minutes buys eight, and 60 minutes is
# therefore not a bound, it is a guess with two digits.
#
# The number becomes derivable the moment rollback uses the same wave width
# as the verify: 32 failures then cost 4 waves = 20 minutes instead of 160,
# and 45 covers verify plus rollback at full width. That change is to the
# recovery path and is not folded into a timeout edit.
timeout-minutes: 30
steps:
- name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
git -C "$repo" reset --hard origin/main
cd "$repo"
- name: Verify workloads and roll back on failure
shell: bash
run: |
set -euo pipefail
./.gitea/workflows/ssh-run.sh verify-k8s
is_disabled() {
local target="$1"
if [ -f "$target" ]; then
target="$(dirname "$target")"
fi
while true; do
if [ -f "$target/DISABLED" ]; then
return 0
fi
if [ "$target" = "$repo" ]; then
break
fi
target="$(dirname "$target")"
case "$target" in
"$repo"/*) ;;
*) break ;;
esac
done
return 1
}
# Asks the public route of every active service whether it is actually
# serving, which the rollout check above structurally cannot: a pod can
# converge and still be crash-looping, or be listening on a port no Service
# points at, or answer 500.
#
# `always()` for the same reason verify-k8s has it, and it runs after that job
# specifically because a rollback is when a route most needs re-checking. The
# needs is a barrier, not a filter: whether verify-k8s passed, failed or was
# cancelled, the probes are what say whether the cluster is serving, and
# suppressing them on a rollback would hide the one run where the answer
# matters most.
smoke:
needs: [verify-k8s]
if: always() && needs.verify-k8s.result != 'skipped'
runs-on: [self-hosted, linux, arch, homelab, prod]
timeout-minutes: 10
steps:
- name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
collect_k8s() {
git ls-files -- "$1" \
| grep -E '\.ya?ml$' \
| grep -Ev '/routing/|/overlays/' \
| grep -Ev '(^|/)(kustomization\.ya?ml|.*\.example\.ya?ml|.*values\.ya?ml|patch-.*\.ya?ml)$' \
| grep -Ev '(^|/)[^/]*secret[^/]*\.ya?ml$' \
| sort
}
collect_k8s_inactive() {
collect_k8s "$1" \
| grep -E '(^|/)namespace\.ya?ml$|/routing/'
}
kustomize_overlay() {
if [ -f "$1/overlays/prod/kustomization.yaml" ]; then
echo "$1/overlays/prod"
elif [ -f "$1/base/kustomization.yaml" ]; then
echo "$1/base"
fi
}
mapfile -t k8s_dirs < <(
git ls-files '*.yaml' '*.yml' \
| grep -E '(^|/)k8s/' \
| sed -E 's#((^|.*/)k8s)/.*#\1#' \
| sort -u
)
k8s_manifests=()
kustomize_apps=()
for kd_rel in "${k8s_dirs[@]}"; do
kd="$repo/$kd_rel"
if is_disabled "$kd"; then
echo "skip (DISABLED): $kd_rel"
continue
fi
if [ -f "$kd/active" ]; then
overlay="$(kustomize_overlay "$kd" || true)"
if [ -n "${overlay:-}" ]; then
echo "kustomize app: ${overlay#$repo/}"
kustomize_apps+=("$overlay")
else
while IFS= read -r f; do
[ -n "$f" ] && k8s_manifests+=("$repo/$f")
done < <(collect_k8s "$kd_rel" || true)
fi
else
while IFS= read -r f; do
[ -n "$f" ] && k8s_manifests+=("$repo/$f")
done < <(collect_k8s_inactive "$kd_rel" || true)
fi
done
mapfile -t compose_rel < <(
git ls-files '*/compose.yaml' '*/compose.yml' compose.yaml compose.yml | sort
)
compose_stacks=()
for cf_rel in "${compose_rel[@]}"; do
cf="$repo/$cf_rel"
if is_disabled "$cf"; then
echo "skip (DISABLED): $cf_rel"
continue
fi
if [ -f "$(dirname "$cf")/k8s/active" ]; then
echo "skip (k8s-managed): $cf_rel"
continue
fi
compose_stacks+=("$cf")
done
echo "== Validate compose stacks =="
for cf in "${compose_stacks[@]}"; do
echo " config: $cf"
docker compose -f "$cf" config --quiet
done
echo "== Validate k8s manifests (kubectl dry-run=client) =="
for m in "${k8s_manifests[@]}"; do
echo " apply --dry-run=client $m"
kubectl apply --dry-run=client -f "$m" >/dev/null
done
for k in "${kustomize_apps[@]}"; do
echo " apply -k --dry-run=client $k"
kubectl apply -k "$k" --dry-run=client >/dev/null
done
echo "== Validate k8s manifests (kubectl dry-run=server) =="
for m in "${k8s_manifests[@]}"; do
echo " apply --dry-run=server $m"
kubectl apply --dry-run=server -f "$m" >/dev/null
done
for k in "${kustomize_apps[@]}"; do
echo " apply -k --dry-run=server $k"
kubectl apply -k "$k" --dry-run=server >/dev/null
done
echo "== Checking referenced Secrets exist =="
echo " (deploy never applies *secret*.yaml; create missing ones from the laptop)"
ref_secrets=()
if [ "${#k8s_manifests[@]}" -gt 0 ]; then
while IFS= read -r s; do
[ -n "$s" ] && ref_secrets+=("$s")
done < <(
{
grep -h -A1 -E 'secretRef:|secretKeyRef:' "${k8s_manifests[@]}" 2>/dev/null || true
grep -h -E 'secretName:' "${k8s_manifests[@]}" 2>/dev/null || true
} | grep -E 'name:' | sed -E 's/.*name:[[:space:]]*//' | tr -d '"'"'"' "'"'" | sed -E 's/[[:space:]]*#.*//' | awk 'NF' | sort -u || true
)
fi
missing_secrets=()
all_secrets="$(kubectl get secrets -A --no-headers -o custom-columns=:metadata.name 2>/dev/null || true)"
for s in "${ref_secrets[@]}"; do
if printf '%s\n' "$all_secrets" | grep -qx "$s"; then
echo " ok: $s"
else
echo " MISSING: $s"
missing_secrets+=("$s")
fi
done
if [ "${#missing_secrets[@]}" -gt 0 ]; then
echo "ERROR: ${#missing_secrets[@]} referenced Secret(s) not found in the cluster:"
printf ' - %s\n' "${missing_secrets[@]}"
echo "Create them manually from the laptop, e.g.:"
echo " kubectl apply -f SERVICE/k8s/secrets.yaml # see SERVICE/k8s/secrets.yaml.example"
exit 1
fi
echo "== Applying Kubernetes manifests =="
ns_files=()
other_files=()
for m in "${k8s_manifests[@]}"; do
case "$m" in
*/namespace.y?ml) ns_files+=("$m") ;;
*) other_files+=("$m") ;;
esac
done
prune_opts=()
if [ "${APPLY_PRUNE:-false}" = "true" ]; then
prune_opts=(--prune -l app.kubernetes.io/managed-by=homelab-deploy)
fi
if [ "${#ns_files[@]}" -gt 0 ]; then
echo " namespaces first: ${ns_files[*]}"
kubectl apply -f "${ns_files[@]}"
fi
if [ -f "$repo/prometheus-stack/k8s/active" ] && ! is_disabled "$repo/prometheus-stack/k8s"; then
if [ ! -f "$repo/prometheus-stack/k8s/grafana-values.yaml" ]; then
echo "ERROR: prometheus-stack/k8s/grafana-values.yaml (gitignored) missing on workstation, restore it first."
exit 1
fi
echo "== Upgrading kube-prometheus-stack =="
helm upgrade --install prometheus-stack prometheus-community/kube-prometheus-stack \
--namespace prometheus \
--version 86.2.3 \
--values "$repo/prometheus-stack/k8s/grafana-values.yaml" \
--wait --timeout 10m
fi
if [ -f "$repo/loki/k8s/active" ] && ! is_disabled "$repo/loki/k8s"; then
echo "== Upgrading loki/alloy =="
helm repo add grafana https://grafana.github.io/helm-charts >/dev/null 2>&1 || true
helm repo update grafana >/dev/null 2>&1 || true
helm upgrade --install loki grafana/loki \
--version 7.3.0 \
--namespace prometheus \
--values "$repo/loki/k8s/loki-values.yaml" \
--wait --timeout 10m
helm upgrade --install alloy grafana/alloy \
--version 1.12.1 \
--namespace prometheus \
--values "$repo/loki/k8s/alloy-values.yaml" \
--wait --timeout 10m
fi
if [ "${#other_files[@]}" -gt 0 ]; then
echo " resources: ${other_files[*]}"
kubectl apply "${prune_opts[@]}" -f "${other_files[@]}"
fi
for k in "${kustomize_apps[@]}"; do
echo "== Applying kustomize app: ${k#$repo/} =="
kubectl apply -k "$k"
done
if [ -f "$repo/userbot/k8s/active" ] && ! is_disabled "$repo/userbot"; then
echo "== userbot panel hook =="
if kubectl get secret userbot-common-secrets -n userbot >/dev/null 2>&1; then
echo " userbot-common-secrets already present in userbot ns, not touching"
elif kubectl get secret userbot-common-secrets -n default >/dev/null 2>&1; then
echo " bootstrapping userbot-common-secrets into userbot ns"
kubectl get secret userbot-common-secrets -n default -o json \
| jq 'del(.metadata.annotations,.metadata.creationTimestamp,.metadata.resourceVersion,.metadata.uid,.metadata.managedFields) | .metadata.namespace = "userbot"' \
| kubectl apply -f -
else
echo " WARNING: userbot-common-secrets missing in both default and userbot ns; create it manually from the laptop"
fi
kubectl rollout restart deployment/userbot-panel -n userbot
kubectl rollout status deployment/userbot-panel -n userbot --timeout=180s
fi
echo "== Redeploying docker compose stacks =="
for cf in "${compose_stacks[@]}"; do
echo " compose: $cf"
if grep -Eq '^\s+pull_policy:\s*build\b' "$cf"; then
docker compose -f "$cf" build
docker compose -f "$cf" push
fi
docker compose -f "$cf" up -d --pull always --remove-orphans
done
EOF
- name: Probe the public route of every active service
shell: bash
run: |
set -euo pipefail
./.gitea/workflows/ssh-run.sh smoke
+233
View File
@@ -0,0 +1,233 @@
#!/usr/bin/env bash
# Installs the pinned CI tools into "$TOOLS_DIR/bin" and echoes that directory
# on stdout, so callers can do:
#
# export PATH="$(bash .gitea/workflows/install-ci-tools.sh kubeconform shellcheck):$PATH"
#
# Versions come from tool-versions.env next to this script and are kept fresh by
# Renovate. Re-running is cheap: an already-installed tool at the pinned version
# is left alone.
set -euo pipefail
here="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
# shellcheck source=tool-versions.env
. "$here/tool-versions.env"
TOOLS_DIR="${TOOLS_DIR:-${RUNNER_TEMP:-/tmp}/homelab-tools}"
BIN_DIR="$TOOLS_DIR/bin"
mkdir -p "$BIN_DIR"
arch="$(uname -m)"
# Upstream projects disagree on arch spelling: kubeconform and actionlint use
# Go names (amd64/arm64), shellcheck uses uname names (x86_64/aarch64), node
# uses neither (x64/arm64), and hadolint mixes the two in a single release
# (x86_64 but arm64).
case "$arch" in
x86_64 | amd64)
goarch=amd64
sharch=x86_64
nodearch=x64
hadolintarch=x86_64
;;
aarch64 | arm64)
goarch=arm64
sharch=aarch64
nodearch=arm64
hadolintarch=arm64
;;
*)
echo "install-ci-tools: unsupported architecture: $arch" >&2
exit 1
;;
esac
fetch() {
# fetch <url> <dest>
if command -v curl >/dev/null 2>&1; then
curl -sSLf --retry 3 -o "$2" "$1"
elif command -v wget >/dev/null 2>&1; then
wget -q -O "$2" "$1"
else
echo "install-ci-tools: neither curl nor wget is available" >&2
exit 1
fi
}
# installed_version <command>
# Prints the version of an already-installed tool, or nothing. Each tool spells
# its version flag differently, hence the case.
installed_version() {
local out
case "$1" in
kubeconform) out="$("$1" -v 2>/dev/null | head -1 || true)" ;;
*) out="$("$1" --version 2>/dev/null | head -1 || true)" ;;
esac
printf '%s' "$out"
}
# at_version <command> <expected>
at_version() {
case "$(installed_version "$1")" in
*"$2"*) return 0 ;;
*) return 1 ;;
esac
}
install_kubeconform() {
if at_version kubeconform "v${KUBECONFORM_VERSION}"; then
return 0
fi
local tmp
tmp="$(mktemp -d)"
fetch "https://github.com/yannh/kubeconform/releases/download/v${KUBECONFORM_VERSION}/kubeconform-linux-${goarch}.tar.gz" \
"$tmp/kubeconform.tar.gz"
tar -xzf "$tmp/kubeconform.tar.gz" -C "$tmp" kubeconform
install -m 0755 "$tmp/kubeconform" "$BIN_DIR/kubeconform"
rm -rf "$tmp"
}
install_shellcheck() {
if at_version shellcheck "${SHELLCHECK_VERSION}"; then
return 0
fi
local tmp
tmp="$(mktemp -d)"
fetch "https://github.com/koalaman/shellcheck/releases/download/v${SHELLCHECK_VERSION}/shellcheck-v${SHELLCHECK_VERSION}.linux.${sharch}.tar.xz" \
"$tmp/shellcheck.tar.xz"
tar -xJf "$tmp/shellcheck.tar.xz" -C "$tmp" --strip-components=1 "shellcheck-v${SHELLCHECK_VERSION}/shellcheck"
install -m 0755 "$tmp/shellcheck" "$BIN_DIR/shellcheck"
rm -rf "$tmp"
}
install_uv() {
if at_version uv "${UV_VERSION}"; then
return 0
fi
local tmp
tmp="$(mktemp -d)"
# uv release tags carry no leading v, unlike every other tool installed here.
fetch "https://github.com/astral-sh/uv/releases/download/${UV_VERSION}/uv-${sharch}-unknown-linux-gnu.tar.gz" \
"$tmp/uv.tar.gz"
tar -xzf "$tmp/uv.tar.gz" -C "$tmp" --strip-components=1 "uv-${sharch}-unknown-linux-gnu/uv"
install -m 0755 "$tmp/uv" "$BIN_DIR/uv"
rm -rf "$tmp"
}
install_hadolint() {
if at_version hadolint "${HADOLINT_VERSION}"; then
return 0
fi
# A bare binary, no archive: hadolint ships one file per platform.
fetch "https://github.com/hadolint/hadolint/releases/download/v${HADOLINT_VERSION}/hadolint-linux-${hadolintarch}" \
"$BIN_DIR/hadolint"
chmod 0755 "$BIN_DIR/hadolint"
}
# ruff and yamllint both come from PyPI as wheels, which uv unpacks for us.
install_uv_tool() {
# <package> <pinned version>
if at_version "$1" "$2"; then
return 0
fi
install_uv
UV_TOOL_BIN_DIR="$BIN_DIR" uv tool install --force "$1==$2" >/dev/null
}
install_ruff() {
install_uv_tool ruff "${RUFF_VERSION}"
}
install_yamllint() {
install_uv_tool yamllint "${YAMLLINT_VERSION}"
}
install_pip_audit() {
install_uv_tool pip-audit "${PIP_AUDIT_VERSION}"
}
install_prettier() {
if at_version prettier "${PRETTIER_VERSION}"; then
return 0
fi
# Not a standalone binary: prettier's entry point requires ../package.json
# relative to its own real path, so the package directory has to survive
# next to it. Hence a versioned directory plus a relative symlink, rather
# than copying the one file out as the other installers do.
local dir="$BIN_DIR/prettier-${PRETTIER_VERSION}"
if [ ! -f "$dir/package/package.json" ]; then
rm -rf "$dir"
mkdir -p "$dir"
fetch "https://registry.npmjs.org/prettier/-/prettier-${PRETTIER_VERSION}.tgz" "$dir/prettier.tgz"
tar -xzf "$dir/prettier.tgz" -C "$dir"
rm -f "$dir/prettier.tgz"
# npm strips the exec bit from bin/ on the way into the tarball.
chmod 0755 "$dir/package/bin/prettier.cjs"
fi
# Relative, so the whole tree stays valid if TOOLS_DIR is relocated.
ln -sfn "prettier-${PRETTIER_VERSION}/package/bin/prettier.cjs" "$BIN_DIR/prettier"
}
install_node() {
# npm gets checked by running it, not by looking it up: what matters is that
# it answers, so a stub, a half-removed Arch package or a name that resolves
# to something broken all have to read as "not installed". The runner's npm
# is a symlink into /usr/lib/node_modules/npm, which is exactly the kind of
# thing that disappears between runs.
if at_version node "v${NODE_VERSION}" && [ -n "$(installed_version npm)" ]; then
return 0
fi
# Same shape as prettier above: the tarball's bin/npm and bin/npx are links
# into lib/node_modules, so the whole tree has to survive next to them.
local dir="$BIN_DIR/node-${NODE_VERSION}"
if [ ! -x "$dir/bin/node" ]; then
rm -rf "$dir"
mkdir -p "$dir"
fetch "https://nodejs.org/dist/v${NODE_VERSION}/node-v${NODE_VERSION}-linux-${nodearch}.tar.xz" \
"$dir/node.tar.xz"
tar -xJf "$dir/node.tar.xz" -C "$dir" --strip-components=1 "node-v${NODE_VERSION}-linux-${nodearch}"
rm -f "$dir/node.tar.xz"
fi
# Relative, so the whole tree stays valid if TOOLS_DIR is relocated.
for bin in node npm npx; do
ln -sfn "node-${NODE_VERSION}/bin/${bin}" "$BIN_DIR/${bin}"
done
}
install_actionlint() {
if at_version actionlint "${ACTIONLINT_VERSION}"; then
return 0
fi
local tmp
tmp="$(mktemp -d)"
fetch "https://github.com/rhysd/actionlint/releases/download/v${ACTIONLINT_VERSION}/actionlint_${ACTIONLINT_VERSION}_linux_${goarch}.tar.gz" \
"$tmp/actionlint.tar.gz"
tar -xzf "$tmp/actionlint.tar.gz" -C "$tmp" actionlint
install -m 0755 "$tmp/actionlint" "$BIN_DIR/actionlint"
rm -rf "$tmp"
}
wanted=("$@")
if [ "${#wanted[@]}" -eq 0 ]; then
wanted=(kubeconform shellcheck actionlint prettier ruff yamllint hadolint)
fi
for tool in "${wanted[@]}"; do
case "$tool" in
kubeconform) install_kubeconform ;;
shellcheck) install_shellcheck ;;
actionlint) install_actionlint ;;
prettier) install_prettier ;;
ruff) install_ruff ;;
yamllint) install_yamllint ;;
pip-audit) install_pip_audit ;;
hadolint) install_hadolint ;;
node) install_node ;;
uv) install_uv ;;
*)
echo "install-ci-tools: unknown tool: $tool" >&2
exit 1
;;
esac
done
printf '%s\n' "$BIN_DIR"
+42 -18
View File
@@ -7,33 +7,58 @@ on:
- main
workflow_dispatch:
permissions:
contents: read
jobs:
validate-renovate:
runs-on: [self-hosted, linux, arch, homelab]
timeout-minutes: 20
steps:
- name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
- name: Validate Renovate Compose draft
# renovate/k8s/cronjob.yaml is the single source of truth for the image tag,
# so the same version that runs in the cluster is the one validated here.
- name: Resolve the deployed Renovate image
id: image
shell: bash
run: |
set -euo pipefail
trap 'rm -f renovate/.env' EXIT
printf '%s\n' \
'RENOVATE_ENDPOINT=https://gitea.example/api/v1' \
'RENOVATE_TOKEN=test-token' \
'RENOVATE_REPOSITORIES=forust/homelab' \
> renovate/.env
docker compose -f renovate/renovate-compose.yaml config --quiet
image="$(sed -n 's|.*image:[[:space:]]*\(renovate/renovate:[^[:space:]]*\).*|\1|p' \
renovate/k8s/cronjob.yaml | head -1)"
if [ -z "$image" ]; then
echo "::error::no renovate/renovate image found in renovate/k8s/cronjob.yaml"
exit 1
fi
echo "using $image"
echo "image=$image" >> "$GITHUB_OUTPUT"
- name: Validate Kubernetes manifests
- name: Validate Renovate repository config
shell: bash
run: |
set -euo pipefail
docker run --rm \
-v "$PWD:/work" \
-w /work \
ghcr.io/yannh/kubeconform:latest \
-v "$PWD/renovate:/opt/renovate:ro" \
-e RENOVATE_CONFIG_FILE=/opt/renovate/renovate.json \
"${{ steps.image.outputs.image }}" \
renovate-config-validator /opt/renovate/renovate.json
# The CronJob cannot read the repository, so renovate/k8s/configmap.yaml
# carries an inlined copy of the config. Fail if it no longer matches.
- name: Check the generated Renovate ConfigMap
shell: bash
run: |
set -euo pipefail
./.gitea/workflows/sync-renovate-configmap.sh --check
- name: Validate Renovate Kubernetes manifests
shell: bash
run: |
set -euo pipefail
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh kubeconform)"
export PATH="$tools_dir:$PATH"
kubeconform \
-strict \
-ignore-missing-schemas \
-summary \
@@ -41,12 +66,11 @@ jobs:
renovate/k8s/configmap.yaml \
renovate/k8s/cronjob.yaml
- name: Validate Renovate repository config
- name: Validate Renovate Compose file
shell: bash
run: |
set -euo pipefail
docker run --rm \
-v "$PWD:/work" \
-w /work \
renovate/renovate:44.103.0 \
renovate-config-validator renovate.json
source .gitea/workflows/compose-lint.sh
mapfile -t safe_flags < <(compose_safe_flags)
validate_compose_file renovate/renovate-compose.yaml \
${safe_flags[@]+"${safe_flags[@]}"}
+29 -6
View File
@@ -21,6 +21,11 @@ on:
default: false
type: boolean
# Renovate writes through its own bot PAT, passed in as RENOVATE_TOKEN, so the
# Actions token is only ever used to read the checkout.
permissions:
contents: read
concurrency:
group: renovate-run
cancel-in-progress: false
@@ -28,18 +33,36 @@ concurrency:
jobs:
run-renovate:
runs-on: [self-hosted, linux, arch, homelab]
timeout-minutes: 60
steps:
- name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
# renovate/k8s/cronjob.yaml is the single source of truth for the image tag.
# Reading it here means this workflow validates and runs the exact version
# that is deployed, instead of a copy that silently goes stale.
- name: Resolve the deployed Renovate image
id: image
shell: bash
run: |
set -euo pipefail
image="$(sed -n 's|.*image:[[:space:]]*\(renovate/renovate:[^[:space:]]*\).*|\1|p' \
renovate/k8s/cronjob.yaml | head -1)"
if [ -z "$image" ]; then
echo "::error::no renovate/renovate image found in renovate/k8s/cronjob.yaml"
exit 1
fi
echo "using $image"
echo "image=$image" >> "$GITHUB_OUTPUT"
- name: Validate Renovate config
shell: bash
run: |
set -euo pipefail
docker run --rm \
-v "$PWD/renovate/config.js:/opt/renovate/config.js:ro" \
-e RENOVATE_CONFIG_FILE=/opt/renovate/config.js \
renovate/renovate:44.103.0 \
-v "$PWD/renovate/renovate.json:/opt/renovate/renovate.json:ro" \
-e RENOVATE_CONFIG_FILE=/opt/renovate/renovate.json \
"${{ steps.image.outputs.image }}" \
renovate-config-validator
- name: Run Renovate
@@ -56,14 +79,14 @@ jobs:
: "${RENOVATE_TOKEN:?missing RENOVATE_TOKEN secret — add a renovate-bot PAT in repo/org Actions secrets}"
docker run --rm \
-v "$PWD/renovate/config.js:/opt/renovate/config.js:ro" \
-v "$PWD/renovate/renovate.json:/opt/renovate/renovate.json:ro" \
-e RENOVATE_PLATFORM=gitea \
-e RENOVATE_ENDPOINT=https://gitea.forust.xyz/api/v1 \
-e RENOVATE_TOKEN="$RENOVATE_TOKEN" \
-e RENOVATE_GITHUB_COM_TOKEN="${RENOVATE_GITHUB_COM_TOKEN:-}" \
-e RENOVATE_REPOSITORIES="${RENOVATE_REPOSITORIES:-forust/homelab}" \
-e RENOVATE_DRY_RUN="${RENOVATE_DRY_RUN:-}" \
-e RENOVATE_CONFIG_FILE=/opt/renovate/config.js \
-e RENOVATE_CONFIG_FILE=/opt/renovate/renovate.json \
-e RENOVATE_BASE_DIR=/tmp/renovate \
-e LOG_LEVEL="${LOG_LEVEL:-info}" \
renovate/renovate:44.103.0
"${{ steps.image.outputs.image }}"
+59
View File
@@ -0,0 +1,59 @@
#!/usr/bin/env bash
# usage: ssh-run.sh <stage>
# Runs one deploy-lib.sh stage on the workstation over SSH.
set -euo pipefail
: "${DEPLOY_HOST:?missing DEPLOY_HOST}"
: "${DEPLOY_USER:?missing DEPLOY_USER}"
: "${DEPLOY_KEY:?missing DEPLOY_SSH_KEY}"
deploy_port="${DEPLOY_PORT:-22}"
deploy_path="${DEPLOY_PATH:-/srv/homelab}"
deploy_path="$(printf '%s' "$deploy_path" | tr -d '\"' | tr -d '\r' | xargs)"
# The private key is written to a per-run directory that is removed on exit, so a
# failed or cancelled job cannot leave deploy credentials in the runner's temp
# directory. Do not use a fixed path: apply-k8s and apply-compose run in parallel.
key_dir="$(mktemp -d "${RUNNER_TEMP:-/tmp}/homelab-deploy-key.XXXXXXXX")"
trap 'rm -rf "$key_dir"' EXIT INT TERM
ssh_key="$key_dir/deploy_key"
printf '%s\n' "$DEPLOY_KEY" > "$ssh_key"
chmod 600 "$ssh_key"
# A connection that died silently used to hang until the job timeout, and the
# stage was never re-run: one flaky TCP session cost a whole 45-minute apply.
# ServerAlive* bounds how long a dead peer goes unnoticed, ConnectTimeout bounds
# setup. Only exit 255 - ssh's own transport failures - is retried. A stage that
# fails on its own merits exits with the remote's status, so a real failure
# still surfaces its own log instead of burning three attempts. The stages are
# declarative applies, so re-running one that had already committed is harmless.
ssh_opts=(
-i "$ssh_key" -p "$deploy_port"
-o BatchMode=yes -o StrictHostKeyChecking=accept-new
-o ConnectTimeout=15
-o ServerAliveInterval=15 -o ServerAliveCountMax=4
)
rc=0
for attempt in 1 2 3; do
if [ "$attempt" -gt 1 ]; then
echo ":: warning::ssh transport failed, retrying (${attempt}/3)"
sleep $((attempt * 5))
fi
rc=0
ssh "${ssh_opts[@]}" "${DEPLOY_USER}@${DEPLOY_HOST}" \
env "REPO=$deploy_path" "APPLY_PRUNE=${APPLY_PRUNE:-false}" \
"DEPLOY_SHA=${DEPLOY_SHA:-}" "DEPLOY_SNAPSHOT_DIR=${DEPLOY_SNAPSHOT_DIR:-}" \
"STAGE=$1" bash -se <<'EOF' || rc=$?
source "$REPO/.gitea/workflows/deploy-lib.sh"
run_stage "$STAGE"
EOF
[ "$rc" -eq 0 ] && break
[ "$rc" -ne 255 ] && break
done
if [ "$rc" -ne 0 ]; then
echo ":: error::stage $1 failed over ssh (exit $rc)"
fi
exit "$rc"
+55
View File
@@ -0,0 +1,55 @@
#!/usr/bin/env bash
# Regenerates renovate/k8s/configmap.yaml from renovate/renovate.json.
#
# renovate/renovate.json is the single source of truth: the CronJob, the Compose
# file and the renovate-run workflow all mount that exact file. A ConfigMap cannot
# read a file from the repository, so the same bytes are inlined here as a literal
# block. This script keeps the copy honest:
#
# .gitea/workflows/sync-renovate-configmap.sh # rewrite in place
# .gitea/workflows/sync-renovate-configmap.sh --check # fail if out of date
#
# renovate-ci runs the --check form on every PR and push, so a config change that
# forgets to regenerate the ConfigMap cannot be merged.
set -euo pipefail
here="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
repo="$(git -C "$here" rev-parse --show-toplevel)"
src="$repo/renovate/renovate.json"
dst="$repo/renovate/k8s/configmap.yaml"
[ -f "$src" ] || {
echo "missing $src" >&2
exit 1
}
render() {
cat <<'HEADER'
# GENERATED FILE - do not edit by hand.
# Source: renovate/renovate.json
# Regenerate: .gitea/workflows/sync-renovate-configmap.sh
# Verify: .gitea/workflows/sync-renovate-configmap.sh --check
apiVersion: v1
kind: ConfigMap
metadata:
name: renovate-config
namespace: renovate
data:
renovate.json: |
HEADER
sed 's/^/ /' "$src"
}
if [ "${1:-}" = "--check" ]; then
if ! diff -u "$dst" <(render) >/dev/null 2>&1; then
echo "ERROR: $dst is out of sync with renovate/renovate.json"
echo "Run: .gitea/workflows/sync-renovate-configmap.sh"
diff -u "$dst" <(render) || true
exit 1
fi
echo "renovate/k8s/configmap.yaml is in sync with renovate/renovate.json"
exit 0
fi
render >"$dst"
echo "wrote $dst"
+33
View File
@@ -0,0 +1,33 @@
# Pinned versions of the CI tools installed by install-ci-tools.sh.
# Renovate keeps these up to date (see customManagers in renovate/renovate.json).
#
# Every version here except NODE_VERSION matches what was already installed on
# the runner, so pinning them changes what CI does not at all. It changes what
# CI does when the runner is rebuilt with something else: today
# install-ci-tools.sh finds the pinned version already on PATH and installs
# nothing, and a runner that drifts gets the pinned one installed over it.
#
# The renovate image version is NOT pinned here: renovate/k8s/cronjob.yaml is the
# single source of truth and the workflows read the tag from it, so there is
# nothing to drift.
ACTIONLINT_VERSION="1.7.7"
SHELLCHECK_VERSION="0.11.0"
KUBECONFORM_VERSION="0.8.0"
PRETTIER_VERSION="3.8.1"
RUFF_VERSION="0.16.8"
YAMLLINT_VERSION="1.38.0"
HADOLINT_VERSION="2.14.0"
# pip-audit reads the advisory database over the network, so a floating version
# would make the same commit report different things on different days. Pin it
# like the rest: the advisories themselves are the moving part, not the tool.
PIP_AUDIT_VERSION="2.10.1"
# uv builds the throwaway venv the pytest job runs in, and unpacks the PyPI
# wheels for ruff, yamllint and pip-audit.
UV_VERSION="0.12.17"
# node runs `npm ci` for the frontend tests and the npm audit, and it is the one
# pin here that does NOT come from the runner: the runner's system node is a
# rolling Arch package (it was node 26 with no npm at all when this was pinned),
# and the panel image is node:22-alpine. Pinned to the image's major on purpose,
# so the tree that gets tested is the tree that gets built. Renovate keeps this
# in step with the Dockerfile's node: tag via the "node runtime" group.
NODE_VERSION="22.23.3"
+5
View File
@@ -21,6 +21,9 @@ checkmk/checkmk/*
downtify/Downtify_downloads
headscale/config/*
headscale/data/*
# NetBird local hostnames and generated secrets
netbird/.env
netbird/secrets/
searxng/core-config/*
# Steaming services files
@@ -93,6 +96,8 @@ replacements.txt
# Temp files
edu_master/temp/
temp/*
# Local-only tooling scratch space (pinned CI tools, verification scripts)
tmp/
# Environment
.env
+1 -1
View File
@@ -73,7 +73,7 @@ spec:
memory: "1.5Gi"
cpu: "300m"
requests:
memory: "500Mi"
memory: "1Gi"
cpu: "50m"
ports:
- containerPort: 3000
+3 -3
View File
@@ -52,7 +52,7 @@ spec:
- containerPort: 9000
resources:
requests:
memory: "700Mi"
memory: "768Mi"
cpu: "300m"
limits:
memory: "1.5Gi"
@@ -86,8 +86,8 @@ spec:
name: authentik-secrets
resources:
requests:
memory: "512Mi"
memory: "320Mi"
cpu: "300m"
limits:
memory: "1Gi"
memory: "768Mi"
cpu: "700m"
+2 -2
View File
@@ -22,10 +22,10 @@ spec:
imagePullPolicy: Always
resources:
requests:
memory: "20Mi"
memory: "32Mi"
cpu: "30m"
limits:
memory: "64Mi"
memory: "128Mi"
cpu: "50m"
envFrom:
- secretRef:
+3 -3
View File
@@ -16,7 +16,7 @@ spec:
spec:
containers:
- name: cloudflared
image: cloudflare/cloudflared:2026.9.1
image: cloudflare/cloudflared:2026.9.3
imagePullPolicy: IfNotPresent
args:
- tunnel
@@ -30,8 +30,8 @@ spec:
key: TUNNEL_TOKEN
resources:
requests:
memory: "32Mi"
memory: "128Mi"
cpu: "30m"
limits:
memory: "128Mi"
memory: "256Mi"
cpu: "200m"
+1 -1
View File
@@ -54,7 +54,7 @@ services:
- "traefik.http.routers.bentopdf.tls.certresolver=letsencrypt"
- "traefik.http.routers.bentopdf.tls=true"
# Local router
- "traefik.http.routers.bentopdf-local.rule=Host(`pdf.wokstation.internal`)"
- "traefik.http.routers.bentopdf-local.rule=Host(`pdf.workstation.internal`)"
- "traefik.http.routers.bentopdf-local.entrypoints=websecure"
- "traefik.http.routers.bentopdf-local.tls=true"
# Dev router
+3 -2
View File
@@ -31,12 +31,13 @@ spec:
name: bentopdf
ports:
- containerPort: 8080
# p95 4M, max 11M over 7 days. Was 50Mi/700Mi.
resources:
requests:
memory: "50Mi"
memory: "32Mi"
cpu: "50m"
ephemeral-storage: "100Mi"
limits:
memory: "700Mi"
memory: "128Mi"
cpu: "700m"
ephemeral-storage: "5Gi"
+3 -2
View File
@@ -38,13 +38,14 @@ spec:
volumeMounts:
- mountPath: /data
name: data
# p95 85M, max 136M over 7 days, spikes while converting. Was 250Mi/1.5Gi.
resources:
requests:
memory: "250Mi"
memory: "128Mi"
cpu: "100m"
limits:
cpu: "1500m"
memory: "1.5Gi"
memory: "512Mi"
volumes:
- name: data
persistentVolumeClaim:
+56 -1
View File
@@ -1,3 +1,17 @@
# crowdsec/k8s is NOT managed by deploy.yaml - apply this by hand, and apply it
# together with a restart:
# kubectl apply -f crowdsec/k8s/crowdsec-middleware.yaml
# kubectl -n traefik rollout restart deploy/traefik
#
# The restart is not optional. In stream mode the plugin runs a package-level
# ticker goroutine (handleStreamTicker over the isCrowdsecStreamHealthy and
# updateFailure globals) that no reconfiguration stops. Applying a change
# wedges the instance: every route referencing it answers 404 and traefik logs
# 'invalid middleware crowdsec-crowdsec-bouncer@kubernetescrd' until the pod is
# replaced. Re-applying the previous config does NOT recover it, and the config
# is not the cause - a valid CIDR cannot fail NewChecker, which is a plain
# net.ParseCIDR. Only a new pod clears it. Measured cost: ~35s down for all
# 20 hosts behind this middleware.
apiVersion: traefik.io/v1alpha1
kind: Middleware
metadata:
@@ -8,7 +22,48 @@ spec:
crowdsec-bouncer:
enabled: true
LogLevel: INFO
CrowdsecMode: live
# `live` blocked on a `GET /v1/decisions` per request, so a burst
# saturated the LAPI and the plugin 403'd IPs that were never banned.
# v1.3.3 ignores UpdateMaxFailure in `live`, so fail-open is only
# reachable in stream mode, which polls into a cache instead - no
# per-request call to saturate. 15s rather than the 60s default: the
# deploy runner shares one public IP with the house, so this bounds
# both how late a ban lands and how long a lifted one lingers.
CrowdsecMode: stream
UpdateIntervalSeconds: 15
# -1 = never block because the LAPI is unreachable. In v1.3.3
# handleStreamTicker only clears isCrowdsecStreamHealthy when
# updateMaxFailure != -1, and ServeHTTP 403s once it is false, so this
# makes a CrowdSec outage mean "no protection", not "every site 403".
UpdateMaxFailure: -1
CrowdsecLapiScheme: http
CrowdsecLapiHost: crowdsec-service.crowdsec.svc.cluster.local:8080
CrowdsecLapiKeyFile: "/etc/traefik/secrets/traefik-api-key"
# Bypasses the bouncer and the decision cache, no LAPI round-trip.
# Keep in sync with forust/local-network in crowdsec-values.yaml.
ClientTrustedIPs:
- "127.0.0.0/8"
- "10.0.0.0/8"
- "172.16.0.0/12"
- "192.168.0.0/16"
- "100.64.0.0/10"
- "169.254.0.0/16"
- "fc00::/7"
- "fe80::/10"
# The mobile operator range from forust/mobile-whitelist, repeated
# deliberately rather than relying on the parser whitelist alone.
# That whitelist drops the event before it reaches a bucket, so no
# decision is ever created - but it is one config away from not
# firing, and the bouncer would then enforce a ban that was never
# justified. This is the last line: even a decision that exists for
# any reason is not served against the phone.
- "84.245.64.0/18"
# The name is HTTPTimeoutSeconds, an int in seconds (min 1) - there is
# no CrowdsecLapiTimeout, and an unrecognised key is silently dropped,
# which is how this sat at the 10s default. Nothing rides on it per
# request any more, so this only bounds the stream pull - and too low
# is the dangerous direction: the LAPI needs ~2s to answer
# /v1/decisions/stream, and a pull that times out leaves the ban cache
# frozen at its startup contents ("failed sending new decisions"),
# i.e. new bans silently never apply. Keep it above the pull latency.
HTTPTimeoutSeconds: 10
+92 -4
View File
@@ -58,9 +58,37 @@ config:
reason: "Mobile IP whitelist"
cidr:
- "84.245.64.0/18"
# CrowdSec's own guidance: CIDR allowlisting belongs at the parser stage.
# A parser whitelist discards the event before it reaches a bucket, so
# these addresses never produce an overflow and never become a decision.
# A postoverflow whitelist is checked only *after* the ban exists, and
# the bouncer answers 403 for as long as it does - which is a window we
# do not want the deploy sitting in.
local-network.yaml: |
name: forust/local-network
description: "Whitelist loopback, private and VPN networks"
whitelist:
reason: "Local network"
cidr:
- "127.0.0.0/8"
- "10.0.0.0/8"
- "172.16.0.0/12"
- "192.168.0.0/16"
# CGNAT range (RFC 6598). The workstation and the k0s node live
# here on WireGuard, and 100.64.0.0/10 is not covered by the
# RFC 1918 blocks above.
- "100.64.0.0/10"
- "169.254.0.0/16"
- "fc00::/7"
- "fe80::/10"
postoverflows:
s01-whitelist:
# The one whitelist that has to stay here: resolving a hostname is a
# network call, and the docs put expensive lookups in postoverflows on
# purpose - it runs only when a bucket actually overflows.
# ddns.forust.xyz is the public home address, not a private one, so
# forust/local-network does not cover it.
home-dynamic-ip.yaml: |
name: forust/home-dynamic-ip
description: "Whitelist home dynamic IP"
@@ -69,6 +97,59 @@ config:
expression:
- evt.Overflow.Alert.Source.IP in LookupHost("ddns.forust.xyz")
# LAPI-only main config override, merged over config.yaml. NOTE: the
# chart's own default for this key is REPLACED, not merged, so its
# auto_registration block is repeated verbatim below - drop it and the
# agent can no longer register itself.
config.yaml.local: |
api:
server:
auto_registration: # Activate if not using TLS for authentication
enabled: true
token: "${REGISTRATION_TOKEN}" # /!\ Do not modify this variable (auto-generated and handled by the chart)
allowed_ranges: # /!\ Make sure to adapt to the pod IP ranges used by your cluster
- "127.0.0.1/32"
- "192.168.0.0/16"
- "10.0.0.0/8"
- "172.16.0.0/12"
# This homelab has no egress to console.crowdsec.cloud: DNS does
# not resolve. The LAPI kept trying anyway ("Signal push: N
# signals to push", "capi metrics: sending" every 10s) and each
# attempt sat on a resolver timeout WHILE HOLDING A WRITE
# TRANSACTION, which is what kept stalling per-request decision
# lookups even with WAL enabled. Nothing to share and nothing to
# pull - turn the Central API off instead of letting it block the
# only database writer we have.
online_client:
sharing: false
pull:
community: false
blocklists: false
disable_usage_metrics_export: true
db_config:
# SQLite without WAL serialises every reader behind the writer's
# rollback journal, and the LAPI writes constantly: the agent pushes
# Traefik alerts read from Loki, the metrics collector counts
# decisions, the bouncer touches "last pull" on every request.
# Symptom: decision lookups taking 10-30s (and a second connection
# that could not even open the database) while the LAPI sat at 28m
# CPU - the process was blocked in fsync, not computing. Every
# bouncer-protected request then blew through the plugin timeout and
# fail-closed with 403, on every site at once.
# The PVC is local-path-retain (hostPath), not a network share, so
# WAL is safe here; the crowdsec docs recommend it for exactly this
# ("allowing more concurrency in SQLite that will improve
# performances in most scenarios").
use_wal: true
# Keeps the alert table bounded. At the 5000/7d default the file
# reached 54MB in 15 days off the Traefik access log alone, and the
# metrics collector counts decisions on a timer; a smaller working
# set means fewer full scans. Crowdsec only prunes - SQLite never
# shrinks the file, so the size stays until a manual VACUUM.
flush:
max_items: 1000
max_age: 24h
lapi:
env:
- name: COLLECTIONS
@@ -90,13 +171,20 @@ lapi:
enabled: true
size: 1Gi
storageClassName: local-path-retain
# LAPI answers a blocking /v1/decisions lookup for EVERY bouncer-protected
# request (whole Traefik front door), so it is the hot path of the proxy.
# At 400m/500Mi it went CPU-throttled and idle lookups measured 1.3-7.4s,
# which pushed requests into the bouncer's fail-closed 403.
# Single replica on purpose: LAPI is stateful (BoltDB on the `data` PVC,
# credentials on the `config` PVC) - two replicas sharing those RWO
# volumes would corrupt the decision store. Scale up CPU, not replicas.
resources:
limits:
cpu: 400m
memory: 500Mi
cpu: 1500m
memory: 1Gi
requests:
cpu: 50m
memory: 150Mi
cpu: 250m
memory: 500Mi
service:
type: ClusterIP
storeLAPICscliCredentialsInSecret: true
+7
View File
@@ -32,6 +32,13 @@
# on their own - same name + same password);
# 4. prune bouncer entries idle for 30d.
#
# It used to also delete LePresidente/http-generic-403-bf decisions hourly.
# That was a workaround for the bouncer failing closed on a slow LAPI and
# 403-ing the deploy runner into a 4h ban. The bouncer now polls decisions
# into a cache and never blocks on an unreachable LAPI, so it cannot
# manufacture those 403s any more, and the scenario only fires against real
# scanners - deleting their decisions hourly was undoing a working ban.
#
# Manual apply (crowdsec/k8s is NOT managed by deploy.yaml):
# kubectl apply -f crowdsec/k8s/janitor-cronjob.yaml
# Force a run:
+1 -1
View File
@@ -1,6 +1,6 @@
services:
dockmon:
image: darthnorse/dockmon:2.4.5
image: darthnorse/dockmon:2.5.0
container_name: dockmon
restart: unless-stopped
# ports:
+1 -1
View File
@@ -29,7 +29,7 @@ spec:
spec:
containers:
- name: dockmon
image: darthnorse/dockmon:2.4.5
image: darthnorse/dockmon:2.5.0
ports:
- containerPort: 443
volumeMounts:
+1 -1
View File
@@ -1,7 +1,7 @@
services:
downtify:
container_name: downtify
image: ghcr.io/henriquesebastiao/downtify:2.13.0
image: ghcr.io/henriquesebastiao/downtify:3.1.0
restart: unless-stopped
# ports:
# - '7077:8000'
+1 -1
View File
@@ -27,7 +27,7 @@ spec:
spec:
containers:
- name: downtify
image: ghcr.io/henriquesebastiao/downtify:2.13.0
image: ghcr.io/henriquesebastiao/downtify:3.1.0
ports:
- containerPort: 8000
volumeMounts:
+3 -3
View File
@@ -1,6 +1,6 @@
services:
redis:
image: redis:8.10.1-alpine
image: redis:8.10.2-alpine
restart: unless-stopped
volumes:
- redis-data:/data
@@ -17,7 +17,7 @@ services:
session-keeper:
build: ./phpsessid-bot
image: gcr.forust.xyz/forust/session-keeper:latest
image: gcr.forust.xyz/forust/session-keeper:prod
pull_policy: build
env_file: .env
restart: unless-stopped
@@ -33,7 +33,7 @@ services:
webinar-checker:
build: ./webinar-checker
image: gcr.forust.xyz/forust/webinar-checker:latest
image: gcr.forust.xyz/forust/webinar-checker:prod
pull_policy: build
env_file: .env
restart: unless-stopped
+20 -1
View File
@@ -11,9 +11,14 @@ spec:
rules:
# No successful webinar check for 5m (~2-3 missed 2-min checks).
# Catches: playwright hangs/timeouts, version skew, site changes, hung job.
# The last_success > 0 guard is mandatory: checker.py initialises
# last_success to 0, so without it `time() - 0` equals the current epoch
# and humanizeDuration renders ~20722d on every pod restart. Keep the
# duration expression on the left so $value stays the real gap.
- alert: WebinarCheckerNoSuccessfulCheck
expr: |
(time() - webinar_check_last_success_timestamp_seconds > 300)
((time() - webinar_check_last_success_timestamp_seconds) > 300)
and (webinar_check_last_success_timestamp_seconds > 0)
and (webinar_check_last_run_timestamp_seconds > 0)
for: 2m
labels:
@@ -22,6 +27,20 @@ spec:
summary: "Webinar checker has no successful check for 5m"
description: "edu-master/webinar-checker: last successful webinar check was {{ $value | humanizeDuration }} ago. Checks are failing or hanging (see consecutive failures alert). Notifications about new webinars are NOT being sent."
# Checks are running but none has ever succeeded since pod start.
# Split out from the rule above so a zeroed gauge never feeds
# humanizeDuration.
- alert: WebinarCheckerNeverSucceeded
expr: |
(webinar_check_last_success_timestamp_seconds == 0)
and (webinar_check_last_run_timestamp_seconds > 0)
for: 10m
labels:
severity: critical
annotations:
summary: "Webinar checker has never completed a successful check"
description: 'edu-master/webinar-checker: checks have been running for 10m but not one has ever succeeded since the pod started, so every check is failing. Check pod logs (Loki: {namespace="edu-master", container="webinar-checker"}).'
# Fast path: 3 consecutive failures (~6+ min at 2-min interval).
- alert: WebinarCheckerConsecutiveFailures
expr: |
+9
View File
@@ -20,6 +20,15 @@ spec:
# renovate: datasource=docker depName=mcr.microsoft.com/playwright versioning=docker
image: mcr.microsoft.com/playwright:v1.56.0-jammy
imagePullPolicy: IfNotPresent
# p95 412M, max 478M over 7 days, no limit before. Request is set at p95
# so the pod is not an eviction candidate; the limit stays above 2x the
# request because browser page lifetimes are unpredictable.
resources:
requests:
cpu: "200m"
memory: "416Mi"
limits:
memory: "1Gi"
command:
- npx
- -y
+3 -3
View File
@@ -18,7 +18,7 @@ spec:
spec:
containers:
- name: redis
image: redis:8.10.1-alpine
image: redis:8.10.2-alpine
imagePullPolicy: IfNotPresent
ports:
- containerPort: 6379
@@ -28,10 +28,10 @@ spec:
resources:
requests:
cpu: 25m
memory: 64Mi
memory: 32Mi
limits:
cpu: 250m
memory: 256Mi
memory: 128Mi
readinessProbe:
exec:
command: ["redis-cli", "ping"]
+4 -5
View File
@@ -17,7 +17,7 @@ spec:
spec:
initContainers:
- name: wait-redis
image: redis:8.10.1-alpine
image: redis:8.10.2-alpine
command:
- /bin/sh
- -ec
@@ -31,18 +31,17 @@ spec:
echo "redis is ready"
containers:
- name: session-keeper
image: gcr.forust.xyz/forust/session-keeper:latest
imagePullPolicy: Always
image: gcr.forust.xyz/forust/session-keeper:prod
envFrom:
- secretRef:
name: edu-master-secrets
resources:
requests:
cpu: 25m
memory: 96Mi
memory: 32Mi
limits:
cpu: 250m
memory: 256Mi
memory: 128Mi
readinessProbe:
exec:
command: ["/bin/sh", "-ec", "redis-cli -h redis EXISTS EDU_PHPSESSID | grep -q 1"]
+12 -5
View File
@@ -19,7 +19,7 @@ spec:
# redis healthy -> session-keeper healthy (EXISTS EDU_PHPSESSID) -> playwright started
initContainers:
- name: wait-deps
image: redis:8.10.1-alpine
image: redis:8.10.2-alpine
command:
- /bin/sh
- -ec
@@ -45,12 +45,19 @@ spec:
echo "playwright ok"
containers:
- name: webinar-checker
image: gcr.forust.xyz/forust/webinar-checker:latest
imagePullPolicy: Always
image: gcr.forust.xyz/forust/webinar-checker:prod
ports:
- name: metrics
containerPort: 8000
protocol: TCP
readinessProbe:
httpGet:
path: /health
port: metrics
periodSeconds: 10
timeoutSeconds: 3
failureThreshold: 12
initialDelaySeconds: 10
envFrom:
- secretRef:
name: edu-master-secrets
@@ -60,7 +67,7 @@ spec:
resources:
requests:
cpu: "50m"
memory: "128Mi"
memory: "192Mi"
limits:
cpu: "600m"
memory: "512Mi"
memory: "384Mi"
+1 -1
View File
@@ -1,7 +1,7 @@
services:
errorpage:
build: .
image: gcr.forust.xyz/forust/error-pages:latest
image: gcr.forust.xyz/forust/error-pages:prod
pull_policy: build
container_name: error-pages
restart: unless-stopped
File renamed without changes.
+15 -1
View File
@@ -27,7 +27,21 @@ spec:
spec:
containers:
- name: error-pages
image: gcr.forust.xyz/forust/error-pages:latest
image: gcr.forust.xyz/forust/error-pages:prod
# p95 6M, max 10M, no limit before.
resources:
requests:
cpu: "10m"
memory: "32Mi"
limits:
memory: "128Mi"
ports:
- containerPort: 80
readinessProbe:
httpGet:
path: /404.html
port: 80
periodSeconds: 10
timeoutSeconds: 2
failureThreshold: 3
---
+6
View File
@@ -17,6 +17,12 @@ data:
GITEA__mailer__ENABLED: "false"
# No code/issue search needed: bleve reindexes the whole issue index on
# every pod restart (cron.rebuild_issue_indexer RUN_AT_START) and hammers
# the rotational disk for an hour. "db" serves issue search from postgres.
GITEA__indexer__ISSUE_INDEXER_TYPE: "db"
GITEA__indexer__REPO_INDEXER_ENABLED: "false"
GITEA__log__logger.access.MODE: "console, file"
USER_UID: "1000"
USER_GID: "1000"
+2 -2
View File
@@ -47,10 +47,10 @@ spec:
mountPath: /data
resources:
requests:
memory: "512Mi"
memory: "320Mi"
cpu: "300m"
limits:
memory: "1.5Gi"
memory: "1Gi"
cpu: "1300m"
volumes:
- name: gitea-data
+7 -3
View File
@@ -15,11 +15,15 @@ spec:
services:
- name: gitea-service
port: 3000
# Registry route: NO crowdsec-bouncer.
# A deploy burst (runner Action API polls, `docker manifest inspect` per
# own image, containerd pulls, smoke probes) fires hundreds of parallel
# registry calls, and a ban on the runner breaks every later job. This
# route only serves authenticated OCI traffic - registry tokens and
# basic-auth are already handled by gitea - and scanners get nothing
# useful from /v2, so there is no bruteforce surface to protect here.
- match: Host(`gcr.forust.xyz`) && PathPrefix(`/v2`)
kind: Rule
middlewares:
- name: crowdsec-bouncer
namespace: crowdsec
services:
- name: gitea-service
port: 3000
+2 -2
View File
@@ -57,10 +57,10 @@ spec:
resources:
requests:
cpu: "50m"
memory: "64Mi"
memory: "32Mi"
limits:
cpu: "200m"
memory: "256Mi"
memory: "128Mi"
volumes:
- name: glance-config
configMap:
File renamed without changes.
+1 -1
View File
@@ -1,6 +1,6 @@
services:
headscale:
image: headscale/headscale:0.29.3
image: headscale/headscale:v0.29.4
restart: unless-stopped
container_name: headscale-server
command: serve
File renamed without changes.
+2 -2
View File
@@ -3,7 +3,7 @@ services:
build:
context: .
dockerfile: Dockerfile.forust
image: gcr.forust.xyz/forust/forust-homepage:latest
image: gcr.forust.xyz/forust/forust-homepage:prod
pull_policy: build
# ports:
# - "8085:80"
@@ -35,7 +35,7 @@ services:
build:
context: .
dockerfile: Dockerfile.xdfnx
image: gcr.forust.xyz/forust/xdfnx-homepage:latest
image: gcr.forust.xyz/forust/xdfnx-homepage:prod
pull_policy: build
restart: unless-stopped
# ports:
+20 -8
View File
@@ -27,16 +27,22 @@ spec:
spec:
containers:
- name: forust-homepage
image: gcr.forust.xyz/forust/forust-homepage:latest
imagePullPolicy: Always
image: gcr.forust.xyz/forust/forust-homepage:prod
ports:
- containerPort: 80
readinessProbe:
httpGet:
path: /
port: 80
periodSeconds: 10
timeoutSeconds: 2
failureThreshold: 3
resources:
requests:
memory: "10Mi"
memory: "32Mi"
cpu: "20m"
limits:
memory: "100Mi"
memory: "128Mi"
cpu: "50m"
---
apiVersion: v1
@@ -68,14 +74,20 @@ spec:
spec:
containers:
- name: xdfnx-homepage
image: gcr.forust.xyz/forust/xdfnx-homepage:latest
imagePullPolicy: Always
image: gcr.forust.xyz/forust/xdfnx-homepage:prod
ports:
- containerPort: 80
readinessProbe:
httpGet:
path: /
port: 80
periodSeconds: 10
timeoutSeconds: 2
failureThreshold: 3
resources:
requests:
memory: "10Mi"
memory: "32Mi"
cpu: "20m"
limits:
memory: "100Mi"
memory: "128Mi"
cpu: "50m"
+24
View File
@@ -0,0 +1,24 @@
# You can find documentation for all the supported env variables at https://docs.immich.app/install/environment-variables
# The location where your uploaded files are stored. The k8s manifests bind
# mount /mnt/immich/library, which is the sdc9 partition - the same place, so
# the two deployment paths look at one library.
UPLOAD_LOCATION=/mnt/immich/library
# The location where your database files are stored. Network shares are not supported for the database
DB_DATA_LOCATION=./postgres
# To set a timezone, uncomment the next line and change Etc/UTC to a TZ identifier from this list: https://en.wikipedia.org/wiki/List_of_tz_database_time_zones#List
# TZ=Etc/UTC
# The Immich version to use. You can pin this to a specific version like "v2.1.0"
IMMICH_VERSION=v3
# Connection secret for postgres. You should change it to a random password
# Please use only the characters `A-Za-z0-9`, without special characters or spaces
DB_PASSWORD=postgres
# The values below this line do not need to be changed
###################################################################################
DB_USERNAME=postgres
DB_DATABASE_NAME=immich
+63
View File
@@ -0,0 +1,63 @@
name: immich
services:
immich-server:
container_name: immich_server
image: ghcr.io/immich-app/immich-server:v3
volumes:
- ${UPLOAD_LOCATION}:/data
- /etc/localtime:/etc/localtime:ro
env_file:
- .env
ports:
- "2283:2283"
depends_on:
- redis
- database
restart: always
healthcheck:
disable: false
immich-machine-learning:
container_name: immich_machine_learning
# For hardware acceleration, add one of -[armnn, cuda, rocm, openvino, rknn] to the image tag.
# Example tag: ${IMMICH_VERSION:-release}-cuda
image: ghcr.io/immich-app/immich-machine-learning:${IMMICH_VERSION:-release}
# extends: # uncomment this section for hardware acceleration - see https://docs.immich.app/features/ml-hardware-acceleration
# file: hwaccel.ml.yml
# service: cpu # set to one of [armnn, cuda, rocm, openvino, openvino-wsl, rknn] for accelerated inference - use the `-wsl` version for WSL2 where applicable
volumes:
- model-cache:/cache
env_file:
- .env
restart: always
healthcheck:
disable: false
redis:
container_name: immich_redis
image: docker.io/valkey/valkey:9@sha256:418652cfb58ef879d4978c33553735d7147016032d5aefaa14c828e611eb9dfd
healthcheck:
test: redis-cli ping | grep -q PONG || exit 1
restart: always
database:
container_name: immich_postgres
image: ghcr.io/immich-app/postgres:14-vectorchord0.4.3-pgvectors0.2.0@sha256:bcf63357191b76a916ae5eb93464d65c07511da41e3bf7a8416db519b40b1c23
environment:
POSTGRES_PASSWORD: ${DB_PASSWORD}
POSTGRES_USER: ${DB_USERNAME}
POSTGRES_DB: ${DB_DATABASE_NAME}
POSTGRES_INITDB_ARGS: "--data-checksums"
# Uncomment the DB_STORAGE_TYPE: 'HDD' var if your database isn't stored on SSDs
# DB_STORAGE_TYPE: 'HDD'
volumes:
# Do not edit the next line. If you want to change the database storage location on your system, edit the value of DB_DATA_LOCATION in the .env file
- ${DB_DATA_LOCATION}:/var/lib/postgresql/data
shm_size: 128mb
restart: always
healthcheck:
disable: false
volumes:
model-cache:
File renamed without changes.
+28
View File
@@ -0,0 +1,28 @@
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: immich-prod-tls
namespace: immich
spec:
secretName: immich-prod-tls
dnsNames:
- immich.forust.xyz
issuerRef:
name: letsencrypt-prod
kind: ClusterIssuer
---
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: internal-wildcard-tls
namespace: immich
spec:
secretName: internal-wildcard-tls
dnsNames:
- "*.workstation.internal"
- "*.gigaforust.internal"
- workstation.internal
- gigaforust.internal
issuerRef:
name: internal-ca
kind: ClusterIssuer
+24
View File
@@ -0,0 +1,24 @@
apiVersion: v1
kind: ConfigMap
metadata:
name: immich-config
namespace: immich
data:
TZ: "Europe/Bratislava"
# The database in this namespace, not the shared one in the database
# namespace: v3 needs VectorChord, and only the dedicated image carries it.
DB_HOSTNAME: "immich-postgres"
DB_PORT: "5432"
DB_USERNAME: "immich"
DB_DATABASE_NAME: "immich"
DB_SSL_MODE: "disable"
DB_VECTOR_EXTENSION: "vectorchord"
REDIS_HOSTNAME: "immich-valkey"
REDIS_PORT: "6379"
# Traefik is the only client of the server, and it is a pod: the address immich
# sees is inside the node's pod CIDR. Without this the server does not trust
# X-Forwarded-For and every request looks like it came from Traefik itself.
IMMICH_TRUSTED_PROXIES: "10.244.0.0/24"
+95
View File
@@ -0,0 +1,95 @@
apiVersion: v1
kind: Service
metadata:
name: immich-service
namespace: immich
spec:
selector:
app: immich
ports:
- name: http
port: 2283
targetPort: 2283
---
apiVersion: apps/v1
kind: Deployment
metadata:
name: immich-deployment
namespace: immich
labels:
app: immich
spec:
replicas: 2
selector:
matchLabels:
app: immich
template:
metadata:
labels:
app: immich
spec:
containers:
- name: immich
image: ghcr.io/immich-app/immich-server:v3
envFrom:
- configMapRef:
name: immich-config
- secretRef:
name: immich-secrets
ports:
- name: http
containerPort: 2283
volumeMounts:
- name: immich-data
mountPath: /data
# The first boot runs migrations and warms the transcoder, which can
# take minutes, so liveness has to wait on the startup probe.
startupProbe:
httpGet:
path: /api/server/ping
port: http
failureThreshold: 60
periodSeconds: 10
timeoutSeconds: 5
readinessProbe:
httpGet:
path: /api/server/ping
port: http
periodSeconds: 10
timeoutSeconds: 5
livenessProbe:
httpGet:
path: /api/server/ping
port: http
initialDelaySeconds: 30
periodSeconds: 30
timeoutSeconds: 5
# Only the request is scheduled against, and the node is already
# oversubscribed (5.58 of 6 cores requested) while actually running
# at about 1.5. So the request states what this sits at while idle -
# tens of millicores - and the limit leaves room for the burst that
# matters: thumbnails, transcodes and metadata extraction.
#
# The limit used to be 2Gi, but the server OOMKilled on boot while
# chewing through a backlog of unprocessed assets (API +Workers in
# one container spike well past idle).
resources:
requests:
cpu: "100m"
memory: "512Mi"
limits:
cpu: "1500m"
memory: "4Gi"
volumes:
- name: immich-data
# The library lives on the node's own disk, not in a PVC. A PVC here
# meant declaring a size up front for data that does not exist yet,
# on a provisioner that cannot grow it, and the only copy of the
# photos was one `kubectl delete namespace` away.
#
# Directory, not DirectoryOrCreate, on purpose: if sdc9 is not
# mounted, this must fail loudly instead of quietly writing the
# library onto the root filesystem.
hostPath:
path: /mnt/immich/library
type: Directory
+36
View File
@@ -0,0 +1,36 @@
apiVersion: traefik.io/v1alpha1
kind: IngressRoute
metadata:
name: immich-prod
namespace: immich
spec:
entryPoints:
- websecure
routes:
- match: Host(`immich.forust.xyz`)
kind: Rule
middlewares:
- name: crowdsec-bouncer
namespace: crowdsec
services:
- name: immich-service
port: 2283
tls:
secretName: immich-prod-tls
---
apiVersion: traefik.io/v1alpha1
kind: IngressRoute
metadata:
name: immich-local
namespace: immich
spec:
entryPoints:
- websecure
routes:
- match: Host(`immich.workstation.internal`) || Host(`immich.gigaforust.internal`)
kind: Rule
services:
- name: immich-service
port: 2283
tls:
secretName: internal-wildcard-tls
+92
View File
@@ -0,0 +1,92 @@
apiVersion: v1
kind: Service
metadata:
name: immich-machine-learning
namespace: immich
spec:
selector:
app: immich-machine-learning
ports:
- name: http
port: 3003
targetPort: 3003
---
apiVersion: apps/v1
kind: Deployment
metadata:
name: immich-machine-learning-deployment
namespace: immich
labels:
app: immich-machine-learning
spec:
replicas: 1
selector:
matchLabels:
app: immich-machine-learning
template:
metadata:
labels:
app: immich-machine-learning
spec:
containers:
- name: immich-machine-learning
image: ghcr.io/immich-app/immich-machine-learning:v3
envFrom:
- configMapRef:
name: immich-config
- secretRef:
name: immich-secrets
ports:
- name: http
containerPort: 3003
volumeMounts:
- name: model-cache
mountPath: /cache
# The first request pulls a model over the internet, so a cold start
# is slower than a container start.
startupProbe:
httpGet:
path: /ping
port: http
failureThreshold: 60
periodSeconds: 5
timeoutSeconds: 5
readinessProbe:
httpGet:
path: /ping
port: http
periodSeconds: 10
timeoutSeconds: 5
livenessProbe:
httpGet:
path: /ping
port: http
initialDelaySeconds: 30
periodSeconds: 30
timeoutSeconds: 5
# Same reasoning as the server: the request covers the idle cost
# only, because the node has no spare cores to schedule against.
# Recognition is the burst - a busy import wants both cores.
resources:
requests:
cpu: "100m"
memory: "1Gi"
limits:
cpu: "2000m"
memory: "3Gi"
volumes:
- name: model-cache
persistentVolumeClaim:
claimName: immich-model-cache-pvc
---
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: immich-model-cache-pvc
namespace: immich
spec:
accessModes:
- ReadWriteOnce
resources:
requests:
storage: 2Gi
+5
View File
@@ -0,0 +1,5 @@
# yaml-language-server: $schema=kubernetes
apiVersion: v1
kind: Namespace
metadata:
name: immich
+138
View File
@@ -0,0 +1,138 @@
# Immich's own database, separate from the shared postgres in the database
# namespace. It has to be separate: v3 checks the VectorChord version at startup
# and refuses to boot without it, VectorChord needs its .so in
# shared_preload_libraries, and that can only be read when postmaster starts.
# So the shared instance would have to be rebuilt on a custom image carrying
# vchord and restarted - for every consumer of it (authentik, gitea, netbox,
# netronome, penpot, statuspage). Not worth it for one photo library.
apiVersion: v1
kind: Service
metadata:
name: immich-postgres
namespace: immich
labels:
app: immich-postgres
spec:
selector:
app: immich-postgres
ports:
- name: postgres
port: 5432
targetPort: postgres
---
apiVersion: apps/v1
kind: StatefulSet
metadata:
name: immich-postgres
namespace: immich
labels:
app: immich-postgres
spec:
serviceName: immich-postgres
replicas: 1
selector:
matchLabels:
app: immich-postgres
template:
metadata:
labels:
app: immich-postgres
spec:
containers:
- name: postgres
# v3.x expects vchord for its vector work and vectors (pgvecto.rs)
# for some index types. This image ships both and preloads them, plus
# its own shared_buffers and wal settings, through
# /etc/postgresql/postgresql.conf - which its entrypoint reaches via
# `postgres -c config_file=...` in the image CMD.
#
# So there is deliberately no `command:` here. Overriding it replaces
# that config_file, and it also loses the step where the entrypoint
# drops from root to the postgres user: postmaster then starts as
# root and refuses to run.
image: ghcr.io/immich-app/postgres:14-vectorchord0.4.3-pgvectors0.2.0@sha256:bcf63357191b76a916ae5eb93464d65c07511da41e3bf7a8416db519b40b1c23
env:
- name: POSTGRES_USER
value: immich
- name: POSTGRES_DB
value: immich
- name: POSTGRES_PASSWORD
valueFrom:
secretKeyRef:
name: immich-secrets
key: DB_PASSWORD
# Only read when the data directory is empty, so the checksums are
# decided here and never again.
- name: POSTGRES_INITDB_ARGS
value: --data-checksums
# The postgres-data volume lives on sdc, which is rotational. The
# SSD template is the default; HDD only changes the planner costs
# (effective_io_concurrency, random_page_cost), nothing structural.
- name: DB_STORAGE_TYPE
value: HDD
- name: TZ
valueFrom:
configMapKeyRef:
name: immich-config
key: TZ
ports:
- name: postgres
containerPort: 5432
volumeMounts:
- name: postgres-data
mountPath: /var/lib/postgresql/data
# The upstream compose file asks docker for 128mb of shm. Kubernetes
# gives every container 64mb, which is not what postmaster expects
# for parallel query workers and the WAL writer.
- name: shm
mountPath: /dev/shm
# Probes use a generous timeout on purpose: the data lives on a
# rotational disk on a loaded single node, and pg_isready can take
# seconds during WAL recovery. A 1s timeout kills the container
# mid-recovery and restarts the spiral.
startupProbe:
exec:
command: ["sh", "-c", "pg_isready -U immich -d immich"]
failureThreshold: 60
periodSeconds: 5
timeoutSeconds: 5
readinessProbe:
exec:
command: ["sh", "-c", "pg_isready -U immich -d immich"]
periodSeconds: 10
timeoutSeconds: 5
livenessProbe:
exec:
command: ["sh", "-c", "pg_isready -U immich -d immich"]
initialDelaySeconds: 30
periodSeconds: 20
timeoutSeconds: 5
# The image template sets shared_buffers to 512MB, and the vchord and
# vectors workers are Rust binaries with a real RSS footprint on top
# of postmaster, checkpointer and friends. 1Gi was enough to start
# the server but the vectors worker kept dying in it, so the limit
# sits at 2Gi. The request stays at the idle cost.
resources:
requests:
cpu: "50m"
memory: "256Mi"
limits:
cpu: "1000m"
memory: "2Gi"
volumes:
- name: shm
emptyDir:
medium: Memory
sizeLimit: 128Mi
volumeClaimTemplates:
- metadata:
name: postgres-data
spec:
accessModes: ["ReadWriteOnce"]
# Retain: this is the metadata for a library that only exists in one
# place, and local-path cannot expand a bound volume, so this size has
# to hold until the library is rebuilt or dumped elsewhere.
storageClassName: local-path-retain
resources:
requests:
storage: 32Gi
+13
View File
@@ -0,0 +1,13 @@
apiVersion: v1
kind: Secret
metadata:
name: immich-secrets
namespace: immich
type: Opaque
stringData:
# Creates the immich superuser in this namespace's own postgres on first
# boot, and is the same value the server connects with. Nothing outside the
# immich namespace needs it. Letters and digits only: immich reads this into
# a connection string.
DB_PASSWORD: "changeme"
REDIS_PASSWORD: "changeme"
+84
View File
@@ -0,0 +1,84 @@
apiVersion: v1
kind: Service
metadata:
name: immich-valkey
namespace: immich
labels:
app: immich-valkey
spec:
clusterIP: None
selector:
app: immich-valkey
ports:
- name: valkey
port: 6379
targetPort: valkey
---
apiVersion: apps/v1
kind: StatefulSet
metadata:
name: immich-valkey
namespace: immich
labels:
app: immich-valkey
spec:
serviceName: immich-valkey
replicas: 1
selector:
matchLabels:
app: immich-valkey
template:
metadata:
labels:
app: immich-valkey
spec:
containers:
- name: valkey
image: docker.io/valkey/valkey:9.1.2-alpine
command:
- sh
- -c
- valkey-server --appendonly yes --save 30 1 --loglevel warning --requirepass "$REDIS_PASSWORD"
envFrom:
- secretRef:
name: immich-secrets
ports:
- name: valkey
containerPort: 6379
volumeMounts:
- name: valkey-data
mountPath: /data
# Same reasoning as postgres: 1s probe timeouts flap on a loaded
# single node with rotational storage.
startupProbe:
exec:
command: ["sh", "-c", 'valkey-cli --pass "$REDIS_PASSWORD" ping | grep -q PONG']
failureThreshold: 20
periodSeconds: 5
timeoutSeconds: 5
readinessProbe:
exec:
command: ["sh", "-c", 'valkey-cli --pass "$REDIS_PASSWORD" ping | grep -q PONG']
periodSeconds: 10
timeoutSeconds: 5
livenessProbe:
exec:
command: ["sh", "-c", 'valkey-cli --pass "$REDIS_PASSWORD" ping | grep -q PONG']
initialDelaySeconds: 20
periodSeconds: 20
timeoutSeconds: 5
resources:
requests:
cpu: "25m"
memory: "64Mi"
limits:
cpu: "250m"
memory: "256Mi"
volumeClaimTemplates:
- metadata:
name: valkey-data
spec:
accessModes: ["ReadWriteOnce"]
resources:
requests:
storage: 1Gi
+1 -1
View File
@@ -36,7 +36,7 @@ services:
- proxy
- kener
redis:
image: redis:8.10.1-alpine
image: redis:8.10.2-alpine
container_name: kener-redis
restart: unless-stopped
volumes:
+1 -1
View File
@@ -30,7 +30,7 @@ spec:
spec:
containers:
- name: redis
image: redis:8.10.1-alpine
image: redis:8.10.2-alpine
ports:
- containerPort: 6379
volumeMounts:
+18 -4
View File
@@ -7,18 +7,32 @@
controller:
type: daemonset
# config-reloader sidecar: p95 33M, max 43M. The chart keeps it at the top level,
# not under `alloy:`.
configReloader:
resources:
requests:
memory: "128Mi"
cpu: "50m"
memory: "32Mi"
cpu: "10m"
limits:
memory: "512Mi"
cpu: "500m"
memory: "128Mi"
image:
tag: "v1.19.2"
alloy:
# p95 275M, max 287M. Alloy tails every pod log and ships it to Loki, so it sits
# on the same IronWolf read path the node is I/O bound on. Request is set at p95.
# The chart key is `alloy.resources`. `controller.resources` is ignored silently,
# which is why this pod shipped with no limits at all.
resources:
requests:
memory: "288Mi"
cpu: "50m"
limits:
memory: "512Mi"
configMap:
create: true
content: |
+27 -4
View File
@@ -39,6 +39,16 @@ loki:
local:
directory: /var/loki/rules
# p95 84M, max 85M for the rules sidecar that shares the singleBinary pod.
# The chart exposes it as `sidecar.resources`, shared with any other sidecar.
sidecar:
resources:
requests:
memory: "96Mi"
cpu: "10m"
limits:
memory: "192Mi"
singleBinary:
replicas: 1
persistence:
@@ -47,10 +57,10 @@ singleBinary:
storageClass: local-path-retain
resources:
requests:
memory: "512Mi"
memory: "256Mi"
cpu: "200m"
limits:
memory: "2Gi"
memory: "1Gi"
cpu: "1000m"
# Zeroed: unused in SingleBinary mode (chart validation requires it).
@@ -63,12 +73,25 @@ backend:
gateway:
replicas: 1
# Single node: chart default is required podAntiAffinity on hostname +
# RollingUpdate 25%/25% (effective maxUnavailable=0 at replicas=1).
# That deadlocks the rollout: the new pod stays Unschedulable while the
# old one lives, and the old one never leaves while the new one is not
# Ready. Null clears the default (an empty map would deep-merge with it
# and keep the required rule); maxUnavailable=1 allows a brief gateway
# outage during rollouts instead of a stuck deploy.
affinity: null
deploymentStrategy:
type: RollingUpdate
rollingUpdate:
maxSurge: 1
maxUnavailable: 1
resources:
requests:
memory: "64Mi"
memory: "32Mi"
cpu: "50m"
limits:
memory: "256Mi"
memory: "128Mi"
cpu: "300m"
monitoring:
+1 -1
View File
@@ -1,6 +1,6 @@
services:
metube:
image: ghcr.io/alexta69/metube:2026.09.15
image: ghcr.io/alexta69/metube:2026.09.27
container_name: metube
restart: unless-stopped
# ports:
+4 -3
View File
@@ -27,7 +27,7 @@ spec:
spec:
containers:
- name: metube
image: ghcr.io/alexta69/metube:2026.09.15
image: ghcr.io/alexta69/metube:2026.09.27
envFrom:
- configMapRef:
name: metube-config
@@ -36,12 +36,13 @@ spec:
volumeMounts:
- name: downloads
mountPath: /downloads
# p95 72M, max 80M over 7 days. Was 600Mi/2Gi.
resources:
requests:
memory: "600Mi"
memory: "96Mi"
cpu: "400m"
limits:
memory: "2Gi"
memory: "384Mi"
cpu: "1700m"
volumes:
- name: downloads
+1 -1
View File
@@ -1,6 +1,6 @@
services:
n8n:
image: docker.n8n.io/n8nio/n8n:2.40.3
image: docker.n8n.io/n8nio/n8n:2.41.3
container_name: n8n
restart: unless-stopped
environment:
+1 -1
View File
@@ -27,7 +27,7 @@ spec:
spec:
containers:
- name: n8n
image: docker.n8n.io/n8nio/n8n:2.40.3
image: docker.n8n.io/n8nio/n8n:2.41.3
envFrom:
- configMapRef:
name: n8n-config
+13
View File
@@ -0,0 +1,13 @@
# Public hostname advertised to NetBird clients and used for TLS/OAuth.
NETBIRD_DOMAIN=nb.forust.xyz
# Internal-only aliases routed by the existing Traefik instance.
NETBIRD_LOCAL_DOMAIN=netbird.workstation.internal
NETBIRD_DEV_DOMAIN=netbird.gigaforust.internal
NETBIRD_PROXY_SUBNET=auto
NETBIRD_CLIENT_HOSTNAME=hostname
# Add a dashboard-generated setup key before starting client.compose.yaml.
# NB_SETUP_KEY=
+123
View File
@@ -0,0 +1,123 @@
# NetBird
Self-hosted NetBird with the combined management, signal, relay, and STUN server. The dashboard and server run behind the repository's existing external Traefik instance on the Docker `proxy` network. Only STUN UDP `3478` is published directly.
The deployment uses SQLite for a single-instance homelab server. The persistent `netbird_data` volume and the datastore encryption key are both required to recover the installation.
## Files
- `compose.yaml`: dashboard and combined server; selected by the marker-driven deploy workflow through `active`.
- `config.template.yaml`: non-secret server configuration rendered at startup.
- `entrypoint.sh`: injects Docker secrets into an in-memory runtime configuration.
- `client.compose.yaml`: optional host-network peer using a dashboard-generated setup key.
- `.env`: ignored local hostnames, the detected Traefik Docker-network subnet, and optional client setup key.
- `secrets/`: ignored relay secret and datastore encryption key.
## First deployment
Run these commands on the Docker host before merging the activating branch. The deploy preflight resets tracked files but preserves ignored local state.
```bash
cd /srv/homelab/netbird
./setup.sh
$EDITOR .env
docker compose config --quiet
docker compose up -d
```
Review the values in `.env` before starting. The example public hostname is `netbird.forust.xyz`; change it if a different public domain was selected. `setup.sh` replaces `NETBIRD_PROXY_SUBNET=auto` with the first IPv4 subnet of the external Docker `proxy` network. Keep that value synchronized with the network; set an explicit CIDR instead if the network is managed elsewhere.
`setup.sh` is idempotent and never replaces existing secrets. Do not delete or regenerate `secrets/datastore-encryption-key` after the first successful start unless all encrypted setup keys and API tokens are intentionally being invalidated.
## Network prerequisites
- Point the public hostname directly to the Docker host. Do not proxy UDP `3478` through Cloudflare or another CDN.
- Allow inbound TCP `80`, TCP `443`, and UDP `3478` through the host firewall and upstream router.
- Ensure the external `proxy` Docker network exists and Traefik uses its `websecure` entrypoint and `letsencrypt` resolver. `NETBIRD_PROXY_SUBNET` must describe that network; it is used to trust only forwarded client addresses from Traefik.
- Ensure the internal names in `.env` resolve where the local and development aliases are needed.
- Keep Traefik's `websecure` read timeout disabled for long-lived gRPC and WebSocket sessions. This repository configures `--entrypoints.websecure.transport.respondingTimeouts.readTimeout=0` in `traefik/compose.yaml`.
After startup, verify OIDC discovery through the public TLS endpoint:
```bash
curl -fsS "https://${NETBIRD_DOMAIN}/oauth2/.well-known/openid-configuration"
```
Open `https://${NETBIRD_DOMAIN}` immediately and complete the initial owner setup. Treat the initial setup flow as public until the owner exists.
## Optional host client
The client intentionally lives in a separate Compose project. Normal server deploys use `--remove-orphans`, so keeping the client in the server project would cause it to be removed.
1. Create a reusable or ephemeral setup key in the NetBird dashboard.
2. Put `NB_SETUP_KEY=<key>` in the ignored `netbird/.env` file.
3. Set `NETBIRD_CLIENT_HOSTNAME` to this machine's desired peer name.
4. Start and inspect the client:
```bash
cd /srv/homelab/netbird
docker compose -f client.compose.yaml config --quiet
docker compose -f client.compose.yaml up -d
docker compose -f client.compose.yaml exec netbird-client netbird status
```
The client uses host networking and requires `NET_ADMIN`, `SYS_ADMIN`, `SYS_RESOURCE`, and `/dev/net/tun`. Remove it without affecting the server stack:
```bash
docker compose -f client.compose.yaml down
```
## Operations
Inspect status and logs:
```bash
docker compose ps
docker compose logs --tail=200 netbird-server dashboard
```
Stop or remove containers without deleting data:
```bash
docker compose down
```
Do not add `-v` to `docker compose down`; it would delete the NetBird datastore.
## Backup and restore
Back up both the persistent volume and the ignored secret files. For a consistent SQLite backup, briefly stop the server first and store the resulting archive and `datastore-encryption-key` in an encrypted backup:
```bash
cd /srv/homelab/netbird
mkdir -p backups
docker compose stop netbird-server
docker run --rm \
-v netbird_data:/data:ro \
-v "$PWD/backups:/backup" \
busybox:1.37.0 \
tar -C /data -czf "/backup/netbird-data-$(date -u +%Y%m%dT%H%M%SZ).tar.gz" .
docker compose start netbird-server
```
Also securely back up:
- `secrets/datastore-encryption-key` — required to decrypt stored secrets.
- `secrets/relay-auth-secret` — keeps issued relay credentials valid across restoration.
- `netbird/.env` — optional, but it records the public and internal hostnames.
Test a restore in an isolated Docker host before relying on a backup.
## Upgrade
1. Take and verify a backup.
2. Review NetBird release notes for server, client, and dashboard compatibility.
3. Update the pinned tags in `compose.yaml`; update `client.compose.yaml` separately when deploying the client.
4. Pull and recreate the selected services:
```bash
docker compose pull
docker compose up -d
```
The image tags are intentionally pinned instead of using `latest`, matching this repository's pull-on-deploy policy.
+31
View File
@@ -0,0 +1,31 @@
name: netbird-client
services:
netbird-client:
image: netbirdio/netbird:0.79.0
container_name: netbird-client
hostname: "${NETBIRD_CLIENT_HOSTNAME:?Set NETBIRD_CLIENT_HOSTNAME in netbird/.env}"
restart: unless-stopped
cap_add:
- NET_ADMIN
- SYS_ADMIN
- SYS_RESOURCE
devices:
- /dev/net/tun
network_mode: host
environment:
NB_SETUP_KEY: "${NB_SETUP_KEY:?Set NB_SETUP_KEY in netbird/.env after creating a peer setup key}"
NB_MANAGEMENT_URL: "https://${NETBIRD_DOMAIN:?Set NETBIRD_DOMAIN in netbird/.env}"
volumes:
- netbird-client:/var/lib/netbird
healthcheck:
test: ["CMD", "/usr/local/bin/netbird", "status", "--check", "live"]
interval: 30s
timeout: 5s
retries: 5
start_period: 30s
stop_grace_period: 30s
volumes:
netbird-client:
name: netbird-client
+152
View File
@@ -0,0 +1,152 @@
name: netbird
services:
netbird-server:
image: netbirdio/netbird-server:0.79.0
container_name: netbird-server
restart: unless-stopped
environment:
NETBIRD_DOMAIN: "${NETBIRD_DOMAIN:?Set NETBIRD_DOMAIN in netbird/.env}"
NETBIRD_PROXY_SUBNET: "${NETBIRD_PROXY_SUBNET:?Set NETBIRD_PROXY_SUBNET in netbird/.env (run setup.sh)}"
entrypoint:
- /bin/sh
- /opt/netbird/entrypoint.sh
command:
- --config
- /run/netbird/config.yaml
ports:
- "3478:3478/udp"
volumes:
- netbird_data:/var/lib/netbird
- ./config.template.yaml:/opt/netbird/config.template.yaml:ro
- ./entrypoint.sh:/opt/netbird/entrypoint.sh:ro
secrets:
- relay_auth_secret
- datastore_encryption_key
tmpfs:
- /run/netbird:mode=0700
healthcheck:
test:
- CMD
- bash
- -ec
- exec 3<>/dev/tcp/127.0.0.1/80
interval: 30s
timeout: 5s
retries: 5
start_period: 30s
stop_grace_period: 30s
labels:
- "traefik.enable=true"
- "traefik.http.services.netbird-server.loadbalancer.server.port=80"
- "traefik.http.services.netbird-server-h2c.loadbalancer.server.port=80"
- "traefik.http.services.netbird-server-h2c.loadbalancer.server.scheme=h2c"
# gRPC routers
# Prod Router
- "traefik.http.routers.netbird-grpc.rule=Host(`${NETBIRD_DOMAIN:?Set NETBIRD_DOMAIN in netbird/.env}`) && (PathPrefix(`/signalexchange.SignalExchange/`) || PathPrefix(`/management.ManagementService/`) || PathPrefix(`/management.ProxyService/`))"
- "traefik.http.routers.netbird-grpc.entrypoints=websecure"
- "traefik.http.routers.netbird-grpc.service=netbird-server-h2c"
- "traefik.http.routers.netbird-grpc.priority=100"
- "traefik.http.routers.netbird-grpc.tls=true"
- "traefik.http.routers.netbird-grpc.tls.certresolver=letsencrypt"
# Local Router
- "traefik.http.routers.netbird-grpc-local.rule=Host(`${NETBIRD_LOCAL_DOMAIN:?Set NETBIRD_LOCAL_DOMAIN in netbird/.env}`) && (PathPrefix(`/signalexchange.SignalExchange/`) || PathPrefix(`/management.ManagementService/`) || PathPrefix(`/management.ProxyService/`))"
- "traefik.http.routers.netbird-grpc-local.entrypoints=websecure"
- "traefik.http.routers.netbird-grpc-local.service=netbird-server-h2c"
- "traefik.http.routers.netbird-grpc-local.priority=100"
- "traefik.http.routers.netbird-grpc-local.tls=true"
# Dev Router
- "traefik.http.routers.netbird-grpc-dev.rule=Host(`${NETBIRD_DEV_DOMAIN:?Set NETBIRD_DEV_DOMAIN in netbird/.env}`) && (PathPrefix(`/signalexchange.SignalExchange/`) || PathPrefix(`/management.ManagementService/`) || PathPrefix(`/management.ProxyService/`))"
- "traefik.http.routers.netbird-grpc-dev.entrypoints=websecure"
- "traefik.http.routers.netbird-grpc-dev.service=netbird-server-h2c"
- "traefik.http.routers.netbird-grpc-dev.priority=100"
- "traefik.http.routers.netbird-grpc-dev.tls=true"
# Backend routers
# Prod Router
- "traefik.http.routers.netbird-backend.rule=Host(`${NETBIRD_DOMAIN:?Set NETBIRD_DOMAIN in netbird/.env}`) && (PathPrefix(`/relay`) || PathPrefix(`/ws-proxy/`) || PathPrefix(`/api`) || PathPrefix(`/oauth2`))"
- "traefik.http.routers.netbird-backend.entrypoints=websecure"
- "traefik.http.routers.netbird-backend.service=netbird-server"
- "traefik.http.routers.netbird-backend.priority=100"
- "traefik.http.routers.netbird-backend.tls=true"
- "traefik.http.routers.netbird-backend.tls.certresolver=letsencrypt"
# Local Router
- "traefik.http.routers.netbird-backend-local.rule=Host(`${NETBIRD_LOCAL_DOMAIN:?Set NETBIRD_LOCAL_DOMAIN in netbird/.env}`) && (PathPrefix(`/relay`) || PathPrefix(`/ws-proxy/`) || PathPrefix(`/api`) || PathPrefix(`/oauth2`))"
- "traefik.http.routers.netbird-backend-local.entrypoints=websecure"
- "traefik.http.routers.netbird-backend-local.service=netbird-server"
- "traefik.http.routers.netbird-backend-local.priority=100"
- "traefik.http.routers.netbird-backend-local.tls=true"
# Dev Router
- "traefik.http.routers.netbird-backend-dev.rule=Host(`${NETBIRD_DEV_DOMAIN:?Set NETBIRD_DEV_DOMAIN in netbird/.env}`) && (PathPrefix(`/relay`) || PathPrefix(`/ws-proxy/`) || PathPrefix(`/api`) || PathPrefix(`/oauth2`))"
- "traefik.http.routers.netbird-backend-dev.entrypoints=websecure"
- "traefik.http.routers.netbird-backend-dev.service=netbird-server"
- "traefik.http.routers.netbird-backend-dev.priority=100"
- "traefik.http.routers.netbird-backend-dev.tls=true"
networks:
- proxy
dashboard:
image: netbirdio/dashboard:v2.93.0
container_name: netbird-dashboard
restart: unless-stopped
environment:
NETBIRD_MGMT_API_ENDPOINT: "https://${NETBIRD_DOMAIN:?Set NETBIRD_DOMAIN in netbird/.env}"
NETBIRD_MGMT_GRPC_API_ENDPOINT: "https://${NETBIRD_DOMAIN:?Set NETBIRD_DOMAIN in netbird/.env}"
AUTH_AUDIENCE: netbird-dashboard
AUTH_CLIENT_ID: netbird-dashboard
AUTH_CLIENT_SECRET: ""
AUTH_AUTHORITY: "https://${NETBIRD_DOMAIN:?Set NETBIRD_DOMAIN in netbird/.env}/oauth2"
AUTH_SUPPORTED_SCOPES: openid profile email groups
AUTH_REDIRECT_URI: /nb-auth
AUTH_SILENT_REDIRECT_URI: /nb-silent-auth
USE_AUTH0: "false"
LETSENCRYPT_DOMAIN: none
depends_on:
netbird-server:
condition: service_healthy
healthcheck:
test: ["CMD", "curl", "--fail", "--silent", "--show-error", "http://127.0.0.1/"]
interval: 30s
timeout: 5s
retries: 5
start_period: 15s
labels:
- "traefik.enable=true"
- "traefik.http.services.netbird-dashboard.loadbalancer.server.port=80"
# Dashboard catch-all routers
# Prod Router
- "traefik.http.routers.netbird-dashboard.rule=Host(`${NETBIRD_DOMAIN:?Set NETBIRD_DOMAIN in netbird/.env}`)"
- "traefik.http.routers.netbird-dashboard.entrypoints=websecure"
- "traefik.http.routers.netbird-dashboard.service=netbird-dashboard"
- "traefik.http.routers.netbird-dashboard.priority=1"
- "traefik.http.routers.netbird-dashboard.tls=true"
- "traefik.http.routers.netbird-dashboard.tls.certresolver=letsencrypt"
# Local Router
- "traefik.http.routers.netbird-dashboard-local.rule=Host(`${NETBIRD_LOCAL_DOMAIN:?Set NETBIRD_LOCAL_DOMAIN in netbird/.env}`)"
- "traefik.http.routers.netbird-dashboard-local.entrypoints=websecure"
- "traefik.http.routers.netbird-dashboard-local.service=netbird-dashboard"
- "traefik.http.routers.netbird-dashboard-local.priority=1"
- "traefik.http.routers.netbird-dashboard-local.tls=true"
# Dev Router
- "traefik.http.routers.netbird-dashboard-dev.rule=Host(`${NETBIRD_DEV_DOMAIN:?Set NETBIRD_DEV_DOMAIN in netbird/.env}`)"
- "traefik.http.routers.netbird-dashboard-dev.entrypoints=websecure"
- "traefik.http.routers.netbird-dashboard-dev.service=netbird-dashboard"
- "traefik.http.routers.netbird-dashboard-dev.priority=1"
- "traefik.http.routers.netbird-dashboard-dev.tls=true"
networks:
- proxy
networks:
proxy:
external: true
volumes:
netbird_data:
name: netbird_data
secrets:
relay_auth_secret:
file: ./secrets/relay-auth-secret
datastore_encryption_key:
file: ./secrets/datastore-encryption-key
+26
View File
@@ -0,0 +1,26 @@
server:
listenAddress: ":80"
exposedAddress: "https://__NETBIRD_DOMAIN__:443"
stunPorts:
- 3478
metricsPort: 9090
healthcheckAddress: ":9000"
logLevel: info
logFile: console
authSecret: "__NETBIRD_AUTH_SECRET__"
dataDir: "/var/lib/netbird"
disableAnonymousMetrics: true
auth:
issuer: "https://__NETBIRD_DOMAIN__/oauth2"
signKeyRefreshEnabled: true
dashboardRedirectURIs:
- "https://__NETBIRD_DOMAIN__/nb-auth"
- "https://__NETBIRD_DOMAIN__/nb-silent-auth"
reverseProxy:
trustedHTTPProxies:
- "__NETBIRD_PROXY_SUBNET__"
trustedPeers:
- "__NETBIRD_PROXY_SUBNET__"
store:
engine: sqlite
encryptionKey: "__NETBIRD_ENCRYPTION_KEY__"
View File
Whitespace-only changes.
+28
View File
@@ -0,0 +1,28 @@
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: netbird-prod-tls
namespace: netbird
spec:
secretName: netbird-prod-tls
dnsNames:
- nb.forust.xyz
issuerRef:
name: letsencrypt-prod
kind: ClusterIssuer
---
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: internal-wildcard-tls
namespace: netbird
spec:
secretName: internal-wildcard-tls
dnsNames:
- "*.workstation.internal"
- "*.gigaforust.internal"
- workstation.internal
- gigaforust.internal
issuerRef:
name: internal-ca
kind: ClusterIssuer
+160
View File
@@ -0,0 +1,160 @@
apiVersion: v1
kind: ConfigMap
metadata:
name: netbird-config
namespace: netbird
data:
# Public hostname, rendered into the server config by entrypoint.sh.
NETBIRD_DOMAIN: "nb.forust.xyz"
NETBIRD_PROXY_SUBNET: "10.244.0.0/16"
NETBIRD_MGMT_API_ENDPOINT: "https://nb.forust.xyz"
NETBIRD_MGMT_GRPC_API_ENDPOINT: "https://nb.forust.xyz"
AUTH_AUDIENCE: "netbird-dashboard"
AUTH_CLIENT_ID: "netbird-dashboard"
AUTH_CLIENT_SECRET: ""
AUTH_AUTHORITY: "https://nb.forust.xyz/oauth2"
AUTH_SUPPORTED_SCOPES: "openid profile email groups"
AUTH_REDIRECT_URI: "/nb-auth"
AUTH_SILENT_REDIRECT_URI: "/nb-silent-auth"
USE_AUTH0: "false"
LETSENCRYPT_DOMAIN: "none"
config.template.yaml: |
server:
listenAddress: ":80"
exposedAddress: "https://__NETBIRD_DOMAIN__:443"
stunPorts:
- 3478
metricsPort: 9090
healthcheckAddress: ":9000"
logLevel: info
logFile: console
authSecret: "__NETBIRD_AUTH_SECRET__"
dataDir: "/var/lib/netbird"
disableAnonymousMetrics: true
auth:
issuer: "https://__NETBIRD_DOMAIN__/oauth2"
signKeyRefreshEnabled: true
dashboardRedirectURIs:
- "https://__NETBIRD_DOMAIN__/nb-auth"
- "https://__NETBIRD_DOMAIN__/nb-silent-auth"
reverseProxy:
trustedHTTPProxies:
- "__NETBIRD_PROXY_SUBNET__"
trustedPeers:
- "__NETBIRD_PROXY_SUBNET__"
store:
engine: sqlite
encryptionKey: "__NETBIRD_ENCRYPTION_KEY__"
entrypoint.sh: |
#!/bin/sh
set -eu
umask 077
TEMPLATE_PATH=/opt/netbird/config.template.yaml
RENDERED_PATH=/run/netbird/config.yaml
RELAY_SECRET_PATH=/run/secrets/relay_auth_secret
ENCRYPTION_KEY_PATH=/run/secrets/datastore_encryption_key
is_valid_proxy_subnet() {
candidate="$1"
case "$candidate" in
0.0.0.0/0)
return 1
;;
*/*)
address="${candidate%%/*}"
prefix="${candidate#*/}"
;;
*)
return 1
;;
esac
case "$prefix" in
0|[1-9]|[1-2][0-9]|3[0-2]) ;;
*)
return 1
;;
esac
old_ifs="$IFS"
IFS=.
# shellcheck disable=SC2086
set -- $address
IFS="$old_ifs"
[ "$#" -eq 4 ] || return 1
for octet do
case "$octet" in
0|[1-9]|[1-9][0-9]|1[0-9][0-9]|2[0-4][0-9]|25[0-5]) ;;
*)
return 1
;;
esac
done
}
read_secret() {
secret_path="$1"
if [ ! -r "$secret_path" ]; then
echo "Required secret is not readable: $secret_path" >&2
exit 1
fi
secret_value="$(cat "$secret_path")"
if [ -z "$secret_value" ]; then
echo "Required secret is empty: $secret_path" >&2
exit 1
fi
printf '%s' "$secret_value"
}
if [ -z "${NETBIRD_DOMAIN:-}" ]; then
echo "NETBIRD_DOMAIN must be set" >&2
exit 1
fi
case "$NETBIRD_DOMAIN" in
*[!A-Za-z0-9.-]*)
echo "NETBIRD_DOMAIN contains unsupported characters" >&2
exit 1
;;
esac
if [ -z "${NETBIRD_PROXY_SUBNET:-}" ] || [ "$NETBIRD_PROXY_SUBNET" = "auto" ]; then
echo "NETBIRD_PROXY_SUBNET must be an explicit IPv4 CIDR; run netbird/setup.sh first" >&2
exit 1
fi
if ! is_valid_proxy_subnet "$NETBIRD_PROXY_SUBNET"; then
echo "NETBIRD_PROXY_SUBNET must be a non-default IPv4 CIDR, for example 172.20.0.0/16" >&2
exit 1
fi
if [ "$#" -ne 2 ] || [ "$1" != "--config" ] || [ "$2" != "$RENDERED_PATH" ]; then
echo "Expected: --config $RENDERED_PATH" >&2
exit 1
fi
relay_secret="$(read_secret "$RELAY_SECRET_PATH")"
encryption_key="$(read_secret "$ENCRYPTION_KEY_PATH")"
mkdir -p "$(dirname "$RENDERED_PATH")"
sed \
-e "s|__NETBIRD_DOMAIN__|${NETBIRD_DOMAIN}|g" \
-e "s|__NETBIRD_AUTH_SECRET__|${relay_secret}|g" \
-e "s|__NETBIRD_ENCRYPTION_KEY__|${encryption_key}|g" \
-e "s|__NETBIRD_PROXY_SUBNET__|${NETBIRD_PROXY_SUBNET}|g" \
"$TEMPLATE_PATH" >"$RENDERED_PATH"
if grep -q '__NETBIRD_' "$RENDERED_PATH"; then
echo "Rendered NetBird configuration still contains unresolved placeholders" >&2
exit 1
fi
exec /go/bin/netbird-server "$@"
+86
View File
@@ -0,0 +1,86 @@
apiVersion: traefik.io/v1alpha1
kind: IngressRoute
metadata:
name: netbird-prod
namespace: netbird
spec:
entryPoints:
- websecure
routes:
# NO crowdsec-bouncer on the API routes. These are the mesh client's own
# endpoints: gRPC-gateway management calls plus signal/relay long-polling,
# authenticated by NetBird's token rather than by a login form. A ban here
# is self-defeating - the client needs the mesh to reach anything else, so
# CrowdSec banning it locks the peer out of the network it needs to
# function. It also backfires: a banned peer keeps retrying, every retry
# is another 403, and LePresidente/http-generic-403-bf turns five 403s in
# ten seconds into a 4h ban, so one 403 loop kept re-arming the ban.
# netbird-local below has always been exempt; this makes prod match.
- match: Host(`nb.forust.xyz`) && (PathPrefix(`/signalexchange.SignalExchange/`) || PathPrefix(`/management.ManagementService/`) || PathPrefix(`/management.ProxyService/`))
kind: Rule
priority: 100
services:
- name: netbird-server-service
port: 80
scheme: h2c
- match: Host(`nb.forust.xyz`) && (PathPrefix(`/relay`) || PathPrefix(`/ws-proxy/`) || PathPrefix(`/api`) || PathPrefix(`/oauth2`))
kind: Rule
priority: 100
services:
- name: netbird-server-service
port: 80
- match: Host(`nb.forust.xyz`)
kind: Rule
priority: 1
middlewares:
- name: crowdsec-bouncer
namespace: crowdsec
services:
- name: netbird-dashboard-service
port: 80
tls:
secretName: netbird-prod-tls
---
apiVersion: traefik.io/v1alpha1
kind: IngressRoute
metadata:
name: netbird-local
namespace: netbird
spec:
entryPoints:
- websecure
routes:
- match: (Host(`netbird.workstation.internal`) || Host(`netbird.gigaforust.internal`)) && (PathPrefix(`/signalexchange.SignalExchange/`) || PathPrefix(`/management.ManagementService/`) || PathPrefix(`/management.ProxyService/`))
kind: Rule
priority: 100
services:
- name: netbird-server-service
port: 80
scheme: h2c
- match: (Host(`netbird.workstation.internal`) || Host(`netbird.gigaforust.internal`)) && (PathPrefix(`/relay`) || PathPrefix(`/ws-proxy/`) || PathPrefix(`/api`) || PathPrefix(`/oauth2`))
kind: Rule
priority: 100
services:
- name: netbird-server-service
port: 80
- match: Host(`netbird.workstation.internal`) || Host(`netbird.gigaforust.internal`)
kind: Rule
priority: 1
services:
- name: netbird-dashboard-service
port: 80
tls:
secretName: internal-wildcard-tls
---
apiVersion: traefik.io/v1alpha1
kind: IngressRouteUDP
metadata:
name: netbird-stun
namespace: netbird
spec:
entryPoints:
- netbird-stun
routes:
- services:
- name: netbird-server-service
port: 3478
+4
View File
@@ -0,0 +1,4 @@
apiVersion: v1
kind: Namespace
metadata:
name: netbird
+182
View File
@@ -0,0 +1,182 @@
apiVersion: v1
kind: Service
metadata:
name: netbird-server-service
namespace: netbird
spec:
selector:
app: netbird-server
ports:
- port: 80
name: http
targetPort: 80
protocol: TCP
- port: 3478
name: stun
targetPort: 3478
protocol: UDP
---
apiVersion: v1
kind: Service
metadata:
name: netbird-dashboard-service
namespace: netbird
spec:
selector:
app: netbird-dashboard
ports:
- port: 80
name: http
targetPort: 80
---
apiVersion: apps/v1
kind: Deployment
metadata:
name: netbird-server-deployment
namespace: netbird
spec:
replicas: 1
selector:
matchLabels:
app: netbird-server
template:
metadata:
labels:
app: netbird-server
spec:
containers:
- name: netbird-server
image: netbirdio/netbird-server:0.79.0
command: ["/bin/sh", "/opt/netbird/entrypoint.sh", "--config", "/run/netbird/config.yaml"]
envFrom:
- configMapRef:
name: netbird-config
ports:
- containerPort: 80
name: http
protocol: TCP
- containerPort: 3478
name: stun
protocol: UDP
volumeMounts:
- name: netbird-data
mountPath: /var/lib/netbird
- name: netbird-files
mountPath: /opt/netbird
readOnly: true
- name: netbird-secrets
mountPath: /run/secrets/relay_auth_secret
subPath: relay_auth_secret
readOnly: true
- name: netbird-secrets
mountPath: /run/secrets/datastore_encryption_key
subPath: datastore_encryption_key
readOnly: true
- name: netbird-run
mountPath: /run/netbird
readinessProbe:
tcpSocket:
port: 80
initialDelaySeconds: 30
periodSeconds: 30
timeoutSeconds: 5
failureThreshold: 5
livenessProbe:
tcpSocket:
port: 80
initialDelaySeconds: 60
periodSeconds: 30
timeoutSeconds: 5
failureThreshold: 5
# p95 97M, max 102M over 7 days. Was 256Mi/1Gi.
resources:
requests:
memory: "128Mi"
cpu: "250m"
limits:
memory: "384Mi"
cpu: "1000m"
volumes:
- name: netbird-data
persistentVolumeClaim:
claimName: netbird-pvc
- name: netbird-files
configMap:
name: netbird-config
defaultMode: 0755
items:
- key: config.template.yaml
path: config.template.yaml
- key: entrypoint.sh
path: entrypoint.sh
- name: netbird-secrets
secret:
secretName: netbird-secrets
items:
- key: relay_auth_secret
path: relay_auth_secret
- key: datastore_encryption_key
path: datastore_encryption_key
- name: netbird-run
emptyDir:
medium: Memory
---
apiVersion: apps/v1
kind: Deployment
metadata:
name: netbird-dashboard-deployment
namespace: netbird
spec:
replicas: 1
selector:
matchLabels:
app: netbird-dashboard
template:
metadata:
labels:
app: netbird-dashboard
spec:
containers:
- name: dashboard
image: netbirdio/dashboard:v2.93.0
envFrom:
- configMapRef:
name: netbird-config
ports:
- containerPort: 80
name: http
readinessProbe:
httpGet:
path: /
port: 80
initialDelaySeconds: 15
periodSeconds: 30
timeoutSeconds: 5
failureThreshold: 5
livenessProbe:
httpGet:
path: /
port: 80
initialDelaySeconds: 30
periodSeconds: 30
timeoutSeconds: 5
failureThreshold: 5
resources:
requests:
memory: "32Mi"
cpu: "50m"
limits:
memory: "128Mi"
cpu: "300m"
---
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: netbird-pvc
namespace: netbird
spec:
accessModes:
- ReadWriteOnce
resources:
requests:
storage: 2Gi
+11
View File
@@ -0,0 +1,11 @@
apiVersion: v1
kind: Secret
metadata:
name: netbird-secrets
namespace: netbird
type: Opaque
stringData:
# hex, 64 chars: openssl rand -hex 32
relay_auth_secret: "REPLACE_ME"
# base64, 44 chars: openssl rand -base64 32
datastore_encryption_key: "REPLACE_ME"
+32
View File
@@ -0,0 +1,32 @@
POSTGRES_DB=netbox
POSTGRES_USER=netbox
POSTGRES_PASSWORD=CHANGE_ME_POSTGRES_PASSWORD
DB_NAME=netbox
DB_USER=netbox
DB_PASSWORD=CHANGE_ME_POSTGRES_PASSWORD
DB_HOST=postgres
DB_PORT=5432
DB_SSLMODE=disable
REDIS_HOST=redis
REDIS_PORT=6379
REDIS_PASSWORD=CHANGE_ME_REDIS_PASSWORD
REDIS_DATABASE=0
REDIS_CACHE_HOST=redis-cache
REDIS_CACHE_PORT=6379
REDIS_CACHE_PASSWORD=CHANGE_ME_REDIS_CACHE_PASSWORD
REDIS_CACHE_DATABASE=1
ALLOWED_HOSTS=localhost,127.0.0.1,[::1],netbox.forust.xyz,netbox.workstation.internal
CSRF_TRUSTED_ORIGINS=https://netbox.forust.xyz,https://netbox.workstation.internal
SECRET_KEY=CHANGE_ME_DJANGO_SECRET_KEY
API_TOKEN_PEPPER_1=CHANGE_ME_API_TOKEN_PEPPER
TIME_ZONE=Europe/Bratislava
TZ=Europe/Bratislava
SKIP_SUPERUSER=false
SUPERUSER_NAME=admin
SUPERUSER_EMAIL=admin@example.com
SUPERUSER_PASSWORD=CHANGE_ME_SUPERUSER_PASSWORD
+96
View File
@@ -0,0 +1,96 @@
# NetBox
NetBox for homelab documentation and visualization. Two runtimes are available:
| Runtime | Manifest | Purpose |
| ------- | -------------- | -------------------------------------------------------------- |
| Docker | `compose.yaml` | Local stand on `127.0.0.1:8000` (no public exposure) |
| k8s | `k8s/` | Homelab service on `netbox.forust.xyz` (and the internal name) |
Both use the same image (`netboxcommunity/netbox:v4.7-5.1.1`) and Valkey for tasks
plus a second logical database for caching. The Docker stand keeps its own
PostgreSQL container, while the k8s deployment uses the shared `database` cluster
(`postgres.database.svc.cluster.local:5432`, role/database `netbox`); only Valkey
stays a per-service StatefulSet.
## Docker Compose
```bash
cp .env.example .env
# replace CHANGE_ME
docker compose up -d
```
The UI is available at <http://localhost:8000>. The port is bound to `127.0.0.1`
intentionally, so this stand is not exposed on the LAN or public interfaces.
The `netbox` service is also attached to the external `proxy` network and carries
Traefik labels for `netbox.forust.xyz` and `netbox.workstation.internal`. Those
labels only take effect while the Docker Traefik stack is running; it is currently
stopped, and the live ingress path in this homelab is the k8s Traefik.
Inspect startup and health with:
```bash
docker compose ps
docker compose logs -f netbox
```
Stop it with `docker compose down`; data is kept in the named volumes
`netbox-postgres`, `netbox-media-files`, `netbox-reports-files`,
`netbox-scripts-files` and `netbox-redis-data`.
## Kubernetes
`k8s/` is deployed in the homelab cluster and serves `netbox.forust.xyz` publicly
plus `netbox.workstation.internal` / `netbox.gigaforust.internal` internally. To
rebuild it from scratch:
```bash
# 1. shared PostgreSQL: the password lives in the shared secret, NetBox keeps a copy
kubectl -n database patch secret postgres-shared-secrets \
--type merge -p '{"stringData":{"NETBOX_DB_PASSWORD":"<same value>"}}'
kubectl -n database exec postgres17-0 -- psql -U postgres -d postgres \
-c 'CREATE ROLE netbox LOGIN PASSWORD ...' -c 'CREATE DATABASE netbox OWNER netbox'
# 2. secrets first: the deploy workflow never applies *secret*.yaml
cp k8s/secrets.yaml.example k8s/secrets.yaml # replace CHANGE_ME
kubectl apply -f k8s/secrets.yaml
# 3. manifests
kubectl apply -f k8s/
```
The shared cluster is reached at `postgres.database.svc.cluster.local:5432`. Its
NetworkPolicy (`postgres/k8s/network-policy.yaml`) must list the `netbox` namespace
or connections are dropped, and `postgres/initdb/01-create-databases.sh` already
creates the role and database on a fresh data directory. NetBox has no PostgreSQL
StatefulSet of its own — only `netbox-valkey`.
`netbox.forust.xyz` resolves to this host (`78.98.72.122`) through the `DOMAINS`
list in the `default/cfddns` secret. cert-manager issues `netbox-prod-tls` with the
`letsencrypt-prod` issuer, the internal route uses `internal-wildcard-tls`.
Resources are permanent again now that the first-boot migrations are complete:
the web container reserves `100m`/`512Mi` and is capped at `2` CPU/`2Gi`, the
worker reserves `50m`/`256Mi` and is capped at `1` CPU/`1Gi`, and Valkey reserves
`25m`/`64Mi` and is capped at `250m`/`256Mi`. The deliberately generous CPU caps
leave enough headroom for future schema migrations without letting one process
consume the whole node.
The first start applies ~810 migrations, each in its own transaction with DDL and
a commit; every later start is a no-op. The startup probe allows 15 minutes and
`progressDeadlineSeconds` is 1800 for the same reason. Probes run inside the pod
and explicitly set `Host: netbox.forust.xyz`; a kubelet `httpGet.host` field would
replace the probe destination with that public hostname and bypass the pod.
## Secrets
- `netbox/.env` (compose) and `netbox/k8s/secrets.yaml` (k8s) are gitignored. Only
`.env.example` and `k8s/secrets.yaml.example` are committed.
- `netbox/configuration/configuration.py` is env-driven: hosts, database, Redis and
the Django keys all come from the environment, so the same settings file works in
both runtimes. The k8s copy lives in the `netbox-settings` ConfigMap
(`k8s/settings.yaml`) and must be kept in sync with the file.
- Rotating `SECRET_KEY` invalidates all sessions; rotating `API_TOKEN_PEPPER_1`
invalidates every API token.
+137
View File
@@ -0,0 +1,137 @@
services:
netbox:
image: docker.io/netboxcommunity/netbox:v4.7-5.1.1
container_name: netbox
restart: unless-stopped
user: "netbox:root"
ports:
- "127.0.0.1:8000:8080"
env_file:
- .env
environment:
GRANIAN_WORKERS: "2"
depends_on:
postgres:
condition: service_healthy
redis:
condition: service_healthy
redis-cache:
condition: service_healthy
volumes:
- ./configuration:/etc/netbox/config:z,ro
- netbox-media-files:/opt/netbox/netbox/media
- netbox-reports-files:/opt/netbox/netbox/reports
- netbox-scripts-files:/opt/netbox/netbox/scripts
networks:
- default
- proxy
labels:
- "traefik.enable=true"
- "traefik.http.services.netbox.loadbalancer.server.port=8080"
# Prod Router
- "traefik.http.routers.netbox.rule=Host(`netbox.forust.xyz`)"
- "traefik.http.routers.netbox.entrypoints=websecure"
- "traefik.http.routers.netbox.middlewares=security-headers@file"
- "traefik.http.routers.netbox.tls.certresolver=letsencrypt"
# Local Router
- "traefik.http.routers.netbox-local.rule=Host(`netbox.workstation.internal`)"
- "traefik.http.routers.netbox-local.entrypoints=websecure"
- "traefik.http.routers.netbox-local.tls=true"
healthcheck:
test: ["CMD", "/opt/netbox/health.sh"]
start_period: 600s
timeout: 5s
interval: 15s
retries: 10
netbox-worker:
image: docker.io/netboxcommunity/netbox:v4.7-5.1.1
container_name: netbox-worker
restart: unless-stopped
user: "netbox:root"
command:
- /opt/netbox/venv/bin/python
- /opt/netbox/netbox/manage.py
- rqworker
env_file:
- .env
depends_on:
netbox:
condition: service_healthy
volumes:
- ./configuration:/etc/netbox/config:z,ro
- netbox-media-files:/opt/netbox/netbox/media
- netbox-reports-files:/opt/netbox/netbox/reports
- netbox-scripts-files:/opt/netbox/netbox/scripts
healthcheck:
test: ["CMD-SHELL", "ps -ef | grep -q '[r]qworker'"]
start_period: 30s
timeout: 5s
interval: 15s
retries: 10
postgres:
image: docker.io/postgres:18.6-alpine
container_name: netbox-postgres
restart: unless-stopped
environment:
POSTGRES_DB: "${POSTGRES_DB:?POSTGRES_DB must be set}"
POSTGRES_USER: "${POSTGRES_USER:?POSTGRES_USER must be set}"
POSTGRES_PASSWORD: "${POSTGRES_PASSWORD:?POSTGRES_PASSWORD must be set}"
volumes:
- netbox-postgres:/var/lib/postgresql
healthcheck:
test: ["CMD-SHELL", 'pg_isready -q -t 2 -d "$${POSTGRES_DB}" -U "$${POSTGRES_USER}"']
start_period: 20s
timeout: 5s
interval: 10s
retries: 10
redis:
image: docker.io/valkey/valkey:9.1.2-alpine
container_name: netbox-redis
restart: unless-stopped
command:
- sh
- -c
- valkey-server --appendonly yes --requirepass "$$REDIS_PASSWORD"
environment:
REDIS_PASSWORD: "${REDIS_PASSWORD:?REDIS_PASSWORD must be set}"
volumes:
- netbox-redis-data:/data
healthcheck:
test: ["CMD-SHELL", 'valkey-cli --pass "$${REDIS_PASSWORD}" ping | grep -q PONG']
start_period: 5s
timeout: 5s
interval: 5s
retries: 10
redis-cache:
image: docker.io/valkey/valkey:9.1.2-alpine
container_name: netbox-redis-cache
restart: unless-stopped
command:
- sh
- -c
- valkey-server --requirepass "$$REDIS_CACHE_PASSWORD"
environment:
REDIS_CACHE_PASSWORD: "${REDIS_CACHE_PASSWORD:?REDIS_CACHE_PASSWORD must be set}"
healthcheck:
test: ["CMD-SHELL", 'valkey-cli --pass "$${REDIS_CACHE_PASSWORD}" ping | grep -q PONG']
start_period: 5s
timeout: 5s
interval: 5s
retries: 10
volumes:
netbox-media-files:
netbox-reports-files:
netbox-scripts-files:
netbox-postgres:
netbox-redis-data:
networks:
default:
proxy:
external: true
+48
View File
@@ -0,0 +1,48 @@
import os
def _csv(name, default=''):
return [item.strip() for item in os.environ.get(name, default).split(',') if item.strip()]
ALLOWED_HOSTS = _csv('ALLOWED_HOSTS', 'localhost,127.0.0.1,[::1]')
CSRF_TRUSTED_ORIGINS = _csv('CSRF_TRUSTED_ORIGINS')
USE_X_FORWARDED_HOST = True
SECURE_PROXY_SSL_HEADER = ('HTTP_X_FORWARDED_PROTO', 'https')
DATABASES = {
'default': {
'NAME': os.environ['DB_NAME'],
'USER': os.environ['DB_USER'],
'PASSWORD': os.environ['DB_PASSWORD'],
'HOST': os.environ['DB_HOST'],
'PORT': os.environ.get('DB_PORT', '5432'),
'OPTIONS': {'sslmode': os.environ.get('DB_SSLMODE', 'disable')},
'CONN_MAX_AGE': int(os.environ.get('DB_CONN_MAX_AGE', '300')),
}
}
REDIS = {
'tasks': {
'HOST': os.environ['REDIS_HOST'],
'PORT': int(os.environ.get('REDIS_PORT', '6379')),
'PASSWORD': os.environ['REDIS_PASSWORD'],
'DATABASE': int(os.environ.get('REDIS_DATABASE', '0')),
'SSL': False,
},
'caching': {
'HOST': os.environ['REDIS_CACHE_HOST'],
'PORT': int(os.environ.get('REDIS_CACHE_PORT', '6379')),
'PASSWORD': os.environ['REDIS_CACHE_PASSWORD'],
'DATABASE': int(os.environ.get('REDIS_CACHE_DATABASE', '1')),
'SSL': False,
},
}
SECRET_KEY = os.environ['SECRET_KEY']
API_TOKEN_PEPPERS = {1: os.environ['API_TOKEN_PEPPER_1']}
TIME_ZONE = os.environ.get('TIME_ZONE', 'UTC')
MEDIA_ROOT = '/opt/netbox/netbox/media'
REPORTS_ROOT = '/opt/netbox/netbox/reports'
SCRIPTS_ROOT = '/opt/netbox/netbox/scripts'
CENSUS_REPORTING_ENABLED = False
View File
Whitespace-only changes.
+28
View File
@@ -0,0 +1,28 @@
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: netbox-prod-tls
namespace: netbox
spec:
secretName: netbox-prod-tls
dnsNames:
- netbox.forust.xyz
issuerRef:
name: letsencrypt-prod
kind: ClusterIssuer
---
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: internal-wildcard-tls
namespace: netbox
spec:
secretName: internal-wildcard-tls
dnsNames:
- "*.workstation.internal"
- "*.gigaforust.internal"
- workstation.internal
- gigaforust.internal
issuerRef:
name: internal-ca
kind: ClusterIssuer
+21
View File
@@ -0,0 +1,21 @@
apiVersion: v1
kind: ConfigMap
metadata:
name: netbox-config
namespace: netbox
data:
DB_HOST: "postgres.database.svc.cluster.local"
DB_PORT: "5432"
DB_SSLMODE: "disable"
REDIS_HOST: "netbox-valkey"
REDIS_PORT: "6379"
REDIS_DATABASE: "0"
REDIS_CACHE_HOST: "netbox-valkey"
REDIS_CACHE_PORT: "6379"
REDIS_CACHE_DATABASE: "1"
TIME_ZONE: "Europe/Bratislava"
TZ: "Europe/Bratislava"
GRANIAN_WORKERS: "2"
ALLOWED_HOSTS: "netbox.forust.xyz,netbox.workstation.internal,netbox.gigaforust.internal"
CSRF_TRUSTED_ORIGINS: "https://netbox.forust.xyz,https://netbox.workstation.internal,https://netbox.gigaforust.internal"
SKIP_SUPERUSER: "false"
+36
View File
@@ -0,0 +1,36 @@
apiVersion: traefik.io/v1alpha1
kind: IngressRoute
metadata:
name: netbox-prod
namespace: netbox
spec:
entryPoints:
- websecure
routes:
- match: Host(`netbox.forust.xyz`)
kind: Rule
middlewares:
- name: crowdsec-bouncer
namespace: crowdsec
services:
- name: netbox-service
port: 8080
tls:
secretName: netbox-prod-tls
---
apiVersion: traefik.io/v1alpha1
kind: IngressRoute
metadata:
name: netbox-local
namespace: netbox
spec:
entryPoints:
- websecure
routes:
- match: Host(`netbox.workstation.internal`) || Host(`netbox.gigaforust.internal`)
kind: Rule
services:
- name: netbox-service
port: 8080
tls:
secretName: internal-wildcard-tls
+4
View File
@@ -0,0 +1,4 @@
apiVersion: v1
kind: Namespace
metadata:
name: netbox
+213
View File
@@ -0,0 +1,213 @@
apiVersion: v1
kind: Service
metadata:
name: netbox-service
namespace: netbox
spec:
selector:
app: netbox
ports:
- name: http
port: 8080
targetPort: http
---
apiVersion: apps/v1
kind: Deployment
metadata:
name: netbox-deployment
namespace: netbox
labels:
app: netbox
spec:
replicas: 1
progressDeadlineSeconds: 300
selector:
matchLabels:
app: netbox
strategy:
# ReadWriteOnce PVC
type: Recreate
template:
metadata:
labels:
app: netbox
spec:
containers:
- name: netbox
image: docker.io/netboxcommunity/netbox:v4.7-5.1.1
ports:
- name: http
containerPort: 8080
envFrom:
- configMapRef:
name: netbox-config
- secretRef:
name: netbox-secrets
volumeMounts:
- name: netbox-config
mountPath: /etc/netbox/config
readOnly: true
- name: netbox-media
mountPath: /opt/netbox/netbox/media
- name: netbox-reports
mountPath: /opt/netbox/netbox/reports
- name: netbox-scripts
mountPath: /opt/netbox/netbox/scripts
startupProbe:
exec:
command:
- /usr/bin/curl
- --fail
- --silent
- --show-error
- --max-time
- "4"
- --header
- "Host: netbox.forust.xyz"
- http://127.0.0.1:8080/login/
failureThreshold: 90
periodSeconds: 10
readinessProbe:
exec:
command:
- /usr/bin/curl
- --fail
- --silent
- --show-error
- --max-time
- "4"
- --header
- "Host: netbox.forust.xyz"
- http://127.0.0.1:8080/login/
periodSeconds: 10
livenessProbe:
exec:
command:
- /usr/bin/curl
- --fail
- --silent
- --show-error
- --max-time
- "4"
- --header
- "Host: netbox.forust.xyz"
- http://127.0.0.1:8080/login/
initialDelaySeconds: 30
periodSeconds: 30
resources:
requests:
cpu: "100m"
memory: "1Gi"
limits:
cpu: "2"
memory: "2Gi"
volumes:
- name: netbox-config
configMap:
name: netbox-settings
- name: netbox-media
persistentVolumeClaim:
claimName: netbox-media-pvc
- name: netbox-reports
persistentVolumeClaim:
claimName: netbox-reports-pvc
- name: netbox-scripts
persistentVolumeClaim:
claimName: netbox-scripts-pvc
---
apiVersion: apps/v1
kind: Deployment
metadata:
name: netbox-worker-deployment
namespace: netbox
labels:
app: netbox-worker
spec:
replicas: 1
progressDeadlineSeconds: 300
selector:
matchLabels:
app: netbox-worker
strategy:
type: Recreate
template:
metadata:
labels:
app: netbox-worker
spec:
containers:
- name: netbox-worker
image: docker.io/netboxcommunity/netbox:v4.7-5.1.1
command:
- /opt/netbox/venv/bin/python
- netbox/manage.py
- rqworker
workingDir: /opt/netbox
envFrom:
- configMapRef:
name: netbox-config
- secretRef:
name: netbox-secrets
volumeMounts:
- name: netbox-config
mountPath: /etc/netbox/config
readOnly: true
- name: netbox-media
mountPath: /opt/netbox/netbox/media
- name: netbox-reports
mountPath: /opt/netbox/netbox/reports
- name: netbox-scripts
mountPath: /opt/netbox/netbox/scripts
resources:
requests:
cpu: "50m"
memory: "256Mi"
limits:
cpu: "1"
memory: "512Mi"
volumes:
- name: netbox-config
configMap:
name: netbox-settings
- name: netbox-media
persistentVolumeClaim:
claimName: netbox-media-pvc
- name: netbox-reports
persistentVolumeClaim:
claimName: netbox-reports-pvc
- name: netbox-scripts
persistentVolumeClaim:
claimName: netbox-scripts-pvc
---
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: netbox-media-pvc
namespace: netbox
spec:
accessModes: ["ReadWriteOnce"]
resources:
requests:
storage: 2Gi
---
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: netbox-reports-pvc
namespace: netbox
spec:
accessModes: ["ReadWriteOnce"]
resources:
requests:
storage: 1Gi
---
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: netbox-scripts-pvc
namespace: netbox
spec:
accessModes: ["ReadWriteOnce"]
resources:
requests:
storage: 1Gi
+18
View File
@@ -0,0 +1,18 @@
apiVersion: v1
kind: Secret
metadata:
name: netbox-secrets
namespace: netbox
type: Opaque
stringData:
DB_NAME: "netbox"
DB_USER: "netbox"
DB_PASSWORD: "CHANGE_ME_POSTGRES_PASSWORD"
REDIS_PASSWORD: "CHANGE_ME_VALKEY_PASSWORD"
REDIS_CACHE_PASSWORD: "CHANGE_ME_VALKEY_PASSWORD"
VALKEY_PASSWORD: "CHANGE_ME_VALKEY_PASSWORD"
SECRET_KEY: "CHANGE_ME_DJANGO_SECRET_KEY"
API_TOKEN_PEPPER_1: "CHANGE_ME_API_TOKEN_PEPPER"
SUPERUSER_NAME: "admin"
SUPERUSER_EMAIL: "admin@example.com"
SUPERUSER_PASSWORD: "CHANGE_ME_SUPERUSER_PASSWORD"
+56
View File
@@ -0,0 +1,56 @@
apiVersion: v1
kind: ConfigMap
metadata:
name: netbox-settings
namespace: netbox
data:
# Sync wit netbox/configuration/configuration.py (the Docker mounts that file).
configuration.py: |
import os
def _csv(name, default=""):
return [item.strip() for item in os.environ.get(name, default).split(",") if item.strip()]
ALLOWED_HOSTS = _csv("ALLOWED_HOSTS", "localhost,127.0.0.1,[::1]")
CSRF_TRUSTED_ORIGINS = _csv("CSRF_TRUSTED_ORIGINS")
USE_X_FORWARDED_HOST = True
SECURE_PROXY_SSL_HEADER = ("HTTP_X_FORWARDED_PROTO", "https")
DATABASES = {
"default": {
"NAME": os.environ["DB_NAME"],
"USER": os.environ["DB_USER"],
"PASSWORD": os.environ["DB_PASSWORD"],
"HOST": os.environ["DB_HOST"],
"PORT": os.environ.get("DB_PORT", "5432"),
"OPTIONS": {"sslmode": os.environ.get("DB_SSLMODE", "disable")},
"CONN_MAX_AGE": int(os.environ.get("DB_CONN_MAX_AGE", "300")),
}
}
REDIS = {
"tasks": {
"HOST": os.environ["REDIS_HOST"],
"PORT": int(os.environ.get("REDIS_PORT", "6379")),
"PASSWORD": os.environ["REDIS_PASSWORD"],
"DATABASE": int(os.environ.get("REDIS_DATABASE", "0")),
"SSL": False,
},
"caching": {
"HOST": os.environ["REDIS_CACHE_HOST"],
"PORT": int(os.environ.get("REDIS_CACHE_PORT", "6379")),
"PASSWORD": os.environ["REDIS_CACHE_PASSWORD"],
"DATABASE": int(os.environ.get("REDIS_CACHE_DATABASE", "1")),
"SSL": False,
},
}
SECRET_KEY = os.environ["SECRET_KEY"]
API_TOKEN_PEPPERS = {1: os.environ["API_TOKEN_PEPPER_1"]}
TIME_ZONE = os.environ.get("TIME_ZONE", "UTC")
MEDIA_ROOT = "/opt/netbox/netbox/media"
REPORTS_ROOT = "/opt/netbox/netbox/reports"
SCRIPTS_ROOT = "/opt/netbox/netbox/scripts"
CENSUS_REPORTING_ENABLED = False
+82
View File
@@ -0,0 +1,82 @@
apiVersion: v1
kind: Service
metadata:
name: netbox-valkey
namespace: netbox
labels:
app: netbox-valkey
spec:
clusterIP: None
selector:
app: netbox-valkey
ports:
- name: valkey
port: 6379
targetPort: valkey
---
apiVersion: apps/v1
kind: StatefulSet
metadata:
name: netbox-valkey
namespace: netbox
labels:
app: netbox-valkey
spec:
serviceName: netbox-valkey
replicas: 1
selector:
matchLabels:
app: netbox-valkey
template:
metadata:
labels:
app: netbox-valkey
spec:
containers:
- name: valkey
image: docker.io/valkey/valkey:9.1.2-alpine
command:
- sh
- -c
- valkey-server --appendonly yes --save 30 1 --loglevel warning --requirepass "$VALKEY_PASSWORD"
env:
- name: VALKEY_PASSWORD
valueFrom:
secretKeyRef:
name: netbox-secrets
key: VALKEY_PASSWORD
ports:
- name: valkey
containerPort: 6379
volumeMounts:
- name: valkey-data
mountPath: /data
startupProbe:
exec:
command: ["sh", "-c", 'valkey-cli --pass "$VALKEY_PASSWORD" ping | grep -q PONG']
failureThreshold: 20
periodSeconds: 5
readinessProbe:
exec:
command: ["sh", "-c", 'valkey-cli --pass "$VALKEY_PASSWORD" ping | grep -q PONG']
periodSeconds: 10
livenessProbe:
exec:
command: ["sh", "-c", 'valkey-cli --pass "$VALKEY_PASSWORD" ping | grep -q PONG']
initialDelaySeconds: 20
periodSeconds: 20
resources:
requests:
cpu: "25m"
memory: "64Mi"
limits:
cpu: "250m"
memory: "256Mi"
volumeClaimTemplates:
- metadata:
name: valkey-data
spec:
accessModes: ["ReadWriteOnce"]
resources:
requests:
storage: 1Gi
+1 -1
View File
@@ -1,6 +1,6 @@
services:
netronome:
image: ghcr.io/autobrr/netronome:v0.14.0
image: ghcr.io/autobrr/netronome:v0.14.1
restart: unless-stopped
container_name: netronome
ports:
+3 -3
View File
@@ -30,7 +30,7 @@ spec:
spec:
containers:
- name: netronome
image: ghcr.io/autobrr/netronome:v0.14.0
image: ghcr.io/autobrr/netronome:v0.14.1
ports:
- name: netronome-port
protocol: TCP
@@ -51,8 +51,8 @@ spec:
key: NETRONOME__DB_PASSWORD
resources:
requests:
memory: "100Mi"
memory: "64Mi"
cpu: "100m"
limits:
memory: "512Mi"
memory: "256Mi"
cpu: "500m"
View File
Whitespace-only changes.
View File
Whitespace-only changes.
+2 -1
View File
@@ -1,7 +1,7 @@
# Shared PostgreSQL
This directory contains the shared PostgreSQL 17 deployment for Authentik,
Gitea, Netronome, and Statuspage. It creates one database and one login role
Gitea, NetBox, Netronome, and Statuspage. It creates one database and one login role
per service. Per-service standalone databases were removed after the
migration (Sep 2026); Penpot stays on its own compose PostgreSQL (archived,
not part of the shared instance).
@@ -12,6 +12,7 @@ not part of the shared instance).
| ---------- | ------------------- | -------------------------------------- |
| Authentik | 2025.10.x | Supported (Authentik requires 14+) |
| Gitea | 1.27.3 | Supported (Gitea requires 12+) |
| NetBox | 4.7.x | Supported (NetBox 4.x requires 13+) |
| Netronome | 0.14.0 | Supported (upstream's example uses 17) |
| Statuspage | custom | Supported |
+2
View File
@@ -3,6 +3,7 @@ set -euo pipefail
: "${AUTHENTIK_DB_PASSWORD:?AUTHENTIK_DB_PASSWORD is required}"
: "${GITEA_DB_PASSWORD:?GITEA_DB_PASSWORD is required}"
: "${NETBOX_DB_PASSWORD:?NETBOX_DB_PASSWORD is required}"
: "${NETRONOME_DB_PASSWORD:?NETRONOME_DB_PASSWORD is required}"
: "${PENPOT_DB_PASSWORD:?PENPOT_DB_PASSWORD is required}"
: "${STATUSPAGE_DB_PASSWORD:?STATUSPAGE_DB_PASSWORD is required}"
@@ -23,6 +24,7 @@ SQL
create_role_and_database authentik authentik "$AUTHENTIK_DB_PASSWORD"
create_role_and_database gitea gitea "$GITEA_DB_PASSWORD"
create_role_and_database netbox netbox "$NETBOX_DB_PASSWORD"
create_role_and_database netronome netronome "$NETRONOME_DB_PASSWORD"
create_role_and_database penpot penpot "$PENPOT_DB_PASSWORD"
create_role_and_database statuspage statuspage "$STATUSPAGE_DB_PASSWORD"
+3
View File
@@ -17,6 +17,9 @@ spec:
- namespaceSelector:
matchLabels:
kubernetes.io/metadata.name: gitea
- namespaceSelector:
matchLabels:
kubernetes.io/metadata.name: netbox
- namespaceSelector:
matchLabels:
kubernetes.io/metadata.name: netronome
+13 -1
View File
@@ -60,16 +60,26 @@ spec:
command: ["pg_isready", "-U", "postgres", "-d", "postgres"]
initialDelaySeconds: 10
periodSeconds: 10
# Generous timeout: on an I/O-bound single node even exec can take
# seconds, and a 1s default kills a healthy postgres mid-recovery.
timeoutSeconds: 5
startupProbe:
exec:
command: ["pg_isready", "-U", "postgres", "-d", "postgres"]
failureThreshold: 30
# Crash recovery on an I/O-starved single node can fsync for 10+
# minutes; killing postgres mid-recovery restarts the fsync from
# zero and loops forever. 90x10s = 15 minutes of grace.
failureThreshold: 90
periodSeconds: 10
livenessProbe:
exec:
command: ["pg_isready", "-U", "postgres", "-d", "postgres"]
initialDelaySeconds: 30
periodSeconds: 20
# Same I/O reasoning as readiness, plus more misses before a kill:
# restarting postgres on a loaded node only makes recovery longer.
timeoutSeconds: 5
failureThreshold: 5
resources:
requests:
memory: "512Mi"
@@ -113,6 +123,7 @@ data:
: "${AUTHENTIK_DB_PASSWORD:?AUTHENTIK_DB_PASSWORD is required}"
: "${GITEA_DB_PASSWORD:?GITEA_DB_PASSWORD is required}"
: "${NETBOX_DB_PASSWORD:?NETBOX_DB_PASSWORD is required}"
: "${NETRONOME_DB_PASSWORD:?NETRONOME_DB_PASSWORD is required}"
: "${PENPOT_DB_PASSWORD:?PENPOT_DB_PASSWORD is required}"
: "${STATUSPAGE_DB_PASSWORD:?STATUSPAGE_DB_PASSWORD is required}"
@@ -133,6 +144,7 @@ data:
create_role_and_database authentik authentik "$AUTHENTIK_DB_PASSWORD"
create_role_and_database gitea gitea "$GITEA_DB_PASSWORD"
create_role_and_database netbox netbox "$NETBOX_DB_PASSWORD"
create_role_and_database netronome netronome "$NETRONOME_DB_PASSWORD"
create_role_and_database penpot penpot "$PENPOT_DB_PASSWORD"
create_role_and_database statuspage statuspage "$STATUSPAGE_DB_PASSWORD"
+1
View File
@@ -8,6 +8,7 @@ stringData:
POSTGRES_ADMIN_PASSWORD: ""
AUTHENTIK_DB_PASSWORD: ""
GITEA_DB_PASSWORD: ""
NETBOX_DB_PASSWORD: ""
NETRONOME_DB_PASSWORD: ""
PENPOT_DB_PASSWORD: ""
STATUSPAGE_DB_PASSWORD: ""
+1 -1
View File
@@ -30,7 +30,7 @@ services:
- proxy
prometheus:
image: prom/prometheus:v3.14.0
image: prom/prometheus:v3.15.0
container_name: prometheus-prometheus
restart: unless-stopped
command:
+27
View File
@@ -27,6 +27,33 @@ spec:
summary: "Pod is crash looping"
description: "Container {{ $labels.container }} in {{ $labels.namespace }}/{{ $labels.pod }} is in CrashLoopBackOff."
- alert: ContainerOOMKilled
expr: max_over_time(kube_pod_container_status_terminated_reason{reason="OOMKilled"}[15m]) >= 1
for: 5m
labels:
severity: warning
annotations:
summary: "Container was OOMKilled"
description: "Container {{ $labels.container }} in {{ $labels.namespace }}/{{ $labels.pod }} was killed for exceeding its memory limit. Raise the limit or reduce the workload."
- alert: ContainerRestartingTooOften
expr: max by (namespace, pod, container) (increase(kube_pod_container_status_restarts_total[30m])) > 3
for: 5m
labels:
severity: warning
annotations:
summary: "Container restarting too often"
description: "Container {{ $labels.container }} in {{ $labels.namespace }}/{{ $labels.pod }} restarted {{ $value }} times in the last 30 minutes."
- alert: PodEvicted
expr: max_over_time(kube_pod_status_reason{reason="Evicted"}[15m]) >= 1
for: 5m
labels:
severity: warning
annotations:
summary: "Pod was evicted"
description: "Pod {{ $labels.namespace }}/{{ $labels.pod }} was evicted, usually for node disk or memory pressure."
- alert: PersistentVolumeClaimFillingUp
expr: kubelet_volume_stats_available_bytes / kubelet_volume_stats_capacity_bytes < 0.15
for: 15m
Loaded 100 of 139 files, more files were not shown because too many files have changed in this diff. Show more