Commit Graph
8 Commits
Author SHA1 Message Date
forust a6af69dca0 fix(k8s): Recreate singletons and trim requests for scheduler headroom
ci / lint-actionlint (push) Successful in 1s
ci / lint-shellcheck (push) Successful in 2s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 2s
ci / lint-compose (push) Successful in 3s
ci / test-backend (push) Successful in 8s
ci / test-frontend (push) Successful in 11s
ci / validate (push) Successful in 3s
renovate-ci / validate-renovate (push) Successful in 13s
ci / lint-dockerfiles (push) Successful in 2s
ci / scan-deps (push) Successful in 18s
ci / build (push) Successful in 1m30s
RollingUpdate with default maxSurge needs a spare pod the single node does not have (99% CPU requested), so multi-workload restarts end Pending and verify times out. Recreate on all replicas:1 Deployments (immich-server and bentopdf keep RollingUpdate at replicas 2). Also trims CPU/memory requests toward measured use (adguard, authentik, gitea, netbox, uptime-kuma, netbird-server) and gives traefik requests/limits so it is no longer BestEffort.
2026-09-28 22:36:46 +02:00
forust f9e4623ade fix(k8s): size the remaining workloads against measured use
Finishes the sizing pass over every workload the deploy actually manages. Each
request is at or above the container's p95 over the last seven days, so nothing
is sized below what it is known to use, and each limit is between 1.6x and 5x
the observed max, which is the figure that decides whether a burst gets an
OOMKill.

Some of these go up, and that is the point. adguard was holding 975M against a
500Mi request and netbox 962M against 512Mi, so both sat permanently above
their own request and were standing eviction candidates on a node that has
about 300M of headroom. Raising a request costs scheduler room; leaving it low
costs the pod its place in the queue when the node gets tight.

Others come down. loki ran with a 2Gi limit on 224M, gitea 1.5Gi on 305M, the
authentik worker 1Gi on 305M, and a tail of single-purpose pods -- redis,
glance, the two homepages, cfddns, session-keeper, the netbird dashboard, the
loki gateway -- each reserved 4x to 16x more than they have ever touched.

prometheus gets the opposite treatment: 768Mi/2560Mi, above its p95, because it
compacts its TSDB in place and that is a burst worth budgeting for rather than
throttling.

Two of these limits are close enough to the observed max to be worth watching
rather than trusting: adguard at 1.5x, and its DNS cache grows monotonically, so
the ceiling is a date, not a margin. That was true before this change too; the
pod sizing does not fix it and the cache needs bounding.

CPU limits are untouched throughout. Leaving postgres alone as well: it sits in
an uncommitted file that belongs to other work in progress.

Verified: every request is at or above p95 and every limit above the observed
max across all 74 containers, and 16/16 local gates pass.
2026-09-28 10:38:19 +02:00
forust 2a4f215546 fix(k8s): bring the over-reserved memory limits down to measured use
Six pods reserved far more memory than they have ever touched. uptime-kuma held
a 3Gi limit against 469M of measured p95, metube 2Gi against 72M, convertx
1.5Gi against 85M, netbird-server 1Gi against 97M, searxng 700Mi against 134M
and bentopdf 700Mi against 4M. Every one of them is a ceiling the scheduler
counts against the node while the memory sits unused.

Requests move down with the limits but never below the measured p95, so none of
these becomes an eviction candidate as a side effect of being right-sized. The
limits keep between 2.2x and 11.6x over the observed max, which is the figure
that decides whether a pod gets OOM-killed during a burst.

Net effect across the six: requests -557M, limits -4.6Gi, all of it ceiling that
was never in use. This is the first change that actually gives memory back.

CPU limits are left exactly as they were. They were not part of the sizing pass,
they are not being hit on a node sitting at 5% CPU, and removing them is a
separate decision from moving memory.

Verified: each limit is above the container's own observed max and each request
is above its p95, and 16/16 local gates pass.
2026-09-28 10:20:24 +02:00
forust 892790822d fix(crowdsec): stop the 403 loop that banned our own VPS and runner
Chasing why apply-k8s kept dying mid-run turned up a self-inflicted
ban loop. 585 of 586 LePresidente/http-generic-403-bf alerts in the LAPI
came from 193.181.211.79 - our own VPS - POSTing
/management.ManagementService/GetServerKey, i.e. the NetBird client's
own management call. netbird-server had never seen a single one of them,
so the 403 was not NetBird's: the bouncer was rejecting the request
before it got there. A banned peer keeps retrying, each retry is another
403, and the scenario turns five 403s in ten seconds into a 4h ban, so
the loop kept re-arming the ban it was serving. The hourly janitor step
that deleted those decisions hourly was masking all of it.

* netbird/k8s/ingress.yaml - drop the bouncer from the mesh API routes
  (gRPC-gateway management, signal, relay, /api, /oauth2). A ban there
  locks a peer out of the network it needs to reach anything else, and
  those endpoints authenticate by NetBird token, not by a login form.
  The dashboard keeps the bouncer; it is a real login surface.
  netbird-local was already exempt, so this makes prod match.

* crowdsec-middleware.yaml - CrowdsecMode stream instead of live. live
  blocked on GET /v1/decisions per request, so a burst saturated the
  LAPI and the plugin 403'd IPs that were never banned. v1.3.3 ignores
  UpdateMaxFailure in live, so fail-open is only reachable in stream;
  -1 now means an unreachable LAPI degrades to "no protection" rather
  than "every site 403". 15s poll instead of the 60s default, because
  the runner shares one public IP with the house.

  Also corrects the key name: HTTPTimeoutSeconds, not
  CrowdsecLapiTimeout, which never existed and was being silently
  dropped, leaving the 10s default. Back at 10, not the 2s f7cd75d
  guessed - a pull that times out leaves the ban cache frozen at its
  startup contents, so new bans would silently never apply.

* crowdsec-values.yaml - CIDR allowlisting moves to parsers/s02-enrich,
  where CrowdSec's docs put it: a parser whitelist drops the event before
  it reaches a bucket, so those addresses never become a decision at
  all. The old postoverflow LAN list was checked only after the ban
  existed, which is the window the deploy kept landing in. Added
  100.64.0.0/10, which the RFC 1918 blocks miss and where the
  workstation, the k0s node and the VPS actually live. The DDNS
  home-IP whitelist stays in postoverflows, because resolving a hostname
  is the expensive check the docs reserve that stage for.

  ClientTrustedIPs mirrors that list so the bouncer skips the LAPI
  round-trip entirely for those addresses.

* janitor-cronjob.yaml - drop step 5. The bouncer can no longer
  manufacture 403s, so the only remaining firings of that scenario are
  real scanners, and deleting their decisions hourly was undoing a
  working ban.

* Also lands the LAPI config.yaml.local (SQLite WAL, Central API off,
  bounded flush) that f7cd75d's comments referenced but never included:
  the LAPI was blocked in fsync on its rollback journal, and the CAPI
  resolver held a write transaction while timing out against a host
  this network cannot reach.

Verified in-cluster: no new 403-bf alerts in the 3.5min after applying,
the management endpoint answers 404 from the backend instead of 403 from
the bouncer in 0.17s, and the VPS client reports Management and Signal
connected with 2/2 relays.
2026-09-27 18:09:07 +02:00
forust 41f18ea993 fix(deploy): move netbird from compose to k8s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 3s
ci / lint-dockerfiles (push) Successful in 1s
ci / validate (push) Successful in 2s
deploy / preflight (push) Successful in 2s
renovate-ci / validate-renovate (push) Successful in 19s
ci / build (push) Successful in 2s
deploy / validate (push) Successful in 1m45s
deploy / apply-k8s (push) Successful in 1m52s
deploy / apply-compose (push) Failing after 24s
netbird runs in-cluster; the root active marker made apply-compose pick up netbird/compose.yaml and fail on gitignored .env vars. Drop the compose marker and enable k8s/active instead.
2026-09-26 17:24:03 +02:00
renovate-bot f1f7dd4a0a chore(deps): update netbirdio/dashboard docker tag to v2.93.0
deploy / validate (push) Skipped
renovate-ci / validate-renovate (push) Skipped
ci / lint-prettier (pull_request) Successful in 4s
ci / lint-ruff (pull_request) Successful in 2s
ci / lint-yaml (pull_request) Successful in 2s
ci / lint-dockerfiles (pull_request) Successful in 1s
ci / validate (pull_request) Successful in 2s
ci / build (pull_request) Skipped
renovate-ci / validate-renovate (pull_request) Successful in 9s
ci / lint-prettier (push) Successful in 2s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 1s
ci / validate (push) Successful in 1s
ci / build (push) Skipped
2026-09-25 22:26:10 +00:00
forust ad4bb8750d chore(deploy): enable netbird
deploy / validate (push) Skipped
renovate-ci / validate-renovate (push) Skipped
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 1s
ci / validate (push) Successful in 2s
ci / build (push) Skipped
ci / lint-prettier (pull_request) Successful in 3s
ci / lint-ruff (pull_request) Successful in 1s
ci / lint-yaml (pull_request) Successful in 2s
ci / lint-dockerfiles (pull_request) Successful in 1s
ci / validate (pull_request) Successful in 1s
ci / build (pull_request) Skipped
renovate-ci / validate-renovate (pull_request) Successful in 8s
2026-09-25 22:16:02 +00:00
forust a4a4bb4cc5 feat(netbird): add tailscale-alternative
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 1s
ci / validate (push) Successful in 2s
ci / build (push) Skipped
deploy / validate (push) Skipped
renovate-ci / validate-renovate (push) Skipped
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-prettier (pull_request) Successful in 3s
ci / lint-ruff (pull_request) Successful in 1s
ci / lint-yaml (pull_request) Successful in 5s
ci / lint-dockerfiles (pull_request) Successful in 1s
ci / validate (pull_request) Successful in 1s
ci / build (pull_request) Skipped
renovate-ci / validate-renovate (pull_request) Successful in 10s
2026-09-25 20:26:11 +02:00