3ecc12300aa5ccb5a39b88e9b78b8345da6f7cd9
9
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
b53ce36d89 | chore(deps): update netbirdio/dashboard docker tag to v2.94.0 | ||
|
|
a6af69dca0 |
fix(k8s): Recreate singletons and trim requests for scheduler headroom
ci / lint-actionlint (push) Successful in 1s
ci / lint-shellcheck (push) Successful in 2s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 2s
ci / lint-compose (push) Successful in 3s
ci / test-backend (push) Successful in 8s
ci / test-frontend (push) Successful in 11s
ci / validate (push) Successful in 3s
renovate-ci / validate-renovate (push) Successful in 13s
ci / lint-dockerfiles (push) Successful in 2s
ci / scan-deps (push) Successful in 18s
ci / build (push) Successful in 1m30s
RollingUpdate with default maxSurge needs a spare pod the single node does not have (99% CPU requested), so multi-workload restarts end Pending and verify times out. Recreate on all replicas:1 Deployments (immich-server and bentopdf keep RollingUpdate at replicas 2). Also trims CPU/memory requests toward measured use (adguard, authentik, gitea, netbox, uptime-kuma, netbird-server) and gives traefik requests/limits so it is no longer BestEffort. |
||
|
|
f9e4623ade |
fix(k8s): size the remaining workloads against measured use
Finishes the sizing pass over every workload the deploy actually manages. Each request is at or above the container's p95 over the last seven days, so nothing is sized below what it is known to use, and each limit is between 1.6x and 5x the observed max, which is the figure that decides whether a burst gets an OOMKill. Some of these go up, and that is the point. adguard was holding 975M against a 500Mi request and netbox 962M against 512Mi, so both sat permanently above their own request and were standing eviction candidates on a node that has about 300M of headroom. Raising a request costs scheduler room; leaving it low costs the pod its place in the queue when the node gets tight. Others come down. loki ran with a 2Gi limit on 224M, gitea 1.5Gi on 305M, the authentik worker 1Gi on 305M, and a tail of single-purpose pods -- redis, glance, the two homepages, cfddns, session-keeper, the netbird dashboard, the loki gateway -- each reserved 4x to 16x more than they have ever touched. prometheus gets the opposite treatment: 768Mi/2560Mi, above its p95, because it compacts its TSDB in place and that is a burst worth budgeting for rather than throttling. Two of these limits are close enough to the observed max to be worth watching rather than trusting: adguard at 1.5x, and its DNS cache grows monotonically, so the ceiling is a date, not a margin. That was true before this change too; the pod sizing does not fix it and the cache needs bounding. CPU limits are untouched throughout. Leaving postgres alone as well: it sits in an uncommitted file that belongs to other work in progress. Verified: every request is at or above p95 and every limit above the observed max across all 74 containers, and 16/16 local gates pass. |
||
|
|
2a4f215546 |
fix(k8s): bring the over-reserved memory limits down to measured use
Six pods reserved far more memory than they have ever touched. uptime-kuma held a 3Gi limit against 469M of measured p95, metube 2Gi against 72M, convertx 1.5Gi against 85M, netbird-server 1Gi against 97M, searxng 700Mi against 134M and bentopdf 700Mi against 4M. Every one of them is a ceiling the scheduler counts against the node while the memory sits unused. Requests move down with the limits but never below the measured p95, so none of these becomes an eviction candidate as a side effect of being right-sized. The limits keep between 2.2x and 11.6x over the observed max, which is the figure that decides whether a pod gets OOM-killed during a burst. Net effect across the six: requests -557M, limits -4.6Gi, all of it ceiling that was never in use. This is the first change that actually gives memory back. CPU limits are left exactly as they were. They were not part of the sizing pass, they are not being hit on a node sitting at 5% CPU, and removing them is a separate decision from moving memory. Verified: each limit is above the container's own observed max and each request is above its p95, and 16/16 local gates pass. |
||
|
|
892790822d |
fix(crowdsec): stop the 403 loop that banned our own VPS and runner
Chasing why apply-k8s kept dying mid-run turned up a self-inflicted
ban loop. 585 of 586 LePresidente/http-generic-403-bf alerts in the LAPI
came from 193.181.211.79 - our own VPS - POSTing
/management.ManagementService/GetServerKey, i.e. the NetBird client's
own management call. netbird-server had never seen a single one of them,
so the 403 was not NetBird's: the bouncer was rejecting the request
before it got there. A banned peer keeps retrying, each retry is another
403, and the scenario turns five 403s in ten seconds into a 4h ban, so
the loop kept re-arming the ban it was serving. The hourly janitor step
that deleted those decisions hourly was masking all of it.
* netbird/k8s/ingress.yaml - drop the bouncer from the mesh API routes
(gRPC-gateway management, signal, relay, /api, /oauth2). A ban there
locks a peer out of the network it needs to reach anything else, and
those endpoints authenticate by NetBird token, not by a login form.
The dashboard keeps the bouncer; it is a real login surface.
netbird-local was already exempt, so this makes prod match.
* crowdsec-middleware.yaml - CrowdsecMode stream instead of live. live
blocked on GET /v1/decisions per request, so a burst saturated the
LAPI and the plugin 403'd IPs that were never banned. v1.3.3 ignores
UpdateMaxFailure in live, so fail-open is only reachable in stream;
-1 now means an unreachable LAPI degrades to "no protection" rather
than "every site 403". 15s poll instead of the 60s default, because
the runner shares one public IP with the house.
Also corrects the key name: HTTPTimeoutSeconds, not
CrowdsecLapiTimeout, which never existed and was being silently
dropped, leaving the 10s default. Back at 10, not the 2s
|
||
|
|
41f18ea993 |
fix(deploy): move netbird from compose to k8s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 3s
ci / lint-dockerfiles (push) Successful in 1s
ci / validate (push) Successful in 2s
deploy / preflight (push) Successful in 2s
renovate-ci / validate-renovate (push) Successful in 19s
ci / build (push) Successful in 2s
deploy / validate (push) Successful in 1m45s
deploy / apply-k8s (push) Successful in 1m52s
deploy / apply-compose (push) Failing after 24s
netbird runs in-cluster; the root active marker made apply-compose pick up netbird/compose.yaml and fail on gitignored .env vars. Drop the compose marker and enable k8s/active instead. |
||
|
|
f1f7dd4a0a |
chore(deps): update netbirdio/dashboard docker tag to v2.93.0
deploy / validate (push) Skipped
renovate-ci / validate-renovate (push) Skipped
ci / lint-prettier (pull_request) Successful in 4s
ci / lint-ruff (pull_request) Successful in 2s
ci / lint-yaml (pull_request) Successful in 2s
ci / lint-dockerfiles (pull_request) Successful in 1s
ci / validate (pull_request) Successful in 2s
ci / build (pull_request) Skipped
renovate-ci / validate-renovate (pull_request) Successful in 9s
ci / lint-prettier (push) Successful in 2s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 1s
ci / validate (push) Successful in 1s
ci / build (push) Skipped
|
||
|
|
ad4bb8750d |
chore(deploy): enable netbird
deploy / validate (push) Skipped
renovate-ci / validate-renovate (push) Skipped
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 1s
ci / validate (push) Successful in 2s
ci / build (push) Skipped
ci / lint-prettier (pull_request) Successful in 3s
ci / lint-ruff (pull_request) Successful in 1s
ci / lint-yaml (pull_request) Successful in 2s
ci / lint-dockerfiles (pull_request) Successful in 1s
ci / validate (pull_request) Successful in 1s
ci / build (pull_request) Skipped
renovate-ci / validate-renovate (pull_request) Successful in 8s
|
||
|
|
a4a4bb4cc5 |
feat(netbird): add tailscale-alternative
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 1s
ci / validate (push) Successful in 2s
ci / build (push) Skipped
deploy / validate (push) Skipped
renovate-ci / validate-renovate (push) Skipped
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-prettier (pull_request) Successful in 3s
ci / lint-ruff (pull_request) Successful in 1s
ci / lint-yaml (pull_request) Successful in 5s
ci / lint-dockerfiles (pull_request) Successful in 1s
ci / validate (pull_request) Successful in 1s
ci / build (pull_request) Skipped
renovate-ci / validate-renovate (pull_request) Successful in 10s
|