Every bouncer-protected request blocks on a synchronous
GET /v1/decisions against the LAPI, and the plugin fails CLOSED with 403
when that lookup exceeds its timeout. The LAPI was capped at 400m/500Mi
on a single replica, so idle lookups measured 1.3-7.4s and a deploy
burst pushed them past the fork's implicit 10s default: gitea answered
403 for ten seconds straight, and containerd turned those 403s on
gcr.forust.xyz into ErrImagePull/ImagePullBackOff on freshly rolled
pods.
Three changes, plus the 403 feedback loop that made it sticky:
* gitea/k8s/ingress.yaml - drop the bouncer from the /v2 registry route.
A deploy fires hundreds of parallel authenticated OCI requests
(runner Action API, manifest inspect per own image, containerd pulls,
smoke probes) and scanners gain nothing from a registry that already
does its own token auth. The web UI route keeps the bouncer.
* crowdsec-values.yaml - LAPI to 1500m/1Gi. Replicas stay at 1 on
purpose: LAPI is stateful (BoltDB plus credentials on two RWO PVCs),
so a second replica would corrupt the decision store.
* crowdsec-middleware.yaml - CrowdsecLapiTimeout: 2s instead of the
implicit 10s, so a slow LAPI costs a fast 403 rather than a 10s hang.
LePresidente/http-generic-403-bf then banned us for our own 403s: five
POSTs answered 403 within 10s earn a 4h ban, and the hairpin-NAT
address 192.168.88.1 that the Gitea Actions runner presents to Traefik
is not covered by the home-dynamic-IP whitelist. That scenario cannot
be dropped per-scenario - it is baked into the hub item
crowdsecurity/http-generic-bf v0.9, and disabling the whole
base-http-scenarios collection would cost ~40 useful detections. So the
janitor now deletes its decisions hourly and a new forust/lan
whitelist postoverflow covers 192.168.88.0/24. First janitor run
removed 165 decisions; none of the remaining ones are local.
Measured after: gitea 200 in 25-148ms (was 403 at 10001ms), a 60-way
parallel burst all 200 with a 422ms max, gcr /v2 back to its 401 auth
challenge, LAPI at 60m CPU with no throttling.
Deploy workflow uses git-tracked manifests, DISABLED flag and kustomize overlays; add webinar-checker metrics with ServiceMonitor and alerts; upgrade shared postgres to 17 with statuspage DB and probes/resources.
All prod IngressRoutes switch tls.certResolver to tls.secretName
backed by per-router Certificates (HTTP-01, letsencrypt-prod).
adguard-prod reuses the shared adguard-certs secret (also feeds
DoT :853); sync CronJob removed as redundant.
Traefik certificatesResolvers removed: its internal
acme-http@internal router hijacks HTTP-01 for every host while
enabled, blocking external solvers. Dormant files (kener,
downtify) converted for consistency but not applied; n8n
untouched per live-only rule.
Move the shared postgres service from 15.19 to 17.6 as the postgres17 StatefulSet with its own PVC, extend the initdb and ingress policy with the statuspage database, and drop the now-unused per-app postgres manifests for authentik, gitea and netronome.
Protect public Traefik routes with CrowdSec HTTP decisions and restore access logging for web traffic analysis.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
services marked k8s/active are applied via kubectl; the rest via docker
compose. inactive services with k8s/ keep only routing manifests
(external Services, EndpointSlices, Ingresses) to reach docker backends.
headscale/nextcloud routing moved to k8s/routing/.
validations: compose config --quiet + kubectl apply --dry-run=client.
namespace manifests applied first. pull_policy:build stacks get
build+push before up so the registry image stays fresh.