ci / lint-compose (push) Successful in 4s
ci / lint-actionlint (push) Successful in 1s
ci / lint-shellcheck (push) Successful in 2s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 0s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 2s
ci / scan-deps (push) Successful in 17s
ci / test-backend (push) Successful in 7s
ci / test-frontend (push) Successful in 10s
ci / validate (push) Successful in 3s
renovate-ci / validate-renovate (push) Successful in 32s
ci / build (push) Successful in 1s
Every bouncer-protected request blocks on a synchronous GET /v1/decisions against the LAPI, and the plugin fails CLOSED with 403 when that lookup exceeds its timeout. The LAPI was capped at 400m/500Mi on a single replica, so idle lookups measured 1.3-7.4s and a deploy burst pushed them past the fork's implicit 10s default: gitea answered 403 for ten seconds straight, and containerd turned those 403s on gcr.forust.xyz into ErrImagePull/ImagePullBackOff on freshly rolled pods. Three changes, plus the 403 feedback loop that made it sticky: * gitea/k8s/ingress.yaml - drop the bouncer from the /v2 registry route. A deploy fires hundreds of parallel authenticated OCI requests (runner Action API, manifest inspect per own image, containerd pulls, smoke probes) and scanners gain nothing from a registry that already does its own token auth. The web UI route keeps the bouncer. * crowdsec-values.yaml - LAPI to 1500m/1Gi. Replicas stay at 1 on purpose: LAPI is stateful (BoltDB plus credentials on two RWO PVCs), so a second replica would corrupt the decision store. * crowdsec-middleware.yaml - CrowdsecLapiTimeout: 2s instead of the implicit 10s, so a slow LAPI costs a fast 403 rather than a 10s hang. LePresidente/http-generic-403-bf then banned us for our own 403s: five POSTs answered 403 within 10s earn a 4h ban, and the hairpin-NAT address 192.168.88.1 that the Gitea Actions runner presents to Traefik is not covered by the home-dynamic-IP whitelist. That scenario cannot be dropped per-scenario - it is baked into the hub item crowdsecurity/http-generic-bf v0.9, and disabling the whole base-http-scenarios collection would cost ~40 useful detections. So the janitor now deletes its decisions hourly and a new forust/lan whitelist postoverflow covers 192.168.88.0/24. First janitor run removed 165 decisions; none of the remaining ones are local. Measured after: gitea 200 in 25-148ms (was 403 at 10001ms), a 60-way parallel burst all 200 with a 422ms max, gcr /v2 back to its 401 auth challenge, LAPI at 60m CPU with no throttling.
215 lines
9.5 KiB
YAML
215 lines
9.5 KiB
YAML
# CrowdSec self-healing: static machine identity + enforcement loops.
|
|
#
|
|
# Problem it fixes: the chart's agent init container runs
|
|
# `cscli lapi register --machine "$POD_NAME" ...`
|
|
# unconditionally. Credentials live in an emptyDir, the machine row lives
|
|
# in LAPI's persistent DB. Any init re-run for an already-known pod name
|
|
# (kubelet restart, node reboot) dies with
|
|
# 403 Forbidden: user '<pod>' already exist
|
|
# and the DaemonSet pod sticks in Init forever. Every DS restart also
|
|
# leaves an orphan machine row that is never cleaned.
|
|
#
|
|
# Design (name-independent):
|
|
# * Agent identity is a STATIC machine `crowdsec-agent-workstation`
|
|
# whose password lives in Secret `crowdsec-agent-credentials`
|
|
# (created once, manually - like all other secrets in this repo).
|
|
# The secret is mounted into agent pods at
|
|
# /tmp_config/local_api_credentials.yaml (see extraVolumeMounts in
|
|
# crowdsec-values.yaml), which is exactly the path the agent's main
|
|
# container copies into place at startup.
|
|
# * The DS init command is patched (strategic merge, by container name)
|
|
# to SKIP registration when that file exists, keeping the legacy
|
|
# register path only as fallback. Detection marker in the patched
|
|
# command: `[ -s /tmp_config`.
|
|
# * This CronJob enforces the desired state hourly, so recovery is
|
|
# automatic even after `helm upgrade` reverts the DS patch or the
|
|
# LAPI database is wiped:
|
|
# 1. patch DS init if it still has the unconditional register
|
|
# (no-op otherwise - no restart churn);
|
|
# 2. prune machines with no heartbeat for 2h (orphan hygiene);
|
|
# 3. ensure the static machine exists, recreating it with the
|
|
# Secret password if missing (agent retry loops reconnect
|
|
# on their own - same name + same password);
|
|
# 4. prune bouncer entries idle for 30d;
|
|
# 5. delete decisions from LePresidente/http-generic-403-bf, a hub
|
|
# scenario that bans an IP for 4h after 5 POST-403s in 10s and
|
|
# therefore bans us for our own bouncer's fail-closed 403s.
|
|
#
|
|
# Manual apply (crowdsec/k8s is NOT managed by deploy.yaml):
|
|
# kubectl apply -f crowdsec/k8s/janitor-cronjob.yaml
|
|
# Force a run:
|
|
# kubectl create job -n crowdsec --from=cronjob/crowdsec-janitor janitor-now
|
|
#
|
|
# Helm upgrades: the janitor's strategic patch puts the DS field under
|
|
# the `kubectl-patch` field manager, so a plain `helm upgrade` FAILS
|
|
# with an SSA conflict on initContainers[].command. Procedure:
|
|
# 1. revert init to chart state (kills the conflict):
|
|
# helm template crowdsec crowdsec/crowdsec --version <ver> \
|
|
# -n crowdsec -f crowdsec/k8s/crowdsec-values.yaml > /tmp/r.yaml
|
|
# python3 -c "import yaml,json; ..." # build revert patch from
|
|
# the rendered DaemonSet init command, then
|
|
# kubectl patch ds crowdsec-agent -n crowdsec \
|
|
# --type strategic -p "\$(cat /tmp/revert_patch.json)"
|
|
# 2. helm upgrade --install crowdsec ... (no --force needed)
|
|
# 3. janitor-now right away (upgrade reverts init; new pods would
|
|
# sit in Init until the next hourly run otherwise).
|
|
#
|
|
# One-time bootstrap (order matters):
|
|
# 1. Create Secret + static machine (see commands in chat).
|
|
# 2. Apply this file, trigger janitor-now, wait for agent 1/1.
|
|
# 3. One-time orphan cleanup:
|
|
# kubectl exec -n crowdsec deploy/crowdsec-lapi -- \
|
|
# cscli machines prune --duration 1h --force
|
|
# 4. Only then `helm upgrade` crowdsec with the extraVolumes values.
|
|
# Upgrade reverts the DS patch; trigger janitor-now right after it
|
|
# (otherwise new pods sit in Init until the next hourly run, then
|
|
# self-heal anyway).
|
|
#
|
|
# Password rotation: update the Secret, delete the machine
|
|
# (`cscli machines delete crowdsec-agent-workstation`), trigger
|
|
# janitor-now (recreates it), then `kubectl rollout restart
|
|
# ds/crowdsec-agent -n crowdsec` (agent reads the file at startup only).
|
|
apiVersion: v1
|
|
kind: ServiceAccount
|
|
metadata:
|
|
name: crowdsec-janitor
|
|
namespace: crowdsec
|
|
labels:
|
|
app.kubernetes.io/part-of: crowdsec
|
|
---
|
|
apiVersion: rbac.authorization.k8s.io/v1
|
|
kind: Role
|
|
metadata:
|
|
name: crowdsec-janitor
|
|
namespace: crowdsec
|
|
labels:
|
|
app.kubernetes.io/part-of: crowdsec
|
|
rules:
|
|
- apiGroups: [""]
|
|
resources: ["pods"]
|
|
verbs: ["get", "list"]
|
|
- apiGroups: [""]
|
|
resources: ["pods/exec"]
|
|
verbs: ["create"]
|
|
- apiGroups: ["apps"]
|
|
resources: ["daemonsets"]
|
|
verbs: ["get", "patch"]
|
|
# `kubectl exec deploy/<name>` resolves deploy -> replicaset -> pod,
|
|
# which needs read access to these (exec itself is pods/exec above).
|
|
- apiGroups: ["apps"]
|
|
resources: ["deployments", "replicasets"]
|
|
verbs: ["get", "list"]
|
|
---
|
|
apiVersion: rbac.authorization.k8s.io/v1
|
|
kind: RoleBinding
|
|
metadata:
|
|
name: crowdsec-janitor
|
|
namespace: crowdsec
|
|
labels:
|
|
app.kubernetes.io/part-of: crowdsec
|
|
subjects:
|
|
- kind: ServiceAccount
|
|
name: crowdsec-janitor
|
|
namespace: crowdsec
|
|
roleRef:
|
|
kind: Role
|
|
name: crowdsec-janitor
|
|
apiGroup: rbac.authorization.k8s.io
|
|
---
|
|
apiVersion: batch/v1
|
|
kind: CronJob
|
|
metadata:
|
|
name: crowdsec-janitor
|
|
namespace: crowdsec
|
|
labels:
|
|
app.kubernetes.io/part-of: crowdsec
|
|
spec:
|
|
schedule: "17 * * * *"
|
|
concurrencyPolicy: Forbid
|
|
successfulJobsHistoryLimit: 3
|
|
failedJobsHistoryLimit: 3
|
|
jobTemplate:
|
|
spec:
|
|
activeDeadlineSeconds: 300
|
|
template:
|
|
metadata:
|
|
labels:
|
|
app.kubernetes.io/part-of: crowdsec
|
|
spec:
|
|
serviceAccountName: crowdsec-janitor
|
|
restartPolicy: OnFailure
|
|
containers:
|
|
- name: janitor
|
|
# Same image the chart itself uses for registration jobs;
|
|
# IfNotPresent so it works while the node is offline
|
|
# (layer cached from the chart install).
|
|
image: alpine/kubectl:latest
|
|
imagePullPolicy: IfNotPresent
|
|
env:
|
|
- name: AGENT_PASSWORD
|
|
valueFrom:
|
|
secretKeyRef:
|
|
name: crowdsec-agent-credentials
|
|
key: password
|
|
command:
|
|
- /bin/sh
|
|
- -c
|
|
- |
|
|
set -eu
|
|
LAPI_EXEC="kubectl exec -n crowdsec deploy/crowdsec-lapi --"
|
|
echo "== 1. enforce patched agent init =="
|
|
CUR=$(kubectl get ds crowdsec-agent -n crowdsec \
|
|
-o jsonpath='{.spec.template.spec.initContainers[0].command[2]}')
|
|
case "$CUR" in
|
|
*'-s /tmp_config'*)
|
|
echo "init already patched"
|
|
;;
|
|
*)
|
|
echo "patching init"
|
|
WAIT='until nc "$LAPI_HOST" "$LAPI_PORT" -z'
|
|
WAIT="$WAIT; do echo waiting for lapi to start; sleep 5; done"
|
|
LINK='ln -s /staging/etc/crowdsec /etc/crowdsec'
|
|
REG='cscli lapi register --machine "$USERNAME"'
|
|
REG="$REG -u \"\$LAPI_URL\" --token \"\$REGISTRATION_TOKEN\""
|
|
CREDS=/tmp_config/local_api_credentials.yaml
|
|
CMD="$WAIT; $LINK; [ -s $CREDS ] || {"
|
|
CMD="$CMD $REG && cp"
|
|
CMD="$CMD /etc/crowdsec/local_api_credentials.yaml $CREDS; }"
|
|
ESC=$(printf '%s' "$CMD" | sed 's/"/\\"/g')
|
|
PATCH='{"spec":{"template":{"spec":{"initContainers":'
|
|
PATCH=$PATCH'[{"name":"wait-for-lapi-and-register",'
|
|
PATCH=$PATCH'"command":["sh","-c","'$ESC'"]}]}}}}'
|
|
kubectl patch ds crowdsec-agent -n crowdsec \
|
|
--type strategic -p "$PATCH"
|
|
;;
|
|
esac
|
|
echo "== 2. prune orphan machines (no heartbeat for 2h) =="
|
|
$LAPI_EXEC cscli machines prune --duration 2h --force
|
|
echo "== 3. ensure static machine exists =="
|
|
if $LAPI_EXEC cscli machines inspect \
|
|
crowdsec-agent-workstation >/dev/null 2>&1; then
|
|
echo "static machine present"
|
|
else
|
|
echo "recreating static machine"
|
|
$LAPI_EXEC cscli machines add crowdsec-agent-workstation \
|
|
--password "$AGENT_PASSWORD" --force
|
|
fi
|
|
echo "== 4. prune stale bouncers (no pull for 30d) =="
|
|
$LAPI_EXEC cscli bouncers prune -d 720h --force
|
|
echo "== 5. drop http-403-bf decisions (4h self-bans) =="
|
|
# `LePresidente/http-generic-403-bf` (hub item
|
|
# crowdsecurity/http-generic-bf v0.9) bans any source IP
|
|
# after 5 POSTs answered 403 within 10s, for 4h. That
|
|
# includes 403s this homelab generates ITSELF (any
|
|
# bouncer fail-closed, any app CSRF/rate-limit 403), and a
|
|
# 4h ban on the runner/home IP silently breaks deploys and
|
|
# browsing. The scenario cannot be removed per-scenario -
|
|
# it is baked into a hub item, and disabling the whole
|
|
# base-http-scenarios collection would drop ~40 useful
|
|
# detections. Instead we keep the detection and drop its
|
|
# decisions hourly; the LAN/home whitelists in
|
|
# crowdsec-values.yaml handle the legit sources, so this
|
|
# only ever hits real scanners (who are re-banned anyway).
|
|
$LAPI_EXEC cscli decisions delete \
|
|
--scenario LePresidente/http-generic-403-bf --all || true
|