feat(ingress): replace traefik crowdsec plugin with firewall bouncer
ci / lint-compose (push) Successful in 9s
ci / lint-actionlint (push) Successful in 4s
ci / lint-shellcheck (push) Successful in 7s
ci / lint-prettier (push) Successful in 12s
ci / lint-ruff (push) Successful in 6s
ci / lint-yaml (push) Successful in 9s
ci / lint-dockerfiles (push) Successful in 5s
ci / validate (push) Successful in 5s
renovate-ci / validate-renovate (push) Successful in 7s
ci / build (push) Failing after 14m22s
ci / lint-compose (push) Successful in 9s
ci / lint-actionlint (push) Successful in 4s
ci / lint-shellcheck (push) Successful in 7s
ci / lint-prettier (push) Successful in 12s
ci / lint-ruff (push) Successful in 6s
ci / lint-yaml (push) Successful in 9s
ci / lint-dockerfiles (push) Successful in 5s
ci / validate (push) Successful in 5s
renovate-ci / validate-renovate (push) Successful in 7s
ci / build (push) Failing after 14m22s
Move L3 enforcement to the host firewall-bouncer (systemd, nftables): drop the Traefik plugin, its secrets volume and the crowdsec Middleware, remove bouncer refs from all IngressRoutes. Disable the http-generic-bf scenario (403-burst bans hurt legit automation under L3 enforcement). Add a Gateway API PoC for homepages prod and CrowdSec PrometheusRule alerts.
This commit is contained in:
1 parent
872f64b887
commit
0859479c0f
27 files changed
+151
-148
No files matched your search
@@ -1,69 +0,0 @@
|
||||
# crowdsec/k8s is NOT managed by deploy.yaml - apply this by hand, and apply it
|
||||
# together with a restart:
|
||||
# kubectl apply -f crowdsec/k8s/crowdsec-middleware.yaml
|
||||
# kubectl -n traefik rollout restart deploy/traefik
|
||||
#
|
||||
# The restart is not optional. In stream mode the plugin runs a package-level
|
||||
# ticker goroutine (handleStreamTicker over the isCrowdsecStreamHealthy and
|
||||
# updateFailure globals) that no reconfiguration stops. Applying a change
|
||||
# wedges the instance: every route referencing it answers 404 and traefik logs
|
||||
# 'invalid middleware crowdsec-crowdsec-bouncer@kubernetescrd' until the pod is
|
||||
# replaced. Re-applying the previous config does NOT recover it, and the config
|
||||
# is not the cause - a valid CIDR cannot fail NewChecker, which is a plain
|
||||
# net.ParseCIDR. Only a new pod clears it. Measured cost: ~35s down for all
|
||||
# 20 hosts behind this middleware.
|
||||
apiVersion: traefik.io/v1alpha1
|
||||
kind: Middleware
|
||||
metadata:
|
||||
name: crowdsec-bouncer
|
||||
namespace: crowdsec
|
||||
spec:
|
||||
plugin:
|
||||
crowdsec-bouncer:
|
||||
enabled: true
|
||||
LogLevel: INFO
|
||||
# `live` blocked on a `GET /v1/decisions` per request, so a burst
|
||||
# saturated the LAPI and the plugin 403'd IPs that were never banned.
|
||||
# v1.3.3 ignores UpdateMaxFailure in `live`, so fail-open is only
|
||||
# reachable in stream mode, which polls into a cache instead - no
|
||||
# per-request call to saturate. 15s rather than the 60s default: the
|
||||
# deploy runner shares one public IP with the house, so this bounds
|
||||
# both how late a ban lands and how long a lifted one lingers.
|
||||
CrowdsecMode: stream
|
||||
UpdateIntervalSeconds: 15
|
||||
# -1 = never block because the LAPI is unreachable. In v1.3.3
|
||||
# handleStreamTicker only clears isCrowdsecStreamHealthy when
|
||||
# updateMaxFailure != -1, and ServeHTTP 403s once it is false, so this
|
||||
# makes a CrowdSec outage mean "no protection", not "every site 403".
|
||||
UpdateMaxFailure: -1
|
||||
CrowdsecLapiScheme: http
|
||||
CrowdsecLapiHost: crowdsec-service.crowdsec.svc.cluster.local:8080
|
||||
CrowdsecLapiKeyFile: "/etc/traefik/secrets/traefik-api-key"
|
||||
# Bypasses the bouncer and the decision cache, no LAPI round-trip.
|
||||
# Keep in sync with forust/local-network in crowdsec-values.yaml.
|
||||
ClientTrustedIPs:
|
||||
- "127.0.0.0/8"
|
||||
- "10.0.0.0/8"
|
||||
- "172.16.0.0/12"
|
||||
- "192.168.0.0/16"
|
||||
- "100.64.0.0/10"
|
||||
- "169.254.0.0/16"
|
||||
- "fc00::/7"
|
||||
- "fe80::/10"
|
||||
# The mobile operator range from forust/mobile-whitelist, repeated
|
||||
# deliberately rather than relying on the parser whitelist alone.
|
||||
# That whitelist drops the event before it reaches a bucket, so no
|
||||
# decision is ever created - but it is one config away from not
|
||||
# firing, and the bouncer would then enforce a ban that was never
|
||||
# justified. This is the last line: even a decision that exists for
|
||||
# any reason is not served against the phone.
|
||||
- "84.245.64.0/18"
|
||||
# The name is HTTPTimeoutSeconds, an int in seconds (min 1) - there is
|
||||
# no CrowdsecLapiTimeout, and an unrecognised key is silently dropped,
|
||||
# which is how this sat at the 10s default. Nothing rides on it per
|
||||
# request any more, so this only bounds the stream pull - and too low
|
||||
# is the dangerous direction: the LAPI needs ~2s to answer
|
||||
# /v1/decisions/stream, and a pull that times out leaves the ban cache
|
||||
# frozen at its startup contents ("failed sending new decisions"),
|
||||
# i.e. new bans silently never apply. Keep it above the pull latency.
|
||||
HTTPTimeoutSeconds: 10
|
||||
@@ -16,6 +16,12 @@ agent:
|
||||
value: crowdsecurity/traefik crowdsecurity/base-http-scenarios
|
||||
- name: DISABLE_COLLECTIONS
|
||||
value: crowdsecurity/sshd
|
||||
# Bans on 401/403 bursts hurt more than they protect: with L3 enforcement
|
||||
# a false positive cuts the IP off everything (SSH included), and past
|
||||
# incidents show legit automation (deploy runner, mesh peers, registry
|
||||
# pulls) tripping this probe. Probing/XSS/SQLi/CVE scenarios stay.
|
||||
- name: DISABLE_SCENARIOS
|
||||
value: crowdsecurity/http-generic-bf
|
||||
metrics:
|
||||
enabled: true
|
||||
serviceMonitor:
|
||||
|
||||
Reference in new issue
Block a user