Files
homelab/reloader/k8s/reloader-values.yaml
T
forust 16aaeb60c1 fix(k8s): set requests and limits on the pods that shipped with neither
Eighteen containers had no memory limit at all, so nothing on the node could
bound them. Three of the values files even claimed to set resources: Helm does
not complain about a key it does not recognise, so the block sat there looking
like a limit while the pod ran unbounded.

alloy is the one that mattered. The chart reads `alloy.resources`; the file had
`controller.resources`, so the DaemonSet that tails every pod log on the node
shipped with nothing at all. `kubeStateMetrics` is the same trap in a different
shape -- that is the condition key, the values live under `kube-state-metrics` --
and `configReloader` in the alloy chart sits at the top level rather than under
`alloy`. Each one is verified by rendering the chart and reading the resources
back off the containers, because a values key that is ignored looks exactly
like one that works.

reloader turned out to be set and still wrong: 64Mi request against a measured
p95 of 73M, so the pod ran permanently above its own request and stayed a
standing eviction candidate. That is the pod that restarts every other pod, so
it is the last one that should be evicted. Raised to 96Mi.

Requests are set at p95 throughout, grafana, playwright and alloy included.
Left at the values first proposed they would have sat below their own p95 and
queued for eviction ahead of everything smaller. CPU limits are deliberately
absent: the node is I/O bound at 5% CPU, and CFS throttling would turn disk
wait into runnable-throttled, which is the failure mode that took the node down.

The prometheus and alertmanager configReloader sidecars are left open: chart
86.2.3 does not template the key, so reaching those two containers needs a
postRenderer.

Verified: all four charts render with the resources landing on the intended
containers, and 16/16 local gates pass.
2026-09-28 10:14:52 +02:00

24 lines
918 B
YAML

# Pinned chart: stakater/reloader 2.2.17 (app v1.4.22).
# Deployed by the deploy workflow, namespace reloader.
# Restarts pods when a ConfigMap or Secret they consume changes. Opt-in per workload
# via the reloader.stakater.com/auto: "true" pod annotation; watchGlobally because
# the workloads that need it are spread across a few dozen namespaces.
reloader:
watchGlobally: true
deployment:
replicas: 1
# The chart defaults to no requests or limits, so the pod is evictable under node
# pressure and the restarts go with it.
# Memory was raised from 64Mi: measured p95 over 7 days is 73M, so the pod was
# running above its own request and sitting in the eviction candidates. This pod
# is the one that restarts every other pod, so it must not be evicted.
resources:
requests:
cpu: "10m"
memory: "96Mi"
limits:
cpu: "100m"
memory: "192Mi"