Finishes the sizing pass over every workload the deploy actually manages. Each request is at or above the container's p95 over the last seven days, so nothing is sized below what it is known to use, and each limit is between 1.6x and 5x the observed max, which is the figure that decides whether a burst gets an OOMKill. Some of these go up, and that is the point. adguard was holding 975M against a 500Mi request and netbox 962M against 512Mi, so both sat permanently above their own request and were standing eviction candidates on a node that has about 300M of headroom. Raising a request costs scheduler room; leaving it low costs the pod its place in the queue when the node gets tight. Others come down. loki ran with a 2Gi limit on 224M, gitea 1.5Gi on 305M, the authentik worker 1Gi on 305M, and a tail of single-purpose pods -- redis, glance, the two homepages, cfddns, session-keeper, the netbird dashboard, the loki gateway -- each reserved 4x to 16x more than they have ever touched. prometheus gets the opposite treatment: 768Mi/2560Mi, above its p95, because it compacts its TSDB in place and that is a burst worth budgeting for rather than throttling. Two of these limits are close enough to the observed max to be worth watching rather than trusting: adguard at 1.5x, and its DNS cache grows monotonically, so the ceiling is a date, not a margin. That was true before this change too; the pod sizing does not fix it and the cache needs bounding. CPU limits are untouched throughout. Leaving postgres alone as well: it sits in an uncommitted file that belongs to other work in progress. Verified: every request is at or above p95 and every limit above the observed max across all 74 containers, and 16/16 local gates pass.
NetBox
NetBox for homelab documentation and visualization. Two runtimes are available:
| Runtime | Manifest | Purpose |
|---|---|---|
| Docker | compose.yaml |
Local stand on 127.0.0.1:8000 (no public exposure) |
| k8s | k8s/ |
Homelab service on netbox.forust.xyz (and the internal name) |
Both use the same image (netboxcommunity/netbox:v4.7-5.1.1) and Valkey for tasks
plus a second logical database for caching. The Docker stand keeps its own
PostgreSQL container, while the k8s deployment uses the shared database cluster
(postgres.database.svc.cluster.local:5432, role/database netbox); only Valkey
stays a per-service StatefulSet.
Docker Compose
cp .env.example .env
# replace CHANGE_ME
docker compose up -d
The UI is available at http://localhost:8000. The port is bound to 127.0.0.1
intentionally, so this stand is not exposed on the LAN or public interfaces.
The netbox service is also attached to the external proxy network and carries
Traefik labels for netbox.forust.xyz and netbox.workstation.internal. Those
labels only take effect while the Docker Traefik stack is running; it is currently
stopped, and the live ingress path in this homelab is the k8s Traefik.
Inspect startup and health with:
docker compose ps
docker compose logs -f netbox
Stop it with docker compose down; data is kept in the named volumes
netbox-postgres, netbox-media-files, netbox-reports-files,
netbox-scripts-files and netbox-redis-data.
Kubernetes
k8s/ is deployed in the homelab cluster and serves netbox.forust.xyz publicly
plus netbox.workstation.internal / netbox.gigaforust.internal internally. To
rebuild it from scratch:
# 1. shared PostgreSQL: the password lives in the shared secret, NetBox keeps a copy
kubectl -n database patch secret postgres-shared-secrets \
--type merge -p '{"stringData":{"NETBOX_DB_PASSWORD":"<same value>"}}'
kubectl -n database exec postgres17-0 -- psql -U postgres -d postgres \
-c 'CREATE ROLE netbox LOGIN PASSWORD ...' -c 'CREATE DATABASE netbox OWNER netbox'
# 2. secrets first: the deploy workflow never applies *secret*.yaml
cp k8s/secrets.yaml.example k8s/secrets.yaml # replace CHANGE_ME
kubectl apply -f k8s/secrets.yaml
# 3. manifests
kubectl apply -f k8s/
The shared cluster is reached at postgres.database.svc.cluster.local:5432. Its
NetworkPolicy (postgres/k8s/network-policy.yaml) must list the netbox namespace
or connections are dropped, and postgres/initdb/01-create-databases.sh already
creates the role and database on a fresh data directory. NetBox has no PostgreSQL
StatefulSet of its own — only netbox-valkey.
netbox.forust.xyz resolves to this host (78.98.72.122) through the DOMAINS
list in the default/cfddns secret. cert-manager issues netbox-prod-tls with the
letsencrypt-prod issuer, the internal route uses internal-wildcard-tls.
Resources are permanent again now that the first-boot migrations are complete:
the web container reserves 100m/512Mi and is capped at 2 CPU/2Gi, the
worker reserves 50m/256Mi and is capped at 1 CPU/1Gi, and Valkey reserves
25m/64Mi and is capped at 250m/256Mi. The deliberately generous CPU caps
leave enough headroom for future schema migrations without letting one process
consume the whole node.
The first start applies ~810 migrations, each in its own transaction with DDL and
a commit; every later start is a no-op. The startup probe allows 15 minutes and
progressDeadlineSeconds is 1800 for the same reason. Probes run inside the pod
and explicitly set Host: netbox.forust.xyz; a kubelet httpGet.host field would
replace the probe destination with that public hostname and bypass the pod.
Secrets
netbox/.env(compose) andnetbox/k8s/secrets.yaml(k8s) are gitignored. Only.env.exampleandk8s/secrets.yaml.exampleare committed.netbox/configuration/configuration.pyis env-driven: hosts, database, Redis and the Django keys all come from the environment, so the same settings file works in both runtimes. The k8s copy lives in thenetbox-settingsConfigMap (k8s/settings.yaml) and must be kept in sync with the file.- Rotating
SECRET_KEYinvalidates all sessions; rotatingAPI_TOKEN_PEPPER_1invalidates every API token.