From 3c4732ae20a47ab49ad8de7aa69aeaab33c97b56 Mon Sep 17 00:00:00 2001 From: mr-forust Date: Tue, 6 Oct 2026 16:07:50 +0200 Subject: [PATCH] docs: document homelab services, deployment, and repository review --- .gitea/README.md | 88 ++++++++++++++++++++ README.md | 163 +++++++++++++++++++++++++++++++++++++ adguardhome/README.md | 22 +++++ authentik/README.md | 22 +++++ cert-manager/README.md | 18 ++++ cfddns/README.md | 22 +++++ checkmk/README.md | 21 +++++ cloudflared/README.md | 21 +++++ converters/README.md | 22 +++++ crowdsec/README.md | 25 ++++++ dockmon/README.md | 21 +++++ docs/repository-review.md | 136 +++++++++++++++++++++++++++++++ downtify/README.md | 20 +++++ edu_master/README.md | 52 ++++++++++++ errorpages/README.md | 21 +++++ gitea/README.md | 25 ++++++ glance/README.md | 26 ++++++ headscale/README.md | 18 ++++ homarr/README.md | 24 ++++++ homepages/README.md | 25 ++++++ immich/README.md | 29 +++++++ kener/README.md | 21 +++++ loki/README.md | 16 ++++ metube/README.md | 21 +++++ n8n/README.md | 21 +++++ netbird/README.md | 145 ++++++++++----------------------- netbox/README.md | 123 ++++++++-------------------- netronome/README.md | 22 +++++ nextcloud/README.md | 17 ++++ penpot/README.md | 12 +++ portainer/README.md | 21 +++++ postgres/README.md | 84 ++++++++++++------- prometheus-stack/README.md | 26 ++++++ rackpeek/README.md | 20 +++++ reloader/README.md | 13 +++ renovate/README.md | 110 +++++++------------------ searxng/README.md | 22 +++++ streaming/README.md | 16 ++++ termix/README.md | 20 +++++ traefik/README.md | 33 ++++++++ uptime-kuma/README.md | 22 +++++ vaultwarden/README.md | 23 ++++++ vpn/README.md | 5 ++ vpn/xui/README.md | 18 ++++ 44 files changed, 1354 insertions(+), 298 deletions(-) create mode 100644 .gitea/README.md create mode 100644 README.md create mode 100644 adguardhome/README.md create mode 100644 authentik/README.md create mode 100644 cert-manager/README.md create mode 100644 cfddns/README.md create mode 100644 checkmk/README.md create mode 100644 cloudflared/README.md create mode 100644 converters/README.md create mode 100644 crowdsec/README.md create mode 100644 dockmon/README.md create mode 100644 docs/repository-review.md create mode 100644 downtify/README.md create mode 100644 edu_master/README.md create mode 100644 errorpages/README.md create mode 100644 gitea/README.md create mode 100644 glance/README.md create mode 100644 headscale/README.md create mode 100644 homarr/README.md create mode 100644 homepages/README.md create mode 100644 immich/README.md create mode 100644 kener/README.md create mode 100644 loki/README.md create mode 100644 metube/README.md create mode 100644 n8n/README.md create mode 100644 netronome/README.md create mode 100644 nextcloud/README.md create mode 100644 penpot/README.md create mode 100644 portainer/README.md create mode 100644 prometheus-stack/README.md create mode 100644 rackpeek/README.md create mode 100644 reloader/README.md create mode 100644 searxng/README.md create mode 100644 streaming/README.md create mode 100644 termix/README.md create mode 100644 traefik/README.md create mode 100644 uptime-kuma/README.md create mode 100644 vaultwarden/README.md create mode 100644 vpn/README.md create mode 100644 vpn/xui/README.md diff --git a/.gitea/README.md b/.gitea/README.md new file mode 100644 index 0000000..856cd9f --- /dev/null +++ b/.gitea/README.md @@ -0,0 +1,88 @@ +# Build and deployment workflows + +Gitea Actions checks this repository, builds its custom images, and deploys +selected services to the workstation. Workflows use the self-hosted runner labels +`linux`, `arch`, and `homelab`; deployment jobs also require `prod`. + +## Checks + +`ci.yaml` runs Compose validation, actionlint, ShellCheck, Prettier, Ruff, +yamllint, hadolint, and kubeconform. Tool versions are pinned in +`workflows/tool-versions.env` and installed by `install-ci-tools.sh`. + +Compose CI checks structure without resolving local environment files or paths. +On the reviewed main commit it only discovers standard filenames; the +`fix/deploy-validation` branch adds the manual Compose entry points too. + +Kubeconform validates known resource schemas. Unknown CRDs are skipped. On main, +CI also attempts server-side dry-runs for marked services; these require an +existing namespace and contact the cluster's admission webhooks. A cluster that +is unreachable produces a warning and skips that CI pass. Deploy validation has +its own dry-run stage. + +`renovate-ci.yaml` validates Renovate settings and checks that its generated +ConfigMap matches `renovate/renovate.json`. + +## Image builds + +CI builds changed custom images for `errorpages`, both `homepages` variants, and +the two `edu_master` Python services. Main builds publish `main`, `prod`, and a +commit tag. Dev builds publish `dev`. Build jobs wait for the lint and manifest +checks. + +Kubernetes deployment resolves the lab's own registry images to digests, preferring +commit-specific tags. Third-party image versions remain declared in the manifests. + +## Deploy selection + +`workflows/deploy-lib.sh` owns the stage logic; `ssh-run.sh` invokes it on the +workstation through SSH. Kubernetes selection uses `k8s/active`; Compose selection +uses an `active` file beside a standard `compose.yaml` or `compose.yml`. +Kustomize overlays are supported, although the current tree primarily contains +plain manifests. + +Secret files, examples, Helm values, and patch files are excluded from plain +manifest selection. Create local Kubernetes Secrets separately in their target +namespaces. The Helm table lists Prometheus, Loki, Alloy, and Reloader, with each +release controlled by its configured marker. Other charts need separate setup. + +## Trigger and required settings + +Automatic deployment follows a successful main CI run when the repository Actions +variable `AUTODEPLOY` is `true`. The manual deploy workflow bypasses that switch +and targets the fetched main branch when no validated commit SHA is provided. +A manual dispatch does not prove that this commit passed CI. + +Configure the Actions secrets `DEPLOY_HOST`, `DEPLOY_USER`, `DEPLOY_SSH_KEY`, and, +where needed, `DEPLOY_PORT` and `DEPLOY_PATH`. Registry publishing uses +`REGISTRY_USERNAME` and `REGISTRY_PASSWORD`. The remote user needs access to Git, +Docker, kubectl, Helm, jq, and the state directory used for snapshots. + +Keep `APPLY_PRUNE` false on the reviewed implementation: its per-file prune loop +is unsafe. `fix/deploy-prune-guard` rejects that option before changes are applied. + +Preflight fetches and resets the remote checkout. It refuses when tracked files +have local changes; ignored local env and Secret files stay in place. Do not use +a development checkout with uncommitted tracked changes as the deployment target. + +## Stages and recovery + +1. Preflight fetches the target commit and checks the remote working tree. +2. Validate selects services, parses Compose, performs Kubernetes dry-runs, and + checks referenced Secrets. +3. Apply Kubernetes records a workload snapshot, upgrades selected Helm releases, + applies resources, and refreshes owned custom images. +4. Apply Compose recreates marked stacks and checks container state. +5. Verify Kubernetes checks changed workloads and attempts rollback for failures. +6. Smoke probes public routes after verification. + +The two apply jobs share a remote lock. Workflow concurrency queues deployments +rather than interrupting an older apply. Snapshots live under +`$XDG_STATE_HOME/homelab-deploy`, or `~/.local/state/homelab-deploy` by default. +They contain the pre-apply workload data and commit identifier. + +Rollback uses workload revisions. It does not restore ConfigMaps, Secrets, +database schemas, or data. Helm-owned workloads are handled through the Helm +upgrade's rollback path; the generic rollback skips them. Compose has no automatic +rollback. See the [review](../docs/repository-review.md) for remaining recovery +limitations, including SSH retries and serial rollback timing. diff --git a/README.md b/README.md new file mode 100644 index 0000000..6ff1bef --- /dev/null +++ b/README.md @@ -0,0 +1,163 @@ +# Homelab + +Configuration for my homelab: Kubernetes manifests, Docker Compose stacks, and the +Gitea Actions that build and deploy them. Most applications have both deployment +formats. Headscale, Nextcloud AIO, and the media stack run on Docker; Kubernetes +provides their ingress through Services and EndpointSlices. + +These files contain this lab's domains, IP addresses, storage paths, and private +registry names. Running them on another machine takes some editing. + +## Start here + +- [Service list](#services) — what each directory contains. +- [Deployment workflow](.gitea/README.md) — selection, validation, and recovery. +- [Repository review](docs/repository-review.md) — confirmed problems and fix branches. +- [Shared PostgreSQL](postgres/README.md), [Traefik](traefik/README.md), and + [cert-manager](cert-manager/README.md) — common dependencies. + +## What gets deployed + +The `active` files are switches for the deploy workflow, not health indicators. + +| File | Effect | +| ---------------------- | ----------------------------------------------------------- | +| `/active` | Include that directory's `compose.yaml` or `compose.yml`. | +| `/k8s/active` | Include its Kubernetes manifests or Kustomize overlay. | +| Both | Run the Compose stack and apply the Kubernetes resources. | +| Neither | Keep the configuration in Git without automatic deployment. | + +`shared-compose.yaml`, `client.compose.yaml`, and `renovate-compose.yaml` are +manual entry points. The deploy script does not discover them. + +Kubernetes selection excludes secret files, examples, Helm values, and patches. +Helm releases listed in `deploy-lib.sh` are upgraded separately. Traefik, +cert-manager, and CrowdSec have additional bootstrap steps; an `active` marker +does not install their charts. + +The table below describes committed configuration. It does not claim that a +service is currently healthy or running. + +## Services + +| Service | Configuration | Selected by markers | +| ----------------------------------------------------------- | ---------------------------- | ------------------- | +| [AdGuard Home](adguardhome/README.md) | Kubernetes + Compose | Kubernetes | +| [Authentik](authentik/README.md) | Kubernetes + Compose | Kubernetes | +| [cert-manager](cert-manager/README.md) | Kubernetes / Helm | Manual | +| [Cloudflare DDNS](cfddns/README.md) | Kubernetes + Compose | Kubernetes | +| [Checkmk](checkmk/README.md) | Kubernetes + Compose | Manual | +| [Cloudflare Tunnel](cloudflared/README.md) | Kubernetes / Helm | Manual | +| [File converters](converters/README.md) | Kubernetes + Compose | Kubernetes | +| [CrowdSec](crowdsec/README.md) | Kubernetes / Helm | Manual | +| [Dockmon](dockmon/README.md) | Kubernetes + Compose | Manual | +| [Downtify](downtify/README.md) | Kubernetes + Compose | Manual | +| [EDU session keeper and Telegram bot](edu_master/README.md) | Kubernetes + Compose | Kubernetes | +| [Error pages](errorpages/README.md) | Kubernetes + Compose | Kubernetes | +| [Gitea](gitea/README.md) | Kubernetes + Compose | Kubernetes | +| [Glance](glance/README.md) | Kubernetes + Compose | Manual | +| [Headscale](headscale/README.md) | Compose + Kubernetes routing | Compose, Kubernetes | +| [Homarr](homarr/README.md) | Kubernetes + Compose | Manual | +| [Homepages](homepages/README.md) | Kubernetes + Compose | Kubernetes | +| [Immich](immich/README.md) | Kubernetes + Compose | Kubernetes | +| [Kener](kener/README.md) | Kubernetes + Compose | Manual | +| [Loki and Alloy](loki/README.md) | Kubernetes / Helm | Kubernetes | +| [MeTube](metube/README.md) | Kubernetes + Compose | Kubernetes | +| [n8n](n8n/README.md) | Kubernetes + Compose | Manual | +| [NetBird](netbird/README.md) | Kubernetes + Compose | Kubernetes | +| [NetBox](netbox/README.md) | Kubernetes + Compose | Kubernetes | +| [Netronome](netronome/README.md) | Kubernetes + Compose | Kubernetes | +| [Nextcloud AIO](nextcloud/README.md) | Compose + Kubernetes routing | Compose, Kubernetes | +| [Penpot](penpot/README.md) | Compose | Manual | +| [Portainer](portainer/README.md) | Kubernetes + Compose | Manual | +| [Shared PostgreSQL](postgres/README.md) | Kubernetes + Compose | Kubernetes | +| [Monitoring stack](prometheus-stack/README.md) | Kubernetes + Compose | Kubernetes | +| [RackPeek](rackpeek/README.md) | Kubernetes + Compose | Kubernetes | +| [Reloader](reloader/README.md) | Kubernetes / Helm | Manual | +| [Renovate](renovate/README.md) | Kubernetes + Compose | Kubernetes | +| [SearXNG](searxng/README.md) | Kubernetes + Compose | Manual | +| [Media stack](streaming/README.md) | Compose + Kubernetes routing | Compose, Kubernetes | +| [Termix](termix/README.md) | Kubernetes + Compose | Manual | +| [Traefik](traefik/README.md) | Kubernetes + Compose | Kubernetes | +| [Uptime Kuma](uptime-kuma/README.md) | Kubernetes + Compose | Kubernetes | +| [Vaultwarden](vaultwarden/README.md) | Kubernetes + Compose | Kubernetes | +| [3x-ui](vpn/xui/README.md) | Kubernetes | Kubernetes | + +## Running a Compose stack + +Use the service README first. Where a service has an env example, copy it inside +that service's directory and replace the placeholders. The root `.env.example` +is an older collection of variables, not a complete configuration for every stack. + +For example, from the repository root: + +```sh +cd netbox +cp .env.example .env +$EDITOR .env +docker compose config --quiet +docker compose up -d +docker compose ps +``` + +Stacks that attach to `proxy` require an existing Docker network of that name and +an appropriate reverse proxy. Published host ports still work independently of +Traefik. Check port conflicts before starting an alternative to a Kubernetes +service: DNS, STUN, and HTTP listeners can share the same host. + +`docker compose down` keeps named volumes. Adding `-v` removes them. + +## Preparing Kubernetes + +The manifests assume Traefik CRDs, cert-manager, and a working storage provisioner. +PrometheusRule and ServiceMonitor resources also need the Prometheus Operator. +Replace the lab's hosts and addresses before using the configuration elsewhere. + +Create a service's namespace, then prepare its ignored Secret from the example. +For example: + +```sh +kubectl apply -f netbox/k8s/namespace.yaml +cp netbox/k8s/secrets.yaml.example netbox/k8s/secrets.yaml +$EDITOR netbox/k8s/secrets.yaml +kubectl apply -f netbox/k8s/secrets.yaml +``` + +The deploy workflow applies the tracked resources for marked services. Avoid +applying an entire `k8s/` directory blindly: some directories contain Helm values, +examples, and alternative routes. For a manual change, apply the selected manifest +explicitly and check the resulting rollout. + +Shared database passwords must agree between the `database` namespace and each +application's Secret. Updating the PostgreSQL Secret does not change an existing +role's password; see the database README. + +## Local checks + +CI pins its tools in `.gitea/workflows/tool-versions.env`. Use the same versions: + +```sh +tools_dir="$(bash .gitea/workflows/install-ci-tools.sh)" +export PATH="$tools_dir:$PATH" +ruff check . +ruff format --check . +actionlint -config-file .gitea/actionlint.yaml .gitea/workflows/*.yaml +.gitea/workflows/sync-renovate-configmap.sh --check +``` + +The [workflow README](.gitea/README.md#checks) lists the rest of the checks. +Structure checks do not establish that local Secrets, mounted files, storage, +or external services are ready. + +## Data and recovery + +State lives outside Git: PVCs, Docker volumes, bind mounts, databases, and ignored +configuration. Keep backups of application data and the keys needed to read it. +An image rollback does not roll back database migrations or ConfigMap contents. + +Many PVCs use the cluster's default StorageClass; monitoring explicitly uses +`local-path`. Check the PV reclaim policy before deleting a PVC or namespace. +The manifests do not provide a repository-wide backup schedule. + +`incident-archive/` contains past incident notes. `.docs/storage-audit-instruction.md` +is a planning document, not evidence that NFS has been installed. diff --git a/adguardhome/README.md b/adguardhome/README.md new file mode 100644 index 0000000..7d4b091 --- /dev/null +++ b/adguardhome/README.md @@ -0,0 +1,22 @@ +# AdGuard Home + +DNS filtering with a web UI, DNS-over-TLS, and certificates from cert-manager. + +The Kubernetes namespace is `adguard`. The workload uses `adguard-pvc` for +configuration and working data, and mounts the `adguard-certs` TLS Secret. +The LoadBalancer Service exposes DNS separately from the web ingress. + +The Compose stack publishes TCP/UDP 53 and TCP 853 on the host. Prepare `conf/` +and `certs/` before starting it. Starting both DNS deployments on the same address +can cause a port conflict. + +## Inspect + +From the repository root: + +```sh +kubectl get pods,svc,pvc -n adguard +kubectl get events -n adguard --sort-by=.metadata.creationTimestamp +``` + +See the [repository README](../README.md) for deployment selection. diff --git a/authentik/README.md b/authentik/README.md new file mode 100644 index 0000000..083be8b --- /dev/null +++ b/authentik/README.md @@ -0,0 +1,22 @@ +# Authentik + +Identity provider with separate server and worker deployments. + +Kubernetes connects to the shared PostgreSQL service in `database`. Set +`AUTHENTIK_DB_PASSWORD` to the same value in both database and application Secrets. +Keep `AUTHENTIK_SECRET_KEY` with the backups. + +Compose uses its own PostgreSQL 15 container and bind-mounted media and templates. +Its image defaults differ from Kubernetes; check both before an upgrade. +The worker mounts the Docker socket for Docker outpost management. + +## Inspect + +From the repository root: + +```sh +kubectl get pods,svc,pvc -n authentik +kubectl get events -n authentik --sort-by=.metadata.creationTimestamp +``` + +See the [repository README](../README.md) for deployment selection. diff --git a/cert-manager/README.md b/cert-manager/README.md new file mode 100644 index 0000000..9e02591 --- /dev/null +++ b/cert-manager/README.md @@ -0,0 +1,18 @@ +# cert-manager + +Public ACME issuers and an internal certificate authority. + +This directory contains chart values and issuer resources, not the controller +installation. Install the cert-manager chart with CRDs and the settings in +`k8s/cert-manager-values.yaml` before applying the issuers. + +`clusterissuer.yaml` defines staging and production Let's Encrypt issuers. +They use HTTP-01 through the Traefik ingress class. Public DNS and inbound HTTP +reachability must work for the requested names before issuance. +`internal-ca.yaml` bootstraps the internal CA. Keep its private-key Secret backed +up; the tracked `.crt` is only a public certificate. + +This directory has no `k8s/active` marker. Apply the issuer files deliberately; +`kubectl apply` does not interpret the Helm values file. + +See the [repository README](../README.md) for deployment selection. diff --git a/cfddns/README.md b/cfddns/README.md new file mode 100644 index 0000000..af4f249 --- /dev/null +++ b/cfddns/README.md @@ -0,0 +1,22 @@ +# Cloudflare DDNS + +Updates the lab DNS records when the public address changes. + +Kubernetes runs in `default` with host networking and reads `cfddns-secrets`. +The Compose stack also uses host networking. Configure the API token and domain +list from the relevant example; keep DNS names consistent with the ingress rules. + +`config.json.example` is a separate configuration example. The current Compose +file does not mount a config.json file. Check configuration against the pinned +DDNS image when changing between environment and file-based settings. + +## Inspect + +From the repository root: + +```sh +kubectl get pods,svc,pvc -n default +kubectl get events -n default --sort-by=.metadata.creationTimestamp +``` + +See the [repository README](../README.md) for deployment selection. diff --git a/checkmk/README.md b/checkmk/README.md new file mode 100644 index 0000000..88efda3 --- /dev/null +++ b/checkmk/README.md @@ -0,0 +1,21 @@ +# Checkmk + +Checkmk Raw monitoring site with web and agent-receiver ingress. + +The site data lives in `checkmk-sites-pvc` on Kubernetes and the `sites` named +volume on Compose. The agent receiver has a separate TCP route; enabling the +web route alone does not expose it. + +Prepare the password in the service env or Secret example. Inspect the Checkmk +container logs during the first site creation. + +## Inspect + +From the repository root: + +```sh +kubectl get pods,svc,pvc -n checkmk +kubectl get events -n checkmk --sort-by=.metadata.creationTimestamp +``` + +See the [repository README](../README.md) for deployment selection. diff --git a/cloudflared/README.md b/cloudflared/README.md new file mode 100644 index 0000000..7027594 --- /dev/null +++ b/cloudflared/README.md @@ -0,0 +1,21 @@ +# Cloudflare Tunnel + +A Kubernetes connector for an existing Cloudflare tunnel. + +The Deployment runs in `default` and reads its token from the ignored Secret +created from `k8s/secret.yaml.example`. Create the tunnel and its hostname rules +in Cloudflare before starting the connector. + +There is no Compose file or `k8s/active` marker. Apply the Secret first, then +`k8s/deployment.yaml` when this tunnel is needed. + +## Inspect + +From the repository root: + +```sh +kubectl get pods,svc,pvc -n default +kubectl get events -n default --sort-by=.metadata.creationTimestamp +``` + +See the [repository README](../README.md) for deployment selection. diff --git a/converters/README.md b/converters/README.md new file mode 100644 index 0000000..0284ab4 --- /dev/null +++ b/converters/README.md @@ -0,0 +1,22 @@ +# File converters + +ConvertX for server-side conversion and BentoPDF for PDF tools. + +ConvertX persists files in `convertx-pvc`; BentoPDF has no persistent volume. +Kubernetes configuration includes a local `config.yaml.example`, excluded from +normal deployment. Copy and apply the real ConfigMap separately where required. + +Compose publishes ConvertX on host port 9992 as well as attaching it to the +proxy network. Replace the authentication settings from `.env.example` before +exposing it outside the lab. + +## Inspect + +From the repository root: + +```sh +kubectl get pods,svc,pvc -n converters +kubectl get events -n converters --sort-by=.metadata.creationTimestamp +``` + +See the [repository README](../README.md) for deployment selection. diff --git a/crowdsec/README.md b/crowdsec/README.md new file mode 100644 index 0000000..4ce98b5 --- /dev/null +++ b/crowdsec/README.md @@ -0,0 +1,25 @@ +# CrowdSec + +Helm values, dashboards, network policy, and a maintenance CronJob. + +Install CrowdSec separately using `k8s/crowdsec-values.yaml`; the deploy +workflow does not have a CrowdSec Helm release entry. There is no `k8s/active` +marker in this directory. + +The LAPI policy and janitor run in `crowdsec`. The dashboard ConfigMaps are in +`prometheus` for Grafana's sidecar. The janitor has its own ServiceAccount and +namespace Role. Review its script and schedule before enabling cleanup. + +Traefik's values state that enforcement moved to a host firewall bouncer. This +repository does not install that host component. + +## Inspect + +From the repository root: + +```sh +kubectl get pods,svc,pvc -n crowdsec +kubectl get events -n crowdsec --sort-by=.metadata.creationTimestamp +``` + +See the [repository README](../README.md) for deployment selection. diff --git a/dockmon/README.md b/dockmon/README.md new file mode 100644 index 0000000..766e46e --- /dev/null +++ b/dockmon/README.md @@ -0,0 +1,21 @@ +# Dockmon + +Docker management UI that talks to the host Docker daemon. + +Both runtimes mount `/var/run/docker.sock`. On Kubernetes the socket belongs +to the node hosting the pod, so this is not a cluster-wide container manager. + +Compose stores application data in a named volume. Kubernetes uses a StatefulSet +with a volume claim template. Its ServersTransport is specific to the upstream +connection; keep it with the ingress resources. + +## Inspect + +From the repository root: + +```sh +kubectl get pods,svc,pvc -n dockmon +kubectl get events -n dockmon --sort-by=.metadata.creationTimestamp +``` + +See the [repository README](../README.md) for deployment selection. diff --git a/docs/repository-review.md b/docs/repository-review.md new file mode 100644 index 0000000..28e0ba9 --- /dev/null +++ b/docs/repository-review.md @@ -0,0 +1,136 @@ +# Repository review + +Reviewed the tracked tree at `cc9c3de` and read the live workstation state on +6 October 2026. Changes are split into documentation and individual fix branches, +all based on that main commit. The original local checkout and its uncommitted +monitoring changes were preserved. No deployment was performed. + +## Confirmed problems with prepared fixes + +| Priority | Problem and consequence | Fix branch | +| -------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------- | +| High | `APPLY_PRUNE=true` is passed to each individual manifest apply. Each invocation sees only that file's desired objects and can delete other resources selected by the shared label. | `fix/deploy-prune-guard` | +| High | Deploy validates Compose with interpolation and env/path resolution disabled. Required settings can pass validation and then fail during apply after other workloads have changed. | `fix/deploy-validation` | +| Medium | Secret validation is text-based and compares names across all namespaces. A Secret elsewhere can hide a missing local Secret; mounted Secrets are also missed. | `fix/deploy-validation` | +| Medium | Compose CI misses `postgres/shared-compose.yaml`, `netbird/client.compose.yaml`, and `renovate/renovate-compose.yaml`. | `fix/deploy-validation` | +| Medium | NetBird Compose mounts `entrypoint.sh`, but it is absent. Its README also calls a missing `setup.sh`; a fresh checkout cannot start this stack as documented. | `fix/netbird-compose-runtime` | +| Medium | Glance's CSS mount uses `glance-config`, whose keys do not include `user.css`. That key is in `glance-assets`; the pod's subPath mount cannot be prepared correctly. | `fix/glance-assets` | +| Medium | The shared PostgreSQL initializer requires `NETBOX_DB_PASSWORD`, but the Compose env example omits it. Following the example leaves first initialization incomplete. | `fix/postgres-env-example` | +| Medium | EDU's Compose env example uses old credential names and full URL variables, while the code reads `KEEPER_*` and paths under `EDU_URL_BASE`. | `fix/session-keeper-reliability` | +| Medium | Session keeper HTTP calls have no timeouts. Its Redis cookie never expires, probes only check existence, and its logs include cookies. A hung or failed refresh can leave a stale session appearing ready. | `fix/session-keeper-reliability` | +| Medium | AdGuard's DoH and SearXNG's Compose rules put Boolean expressions inside `Host(...)`. They are invalid router expressions despite valid YAML. | `fix/compose-router-rules` | + +Traefik matchers should be combined as `Host(a) || Host(b)`; the rule syntax is +described in the [Traefik rules documentation](https://doc.traefik.io/traefik/reference/routing-configuration/http/routing/rules-and-priority/). +The fix retains the DoH path constraint for both hostnames. + +The prune fix deliberately rejects the unsafe option. It does not introduce +automatic deletion under a different implementation. Prune defaults to false, +and no tracked resource currently carries the selector label, so this is a +latent defect rather than evidence of a live deletion incident. + +The session fix bounds HTTP and Redis calls, validates required credentials, +sets a cookie lifetime of two refresh intervals, and marks success only after +publishing the verified cookie. With the default ten-minute interval, an outage +longer than twenty minutes will make the existing Redis-key readiness checks fail. +That is an intentional change from indefinite apparent readiness. + +The deployment fix extracts required pod Secret references from rendered JSON, +checks their namespaces, includes init containers, image-pull credentials, and +mounted/projected Secrets, and honors optional references. Ingress TLS Secrets +issued by cert-manager are not treated as pre-existing pod prerequisites. +It checks existence/access, not every key's contents or application validity. + +## Live workstation observations + +The SSH alias `workstation` is reachable. It has one Ready control-plane node, +Kubernetes `v1.35.4+k0s`, and a Docker daemon alongside containerd. At inspection, +no pods were Pending or in another non-running, non-completed phase. This is a +point-in-time observation, not a complete application health test. + +The deployment checkout at `/srv/homelab` is on main commit `2adf17c`, behind the +reviewed local commit. It has untracked host configuration and a separate +`userbot/` directory. It was not reset or cleaned. + +| Observed difference | Implication | +| ----------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------- | +| VictoriaMetrics and vmalert are running; the Prometheus StatefulSet has zero replicas. | A monitoring migration is already in progress outside committed main. Deploying the old Helm values can overwrite those settings. | +| Homarr, Cloudflared, and Reloader are installed without their current Git active markers. | Installed services and marker-selected services are different inventories. Missing markers do not establish that a service is stopped. | +| Cloudflare DDNS is running in both Docker and Kubernetes. | Confirm which instance should own DNS updates and whether their domain lists overlap before retiring either one. Secret values were not inspected. | +| Traefik's LoadBalancer exposes port 8080 at `192.168.80.2`. | The direct API listener is deployed; its external reachability was not tested. | +| Default `local-path` has reclaim policy Delete, while many existing PVs have been changed to Retain. | Current retention is partly live state. Recreating a claim can get a different policy from the old PV. | +| NetBird, NetBox media/reports/scripts, EDU Redis, Homarr, and VictoriaMetrics have Delete-policy PVs. | Deleting their claims can delete important state. Plan backup and retention changes before namespace cleanup. | + +The monitoring files already modified in the user's local tree correspond to the +live migration. They are excluded from these branches. Reconcile that work before +using this review's baseline to deploy monitoring. + +## Remaining work + +These need recovery design or infrastructure decisions rather than a small +configuration correction: + +- **SSH apply retries can replace the rollback baseline.** `ssh-run.sh` retries + exit 255, including `apply-k8s`; every new invocation publishes a fresh snapshot. + If the first attempt already changed workloads, the retry snapshots that partial + state. Preserve a run-specific original baseline and verify it across retries. +- **Rollback can exceed the job budget.** Verification is parallel, but + `rollback_workloads` is serial with a five-minute limit per workload. The + thirty-minute job budget can expire before recovery finishes. Bound recovery + concurrency and account for both phases before choosing a new timeout. +- **Snapshot collection is allowed to fail.** Generation and workload snapshot + errors are warnings; verify can fall back to all workloads. A snapshot failure + must not permit unrelated workloads to be selected for automatic undo. +- **Rollback uses the previous revision, not the captured revision.** `rollout undo` + without an explicit revision cannot guarantee restoration to the snapshot after + retries or intervening rollouts. First deployments also have no previous revision. +- **Manual deploy dispatch bypasses the CI-success trigger.** Either validate the + target commit's successful CI run or document manual dispatch as an operator + override with its own required checks. +- **Direct Traefik API exposure is unauthenticated.** The latest local commit + explicitly added it for Homarr. Preserve that integration while choosing a + cluster-internal authenticated path or a verified network restriction; do not + simply disable an integration that is already in use. +- **Storage retention and backup are not reproducible as a whole.** Defaults and + several important PV policies are Delete. There is no repository-wide backup + schedule. Existing PVC StorageClass changes require migration rather than an + in-place YAML edit. +- **MeTube downloads are temporary on Kubernetes.** `/downloads` is a 20 GiB + emptyDir. Decide whether pod replacement should discard files or whether it + should use persistent storage. Compose uses a host directory instead. +- **First-time activation needs a bootstrap path.** Deploy validation dry-runs + namespaced resources before the apply stage creates namespaces and installs + selected charts. On a fresh cluster, missing namespaces and CRDs need separate + preparation; activation is not a complete installer. + +## Validation + +Baseline lint checks passed for Python, shell, workflows, YAML, standard Compose +files, and Kubernetes resources with available schemas. Kubeconform found 347 +resources in 174 files: 201 valid, 146 skipped CRDs, zero invalid resources. +That skip count matters: passing schema validation does not validate Traefik rule +strings or other controller-specific behavior. + +Fix validation covers: + +- Compose discovery of manual entry points, rejection of required-variable gaps, + namespace-scoped and optional Secret references, and API/render failures. +- NetBird setup idempotence, preservation of existing keys, file permissions, + runtime rendering, and rejection of invalid trusted proxy CIDRs. +- Session refresh success and failure paths, timeouts, cookie expiry, log redaction, + missing credentials, and nonpositive refresh intervals. +- Correct Glance ConfigMap key selection and PostgreSQL initializer/env alignment. +- YAML and Compose structure for the corrected router rules, compared with the + documented Traefik grammar. They were not exercised on the live proxy. +- Prune rejection before any cluster invocation. + +All seven fix branches and the documentation branch merged together without +conflicts in a disposable validation worktree. The combined tree passed the +CI-equivalent local checks, Markdown formatting/lint and link checks, all 35 +Compose structure checks, and 11 Python regression tests plus the shell +validation regressions. CRD server-side validation and live rollout tests were +not run. + +Runtime tests use fixtures and mocks, not production credentials. Live checks read +workload metadata, storage policies, chart versions, and container state only. +They did not read Secret contents or change services. diff --git a/downtify/README.md b/downtify/README.md new file mode 100644 index 0000000..d7f7758 --- /dev/null +++ b/downtify/README.md @@ -0,0 +1,20 @@ +# Downtify + +Download UI with a persistent downloads directory. + +Compose stores downloads under `Downtify_downloads/`; Kubernetes uses +`downtify-downloads-pvc`. The ingress manifests reference shared infrastructure, +so check certificate and middleware availability before enabling them. + +Back up downloads separately if they need to survive storage replacement. + +## Inspect + +From the repository root: + +```sh +kubectl get pods,svc,pvc -n downtify +kubectl get events -n downtify --sort-by=.metadata.creationTimestamp +``` + +See the [repository README](../README.md) for deployment selection. diff --git a/edu_master/README.md b/edu_master/README.md new file mode 100644 index 0000000..0429ee3 --- /dev/null +++ b/edu_master/README.md @@ -0,0 +1,52 @@ +# EDU session keeper and Telegram bot + +Keeps an EDU login session in Redis and sends Telegram notifications for new webinars. The bot also serves diary and schedule commands. + +`phpsessid-bot/` logs into EDU and publishes `EDU_PHPSESSID` in Redis. +`webinar-checker/` uses that cookie through a remote Playwright browser and stores +subscribers, language preferences, and webinar history in Redis. + +Kubernetes runs in `edu-master`, with Redis data in `redis-data-pvc`. +`service.yaml`, `servicemonitor.yaml`, and `alerts.yaml` expose and monitor the +checker's metrics on port 8000. Its `/health` endpoint reflects recent checks. + +## Configuration + +Use the keys in `k8s/secrets.yaml.example` as the reference. The committed Compose +`.env.example` has stale names until `fix/session-keeper-reliability` is merged. +The code reads: + +| Variable | Purpose | +| ----------------------------------------------------- | ------------------------------------------------ | +| `KEEPER_LOGIN`, `KEEPER_PASSWORD` | EDU login credentials. | +| `KEEPER_INTERVAL` | Session refresh interval in minutes; default 10. | +| `EDU_URL_BASE` | EDU site origin. | +| `EDU_URL_LOGIN`, `EDU_URL_COURSES`, `EDU_URL_WEBINAR` | Paths under that origin. | +| `WEBINAR_TELEGRAM_TOKEN`, `WEBINAR_ADMIN_ID` | Telegram bot and administrator. | +| `WEBINAR_CHECK_INTERVAL` | Checker interval in seconds; default 60. | +| `REDIS_HOST`, `REDIS_PORT` | Redis connection. | +| `PLAYWRIGHT_WS` | Remote browser WebSocket endpoint. | + +Set the keeper keys explicitly in the Compose `.env`. Keep the Playwright Python +package, browser image, server command, and `PLAYWRIGHT_VERSION` file on matching +versions. The two Python images are built and published by CI. + +## Bot use + +Start a private chat with `/start` to subscribe. `/stop`, `/language`, `/diary`, +`/schedule`, and `/setclass` manage subscriptions and school views. The +administrator can manage the whitelist with `/adduser` and `/removeuser`. + +Back up Redis if subscriber settings and notification history matter. Session +cookies and Telegram tokens are credentials; keep them out of shared logs. + +## Inspect + +From the repository root: + +```sh +kubectl get pods,svc,pvc -n edu-master +kubectl get events -n edu-master --sort-by=.metadata.creationTimestamp +``` + +See the [repository README](../README.md) for deployment selection. diff --git a/errorpages/README.md b/errorpages/README.md new file mode 100644 index 0000000..3bbc74a --- /dev/null +++ b/errorpages/README.md @@ -0,0 +1,21 @@ +# Error pages + +Static HTTP error pages served by an Nginx image built in CI. + +Edit the HTML in `html/`; the Dockerfile copies it into the image. +Kubernetes exposes `error-pages-service` in `error-pages` for Traefik's error +middleware. Keep the middleware's namespace and port aligned with that Service. + +For a local build, run `docker build -t homelab-error-pages .` from this directory. +Compose references the private registry image rather than a build context. + +## Inspect + +From the repository root: + +```sh +kubectl get pods,svc,pvc -n error-pages +kubectl get events -n error-pages --sort-by=.metadata.creationTimestamp +``` + +See the [repository README](../README.md) for deployment selection. diff --git a/gitea/README.md b/gitea/README.md new file mode 100644 index 0000000..8974ea4 --- /dev/null +++ b/gitea/README.md @@ -0,0 +1,25 @@ +# Gitea + +Git hosting with HTTP and a separate SSH route. + +Kubernetes uses the shared PostgreSQL service and `gitea-pvc` for repositories +and application data. Match the Gitea database password with the shared database +Secret. SSH is routed through Traefik's TCP entrypoint on 2221. + +Compose uses a separate PostgreSQL 14 database, bind mounts `gitea-data/` and +`gitea-db/`, and publishes host port 2221. It is an alternative deployment with +its own database, not a second frontend for the Kubernetes instance. + +Back up repositories, application configuration, and a consistent database dump +together. Gitea Actions definitions for this repository live in `../.gitea/`. + +## Inspect + +From the repository root: + +```sh +kubectl get pods,svc,pvc -n gitea +kubectl get events -n gitea --sort-by=.metadata.creationTimestamp +``` + +See the [repository README](../README.md) for deployment selection. diff --git a/glance/README.md b/glance/README.md new file mode 100644 index 0000000..29b6e90 --- /dev/null +++ b/glance/README.md @@ -0,0 +1,26 @@ +# Glance + +Dashboard pages for links, service checks, and Docker containers. + +Compose mounts `config/` and `assets/`. The Kubernetes equivalents are embedded +in `k8s/glance-config.yaml`: `glance-config` holds pages and `glance-assets` holds +`user.css`. Update both copies when changing shared content. + +Kubernetes serves the dashboard under `/glance`. Its pod also mounts the node's +Docker socket. It references `glance-secrets` for `ADGUARD_PASSWORD`, but there is +no tracked Secret example; create that Secret in `glance` before starting it. +Compose expects a local `.env` with the same password. + +The CSS mount points at the wrong ConfigMap on the reviewed main commit; +`fix/glance-assets` corrects it. + +## Inspect + +From the repository root: + +```sh +kubectl get pods,svc,pvc -n glance +kubectl get events -n glance --sort-by=.metadata.creationTimestamp +``` + +See the [repository README](../README.md) for deployment selection. diff --git a/headscale/README.md b/headscale/README.md new file mode 100644 index 0000000..924bbc8 --- /dev/null +++ b/headscale/README.md @@ -0,0 +1,18 @@ +# Headscale + +Headscale, Headplane, and a separate web administration UI on Docker. + +Kubernetes only provides routes to the Docker host. Update the addresses in +`k8s/routing/external-service.yaml` if the host moves. + +Copy `config/headscale.yaml.example`, `config/headplane.yaml.example`, and +`config/policy.json.example` to their names without `.example`. Set the public +server URL, DNS settings, Headplane cookie secret, and Headscale public URL. +The example URLs are placeholders. + +Compose publishes Headscale on 18080, its metrics port on 19090, Headplane on +13000, and the other UI on 10080. The data volumes store the Headscale database, +keys, and Headplane state. The embedded DERP configuration needs reachable +addresses; Compose does not publish its UDP 3478 listener. + +See the [repository README](../README.md) for deployment selection. diff --git a/homarr/README.md b/homarr/README.md new file mode 100644 index 0000000..b3ae09c --- /dev/null +++ b/homarr/README.md @@ -0,0 +1,24 @@ +# Homarr + +Dashboard with Kubernetes integration and persistent application state. + +Kubernetes uses the `homarr` ServiceAccount and the read-only ClusterRole in +`k8s/rbac.yaml`. Application data lives in `homarr-pvc`; supply the encryption key +from `k8s/secrets.yaml.example` before the first start and retain it with backups. + +The committed ingress is internal. There is no `k8s/active` marker even though +manifests exist, so the workflow does not select Homarr automatically. + +Compose publishes ports 80 and 81, mounts appdata and the Docker socket, and +expects a local kubeconfig. Check these host ports against Traefik before use. + +## Inspect + +From the repository root: + +```sh +kubectl get pods,svc,pvc -n homarr +kubectl get events -n homarr --sort-by=.metadata.creationTimestamp +``` + +See the [repository README](../README.md) for deployment selection. diff --git a/homepages/README.md b/homepages/README.md new file mode 100644 index 0000000..7ac6d53 --- /dev/null +++ b/homepages/README.md @@ -0,0 +1,25 @@ +# Homepages + +Two static sites: Forust and xdfnx. + +The site sources are in `forust_files/` and `xdfnx_files/`. CI builds each with +its own Dockerfile and publishes it to the private registry. Kubernetes serves +the image contents; Compose overlays the source directories as bind mounts. + +Both Traefik IngressRoute and Gateway API route manifests are committed. +Keep their hostnames and backend Services aligned when changing routes. +Certificate resources cover public and internal hostnames. + +Build either site locally with `docker build -f Dockerfile.forust .` or +`docker build -f Dockerfile.xdfnx .` from this directory. + +## Inspect + +From the repository root: + +```sh +kubectl get pods,svc,pvc -n homepages +kubectl get events -n homepages --sort-by=.metadata.creationTimestamp +``` + +See the [repository README](../README.md) for deployment selection. diff --git a/immich/README.md b/immich/README.md new file mode 100644 index 0000000..cae47c5 --- /dev/null +++ b/immich/README.md @@ -0,0 +1,29 @@ +# Immich + +Photo library with its own vector-enabled PostgreSQL and machine-learning service. + +This database is separate from the shared PostgreSQL instance. Keep the server +and machine-learning versions aligned when upgrading. + +Kubernetes bind-mounts `/mnt/immich/library` from the node. That directory must +already exist and contain the intended library; moving the pod to a different +node does not move the files. PostgreSQL and Valkey use StatefulSet storage, and +the model cache has its own PVC. + +Compose reads `UPLOAD_LOCATION` and `DB_DATA_LOCATION` from `.env`. The example +uses the same library path as Kubernetes. Run one writer against that library; +do not start both deployments as independent instances over the same files. + +Back up the library and a consistent database dump together. The model cache +can be rebuilt; the photo database cannot. + +## Inspect + +From the repository root: + +```sh +kubectl get pods,svc,pvc -n immich +kubectl get events -n immich --sort-by=.metadata.creationTimestamp +``` + +See the [repository README](../README.md) for deployment selection. diff --git a/kener/README.md b/kener/README.md new file mode 100644 index 0000000..55afade --- /dev/null +++ b/kener/README.md @@ -0,0 +1,21 @@ +# Kener + +Status page with Redis and persistent database and upload directories. + +Kubernetes uses `kener-db-pvc`, `kener-uploads-pvc`, and a Redis StatefulSet. +Compose keeps the corresponding directories in named volumes. Set the signing +and other credentials from the env or Secret example. + +The monitors and route settings live in `k8s/config.yaml` and `k8s/ingress.yaml`. +There is no active marker for either runtime. + +## Inspect + +From the repository root: + +```sh +kubectl get pods,svc,pvc -n kener +kubectl get events -n kener --sort-by=.metadata.creationTimestamp +``` + +See the [repository README](../README.md) for deployment selection. diff --git a/loki/README.md b/loki/README.md new file mode 100644 index 0000000..d64bcae --- /dev/null +++ b/loki/README.md @@ -0,0 +1,16 @@ +# Loki and Alloy + +Loki log storage and Alloy collection, both deployed through Helm. + +The deploy library lists separate `loki` and `alloy` releases in `prometheus`, +controlled by this directory's `k8s/active` marker. Chart versions are pinned in +`deploy-lib.sh`; settings live in `loki-values.yaml` and `alloy-values.yaml`. + +Alloy collects Kubernetes logs. Grafana's Loki datasource is configured in the +monitoring stack. Review Loki retention and storage settings before enabling +collection on a new cluster. + +Check releases with `helm list -n prometheus` and inspect collector logs before +assuming that an empty Grafana query means there were no events. + +See the [repository README](../README.md) for deployment selection. diff --git a/metube/README.md b/metube/README.md new file mode 100644 index 0000000..72be212 --- /dev/null +++ b/metube/README.md @@ -0,0 +1,21 @@ +# MeTube + +Web downloader behind Traefik. + +Compose bind-mounts `MeTube_downloads/` on the host. Kubernetes uses a 20 GiB +`emptyDir` for `/downloads`: completed downloads disappear when the pod is +replaced. Download files from the UI promptly if this temporary storage is intended. + +Application settings are in `k8s/config.yaml`. Persisting downloads in Kubernetes +would require changing the volume to a PVC and choosing a storage policy. + +## Inspect + +From the repository root: + +```sh +kubectl get pods,svc,pvc -n metube +kubectl get events -n metube --sort-by=.metadata.creationTimestamp +``` + +See the [repository README](../README.md) for deployment selection. diff --git a/n8n/README.md b/n8n/README.md new file mode 100644 index 0000000..4184843 --- /dev/null +++ b/n8n/README.md @@ -0,0 +1,21 @@ +# n8n + +Workflow automation with persistent application and file storage. + +Kubernetes keeps application state in `n8n-node-pvc` and files in +`n8n-files-pvc`; Compose uses `node-data` and `files` named volumes. +Webhook URLs and proxy settings are committed in the application config. + +There is no active marker. Review the URLs before enabling the stack, and retain +the credential encryption key with the database or application-data backup. + +## Inspect + +From the repository root: + +```sh +kubectl get pods,svc,pvc -n n8n +kubectl get events -n n8n --sort-by=.metadata.creationTimestamp +``` + +See the [repository README](../README.md) for deployment selection. diff --git a/netbird/README.md b/netbird/README.md index f5ef834..9be7dfe 100644 --- a/netbird/README.md +++ b/netbird/README.md @@ -1,123 +1,64 @@ # NetBird -Self-hosted NetBird with the combined management, signal, relay, and STUN server. The dashboard and server run behind the repository's existing external Traefik instance on the Docker `proxy` network. Only STUN UDP `3478` is published directly. +Self-hosted NetBird with the combined management, signal, relay, and STUN server. -The deployment uses SQLite for a single-instance homelab server. The persistent `netbird_data` volume and the datastore encryption key are both required to recover the installation. +Kubernetes runs the server and dashboard in `netbird`. The server uses SQLite +in `netbird-pvc`; `k8s/config.yaml` contains the template and runtime renderer. +Prepare `netbird-secrets` from `k8s/secrets.yaml.example` before the first start. -## Files +## Routing and keys -- `compose.yaml`: dashboard and combined server; selected by the marker-driven deploy workflow through `active`. -- `config.template.yaml`: non-secret server configuration rendered at startup. -- `entrypoint.sh`: injects Docker secrets into an in-memory runtime configuration. -- `client.compose.yaml`: optional host-network peer using a dashboard-generated setup key. -- `.env`: ignored local hostnames, the detected Traefik Docker-network subnet, and optional client setup key. -- `secrets/`: ignored relay secret and datastore encryption key. +The public hostname is set in the ConfigMap and ingress rules. Keep the issuer, +dashboard endpoints, and public routes consistent. HTTP, WebSocket, and gRPC +traffic go through Traefik; the STUN route uses UDP 3478. A CDN's HTTP proxy does +not provide that UDP listener. -## First deployment +`NETBIRD_PROXY_SUBNET` controls which forwarded client addresses are trusted. +Use the actual proxy network CIDR rather than assuming another lab's subnet. +Keep the datastore encryption key with every datastore backup. Regenerating it +can make stored credentials unreadable. -Run these commands on the Docker host before merging the activating branch. The deploy preflight resets tracked files but preserves ignored local state. +## Compose alternative -```bash -cd /srv/homelab/netbird +`compose.yaml` expects `entrypoint.sh`, a local `.env`, and two local files: +`secrets/relay-auth-secret` and `secrets/datastore-encryption-key`. +The reviewed main commit is missing the renderer and setup script. +`fix/netbird-compose-runtime` restores them. Merge that fix before following +these setup commands: + +```sh +cd netbird ./setup.sh $EDITOR .env docker compose config --quiet docker compose up -d ``` -Review the values in `.env` before starting. The example public hostname is `netbird.forust.xyz`; change it if a different public domain was selected. `setup.sh` replaces `NETBIRD_PROXY_SUBNET=auto` with the first IPv4 subnet of the external Docker `proxy` network. Keep that value synchronized with the network; set an explicit CIDR instead if the network is managed elsewhere. - -`setup.sh` is idempotent and never replaces existing secrets. Do not delete or regenerate `secrets/datastore-encryption-key` after the first successful start unless all encrypted setup keys and API tokens are intentionally being invalidated. - -## Network prerequisites - -- Point the public hostname directly to the Docker host. Do not proxy UDP `3478` through Cloudflare or another CDN. -- Allow inbound TCP `80`, TCP `443`, and UDP `3478` through the host firewall and upstream router. -- Ensure the external `proxy` Docker network exists and Traefik uses its `websecure` entrypoint and `letsencrypt` resolver. `NETBIRD_PROXY_SUBNET` must describe that network; it is used to trust only forwarded client addresses from Traefik. -- Ensure the internal names in `.env` resolve where the local and development aliases are needed. -- Keep Traefik's `websecure` read timeout disabled for long-lived gRPC and WebSocket sessions. This repository configures `--entrypoints.websecure.transport.respondingTimeouts.readTimeout=0` in `traefik/compose.yaml`. - -After startup, verify OIDC discovery through the public TLS endpoint: - -```bash -curl -fsS "https://${NETBIRD_DOMAIN}/oauth2/.well-known/openid-configuration" -``` - -Open `https://${NETBIRD_DOMAIN}` immediately and complete the initial owner setup. Treat the initial setup flow as public until the owner exists. +The setup script detects the IPv4 subnet of the external `proxy` network and +preserves existing secrets. Complete the initial owner setup through the public +TLS endpoint after starting the server. ## Optional host client -The client intentionally lives in a separate Compose project. Normal server deploys use `--remove-orphans`, so keeping the client in the server project would cause it to be removed. +`client.compose.yaml` runs a host-network peer in a separate Compose project. +Set `NB_SETUP_KEY` and `NETBIRD_CLIENT_HOSTNAME` in the local `.env`, then run +`docker compose -f client.compose.yaml up -d`. It needs `/dev/net/tun` and elevated +network capabilities. The normal server deployment does not start this client. -1. Create a reusable or ephemeral setup key in the NetBird dashboard. -2. Put `NB_SETUP_KEY=` in the ignored `netbird/.env` file. -3. Set `NETBIRD_CLIENT_HOSTNAME` to this machine's desired peer name. -4. Start and inspect the client: +## Backup -```bash -cd /srv/homelab/netbird -docker compose -f client.compose.yaml config --quiet -docker compose -f client.compose.yaml up -d -docker compose -f client.compose.yaml exec netbird-client netbird status +Back up the SQLite data while the server is stopped, together with the encryption +key, relay secret, and local configuration. Test a restore on an isolated host. +For Compose, the datastore volume has the explicit name `netbird_data`. +Do not use `docker compose down -v` when keeping the installation. + +## Inspect + +From the repository root: + +```sh +kubectl get pods,svc,pvc -n netbird +kubectl get events -n netbird --sort-by=.metadata.creationTimestamp ``` -The client uses host networking and requires `NET_ADMIN`, `SYS_ADMIN`, `SYS_RESOURCE`, and `/dev/net/tun`. Remove it without affecting the server stack: - -```bash -docker compose -f client.compose.yaml down -``` - -## Operations - -Inspect status and logs: - -```bash -docker compose ps -docker compose logs --tail=200 netbird-server dashboard -``` - -Stop or remove containers without deleting data: - -```bash -docker compose down -``` - -Do not add `-v` to `docker compose down`; it would delete the NetBird datastore. - -## Backup and restore - -Back up both the persistent volume and the ignored secret files. For a consistent SQLite backup, briefly stop the server first and store the resulting archive and `datastore-encryption-key` in an encrypted backup: - -```bash -cd /srv/homelab/netbird -mkdir -p backups -docker compose stop netbird-server -docker run --rm \ - -v netbird_data:/data:ro \ - -v "$PWD/backups:/backup" \ - busybox:1.37.0 \ - tar -C /data -czf "/backup/netbird-data-$(date -u +%Y%m%dT%H%M%SZ).tar.gz" . -docker compose start netbird-server -``` - -Also securely back up: - -- `secrets/datastore-encryption-key` — required to decrypt stored secrets. -- `secrets/relay-auth-secret` — keeps issued relay credentials valid across restoration. -- `netbird/.env` — optional, but it records the public and internal hostnames. - -Test a restore in an isolated Docker host before relying on a backup. - -## Upgrade - -1. Take and verify a backup. -2. Review NetBird release notes for server, client, and dashboard compatibility. -3. Update the pinned tags in `compose.yaml`; update `client.compose.yaml` separately when deploying the client. -4. Pull and recreate the selected services: - -```bash -docker compose pull -docker compose up -d -``` - -The image tags are intentionally pinned instead of using `latest`, matching this repository's pull-on-deploy policy. +See the [repository README](../README.md) for deployment selection. diff --git a/netbox/README.md b/netbox/README.md index 3a352f8..5ed70c6 100644 --- a/netbox/README.md +++ b/netbox/README.md @@ -1,96 +1,43 @@ # NetBox -NetBox for homelab documentation and visualization. Two runtimes are available: +Inventory and network documentation with a web process, worker, and Valkey. -| Runtime | Manifest | Purpose | -| ------- | -------------- | -------------------------------------------------------------- | -| Docker | `compose.yaml` | Local stand on `127.0.0.1:8000` (no public exposure) | -| k8s | `k8s/` | Homelab service on `netbox.forust.xyz` (and the internal name) | +Kubernetes uses the shared PostgreSQL service at +`postgres.database.svc.cluster.local:5432`, database and role `netbox`. +The database and application Secrets must contain the same password. +Media, reports, scripts, and Valkey have persistent storage. -Both use the same image (`netboxcommunity/netbox:v4.7-5.1.1`) and Valkey for tasks -plus a second logical database for caching. The Docker stand keeps its own -PostgreSQL container, while the k8s deployment uses the shared `database` cluster -(`postgres.database.svc.cluster.local:5432`, role/database `netbox`); only Valkey -stays a per-service StatefulSet. +Compose has its own PostgreSQL container and Valkey instances. It publishes the +web UI on `127.0.0.1:8000`; its Traefik labels can also expose it while a Docker +proxy is running. Copy `.env.example` to `.env`, replace the credentials, and run +`docker compose config --quiet` before starting it. -## Docker Compose +## First Kubernetes start -```bash -cp .env.example .env -# replace CHANGE_ME -docker compose up -d +Create the namespace and application Secret. Provision the database through the +shared database initializer on a fresh instance, or create the role and database +manually on an existing instance; see [PostgreSQL](../postgres/README.md). +The database NetworkPolicy already includes `netbox`. + +Apply the selected application manifests after the database is ready. Startup +runs schema migrations, so the probes allow a longer first boot. Inspect web and +worker logs before retrying a slow migration. + +## Settings and backup + +`configuration/configuration.py` is the Compose settings file. Its Kubernetes +copy is embedded in `k8s/settings.yaml`; keep them aligned. +Back up the database and media together. Keep `SECRET_KEY` and +`API_TOKEN_PEPPER_1`: changing them invalidates sessions or API tokens. +A container rollback cannot undo a database migration. + +## Inspect + +From the repository root: + +```sh +kubectl get pods,svc,pvc -n netbox +kubectl get events -n netbox --sort-by=.metadata.creationTimestamp ``` -The UI is available at . The port is bound to `127.0.0.1` -intentionally, so this stand is not exposed on the LAN or public interfaces. - -The `netbox` service is also attached to the external `proxy` network and carries -Traefik labels for `netbox.forust.xyz` and `netbox.workstation.internal`. Those -labels only take effect while the Docker Traefik stack is running; it is currently -stopped, and the live ingress path in this homelab is the k8s Traefik. - -Inspect startup and health with: - -```bash -docker compose ps -docker compose logs -f netbox -``` - -Stop it with `docker compose down`; data is kept in the named volumes -`netbox-postgres`, `netbox-media-files`, `netbox-reports-files`, -`netbox-scripts-files` and `netbox-redis-data`. - -## Kubernetes - -`k8s/` is deployed in the homelab cluster and serves `netbox.forust.xyz` publicly -plus `netbox.workstation.internal` / `netbox.gigaforust.internal` internally. To -rebuild it from scratch: - -```bash -# 1. shared PostgreSQL: the password lives in the shared secret, NetBox keeps a copy -kubectl -n database patch secret postgres-shared-secrets \ - --type merge -p '{"stringData":{"NETBOX_DB_PASSWORD":""}}' -kubectl -n database exec postgres17-0 -- psql -U postgres -d postgres \ - -c 'CREATE ROLE netbox LOGIN PASSWORD ...' -c 'CREATE DATABASE netbox OWNER netbox' - -# 2. secrets first: the deploy workflow never applies *secret*.yaml -cp k8s/secrets.yaml.example k8s/secrets.yaml # replace CHANGE_ME -kubectl apply -f k8s/secrets.yaml - -# 3. manifests -kubectl apply -f k8s/ -``` - -The shared cluster is reached at `postgres.database.svc.cluster.local:5432`. Its -NetworkPolicy (`postgres/k8s/network-policy.yaml`) must list the `netbox` namespace -or connections are dropped, and `postgres/initdb/01-create-databases.sh` already -creates the role and database on a fresh data directory. NetBox has no PostgreSQL -StatefulSet of its own — only `netbox-valkey`. - -`netbox.forust.xyz` resolves to this host (`78.98.72.122`) through the `DOMAINS` -list in the `default/cfddns` secret. cert-manager issues `netbox-prod-tls` with the -`letsencrypt-prod` issuer, the internal route uses `internal-wildcard-tls`. - -Resources are permanent again now that the first-boot migrations are complete: -the web container reserves `100m`/`512Mi` and is capped at `2` CPU/`2Gi`, the -worker reserves `50m`/`256Mi` and is capped at `1` CPU/`1Gi`, and Valkey reserves -`25m`/`64Mi` and is capped at `250m`/`256Mi`. The deliberately generous CPU caps -leave enough headroom for future schema migrations without letting one process -consume the whole node. - -The first start applies ~810 migrations, each in its own transaction with DDL and -a commit; every later start is a no-op. The startup probe allows 15 minutes and -`progressDeadlineSeconds` is 1800 for the same reason. Probes run inside the pod -and explicitly set `Host: netbox.forust.xyz`; a kubelet `httpGet.host` field would -replace the probe destination with that public hostname and bypass the pod. - -## Secrets - -- `netbox/.env` (compose) and `netbox/k8s/secrets.yaml` (k8s) are gitignored. Only - `.env.example` and `k8s/secrets.yaml.example` are committed. -- `netbox/configuration/configuration.py` is env-driven: hosts, database, Redis and - the Django keys all come from the environment, so the same settings file works in - both runtimes. The k8s copy lives in the `netbox-settings` ConfigMap - (`k8s/settings.yaml`) and must be kept in sync with the file. -- Rotating `SECRET_KEY` invalidates all sessions; rotating `API_TOKEN_PEPPER_1` - invalidates every API token. +See the [repository README](../README.md) for deployment selection. diff --git a/netronome/README.md b/netronome/README.md new file mode 100644 index 0000000..ce8ec48 --- /dev/null +++ b/netronome/README.md @@ -0,0 +1,22 @@ +# Netronome + +Network monitoring application using the shared PostgreSQL instance on Kubernetes. + +Kubernetes reads application settings from its ConfigMap and Secret. Match the +Netronome role password with `NETRONOME_DB_PASSWORD` in the shared database Secret. +Its namespace is included in the PostgreSQL NetworkPolicy. + +The Compose configuration is a separate deployment; review its local database +settings and env example before starting it. Keep monitoring history in the +database backup. + +## Inspect + +From the repository root: + +```sh +kubectl get pods,svc,pvc -n netronome +kubectl get events -n netronome --sort-by=.metadata.creationTimestamp +``` + +See the [repository README](../README.md) for deployment selection. diff --git a/nextcloud/README.md b/nextcloud/README.md new file mode 100644 index 0000000..54df8e9 --- /dev/null +++ b/nextcloud/README.md @@ -0,0 +1,17 @@ +# Nextcloud AIO + +Nextcloud All-in-One on Docker, with Kubernetes routes to the Docker host. + +The master container manages its own child containers through the Docker +socket. Kubernetes does not run the Nextcloud application; the EndpointSlices +under `k8s/routing/` point to host services. + +Compose publishes the AIO administration interface on 8888. The Apache frontend +uses host port 11000. `NEXTCLOUD_DATADIR` is `/mnt/nextcloud/ncdata`; prepare that +storage before first setup and do not change the path casually afterwards. + +Use AIO's backup and restore tools for the managed application. Keep the master +configuration volume and the data directory with the recovery plan. Do not +remove child containers just because they do not appear as Compose services. + +See the [repository README](../README.md) for deployment selection. diff --git a/penpot/README.md b/penpot/README.md new file mode 100644 index 0000000..4b50da0 --- /dev/null +++ b/penpot/README.md @@ -0,0 +1,12 @@ +# Penpot + +A Compose-only design application with frontend, backend, exporter, database, and cache. + +There is no active marker or Kubernetes deployment here. Configure the public +URL and credentials from `.env.example` before starting `compose.yaml`. + +Penpot has its own PostgreSQL container. The shared database initializer still +contains a Penpot role, but this Compose stack does not use it. +Back up the application assets and database together. + +See the [repository README](../README.md) for deployment selection. diff --git a/portainer/README.md b/portainer/README.md new file mode 100644 index 0000000..43d5a3f --- /dev/null +++ b/portainer/README.md @@ -0,0 +1,21 @@ +# Portainer + +Container management UI backed by the host Docker socket. + +Kubernetes mounts the node's Docker socket and persists application data in +`portainer-data-pvc`. This targets Docker on that node, not Kubernetes workloads. +Compose uses the `portainer_data` volume for its state. + +Review initial administrator setup and route access before exposing the UI. +Neither deployment has an active marker. + +## Inspect + +From the repository root: + +```sh +kubectl get pods,svc,pvc -n portainer +kubectl get events -n portainer --sort-by=.metadata.creationTimestamp +``` + +See the [repository README](../README.md) for deployment selection. diff --git a/postgres/README.md b/postgres/README.md index db166ac..c628b44 100644 --- a/postgres/README.md +++ b/postgres/README.md @@ -1,35 +1,63 @@ # Shared PostgreSQL -This directory contains the shared PostgreSQL 17 deployment for Authentik, -Gitea, NetBox, Netronome, and Statuspage. It creates one database and one login role -per service. Per-service standalone databases were removed after the -migration (Sep 2026); Penpot stays on its own compose PostgreSQL (archived, -not part of the shared instance). +PostgreSQL 17 for the Kubernetes deployments of Authentik, Gitea, NetBox, and Netronome. -## Compatibility baseline +The server runs in `database` as StatefulSet `postgres17`, with data in +`postgres17-data`. Applications connect to +`postgres.database.svc.cluster.local:5432`. The NetworkPolicy allows only the +listed application namespaces; add a new consumer there as well as provisioning +its database. -| Service | Current application | Shared PostgreSQL 17 | -| ---------- | ------------------- | -------------------------------------- | -| Authentik | 2025.10.x | Supported (Authentik requires 14+) | -| Gitea | 1.27.3 | Supported (Gitea requires 12+) | -| NetBox | 4.7.x | Supported (NetBox 4.x requires 13+) | -| Netronome | 0.14.0 | Supported (upstream's example uses 17) | -| Statuspage | custom | Supported | +## Initialization -A major-version change must use a logical dump/restore; changing only the -image tag while keeping a data directory is not supported. +`initdb/01-create-databases.sh` creates roles and databases on an empty data +directory. The Kubernetes copy is embedded in `k8s/postgres.yaml`. +It also provisions Penpot and Statuspage roles, even though those are not active +consumers in the current Kubernetes manifests. -For Compose, copy `.env.example` to `.env`, set all passwords, and start it with -`docker compose -f shared-compose.yaml up -d`. This file is intentionally not -named `compose.yaml`, so the repository deploy workflow does not start a second -database accidentally. -Applications that use this database must also join that external network and use -`homelab-postgres:5432`. +The initializer requires every listed password. Prepare `k8s/secrets.yaml` from +the example before applying the StatefulSet. Existing application Secrets keep +copies of their own database passwords; they must match the corresponding role. -For Kubernetes, create `k8s/secrets.yaml` from the example before applying the -manifests. The `k8s/active` marker makes the normal deploy workflow include the -namespace, StatefulSet, ConfigMap, and NetworkPolicy. Applications use -`postgres.database.svc.cluster.local:5432`. -Migrate each existing database with a tested logical dump/restore before -switching an application. Do not reuse a PostgreSQL 14 or 17 data directory -with PostgreSQL 15. +The init scripts do not run again when an existing data directory is mounted. +Changing a Secret does not rotate the PostgreSQL role password. Rotate the role +with SQL and update the application Secret together. + +## Compose alternative + +From this directory: + +```sh +cp .env.example .env +$EDITOR .env +docker compose -f shared-compose.yaml config --quiet +docker compose -f shared-compose.yaml up -d +``` + +Add `NETBOX_DB_PASSWORD` to `.env` as well: the reviewed env example omits it; +`fix/postgres-env-example` restores the key. Fill every required password. +This stack creates the `homelab-database` Docker network and the +`homelab-postgres` container. Compose applications need to join that network +explicitly to use it; several committed Compose stacks use their own databases. + +The filename is intentional: the automatic deploy discovery does not start this +stack just because the Kubernetes database is active. + +## Backup and upgrades + +Keep database dumps and role definitions, including ownership and grants. +Take a logical backup before changing a major PostgreSQL version. A new image +tag over the existing data directory is not a major-version migration. +Test restores separately before changing application connection settings. +Immich uses its own vector-enabled database and is outside this shared instance. + +## Inspect + +From the repository root: + +```sh +kubectl get pods,svc,pvc -n database +kubectl get events -n database --sort-by=.metadata.creationTimestamp +``` + +See the [repository README](../README.md) for deployment selection. diff --git a/prometheus-stack/README.md b/prometheus-stack/README.md new file mode 100644 index 0000000..3fc4224 --- /dev/null +++ b/prometheus-stack/README.md @@ -0,0 +1,26 @@ +# Monitoring stack + +Prometheus, Grafana, Alertmanager, and application alert rules. + +Kubernetes installs `kube-prometheus-stack` in `prometheus` through the deploy +library. Its chart version is pinned there; `k8s/grafana-values.yaml` contains the +values for the whole stack, despite the filename. + +The values file is tracked. `.gitignore` also lists it, but that does not stop Git +tracking later edits. Keep local credentials in Secrets rather than treating +changes to this file as ignored. + +Prepare Grafana admin and Alertmanager Secrets separately. The Alertmanager +configuration example and Telegram template are in `k8s/`; the main deploy +selection excludes the example config. Certificates and ingress expose Grafana, +and the rule files add service-specific alerts. + +The values use `local-path` PVCs for Grafana, Prometheus, and Alertmanager. +Retention is limited by both time and size. Back up Grafana state and any history +that must survive a storage failure. + +The Compose stack has separate Prometheus and Alertmanager configuration files. +There is no root active marker. This README describes committed main files; +local VictoriaMetrics experiments are not part of that configuration. + +See the [repository README](../README.md) for deployment selection. diff --git a/rackpeek/README.md b/rackpeek/README.md new file mode 100644 index 0000000..baabf9f --- /dev/null +++ b/rackpeek/README.md @@ -0,0 +1,20 @@ +# RackPeek + +Rack inventory UI behind Traefik. + +Kubernetes stores configuration in `rackpeek-pvc`. The Compose alternative uses +its own data mount. Keep rack descriptions and inventory data in the backup. + +Public and internal certificates and routes are in `k8s/`. There are no tracked +Secret examples for this service. + +## Inspect + +From the repository root: + +```sh +kubectl get pods,svc,pvc -n rackpeek +kubectl get events -n rackpeek --sort-by=.metadata.creationTimestamp +``` + +See the [repository README](../README.md) for deployment selection. diff --git a/reloader/README.md b/reloader/README.md new file mode 100644 index 0000000..2ac9b54 --- /dev/null +++ b/reloader/README.md @@ -0,0 +1,13 @@ +# Reloader + +Helm settings for restarting workloads when referenced configuration changes. + +The deploy library knows the `reloader` Helm release and the values file here, +but there is no `k8s/active` marker, so it is not upgraded automatically. + +Reloader only acts on workloads configured for it. A ConfigMap mounted with +`subPath` does not update inside an existing container by itself. Verify that the +application restarted after a configuration change rather than assuming an +apply updated the running process. + +See the [repository README](../README.md) for deployment selection. diff --git a/renovate/README.md b/renovate/README.md index 99039b6..9296c4b 100644 --- a/renovate/README.md +++ b/renovate/README.md @@ -1,101 +1,51 @@ -# Renovate for Gitea +# Renovate -Renovate runs as a Kubernetes CronJob and creates container image update pull -requests in Gitea. It does not deploy changes itself. +Container and chart dependency updates for the Gitea repository. -## Kubernetes +The Kubernetes CronJob runs in `renovate` every six hours with overlapping +CronJob executions forbidden. Prepare the bot PAT from the Secret example. +Give the dedicated Gitea user access to the repositories it should update. -Create a dedicated Gitea user named `renovate-bot`, create a repository access -token, and grant it repository read/write plus issue read/write permissions. -Add `read:packages` if Renovate must inspect private Gitea registry images. - -Create the ignored Secret locally; never commit the PAT: +`renovate.json` is the source configuration. The ConfigMap is a generated copy: ```sh -cp renovate/k8s/secrets.yaml.example renovate/k8s/secrets.yaml -$EDITOR renovate/k8s/secrets.yaml -kubectl apply -f renovate/k8s/namespace.yaml -kubectl apply -f renovate/k8s/secrets.yaml -kubectl apply -f renovate/k8s/configmap.yaml -kubectl apply -f renovate/k8s/cronjob.yaml +.gitea/workflows/sync-renovate-configmap.sh +.gitea/workflows/sync-renovate-configmap.sh --check ``` -The `renovate/k8s/active` marker makes the normal deployment workflow include -the namespace, ConfigMap, and CronJob. The Secret is intentionally excluded -from Git and must be applied separately after every new cluster. +Run those commands from the repository root. The `renovate-ci` workflow checks +that the generated configuration agrees with the source. -Run it immediately instead of waiting for the six-hour schedule. +## Run manually -Two options, both use the same `renovate/renovate.json`: +From the repository root: ```sh kubectl create job --from=cronjob/renovate renovate-manual-$(date +%s) -n renovate +kubectl get jobs,pods -n renovate ``` -or the `renovate-run` Actions workflow (Actions tab → `renovate-run` → -Run workflow). It runs the same image as the CronJob on the self-hosted runner -via Docker — the tag is read out of `renovate/k8s/cronjob.yaml` at run time -rather than hardcoded, so the two cannot drift apart. Required Actions secrets -(repo or org settings): +Alternatively use the `renovate-run` Actions workflow. It reads the image tag +from the CronJob and accepts repository, log-level, and dry-run inputs. Actions +requires `RENOVATE_TOKEN`; `RENOVATE_GITHUB_COM_TOKEN` is optional. +The Actions concurrency group and the CronJob policy are separate, so avoid +starting both against the same repository at once. -- `RENOVATE_TOKEN` — renovate-bot PAT (repository + issue read/write). -- `RENOVATE_GITHUB_COM_TOKEN` — optional, for changelogs and GitHub rate limits. +For Compose, copy `.env.example` to `.env` in this directory and run +`docker compose -f renovate-compose.yaml run --rm renovate`. That file is a +manual entry point and is not selected by the deploy workflow. -Inputs: `repositories` (default `forust/homelab`), `log_level` -(`info`/`debug`). Only one run at a time (concurrency group -`renovate-run`), same as the CronJob `Forbid` policy. +The config also tracks chart versions in `deploy-lib.sh` and tool versions in +`tool-versions.env`. Renovate opens pull requests; the normal CI and deploy +workflows handle changes after merge. -Inspect runs with: +## Inspect + +From the repository root: ```sh -kubectl get cronjob,jobs,pods -n renovate -kubectl logs -n renovate job/ +kubectl get pods,svc,pvc -n renovate +kubectl get events -n renovate --sort-by=.metadata.creationTimestamp ``` -`RENOVATE_GITHUB_COM_TOKEN` is optional but recommended for changelogs and -GitHub API rate limits. Set it in the Kubernetes Secret if available. - -## Compose - -Copy `.env.example` to `.env`, set the PAT, and run: - -```sh -docker compose -f renovate-compose.yaml run --rm renovate -``` - -The Compose file is intentionally named `renovate-compose.yaml`, so the -repository's automatic deployment discovery does not start it accidentally. - -## Configuration - -`renovate/renovate.json` is the single source of truth. The Compose file and the -`renovate-run` workflow mount that file directly. - -A ConfigMap cannot read from the repository, so the CronJob needs the config -inlined. `renovate/k8s/configmap.yaml` is therefore a **generated** copy: - -```sh -.gitea/workflows/sync-renovate-configmap.sh # regenerate after editing -.gitea/workflows/sync-renovate-configmap.sh --check # fail if out of date -``` - -The `renovate-ci` workflow runs the `--check` form on every PR and push, so a -config edit that forgets to regenerate the ConfigMap cannot be merged. - -Beyond images, `customManagers` in the config track: - -- Helm chart versions pinned in `.gitea/workflows/deploy-lib.sh`. The built-in - `helmv3` manager only reads `Chart.yaml` and `helm-values` only reads values - files, so neither sees a version written into a `helm upgrade` command — - these are declared as `custom.regex` managers against the `helm` datasource. -- CI linter versions in `.gitea/workflows/tool-versions.env`. - -The Renovate image tag is deliberately _not_ in `tool-versions.env`: -`renovate/k8s/cronjob.yaml` owns it, and the workflows read it from there. - -## How updates flow - -Renovate scans both `compose.yaml` files and Kubernetes manifests, opens a -branch and PR with image tag changes, and waits for CI. After merge, the -existing deployment workflow applies Kubernetes changes or redeploys Compose -stacks. Renovate never updates running workloads directly. +See the [repository README](../README.md) for deployment selection. diff --git a/searxng/README.md b/searxng/README.md new file mode 100644 index 0000000..ef3d047 --- /dev/null +++ b/searxng/README.md @@ -0,0 +1,22 @@ +# SearXNG + +Search frontend with a separate Valkey cache. + +Kubernetes keeps the application settings in a ConfigMap and starts Valkey as a +StatefulSet. Set the secret from the example before exposing the search endpoint. +There is no active marker. + +Compose expects local configuration under `core-config/`, which is ignored. +Prepare it before starting the stack; a container image alone does not supply +this lab's settings. + +## Inspect + +From the repository root: + +```sh +kubectl get pods,svc,pvc -n searxng +kubectl get events -n searxng --sort-by=.metadata.creationTimestamp +``` + +See the [repository README](../README.md) for deployment selection. diff --git a/streaming/README.md b/streaming/README.md new file mode 100644 index 0000000..e0d3cf6 --- /dev/null +++ b/streaming/README.md @@ -0,0 +1,16 @@ +# Media stack + +Docker services for playback, requests, library management, and downloads. + +Compose runs Jellyfin, Jellyseerr, Sonarr, Radarr, Prowlarr, qBittorrent, and the +other services declared in the file. Kubernetes only routes to host endpoints; +update `k8s/routing/external-service.yaml` when the Docker host or ports change. + +Prepare the paths, user/group IDs, and credentials from `.env.example`. Service +configuration and media/download directories are bind mounts. Preserve their +permissions when moving data, and keep the application databases with backups. + +Review device mounts for hardware acceleration before starting on another host. +Both active markers are present, so normal deploys include Compose and routing. + +See the [repository README](../README.md) for deployment selection. diff --git a/termix/README.md b/termix/README.md new file mode 100644 index 0000000..be8e14b --- /dev/null +++ b/termix/README.md @@ -0,0 +1,20 @@ +# Termix + +Terminal and SSH connection manager with persistent application data. + +Kubernetes stores state in `termix-pvc`; Compose mounts `termix-data/`. +The application config and routes are committed separately under `k8s/`. + +There is no active marker. Review access control and retain the application data +needed to recover saved connections before enabling it. + +## Inspect + +From the repository root: + +```sh +kubectl get pods,svc,pvc -n termix +kubectl get events -n termix --sort-by=.metadata.creationTimestamp +``` + +See the [repository README](../README.md) for deployment selection. diff --git a/traefik/README.md b/traefik/README.md new file mode 100644 index 0000000..c4c8e75 --- /dev/null +++ b/traefik/README.md @@ -0,0 +1,33 @@ +# Traefik + +Ingress for HTTP, gRPC, TCP, and UDP services, with public and internal TLS. + +Kubernetes uses the Helm settings in `k8s/traefik-values.yaml`. The deploy +library applies supporting resources in this directory but does not install or +upgrade the Traefik chart. Bootstrap the chart and CRDs separately. + +The LoadBalancer address is set to `192.168.80.2`. Change it for another network. +Entrypoints include web traffic, Gitea SSH, NetBird STUN, and other lab protocols. +Public certificates come from cert-manager; internal certificates use the lab CA. +The file provider reads `traefik-dynamic` through an additional volume and flags. + +## API access + +The committed chart values enable `api.insecure` and expose TCP 8080 through the +LoadBalancer for Homarr integration. That listener has no Traefik authentication. +Its reachability depends on external network controls. Review those controls +before deploying these values outside the trusted network. + +The normal dashboard IngressRoute is a separate path; protecting that route does +not protect the direct port 8080 listener. + +## Compose alternative + +Compose mounts static and dynamic config, certificates, ACME state, and the +Docker socket. It needs the external `proxy` network. Local file-server routing +and TLS files have `.example` templates; copy only the ones needed for the host. + +Keep ACME state and private keys with backups. Changing ingress values can affect +every service at once, so inspect routes and entrypoints after an upgrade. + +See the [repository README](../README.md) for deployment selection. diff --git a/uptime-kuma/README.md b/uptime-kuma/README.md new file mode 100644 index 0000000..20e9d4a --- /dev/null +++ b/uptime-kuma/README.md @@ -0,0 +1,22 @@ +# Uptime Kuma + +Service checks, status pages, and Prometheus metrics. + +Kubernetes keeps state in `uptime-kuma-pvc` and exposes metrics through a +ServiceMonitor. The metrics credentials come from the local Secret example. +`alerts.yaml` adds Prometheus rules; a running Kuma UI alone does not establish +that Prometheus is scraping it. + +Compose stores state in `data/`. Back up that application database and verify +notification delivery after restoring it. + +## Inspect + +From the repository root: + +```sh +kubectl get pods,svc,pvc -n uptime-kuma +kubectl get events -n uptime-kuma --sort-by=.metadata.creationTimestamp +``` + +See the [repository README](../README.md) for deployment selection. diff --git a/vaultwarden/README.md b/vaultwarden/README.md new file mode 100644 index 0000000..68d03f4 --- /dev/null +++ b/vaultwarden/README.md @@ -0,0 +1,23 @@ +# Vaultwarden + +Password vault server with persistent data and public/internal ingress. + +Kubernetes uses `vaultwarden-pvc` and the public URL from a ConfigMap. +Compose has a separate data volume. Preserve the database, attachments, and keys +as part of the same backup. + +The env example only sets `DOMAIN`; there is no tracked administrator Secret +example. Configure any administrator token separately and keep it out of Git. Verify +sign-in and client synchronization after any update. Do not use a successful +container restart as the only restore check. + +## Inspect + +From the repository root: + +```sh +kubectl get pods,svc,pvc -n vaultwarden +kubectl get events -n vaultwarden --sort-by=.metadata.creationTimestamp +``` + +See the [repository README](../README.md) for deployment selection. diff --git a/vpn/README.md b/vpn/README.md new file mode 100644 index 0000000..d7be566 --- /dev/null +++ b/vpn/README.md @@ -0,0 +1,5 @@ +# VPN services + +[x-ui](xui/README.md) contains the Kubernetes deployment for 3x-ui. This directory +has no shared Compose stack. Active markers are checked at each service's `k8s/` +level, including nested paths. diff --git a/vpn/xui/README.md b/vpn/xui/README.md new file mode 100644 index 0000000..7b57d8e --- /dev/null +++ b/vpn/xui/README.md @@ -0,0 +1,18 @@ +# 3x-ui + +Kubernetes deployment for the 3x-ui management panel in namespace `xui`. +The `k8s/active` marker includes it in normal deploy selection. + +The workload, data mounts, and ports are in `k8s/xui.yaml`; panel routing and TLS +are in the ingress and certificate files. Keep panel access and proxy protocol +ports separate when changing the configuration. + +Back up the application's database and keys before upgrades. Check the actual +host and volume paths before moving the workload to another node. + +```sh +kubectl get pods,svc,pvc -n xui +kubectl get events -n xui --sort-by=.metadata.creationTimestamp +``` + +See the [repository README](../../README.md) for deployment selection.