Compare commits

...
Author SHA1 Message Date
forust 033a69bc1e feat(paperless): add alternative Compose deployment
ci / Workflows (pull_request) Successful in 9s
ci / Compose (pull_request) Successful in 15s
ci / Shell (pull_request) Successful in 19s
ci / Formatting (pull_request) Successful in 24s
ci / Python and tests (pull_request) Successful in 11s
ci / YAML (pull_request) Successful in 10s
ci / Dockerfiles (pull_request) Successful in 6s
ci / Kubernetes (pull_request) Successful in 7s
ci / image-plan (pull_request) Skipped
ci / Image (${{ matrix.name }}) (pull_request) Skipped
ci / build (pull_request) Skipped
2026-10-08 08:11:57 +02:00
forust 8a5d23b560 docs(postgres): format Paperless compatibility table
ci / Compose (pull_request) Successful in 26s
ci / Workflows (pull_request) Successful in 18s
ci / Shell (pull_request) Successful in 52s
ci / Formatting (pull_request) Successful in 27s
ci / Python and tests (pull_request) Successful in 8s
ci / YAML (pull_request) Successful in 17s
ci / Dockerfiles (pull_request) Successful in 13s
ci / Kubernetes (pull_request) Successful in 12s
ci / image-plan (pull_request) Skipped
ci / Image (${{ matrix.name }}) (pull_request) Skipped
ci / build (pull_request) Skipped
2026-10-08 00:34:43 +02:00
forust fb46fb48b5 feat(paperless): deploy Paperless-ngx on shared PostgreSQL
ci / Compose (pull_request) Canceled after 0s
ci / Workflows (pull_request) Canceled after 0s
ci / Shell (pull_request) Canceled after 0s
ci / Formatting (pull_request) Canceled after 0s
ci / Python and tests (pull_request) Canceled after 0s
ci / YAML (pull_request) Canceled after 0s
ci / Dockerfiles (pull_request) Canceled after 0s
ci / Kubernetes (pull_request) Canceled after 0s
ci / image-plan (pull_request) Canceled after 0s
ci / Image (${{ matrix.name }}) (pull_request) Canceled after 0s
ci / build (pull_request) Canceled after 0s
2026-10-08 00:32:54 +02:00
forust 4c7c53e0f2 fix(cicd): align apply timeout with stage budgets
ci / Compose (push) Successful in 11s
ci / Workflows (push) Successful in 6s
ci / Shell (push) Successful in 20s
ci / Formatting (push) Successful in 17s
ci / Python and tests (push) Successful in 7s
ci / YAML (push) Successful in 10s
ci / Dockerfiles (push) Successful in 4s
ci / image-plan (push) Successful in 13s
ci / Image (error-pages) (push) Successful in 15s
ci / Image (forust-homepage) (push) Successful in 15s
ci / Kubernetes (push) Successful in 7s
ci / Image (xdfnx-homepage) (push) Successful in 15s
ci / build (push) Successful in 19s
2026-10-07 23:28:17 +02:00
forust ba934265ac Merge pull request 'docs: record completed EDU ownership handoff' (#115) from docs/record-edu-handoff-complete into main
ci / Compose (push) Successful in 14s
ci / Workflows (push) Successful in 8s
ci / Python and tests (push) Successful in 7s
ci / YAML (push) Successful in 8s
ci / Kubernetes (push) Successful in 8s
ci / Image (forust-homepage) (push) Successful in 13s
ci / Shell (push) Successful in 18s
ci / Formatting (push) Successful in 24s
ci / Dockerfiles (push) Successful in 5s
ci / image-plan (push) Successful in 14s
ci / Image (error-pages) (push) Successful in 18s
ci / Image (xdfnx-homepage) (push) Successful in 17s
ci / build (push) Successful in 20s
Reviewed-on: #115
2026-10-07 21:10:52 +00:00
forust c7155808d9 docs: record completed EDU ownership handoff
ci / Workflows (pull_request) Successful in 13s
ci / Shell (pull_request) Successful in 31s
ci / Compose (pull_request) Successful in 26s
ci / Dockerfiles (pull_request) Successful in 11s
ci / Formatting (pull_request) Successful in 45s
ci / Python and tests (pull_request) Successful in 13s
ci / YAML (pull_request) Successful in 21s
ci / Kubernetes (pull_request) Successful in 13s
ci / image-plan (pull_request) Skipped
ci / Image (${{ matrix.name }}) (pull_request) Skipped
ci / build (pull_request) Skipped
2026-10-07 20:43:12 +02:00
forust fc4d64bdb2 chore(streaming): disable Kubernetes routing
ci / Workflows (push) Successful in 17s
ci / Compose (push) Successful in 29s
ci / Shell (push) Successful in 57s
ci / Formatting (push) Successful in 29s
ci / Python and tests (push) Successful in 14s
ci / YAML (push) Successful in 14s
ci / Dockerfiles (push) Successful in 8s
ci / Kubernetes (push) Successful in 7s
ci / image-plan (push) Successful in 15s
ci / Image (error-pages) (push) Successful in 28s
ci / Image (forust-homepage) (push) Successful in 33s
ci / Image (xdfnx-homepage) (push) Successful in 34s
ci / build (push) Successful in 37s
2026-10-07 18:23:16 +02:00
forust a6f7fc6030 chore(streaming): disable streaming stack
ci / Compose (pull_request) Successful in 14s
ci / Formatting (pull_request) Successful in 21s
ci / Python and tests (pull_request) Successful in 7s
ci / YAML (pull_request) Successful in 8s
ci / Kubernetes (pull_request) Successful in 7s
ci / Workflows (pull_request) Successful in 8s
ci / Shell (pull_request) Successful in 19s
ci / Dockerfiles (pull_request) Successful in 5s
ci / image-plan (pull_request) Skipped
ci / Image (${{ matrix.name }}) (pull_request) Skipped
ci / build (pull_request) Skipped
ci / Compose (push) Successful in 14s
ci / Formatting (push) Successful in 20s
ci / Kubernetes (push) Successful in 7s
ci / image-plan (push) Successful in 11s
ci / Workflows (push) Successful in 8s
ci / Shell (push) Successful in 14s
ci / Python and tests (push) Successful in 8s
ci / YAML (push) Successful in 10s
ci / Dockerfiles (push) Successful in 5s
ci / Image (error-pages) (push) Successful in 13s
ci / Image (forust-homepage) (push) Successful in 15s
ci / Image (xdfnx-homepage) (push) Successful in 15s
ci / build (push) Successful in 16s
2026-10-07 17:48:54 +02:00
forust 11de1d1468 Merge pull request 'fix(cicd): preserve AIO tag in rollback snapshot' (#112) from fix/nextcloud-aio-rollback-tag into main
ci / Workflows (push) Successful in 7s
ci / Image (forust-homepage) (push) Successful in 13s
ci / Image (xdfnx-homepage) (push) Successful in 18s
ci / Compose (push) Successful in 14s
ci / Shell (push) Successful in 21s
ci / Formatting (push) Successful in 24s
ci / Python and tests (push) Successful in 9s
ci / YAML (push) Successful in 8s
ci / Dockerfiles (push) Successful in 5s
ci / Kubernetes (push) Successful in 9s
ci / image-plan (push) Successful in 15s
ci / Image (error-pages) (push) Successful in 17s
ci / build (push) Successful in 18s
Reviewed-on: #112
2026-10-07 15:29:11 +00:00
forust 64962d1a63 fix(cicd): preserve AIO tag in recovery snapshot
ci / Workflows (pull_request) Successful in 6s
ci / Kubernetes (pull_request) Successful in 6s
ci / image-plan (pull_request) Skipped
ci / Image (${{ matrix.name }}) (pull_request) Skipped
ci / build (pull_request) Skipped
ci / Compose (pull_request) Successful in 15s
ci / Shell (pull_request) Successful in 19s
ci / Formatting (pull_request) Successful in 30s
ci / Python and tests (pull_request) Successful in 8s
ci / YAML (pull_request) Successful in 10s
ci / Dockerfiles (pull_request) Successful in 4s
2026-10-07 16:53:14 +02:00
forust b08a0a927d Merge pull request 'fix(cicd): preserve Nextcloud AIO tag' (#111) from fix/preserve-nextcloud-aio-tag into main
ci / Compose (push) Successful in 27s
ci / Workflows (push) Successful in 14s
ci / Shell (push) Successful in 50s
ci / Kubernetes (push) Successful in 5s
ci / image-plan (push) Successful in 12s
ci / Image (forust-homepage) (push) Successful in 14s
ci / Image (xdfnx-homepage) (push) Successful in 15s
ci / Formatting (push) Successful in 24s
ci / Python and tests (push) Successful in 9s
ci / YAML (push) Successful in 11s
ci / Dockerfiles (push) Successful in 5s
ci / Image (error-pages) (push) Successful in 16s
ci / build (push) Successful in 18s
Reviewed-on: #111
2026-10-07 14:46:02 +00:00
forust 8203ba1b0b fix(cicd): preserve Nextcloud AIO image tag
ci / Workflows (pull_request) Successful in 9s
ci / Compose (pull_request) Successful in 14s
ci / Shell (pull_request) Successful in 19s
ci / Formatting (pull_request) Successful in 33s
ci / Python and tests (pull_request) Successful in 16s
ci / Dockerfiles (pull_request) Successful in 14s
ci / Kubernetes (pull_request) Successful in 17s
ci / Image (${{ matrix.name }}) (pull_request) Skipped
ci / build (pull_request) Skipped
ci / YAML (pull_request) Successful in 19s
ci / image-plan (pull_request) Skipped
2026-10-07 14:45:10 +00:00
forust 48033b5495 Merge pull request 'fix(cicd): support Helm 4 release listing' (#110) from fix/helm4-list into main
ci / Workflows (push) Successful in 6s
ci / Shell (push) Successful in 17s
ci / Image (error-pages) (push) Successful in 13s
ci / Image (forust-homepage) (push) Successful in 13s
ci / Compose (push) Successful in 11s
ci / Formatting (push) Successful in 20s
ci / Python and tests (push) Successful in 6s
ci / YAML (push) Successful in 9s
ci / Dockerfiles (push) Successful in 4s
ci / Kubernetes (push) Successful in 7s
ci / image-plan (push) Successful in 11s
ci / Image (xdfnx-homepage) (push) Successful in 13s
ci / build (push) Successful in 16s
Reviewed-on: #110
2026-10-07 13:55:15 +00:00
forust 73d2af73e5 fix(cicd): support Helm 4 release listing
ci / build (pull_request) Skipped
ci / Workflows (pull_request) Successful in 7s
ci / Python and tests (pull_request) Successful in 5s
ci / Compose (pull_request) Successful in 11s
ci / Shell (pull_request) Successful in 17s
ci / Formatting (pull_request) Successful in 17s
ci / YAML (pull_request) Successful in 8s
ci / Dockerfiles (pull_request) Successful in 5s
ci / Kubernetes (pull_request) Successful in 7s
ci / image-plan (pull_request) Skipped
ci / Image (${{ matrix.name }}) (pull_request) Skipped
2026-10-07 15:51:57 +02:00
forust 64253e005e Merge pull request 'fix(cicd): use Gitea artifact v4 backend' (#109) from fix/gitea-v4-artifacts into main
ci / Workflows (push) Successful in 8s
ci / Shell (push) Successful in 22s
ci / YAML (push) Successful in 9s
ci / image-plan (push) Successful in 49s
ci / Compose (push) Successful in 14s
ci / Formatting (push) Successful in 19s
ci / Python and tests (push) Successful in 8s
ci / Dockerfiles (push) Successful in 5s
ci / Kubernetes (push) Successful in 8s
ci / Image (error-pages) (push) Successful in 39s
ci / Image (forust-homepage) (push) Successful in 16s
ci / Image (xdfnx-homepage) (push) Successful in 13s
ci / build (push) Successful in 17s
Reviewed-on: #109
2026-10-07 13:35:36 +00:00
forust 69accd1752 fix(cicd): use Gitea artifact v4 backend
ci / Shell (push) Skipped
ci / Formatting (push) Skipped
ci / Python and tests (push) Skipped
ci / YAML (push) Skipped
ci / Dockerfiles (push) Skipped
ci / Kubernetes (push) Skipped
ci / Compose (pull_request) Successful in 11s
ci / Kubernetes (pull_request) Successful in 7s
ci / Compose (push) Skipped
ci / Workflows (push) Skipped
ci / Workflows (pull_request) Successful in 7s
ci / Shell (pull_request) Successful in 18s
ci / Formatting (pull_request) Successful in 20s
ci / Python and tests (pull_request) Successful in 7s
ci / YAML (pull_request) Successful in 8s
ci / Dockerfiles (pull_request) Successful in 4s
ci / image-plan (pull_request) Skipped
ci / Image (${{ matrix.name }}) (pull_request) Skipped
ci / build (pull_request) Skipped
2026-10-07 15:29:49 +02:00
forust 0cf4b08a95 Merge pull request 'fix(cicd): isolate pull request runner jobs' (#108) from fix/cicd-pr-runner into main
ci / Workflows (push) Successful in 7s
ci / Shell (push) Successful in 21s
ci / Python and tests (push) Successful in 5s
ci / Compose (push) Successful in 15s
ci / Formatting (push) Successful in 17s
ci / Kubernetes (push) Successful in 6s
ci / YAML (push) Successful in 9s
ci / Dockerfiles (push) Successful in 4s
renovate-ci / validate-renovate (push) Successful in 2m21s
ci / image-plan (push) Successful in 18s
ci / Image (error-pages) (push) Successful in 13s
ci / Image (xdfnx-homepage) (push) Successful in 12s
ci / Image (forust-homepage) (push) Successful in 11s
ci / build (push) Successful in 16s
Reviewed-on: #108
2026-10-07 13:03:01 +00:00
forust 86df5d9048 docs(cicd): document user-scoped PR runner
ci / Compose (push) Skipped
ci / Workflows (push) Skipped
ci / Shell (push) Skipped
ci / Formatting (push) Skipped
ci / Python and tests (push) Skipped
ci / Dockerfiles (push) Skipped
ci / Workflows (pull_request) Successful in 14s
ci / Shell (pull_request) Successful in 33s
ci / Formatting (pull_request) Successful in 36s
ci / YAML (pull_request) Successful in 28s
ci / Kubernetes (pull_request) Successful in 11s
ci / YAML (push) Skipped
ci / Kubernetes (push) Skipped
ci / Compose (pull_request) Successful in 23s
ci / Python and tests (pull_request) Successful in 16s
ci / Dockerfiles (pull_request) Successful in 11s
ci / Image (${{ matrix.name }}) (pull_request) Skipped
ci / build (pull_request) Skipped
ci / image-plan (pull_request) Skipped
2026-10-07 14:50:32 +02:00
forust d7441bbbc2 fix(cicd): isolate pull request runner jobs
ci / Compose (push) Skipped
ci / Workflows (push) Skipped
ci / Shell (push) Skipped
ci / Formatting (push) Skipped
ci / Python and tests (push) Skipped
ci / Compose (pull_request) Successful in 37s
ci / Shell (pull_request) Successful in 18s
ci / YAML (push) Skipped
ci / Dockerfiles (push) Skipped
ci / Kubernetes (push) Skipped
ci / Workflows (pull_request) Successful in 8s
ci / Python and tests (pull_request) Successful in 11s
ci / Kubernetes (pull_request) Successful in 8s
ci / Formatting (pull_request) Successful in 18s
ci / YAML (pull_request) Successful in 11s
ci / Dockerfiles (pull_request) Successful in 7s
ci / image-plan (pull_request) Skipped
ci / Image (${{ matrix.name }}) (pull_request) Skipped
ci / build (pull_request) Skipped
2026-10-07 14:39:57 +02:00
forust 2be089e048 Merge pull request 'Collect Headscale, NetBird, Gitea and Immich metrics' (#100) from feat/service-metrics into main
ci / Compose (push) Successful in 11s
ci / Workflows (push) Successful in 7s
ci / YAML (push) Successful in 9s
ci / Dockerfiles (push) Successful in 5s
ci / Shell (push) Successful in 16s
ci / Formatting (push) Successful in 17s
ci / Python and tests (push) Successful in 7s
ci / Kubernetes (push) Successful in 6s
ci / image-plan (push) Successful in 15s
ci / Image (error-pages) (push) Successful in 12s
ci / Image (forust-homepage) (push) Successful in 12s
ci / Image (xdfnx-homepage) (push) Successful in 11s
ci / build (push) Successful in 13s
Reviewed-on: #100
2026-10-07 11:46:16 +00:00
forust 83b2e68371 feat(metrics): collect Headscale, NetBird, Gitea and Immich metrics 2026-10-07 11:46:16 +00:00
forust 597f64cbb0 Merge pull request 'Show CI checks and deployment results' (#99) from codex/ci-visible-checks into main
ci / Workflows (push) Successful in 7s
ci / Formatting (push) Successful in 17s
ci / Python and tests (push) Successful in 7s
ci / YAML (push) Successful in 9s
ci / Compose (push) Successful in 11s
ci / Shell (push) Successful in 16s
ci / Kubernetes (push) Successful in 7s
ci / Dockerfiles (push) Successful in 5s
ci / Image (forust-homepage) (push) Successful in 12s
ci / Image (xdfnx-homepage) (push) Successful in 11s
renovate-ci / validate-renovate (push) Successful in 11s
ci / image-plan (push) Successful in 16s
ci / Image (error-pages) (push) Successful in 41s
ci / build (push) Successful in 14s
Reviewed-on: #99
2026-10-07 11:46:05 +00:00
forust f767f3ce1a docs: update EDU handoff status
ci / Compose (push) Skipped
ci / Workflows (push) Skipped
ci / Shell (push) Skipped
ci / Formatting (push) Skipped
ci / Kubernetes (push) Skipped
ci / Compose (pull_request) Successful in 11s
ci / Shell (pull_request) Successful in 15s
ci / Python and tests (push) Skipped
ci / YAML (push) Skipped
ci / Dockerfiles (push) Skipped
ci / Workflows (pull_request) Successful in 6s
ci / Formatting (pull_request) Successful in 16s
ci / Python and tests (pull_request) Successful in 6s
ci / YAML (pull_request) Successful in 7s
ci / Dockerfiles (pull_request) Successful in 5s
ci / build (pull_request) Skipped
ci / Kubernetes (pull_request) Successful in 6s
ci / image-plan (pull_request) Skipped
ci / Image (${{ matrix.name }}) (pull_request) Skipped
renovate-ci / validate-renovate (pull_request) Successful in 9s
2026-10-07 10:25:17 +02:00
forust 95e4d8f875 ci: add per-image jobs to homelab CI
ci / Compose (push) Skipped
ci / Workflows (push) Skipped
ci / Shell (push) Skipped
ci / Formatting (push) Skipped
ci / Python and tests (push) Skipped
ci / YAML (push) Skipped
ci / Dockerfiles (push) Skipped
ci / Kubernetes (push) Skipped
ci / Compose (pull_request) Successful in 11s
ci / Workflows (pull_request) Successful in 7s
ci / Shell (pull_request) Successful in 15s
ci / Formatting (pull_request) Successful in 15s
ci / Python and tests (pull_request) Successful in 5s
ci / YAML (pull_request) Successful in 7s
ci / Dockerfiles (pull_request) Successful in 5s
ci / Kubernetes (pull_request) Successful in 6s
ci / image-plan (pull_request) Skipped
ci / Image (${{ matrix.name }}) (pull_request) Skipped
ci / build (pull_request) Skipped
renovate-ci / validate-renovate (pull_request) Successful in 11s
2026-10-07 10:19:12 +02:00
forust 1d93588e06 refactor: remove EDU ownership from homelab 2026-10-07 10:19:12 +02:00
forust fb400eea6e Use plain text for the deploy request details
ci / Compose (push) Skipped
ci / Shell (push) Skipped
ci / Formatting (push) Skipped
ci / YAML (push) Skipped
ci / Dockerfiles (push) Skipped
ci / Kubernetes (push) Skipped
ci / YAML (pull_request) Successful in 11s
ci / build (pull_request) Skipped
ci / Workflows (push) Skipped
ci / Python and tests (push) Skipped
ci / Compose (pull_request) Successful in 13s
ci / Workflows (pull_request) Successful in 9s
ci / Shell (pull_request) Successful in 21s
ci / Formatting (pull_request) Successful in 22s
ci / Python and tests (pull_request) Successful in 6s
ci / Dockerfiles (pull_request) Successful in 5s
ci / Kubernetes (pull_request) Successful in 7s
2026-10-06 23:44:24 +02:00
forust 0c76426c17 Show failure details in CI and deploy summaries
ci / Formatting (push) Skipped
ci / Python and tests (push) Skipped
ci / Kubernetes (push) Skipped
ci / Compose (push) Skipped
ci / Shell (push) Skipped
ci / YAML (push) Skipped
ci / Dockerfiles (push) Skipped
ci / Workflows (push) Skipped
ci / Compose (pull_request) Successful in 11s
ci / Workflows (pull_request) Failing after 8s
ci / Formatting (pull_request) Canceled after 0s
ci / Python and tests (pull_request) Canceled after 0s
ci / YAML (pull_request) Canceled after 0s
ci / Dockerfiles (pull_request) Canceled after 0s
ci / Kubernetes (pull_request) Canceled after 0s
ci / build (pull_request) Canceled after 0s
ci / Shell (pull_request) Canceled after 7s
2026-10-06 23:44:00 +02:00
forust d74822cd27 Add CI and deploy summaries
ci / Compose (push) Skipped
ci / Workflows (push) Skipped
ci / Shell (push) Skipped
ci / Formatting (push) Skipped
ci / Python and tests (push) Skipped
ci / YAML (push) Skipped
ci / Kubernetes (push) Skipped
ci / Dockerfiles (push) Skipped
ci / Compose (pull_request) Successful in 14s
ci / Workflows (pull_request) Successful in 7s
ci / Shell (pull_request) Successful in 23s
ci / Formatting (pull_request) Successful in 26s
ci / Python and tests (pull_request) Successful in 8s
ci / YAML (pull_request) Successful in 14s
ci / Dockerfiles (pull_request) Successful in 7s
ci / Kubernetes (pull_request) Successful in 10s
ci / build (pull_request) Skipped
2026-10-06 23:28:56 +02:00
forust b677d553b4 refactor(ci): split checks into visible jobs
ci / YAML (pull_request) Successful in 11s
ci / Dockerfiles (pull_request) Successful in 6s
ci / Kubernetes (pull_request) Successful in 6s
ci / build (pull_request) Skipped
renovate-ci / validate-renovate (push) Skipped
renovate-ci / validate-renovate (pull_request) Skipped
ci / Compose (pull_request) Successful in 10s
ci / Workflows (pull_request) Successful in 5s
ci / Shell (pull_request) Successful in 22s
ci / Formatting (pull_request) Successful in 21s
ci / Python and tests (pull_request) Successful in 8s
2026-10-06 23:22:15 +02:00
forust 898463b759 Merge pull request 'refactor(ci): native VPS runner and durable incremental deploys' (#98) from codex/cicd-runner-deploy into main
ci / checks (push) Successful in 53s
renovate-ci / validate-renovate (push) Successful in 12s
ci / build (push) Successful in 1m44s
Reviewed-on: #98
2026-10-06 21:19:55 +00:00
forust 9a76529be8 refactor(ci): use native runner and durable incremental deploys
renovate-ci / validate-renovate (push) Skipped
ci / checks (pull_request) Successful in 51s
ci / build (pull_request) Skipped
renovate-ci / validate-renovate (pull_request) Successful in 21s
2026-10-06 23:08:47 +02:00
forust 5f9354b9a8 fix(edu): deploy reviewed application release with Redis authentication
renovate-ci / validate-renovate (push) Skipped
ci / lint-compose (push) Successful in 14s
ci / lint-actionlint (push) Successful in 6s
ci / lint-shellcheck (push) Successful in 13s
ci / lint-prettier (push) Successful in 19s
ci / lint-ruff (push) Successful in 9s
ci / lint-yaml (push) Successful in 10s
ci / lint-dockerfiles (push) Successful in 6s
ci / validate (push) Successful in 10s
ci / build (push) Successful in 30s
2026-10-06 22:56:02 +02:00
forust fcd16f6128 Merge pull request 'chore(deps): update renovate/renovate docker tag to v44.140.0' (#96) from renovate/renovate-self-update into main
renovate-ci / validate-renovate (push) Successful in 1m25s
ci / lint-compose (push) Successful in 12s
ci / lint-actionlint (push) Successful in 7s
ci / lint-shellcheck (push) Successful in 16s
ci / lint-prettier (push) Successful in 23s
ci / lint-ruff (push) Successful in 9s
ci / lint-yaml (push) Successful in 11s
ci / lint-dockerfiles (push) Successful in 7s
ci / validate (push) Successful in 6s
ci / build (push) Successful in 25s
Reviewed-on: #96
2026-10-06 17:51:40 +00:00
renovate-bot Bot 2ac2a94bb4 chore(deps): update renovate/renovate docker tag to v44.140.0 2026-10-06 17:51:40 +00:00
forust 3c7e358dd5 Merge pull request 'chore(deps): update lscr.io/linuxserver/qbittorrent docker tag to v20' (#97) from renovate/lscr.io-linuxserver-qbittorrent-20.x into main
renovate-ci / validate-renovate (push) Skipped
ci / lint-compose (push) Successful in 16s
ci / lint-actionlint (push) Successful in 7s
ci / lint-shellcheck (push) Successful in 15s
ci / lint-prettier (push) Successful in 20s
ci / lint-ruff (push) Successful in 9s
ci / lint-yaml (push) Successful in 14s
ci / lint-dockerfiles (push) Successful in 7s
ci / validate (push) Successful in 8s
ci / build (push) Successful in 24s
Reviewed-on: #97
2026-10-06 17:51:25 +00:00
renovate-bot Bot 9be8fb6c69 chore(deps): update lscr.io/linuxserver/qbittorrent docker tag to v20 2026-10-06 17:51:25 +00:00
forust 7fdeacffb8 feat(monitoring): stand down Prometheus server during VM trial
renovate-ci / validate-renovate (push) Skipped
ci / lint-compose (push) Successful in 13s
ci / lint-actionlint (push) Successful in 5s
ci / lint-shellcheck (push) Successful in 10s
ci / lint-prettier (push) Successful in 14s
ci / lint-ruff (push) Successful in 8s
ci / lint-yaml (push) Successful in 10s
ci / lint-dockerfiles (push) Successful in 6s
ci / validate (push) Successful in 7s
ci / build (push) Successful in 19s
vmagent scrapes and remote-writes to VictoriaMetrics, so the Prometheus server scales to 0. Encoded as prometheusSpec.replicas in values instead of a kubectl patch, so helm keeps owning spec.replicas and Helm 4 server-side apply stops conflicting with the kubectl-patch field manager.
2026-10-06 19:29:30 +02:00
forust cef499de73 fix(deploy): skip VMAgent in secrets check before CRD install
renovate-ci / validate-renovate (push) Skipped
ci / lint-compose (push) Successful in 12s
ci / lint-actionlint (push) Successful in 6s
ci / lint-shellcheck (push) Successful in 15s
ci / lint-ruff (push) Successful in 6s
ci / lint-dockerfiles (push) Successful in 6s
ci / lint-prettier (push) Successful in 21s
ci / lint-yaml (push) Successful in 10s
ci / validate (push) Successful in 8s
ci / build (push) Successful in 21s
check_referenced_secrets ran kubectl create on vmagent.yaml even when the VMAgent CRD is not installed yet, failing validate with 'no matches for kind VMAgent'. Apply the same skip_uninstalled_vmagent_crd guard used by both dry-run loops.
2026-10-06 19:19:47 +02:00
forust 1b70a55300 fix(ci): handle malformed push before SHA
renovate-ci / validate-renovate (push) Skipped
ci / lint-compose (push) Successful in 10s
ci / lint-actionlint (push) Successful in 7s
ci / lint-shellcheck (push) Successful in 15s
ci / lint-prettier (push) Failing after 22s
ci / lint-ruff (push) Failing after 2s
ci / lint-yaml (push) Failing after 2s
ci / lint-dockerfiles (push) Failing after 3s
ci / validate (push) Failing after 2s
ci / build (push) Skipped
2026-10-06 18:21:19 +02:00
forust 3f2b4e9acf fix(deploy): skip VMAgent preflight before CRD install
renovate-ci / validate-renovate (push) Skipped
ci / lint-compose (push) Successful in 11s
ci / lint-actionlint (push) Successful in 5s
ci / lint-shellcheck (push) Successful in 10s
ci / lint-prettier (push) Successful in 15s
ci / lint-ruff (push) Successful in 8s
ci / lint-yaml (push) Successful in 10s
ci / lint-dockerfiles (push) Successful in 7s
ci / validate (push) Successful in 10s
ci / build (push) Successful in 25s
2026-10-06 18:18:25 +02:00
forust 6057734a4f fix(renovate): sync generated configmap
renovate-ci / validate-renovate (push) Successful in 9s
ci / lint-compose (push) Successful in 9s
ci / lint-actionlint (push) Successful in 4s
ci / lint-shellcheck (push) Successful in 10s
ci / lint-prettier (push) Successful in 19s
ci / lint-ruff (push) Successful in 8s
ci / lint-yaml (push) Successful in 11s
ci / lint-dockerfiles (push) Successful in 5s
ci / validate (push) Successful in 7s
ci / build (push) Successful in 18s
2026-10-06 18:08:17 +02:00
forust 8eedf8b74b Merge pull request 'feat(monitoring): add VictoriaMetrics trial stack' (#95) from feat/victoria-metrics-migration into main
ci / lint-compose (push) Successful in 11s
ci / lint-actionlint (push) Successful in 2m36s
ci / lint-shellcheck (push) Successful in 15s
ci / lint-prettier (push) Successful in 19s
ci / lint-ruff (push) Successful in 7s
ci / lint-yaml (push) Successful in 10s
ci / lint-dockerfiles (push) Successful in 5s
ci / validate (push) Successful in 7s
renovate-ci / validate-renovate (push) Failing after 9s
ci / build (push) Successful in 19s
Reviewed-on: #95
2026-10-06 16:05:46 +00:00
forust 8c0e36a5c0 feat(monitoring): replace scrape dump with vmagent
ci / lint-prettier (push) Skipped
ci / lint-ruff (push) Skipped
ci / lint-yaml (push) Skipped
ci / lint-dockerfiles (push) Skipped
ci / validate (push) Skipped
renovate-ci / validate-renovate (push) Skipped
ci / lint-compose (pull_request) Successful in 11s
ci / lint-actionlint (pull_request) Successful in 6s
ci / lint-shellcheck (pull_request) Successful in 16s
ci / lint-dockerfiles (pull_request) Successful in 6s
ci / lint-prettier (pull_request) Successful in 16s
ci / lint-ruff (pull_request) Successful in 7s
ci / lint-yaml (pull_request) Successful in 10s
ci / validate (pull_request) Successful in 7s
ci / build (pull_request) Skipped
renovate-ci / validate-renovate (pull_request) Failing after 1m40s
2026-10-06 18:03:59 +02:00
forust 78cd15f12c chore(monitoring): remove vmctl backfill job
renovate-ci / validate-renovate (pull_request) Skipped
ci / lint-compose (pull_request) Successful in 11s
ci / lint-yaml (pull_request) Successful in 11s
ci / lint-prettier (push) Skipped
ci / lint-ruff (push) Skipped
ci / lint-yaml (push) Skipped
ci / lint-dockerfiles (push) Skipped
ci / validate (push) Skipped
renovate-ci / validate-renovate (push) Skipped
ci / lint-actionlint (pull_request) Successful in 5s
ci / lint-shellcheck (pull_request) Successful in 16s
ci / lint-prettier (pull_request) Successful in 15s
ci / lint-ruff (pull_request) Successful in 6s
ci / lint-dockerfiles (pull_request) Successful in 7s
ci / validate (pull_request) Successful in 7s
ci / build (pull_request) Skipped
One-shot Prometheus history backfill is complete; drop the Job and clean up related comments.
2026-10-06 17:53:24 +02:00
forust 18c633c242 feat(monitoring): add VictoriaMetrics trial stack
ci / lint-prettier (push) Skipped
ci / lint-ruff (push) Skipped
ci / lint-yaml (push) Skipped
ci / lint-dockerfiles (push) Skipped
ci / validate (push) Skipped
renovate-ci / validate-renovate (push) Skipped
renovate-ci / validate-renovate (pull_request) Skipped
ci / lint-compose (pull_request) Successful in 12s
ci / lint-actionlint (pull_request) Successful in 6s
ci / lint-shellcheck (pull_request) Successful in 13s
ci / lint-prettier (pull_request) Successful in 21s
ci / lint-ruff (pull_request) Successful in 7s
ci / lint-yaml (pull_request) Successful in 12s
ci / lint-dockerfiles (pull_request) Successful in 7s
ci / validate (pull_request) Successful in 7s
ci / build (pull_request) Skipped
2026-10-06 17:45:44 +02:00
forust 4510394531 Merge pull request 'chore(deps): update renovate/renovate docker tag to v44.139.0' (#81) from renovate/renovate-self-update into main
ci / lint-compose (push) Successful in 11s
ci / lint-actionlint (push) Successful in 6s
ci / lint-shellcheck (push) Successful in 15s
ci / lint-prettier (push) Successful in 19s
ci / lint-ruff (push) Successful in 7s
ci / lint-yaml (push) Successful in 11s
ci / lint-dockerfiles (push) Successful in 6s
ci / validate (push) Successful in 9s
renovate-ci / validate-renovate (push) Successful in 13s
ci / build (push) Successful in 20s
Reviewed-on: #81
2026-10-06 15:33:45 +00:00
renovate-bot Bot 2545312db1 chore(deps): update renovate/renovate docker tag to v44.139.0 2026-10-06 15:33:45 +00:00
forust 8ce0b809c0 Merge pull request 'chore(deps): update all minor updates' (#77) from renovate/all-minor into main
renovate-ci / validate-renovate (push) Skipped
ci / lint-compose (push) Canceled after 0s
ci / lint-actionlint (push) Canceled after 0s
ci / lint-shellcheck (push) Canceled after 0s
ci / lint-prettier (push) Canceled after 0s
ci / lint-ruff (push) Canceled after 0s
ci / lint-yaml (push) Canceled after 0s
ci / lint-dockerfiles (push) Canceled after 0s
ci / validate (push) Canceled after 0s
ci / build (push) Canceled after 0s
Reviewed-on: #77
2026-10-06 15:33:28 +00:00
renovate-bot Bot 506c04e15c chore(deps): update all minor updates 2026-10-06 15:33:28 +00:00
forust bdcbb3d5af Merge pull request 'chore(deps): update all patch updates' (#82) from renovate/all-patch into main
renovate-ci / validate-renovate (push) Skipped
ci / lint-compose (push) Canceled after 0s
ci / lint-actionlint (push) Canceled after 0s
ci / lint-shellcheck (push) Canceled after 0s
ci / lint-prettier (push) Canceled after 0s
ci / lint-ruff (push) Canceled after 0s
ci / lint-yaml (push) Canceled after 0s
ci / lint-dockerfiles (push) Canceled after 0s
ci / validate (push) Canceled after 0s
ci / build (push) Canceled after 0s
Reviewed-on: #82
2026-10-06 15:33:09 +00:00
renovate-bot Bot faac6febb8 chore(deps): update all patch updates 2026-10-06 15:33:09 +00:00
forust 695da1308c Merge pull request 'fix(renovate): restore custom extraction and update groups' (#93) from fix/renovate-extraction-groups into main
ci / lint-compose (push) Successful in 10s
ci / lint-actionlint (push) Successful in 6s
ci / lint-shellcheck (push) Successful in 17s
ci / lint-prettier (push) Successful in 18s
ci / lint-ruff (push) Successful in 8s
ci / lint-yaml (push) Successful in 12s
ci / lint-dockerfiles (push) Successful in 8s
ci / validate (push) Successful in 8s
renovate-ci / validate-renovate (push) Successful in 11s
ci / build (push) Successful in 37s
Reviewed-on: #93
2026-10-06 15:31:25 +00:00
forust 552b22cb66 fix(renovate): include self-update Compose filename 2026-10-06 15:31:25 +00:00
forust 3756e60c95 fix(renovate): restore custom extraction and update groups 2026-10-06 15:31:25 +00:00
forust d0d217ddf4 Merge pull request 'fix(deploy): skip verify and smoke when deployment is disabled' (#94) from fix/deploy-skip-when-disabled into main
renovate-ci / validate-renovate (push) Skipped
ci / lint-compose (push) Canceled after 0s
ci / lint-actionlint (push) Canceled after 0s
ci / lint-shellcheck (push) Canceled after 0s
ci / lint-prettier (push) Canceled after 0s
ci / lint-ruff (push) Canceled after 0s
ci / lint-yaml (push) Canceled after 0s
ci / lint-dockerfiles (push) Canceled after 0s
ci / validate (push) Canceled after 0s
ci / build (push) Canceled after 0s
Reviewed-on: #94
2026-10-06 15:24:17 +00:00
forust 24f84bab2f fix(deploy): skip verification when deployment is disabled
ci / lint-yaml (push) Skipped
ci / lint-dockerfiles (push) Skipped
ci / validate (push) Skipped
renovate-ci / validate-renovate (push) Skipped
ci / lint-prettier (push) Skipped
ci / lint-ruff (push) Skipped
renovate-ci / validate-renovate (pull_request) Skipped
ci / lint-compose (pull_request) Canceled after 0s
ci / lint-actionlint (pull_request) Canceled after 0s
ci / lint-shellcheck (pull_request) Canceled after 0s
ci / lint-prettier (pull_request) Canceled after 0s
ci / lint-ruff (pull_request) Canceled after 0s
ci / lint-yaml (pull_request) Canceled after 0s
ci / lint-dockerfiles (pull_request) Canceled after 0s
ci / validate (pull_request) Canceled after 0s
ci / build (pull_request) Canceled after 0s
2026-10-06 15:24:06 +00:00
forust 0b8a3c16a7 Merge pull request 'fix(glance): mount CSS from the correct ConfigMap' (#87) from fix/glance-assets into main
renovate-ci / validate-renovate (push) Skipped
ci / lint-compose (push) Canceled after 0s
ci / lint-actionlint (push) Canceled after 0s
ci / lint-shellcheck (push) Canceled after 0s
ci / lint-prettier (push) Canceled after 0s
ci / lint-ruff (push) Canceled after 0s
ci / lint-yaml (push) Canceled after 0s
ci / lint-dockerfiles (push) Canceled after 0s
ci / validate (push) Canceled after 0s
ci / build (push) Canceled after 0s
Reviewed-on: #87
2026-10-06 15:17:47 +00:00
forust e9a68aae77 fix(glance): mount CSS from the assets ConfigMap 2026-10-06 15:17:47 +00:00
forust 5359df5ed6 Merge pull request 'fix(postgres): include the required NetBox database password' (#88) from fix/postgres-env-example into main
renovate-ci / validate-renovate (push) Skipped
ci / lint-compose (push) Successful in 14s
ci / lint-actionlint (push) Successful in 8s
ci / lint-shellcheck (push) Successful in 17s
ci / lint-prettier (push) Successful in 24s
ci / lint-ruff (push) Successful in 9s
ci / lint-yaml (push) Successful in 11s
ci / lint-dockerfiles (push) Successful in 5s
ci / validate (push) Successful in 11s
ci / build (push) Canceled after 0s
Reviewed-on: #88
2026-10-06 15:17:19 +00:00
forust 8aea0f0d13 fix(postgres): include required NetBox password in Compose env example 2026-10-06 15:17:19 +00:00
forust 65d4f2d482 Merge pull request 'fix(traefik): correct AdGuard and SearXNG Compose rules' (#90) from fix/compose-router-rules into main
renovate-ci / validate-renovate (push) Skipped
ci / lint-compose (push) Successful in 12s
ci / lint-actionlint (push) Successful in 6s
ci / lint-shellcheck (push) Successful in 13s
ci / lint-prettier (push) Successful in 16s
ci / lint-ruff (push) Successful in 8s
ci / lint-yaml (push) Successful in 11s
ci / lint-dockerfiles (push) Successful in 6s
ci / validate (push) Successful in 10s
ci / build (push) Successful in 23s
Reviewed-on: #90
2026-10-06 15:10:49 +00:00
forust 2fe9b7632f fix(traefik): correct AdGuard and SearXNG Compose router expressions 2026-10-06 15:10:49 +00:00
forust e436d89eef Merge pull request 'fix(netbird): restore Compose setup and runtime renderer' (#86) from fix/netbird-compose-runtime into main
renovate-ci / validate-renovate (push) Skipped
ci / lint-compose (push) Successful in 11s
ci / lint-actionlint (push) Successful in 4s
ci / lint-shellcheck (push) Successful in 15s
ci / lint-prettier (push) Successful in 21s
ci / lint-ruff (push) Successful in 7s
ci / lint-yaml (push) Successful in 11s
ci / lint-dockerfiles (push) Successful in 6s
ci / validate (push) Successful in 7s
ci / build (push) Successful in 20s
Reviewed-on: #86
2026-10-06 15:10:12 +00:00
forust 8729cb5062 fix(netbird): restore Compose setup and server entrypoint
ci / validate (push) Skipped
ci / lint-prettier (push) Skipped
ci / lint-ruff (push) Skipped
ci / lint-yaml (push) Skipped
ci / lint-dockerfiles (push) Skipped
renovate-ci / validate-renovate (push) Skipped
renovate-ci / validate-renovate (pull_request) Skipped
ci / lint-compose (pull_request) Successful in 12s
ci / lint-actionlint (pull_request) Successful in 8s
ci / lint-shellcheck (pull_request) Successful in 21s
ci / lint-prettier (pull_request) Successful in 17s
ci / lint-ruff (pull_request) Successful in 6s
ci / lint-yaml (pull_request) Successful in 9s
ci / lint-dockerfiles (pull_request) Successful in 5s
ci / validate (pull_request) Successful in 6s
ci / build (pull_request) Skipped
2026-10-06 15:08:46 +00:00
forust f1e4a1088d Merge pull request 'fix(ci): deduplicate PR checks and filter deploy triggers' (#92) from fix/ci-trigger-dedup into main
renovate-ci / validate-renovate (push) Successful in 9s
ci / lint-compose (push) Successful in 10s
ci / lint-actionlint (push) Successful in 6s
ci / lint-shellcheck (push) Successful in 12s
ci / lint-prettier (push) Successful in 19s
ci / lint-ruff (push) Successful in 8s
ci / lint-yaml (push) Successful in 12s
ci / lint-dockerfiles (push) Successful in 7s
ci / validate (push) Successful in 7s
ci / build (push) Successful in 19s
Reviewed-on: #92
2026-10-06 15:06:47 +00:00
forust b357ef95d8 fix(deploy): only create automatic runs for main CI
ci / lint-ruff (push) Skipped
ci / lint-yaml (push) Skipped
ci / lint-dockerfiles (push) Skipped
ci / validate (push) Skipped
ci / lint-prettier (push) Skipped
renovate-ci / validate-renovate (push) Skipped
ci / lint-compose (pull_request) Successful in 13s
ci / lint-actionlint (pull_request) Successful in 6s
ci / lint-shellcheck (pull_request) Successful in 11s
ci / lint-prettier (pull_request) Successful in 18s
ci / lint-ruff (pull_request) Successful in 8s
ci / lint-yaml (pull_request) Successful in 13s
ci / lint-dockerfiles (pull_request) Successful in 5s
ci / validate (pull_request) Successful in 9s
ci / build (pull_request) Skipped
renovate-ci / validate-renovate (pull_request) Successful in 12s
2026-10-06 17:00:05 +02:00
forust 67422663b7 fix(ci): avoid duplicate branch and PR runs
ci / validate (push) Skipped
ci / lint-prettier (push) Skipped
ci / lint-ruff (push) Skipped
ci / lint-yaml (push) Skipped
ci / lint-dockerfiles (push) Skipped
renovate-ci / validate-renovate (push) Skipped
ci / lint-compose (pull_request) Canceled after 0s
ci / lint-actionlint (pull_request) Canceled after 0s
ci / lint-shellcheck (pull_request) Canceled after 0s
ci / lint-prettier (pull_request) Canceled after 0s
ci / lint-ruff (pull_request) Canceled after 0s
ci / lint-yaml (pull_request) Canceled after 0s
ci / lint-dockerfiles (pull_request) Canceled after 0s
ci / validate (pull_request) Canceled after 0s
ci / build (pull_request) Canceled after 0s
renovate-ci / validate-renovate (pull_request) Successful in 10s
2026-10-06 16:59:26 +02:00
forust 88705fec88 Merge pull request 'fix(deploy): reject unsafe per-file pruning' (#85) from fix/deploy-prune-guard into main
renovate-ci / validate-renovate (push) Successful in 9s
ci / lint-compose (push) Successful in 10s
ci / lint-actionlint (push) Successful in 4s
ci / lint-shellcheck (push) Successful in 11s
ci / lint-prettier (push) Successful in 15s
ci / lint-ruff (push) Successful in 5s
ci / lint-yaml (push) Successful in 10s
ci / lint-dockerfiles (push) Successful in 5s
ci / validate (push) Successful in 8s
ci / build (push) Successful in 23s
Reviewed-on: #85
2026-10-06 14:57:22 +00:00
forust c098807aa4 fix(deploy): reject destructive per-file pruning before apply
renovate-ci / validate-renovate (push) Skipped
ci / lint-compose (push) Successful in 14s
ci / lint-actionlint (push) Successful in 8s
ci / lint-shellcheck (push) Successful in 13s
ci / lint-prettier (push) Successful in 19s
ci / lint-ruff (push) Successful in 8s
ci / lint-yaml (push) Successful in 12s
ci / lint-dockerfiles (push) Successful in 8s
ci / validate (push) Successful in 10s
ci / build (push) Skipped
ci / lint-compose (pull_request) Successful in 12s
ci / lint-actionlint (pull_request) Successful in 5s
ci / lint-shellcheck (pull_request) Successful in 11s
ci / lint-prettier (pull_request) Successful in 16s
ci / lint-ruff (pull_request) Successful in 8s
ci / lint-yaml (pull_request) Successful in 11s
ci / lint-dockerfiles (pull_request) Successful in 7s
ci / validate (pull_request) Successful in 7s
ci / build (pull_request) Skipped
renovate-ci / validate-renovate (pull_request) Successful in 10s
2026-10-06 14:57:11 +00:00
forust e1eee2d3c7 Merge pull request 'feat(reloader): enable deploy and integrate configuration reloads' (#91) from fix/reloader-integration into main
ci / lint-compose (push) Canceled after 0s
ci / lint-actionlint (push) Canceled after 0s
ci / lint-shellcheck (push) Canceled after 0s
ci / lint-prettier (push) Canceled after 0s
ci / lint-ruff (push) Canceled after 0s
ci / lint-yaml (push) Canceled after 0s
ci / lint-dockerfiles (push) Canceled after 0s
ci / validate (push) Canceled after 0s
ci / build (push) Canceled after 0s
renovate-ci / validate-renovate (push) Successful in 9s
Reviewed-on: #91
2026-10-06 14:55:16 +00:00
108 changed files with 4591 additions and 3714 deletions

No files matched your search

+50
View File
@@ -0,0 +1,50 @@
# EDU ownership handoff
## Status
The EDU ownership handoff is complete. The homelab repository no longer owns
EDU workloads, images, routes, alerts, or deployment selection. The EDU
repository is the only deployment owner: [forust/edu-master](https://git.forust.xyz/forust/edu-master).
Homelab PRs #99 and #105 are merged. PR #105 removed the EDU subtree and its
build, deploy, rollback, verification, route-probe, and registry references.
It also added the serial image build matrix for the homelab services. This
handoff record is the only remaining EDU-specific file in homelab Git.
The dedicated workstation checkout is `/srv/edu-master`, at release
`4f2b2a0e37dc11ac2c75441a15076c178e219d37`. It contains `k8s/active`; root
`active` is absent. The old untracked `/srv/homelab/edu_master` checkout was
moved outside the homelab repository to
`/srv/edu-master-legacy-archive-20261007/edu_master`. Its private files remain
mode `0600` inside an archive directory with mode `0700`. The homelab deploy
checkout has no EDU marker or tracked EDU application/deployment files.
`AUTODEPLOY=false` remains in place for homelab deployment.
## Release evidence
EDU PR #4 merged after its review and CI checks. Main-push CI run 1652 passed
all validation and both image builds. Deploy run 1653 passed for the exact main
SHA above.
The workstation rollout completed for both Deployments. The deployment
verified `/health` and `/live` with HTTP 200, Redis AUTH, session TTL of 1058
seconds, a delivery backlog of zero, and all nine EDU vmalert rules with
matching expressions and healthy evaluation.
The images now run by digest:
- Session keeper: `sha256:998dea51aa3015fd9cabefb0f53b030157a650c3bef72e02fe84f17d5762613d`
- Webinar checker: `sha256:92f3c1fa2bb7f9b4680a9fc76a5b33dfbea8ef3dd9c6490ebc45876fd4c54461`
Redis StatefulSet was unchanged. PVC `redis-data-pvc` remains bound to PV
`pvc-a4f2a79a-363a-4c12-ae91-92cdfc2a0d2e` with capacity 1 GiB. The existing
runtime Secret and Fernet key were preserved during the handoff. Notification
delivery was verified before closeout, as confirmed by the operator. The
deployment did not record downtime.
The release rollback snapshot is
`/home/forust/.local/state/edu-master-deploy/20261007T180541Z-4f2b2a0e37dc11ac2c75441a15076c178e219d37`.
The handoff data snapshot remains at
`/home/forust/.local/state/edu-master-deploy/handoff-20261007T080838Z`.
Both snapshots are outside Git. Do not restore old Redis data unless recovery
requires it. Never delete or recreate the Redis PVC.
+1
View File
@@ -7,4 +7,5 @@ self-hosted-runner:
labels: labels:
- arch - arch
- homelab - homelab
- homelab-pr
- prod - prod
+3
View File
@@ -0,0 +1,3 @@
{
"postgres": ["authentik", "gitea", "immich", "n8n", "netbox", "netronome"]
}
+159
View File
@@ -0,0 +1,159 @@
# Homelab CI/CD
The native Gitea runners run on **vps**; production runs on **workstation**.
Main-branch checks and image builds use `homelab:host`. Pull request and
non-main checks use `homelab-pr:host` under a separate account without Docker
access. The `homelab-pr` runner is registered at User scope for `forust`, so
any repository under that account can schedule jobs that request this label.
Each runner accepts one job at a time; the build waits for every check to pass.
CI and deploy runs also show a summary with
the release SHA, image build or reuse results, deploy mode, selected services,
and image digests. Failed runs keep a summary of completed image builds, stage
results, apply results, and recorded Kubernetes recovery. The final deploy
summary is in the smoke job; earlier jobs show the state observed at that time.
Apply success is separate from health and recovery. Update the installed
workstation controller with `setup-workstation.sh` when no deploy is running.
No job images or Kubernetes credentials are needed on the VPS. Builds use one
pinned BuildKit helper container. CI and deploy are separate workflows.
## Runner installation
Install Docker Engine with Compose and Buildx, Git, Python 3.11+, Bash, curl,
GNU tar/xz, flock and systemd using the host's package manager. Keep the existing
Gitea runner 3.0.2 binary at `/usr/local/bin/gitea-runner`.
From this checkout on the VPS:
```sh
sudo bash .gitea/runner/setup-runner.sh
```
The installer reuses `/var/lib/gitea-runner/.runner` and the existing service.
For a new host, install the same runner binary and register as `gitea-runner`
using the registration token interactively, label `homelab:host`, and working
directory `/var/lib/gitea-runner`; then rerun the installer. Tokens never belong
in this repository or command-line examples.
Pinned tools live in the runner user's `~/.cache/homelab-ci`; CI repairs version
drift there. Installations are locked. Buildx uses only the `homelab-ci` builder,
pushes directly to the registry, and caps retained local cache at 1 GiB with a
2 GiB free-space target. This is not a hard limit on peak build disk usage.
Nothing runs `docker system prune`, removes unrelated images, or deletes volumes.
### Pull request runner
Install the unprivileged host runner on the VPS:
```sh
sudo bash .gitea/runner/setup-pr-runner.sh
```
Get a registration token from the user Actions runner settings. Run the
installer in a terminal. It asks for the token without echoing it, registers the
runner as `homelab-pr` with label `homelab-pr:host`, then enables the service.
The work directory is `/var/lib/gitea-pr-runner`. Confirm that Gitea lists the
runner as User scope before merging the workflow change. An unmatched label can
fall back to the default job image.
Renovate PR validation uses `pull_request_target`, which reads the workflow from
the base branch. It checks out the PR head only after runner selection and runs
that code on `homelab-pr`. Keep this workflow read-only and do not add secrets.
The PR runner has a separate home and tool cache. Do not add it to the `docker`
group or give it access to `/var/run/docker.sock`. It runs repository code from
pull requests, so keep its registration and permissions separate from the
trusted `homelab` runner. This separates users and host permissions, but both
runners still share the VPS kernel and network. Use a disposable VM if PRs from
untrusted external authors must be fully isolated.
## Workstation setup
As the existing SSH deploy user on workstation:
```sh
sudo loginctl enable-linger forust
bash .gitea/runner/setup-workstation.sh
```
The controller uses `/srv/homelab` as the persistent configuration tree and makes
a detached source worktree for each SHA. It never resets `/srv/homelab`, moves
local configuration, renames Compose projects, or changes volume names.
The installer records the current Kubernetes context and cluster UID in
`~/.config/homelab-deploy/environment`. Check these before installing.
Configure Gitea Actions Variables:
- `DEPLOY_HOST`, `DEPLOY_USER`, `DEPLOY_PORT`: the existing VPS-to-workstation SSH endpoint.
- `DEPLOY_KNOWN_HOSTS`: workstation's verified SSH host key entry for that endpoint.
- `AUTODEPLOY`: `false` initially; `true` enables deployment after successful main CI.
Keep `DEPLOY_SSH_KEY`, `REGISTRY_USERNAME` and `REGISTRY_PASSWORD` in Actions
Secrets. Legacy endpoint secrets remain accepted during migration. The Actions
token must have repository read and Actions read access for release downloads.
The deploy user's existing Docker registry authentication remains necessary.
## Releases and deployment
CI publishes `release-<full SHA>` as a Gitea artifact with all three owned image
digests and build input fingerprints. Unchanged images are reused only from a
successful main CI artifact, never from `:prod`. Expired artifacts cause CI to
rebuild images; they block deployment until CI is rerun.
Run deploy from main with `deploy_ref=main` or a checked SHA:
- `full`: required for the first baseline; reconcile all active components.
- `changed`: compare with the last fully successful production deploy.
- `plan`: validate configuration and show selection without changing production resources.
- `refresh_images=true`: explicitly refresh mutable third-party Compose tags.
The manual and automatic paths both require successful CI, a successful build
job and the exact SHA's release artifact. PRs cannot publish images or deploy.
Removed resources are reported and require explicit removal; no automatic prune.
Service dependencies are listed in `.gitea/deploy-dependencies.json`.
A workstation user systemd service holds the deploy lock across validation,
sequential apply, verification and smoke checks. SSH clients only submit/follow:
disconnecting or cancelling the Actions client does not kill production apply.
Retrying the same run ID does not start another apply. `ExecStopPost` recovers
interrupted runs before the unit finishes. Kubernetes rolls back to captured
revisions; configuration and persistent data are not reverted.
## Status and recovery
`--retry` repeats failed verification and smoke checks, never apply. Recovery
keeps a failed deploy out of the successful baseline, even after rollback.
On workstation (replace the numeric ID with Actions run ID and attempt):
```sh
python3 ~/.local/lib/homelab-deploy/controller.py status 123-1
python3 ~/.local/lib/homelab-deploy/controller.py recover 123-1 --retry
journalctl --user -u homelab-deploy@123-1
```
Runs live in `~/.local/state/homelab-deploy/runs`. Compose stores resolved configs
with restricted permissions; these may contain credentials and must never be
uploaded as CI artifacts. Stage logs print the exact manual recovery command
using `compose-before/<stack>.json`, the original project directory and project
name. Compose does not automatically roll back, and Nextcloud AIO's child
containers remain managed by AIO. Preserve its own backups for data recovery.
The controller retains twenty successful/planned runs and preserves failures.
Update the workstation dispatcher only when no deploy is running.
## Validation and migration rollback
```sh
python3 -m unittest discover -s tests -v
bash .gitea/tests/deploy-validation.sh
```
Test on a separate namespace before the initial production `full` run. Check a
failed rollout, interrupted SSH and repeated run ID, and verify that an isolated
service change does not upgrade unrelated Helm releases or Compose stacks.
To roll back the migration, disable autodeploy and finish or recover the remote
run first. Restore the runner config/unit from `.before-<timestamp>` backups,
reload systemd and restart the runner. Restore the prior workflows from Git.
Production data and persistent volumes stay where they were. Do not remove run
state or Compose recovery files until recovery is confirmed.
+11
View File
@@ -0,0 +1,11 @@
[worker.oci]
gc = true
reservedSpace = "256MB"
maxUsedSpace = "1GB"
minFreeSpace = "2GB"
[[worker.oci.gcpolicy]]
reservedSpace = "256MB"
maxUsedSpace = "1GB"
minFreeSpace = "2GB"
all = true
+10
View File
@@ -0,0 +1,10 @@
runner:
file: /var/lib/gitea-runner/.runner
capacity: 1
timeout: 5h
labels:
- homelab:host
cache:
enabled: false
container:
docker_host: unix:///var/run/docker.sock
+18
View File
@@ -0,0 +1,18 @@
[Unit]
Description=Gitea Actions runner
After=network-online.target docker.service
Wants=network-online.target
[Service]
User=gitea-runner
Group=gitea-runner
SupplementaryGroups=docker
WorkingDirectory=/var/lib/gitea-runner
Environment=PATH=/var/lib/gitea-runner/.cache/homelab-ci/bin:/usr/local/bin:/usr/bin:/bin
ExecStart=/usr/local/bin/gitea-runner daemon --config /etc/gitea-runner/config.yaml
Restart=on-failure
RestartSec=5
UMask=0077
[Install]
WantedBy=multi-user.target
+12
View File
@@ -0,0 +1,12 @@
[Unit]
Description=Homelab deploy %i
[Service]
Type=exec
EnvironmentFile=%h/.config/homelab-deploy/environment
ExecStart=/usr/bin/python3 %h/.local/lib/homelab-deploy/controller.py execute %i
ExecStopPost=/usr/bin/python3 %h/.local/lib/homelab-deploy/controller.py recover %i
RuntimeMaxSec=5h
TimeoutStopSec=135min
KillMode=control-group
UMask=0077
+8
View File
@@ -0,0 +1,8 @@
runner:
file: /var/lib/gitea-pr-runner/.runner
capacity: 1
timeout: 5h
labels:
- homelab-pr:host
cache:
enabled: false
+27
View File
@@ -0,0 +1,27 @@
[Unit]
Description=Gitea Actions untrusted pull request runner
After=network-online.target
Wants=network-online.target
[Service]
User=gitea-pr-runner
Group=gitea-pr-runner
WorkingDirectory=/var/lib/gitea-pr-runner
Environment=HOME=/var/lib/gitea-pr-runner
Environment=PATH=/var/lib/gitea-pr-runner/.cache/homelab-ci/bin:/usr/local/bin:/usr/bin:/bin
ExecStart=/usr/local/bin/gitea-runner daemon --config /etc/gitea-pr-runner/config.yaml
Restart=on-failure
RestartSec=5
NoNewPrivileges=yes
PrivateTmp=yes
ProtectSystem=full
ProtectHome=yes
ProtectKernelTunables=yes
ProtectKernelModules=yes
ProtectControlGroups=yes
RestrictSUIDSGID=yes
LockPersonality=yes
UMask=0077
[Install]
WantedBy=multi-user.target
+56
View File
@@ -0,0 +1,56 @@
#!/usr/bin/env bash
# Install a native runner for untrusted PR jobs without Docker access.
set -euo pipefail
here="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
[ "$(id -u)" -eq 0 ] || { echo 'Run with sudo on the runner host' >&2; exit 1; }
for tool in cp cut date getent id install runuser systemctl useradd; do
command -v "$tool" >/dev/null || { echo "Install missing prerequisite: $tool" >&2; exit 1; }
done
command -v /usr/local/bin/gitea-runner >/dev/null || {
echo 'Install gitea-runner 3.0.2 at /usr/local/bin/gitea-runner first' >&2
exit 1
}
id gitea-pr-runner >/dev/null 2>&1 || \
useradd --system --create-home --home-dir /var/lib/gitea-pr-runner --shell /usr/bin/bash gitea-pr-runner
runner_home="$(getent passwd gitea-pr-runner | cut -d: -f6)"
[ "$runner_home" = /var/lib/gitea-pr-runner ] || {
echo 'Unexpected PR runner home; inspect the existing service first' >&2
exit 1
}
case " $(id -nG gitea-pr-runner) " in
*' docker '*)
echo 'The PR runner account must not belong to the docker group' >&2
exit 1
;;
esac
install -d -m 0755 /etc/gitea-pr-runner
stamp="$(date -u +%Y%m%dT%H%M%SZ)"
for existing in /etc/gitea-pr-runner/config.yaml /etc/systemd/system/gitea-pr-runner.service; do
[ ! -f "$existing" ] || cp -p "$existing" "$existing.before-$stamp"
done
install -m 0644 "$here/pr-config.yaml" /etc/gitea-pr-runner/config.yaml
install -m 0644 "$here/pr-runner.service" /etc/systemd/system/gitea-pr-runner.service
if [ ! -f /var/lib/gitea-pr-runner/.runner ]; then
read -r -s -p 'Enter the Gitea repository runner registration token: ' runner_token
printf '\n'
[ -n "$runner_token" ] || { echo 'Runner token is required' >&2; exit 1; }
export GITEA_RUNNER_REGISTRATION_TOKEN="$runner_token"
unset runner_token
runuser --preserve-environment -u gitea-pr-runner -- \
/usr/local/bin/gitea-runner register \
--config /etc/gitea-pr-runner/config.yaml \
--instance https://gitea.forust.xyz \
--name homelab-pr \
--labels homelab-pr:host \
--no-interactive
unset GITEA_RUNNER_REGISTRATION_TOKEN
fi
chmod 0600 /var/lib/gitea-pr-runner/.runner
systemctl daemon-reload
systemctl enable --now gitea-pr-runner.service
systemctl restart gitea-pr-runner.service
echo "PR runner ready. Configuration backups: *.before-$stamp"
+48
View File
@@ -0,0 +1,48 @@
#!/usr/bin/env bash
# Native host runner, with pinned user-space tools and no extra CI images.
set -euo pipefail
here="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
[ "$(id -u)" -eq 0 ] || { echo 'Run with sudo on the runner host' >&2; exit 1; }
for tool in docker curl python3 git tar xz flock runuser systemctl; do
command -v "$tool" >/dev/null || { echo "Install missing prerequisite: $tool" >&2; exit 1; }
done
docker info >/dev/null
docker compose version >/dev/null
docker buildx version >/dev/null
id gitea-runner >/dev/null 2>&1 || useradd --system --create-home --home-dir /var/lib/gitea-runner --shell /usr/bin/bash gitea-runner
# Reuse the established service account and runner registration.
runner_home="$(getent passwd gitea-runner | cut -d: -f6)"
[ "$runner_home" = /var/lib/gitea-runner ] || { echo 'Unexpected runner home; inspect the existing service first' >&2; exit 1; }
runuser -u gitea-runner -- docker info >/dev/null || { echo "The runner user needs access to Docker before setup" >&2; exit 1; }
command -v gitea-runner >/dev/null || { echo 'Install gitea-runner 3.0.2 at /usr/local/bin/gitea-runner first' >&2; exit 1; }
mkdir -p /etc/gitea-runner
stamp="$(date -u +%Y%m%dT%H%M%SZ)"
for existing in /etc/gitea-runner/config.yaml /etc/systemd/system/gitea-runner.service; do
[ ! -f "$existing" ] || cp -p "$existing" "$existing.before-$stamp"
done
scratch="$(mktemp -d)"
trap 'rm -rf "$scratch"' EXIT
chmod 755 "$scratch"
install -m 0644 "$here/../workflows/install-ci-tools.sh" "$here/../workflows/tool-versions.env" "$scratch/"
runuser -u gitea-runner -- bash "$scratch/install-ci-tools.sh"
install -m 0644 "$here/config.yaml" /etc/gitea-runner/config.yaml
python3 - <<'PYLABELS'
import json
from pathlib import Path
registration = Path('/var/lib/gitea-runner/.runner')
if registration.exists():
labels = json.loads(registration.read_text()).get('labels', [])
labels = [label for label in labels if isinstance(label, str) and label.split(':')[0] != 'homelab']
labels.append('homelab:host')
config = Path('/etc/gitea-runner/config.yaml')
config.write_text(config.read_text().replace(' - homelab:host', '\n'.join(' - ' + json.dumps(label) for label in labels)))
PYLABELS
install -m 0644 "$here/gitea-runner.service" /etc/systemd/system/gitea-runner.service
if [ ! -f /var/lib/gitea-runner/.runner ]; then
echo 'Register once as gitea-runner with homelab:host before starting the service.'
exit 0
fi
systemctl daemon-reload
systemctl enable --now gitea-runner.service
systemctl restart gitea-runner.service
echo "Runner ready. Configuration backups: *.before-$stamp"
+32
View File
@@ -0,0 +1,32 @@
#!/usr/bin/env bash
# Run as the existing deploy user on workstation. Never resets the working tree.
set -euo pipefail
here="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
repo="${HOMELAB_REPO:-/srv/homelab}"
for tool in python3 git kubectl helm docker flock timeout; do
command -v "$tool" >/dev/null || { echo "Install missing dependency: $tool" >&2; exit 1; }
done
[ -d "$repo/.git" ] || { echo "Missing deploy checkout: $repo" >&2; exit 1; }
[[ "$repo" =~ ^/[A-Za-z0-9_./-]+$ ]] || { echo 'Deploy path must be absolute and contain no whitespace' >&2; exit 1; }
if [ "$(loginctl show-user "$USER" -p Linger --value)" != yes ]; then
echo "Run once: sudo loginctl enable-linger $USER" >&2
exit 1
fi
config="${XDG_CONFIG_HOME:-$HOME/.config}/homelab-deploy"
mkdir -p "$config" "$HOME/.local/lib/homelab-deploy" "$HOME/.config/systemd/user"
chmod 700 "$config"
if [ ! -f "$config/environment" ]; then
context="$(kubectl config current-context)"
cluster_uid="$(kubectl get namespace kube-system -o jsonpath='{.metadata.uid}')"
printf 'HOMELAB_REPO=%s\nKUBE_CONTEXT=%s\nEXPECTED_CLUSTER_UID=%s\n' "$repo" "$context" "$cluster_uid" >"$config/environment"
chmod 600 "$config/environment"
fi
# Do not replace a dispatcher while an existing deploy uses it.
if systemctl --user list-units 'homelab-deploy@*' --state=running --no-legend | grep -q .; then
echo 'An existing deploy is running; wait before updating the controller' >&2
exit 1
fi
install -m 0755 "$here/../workflows/deploy-controller.py" "$HOME/.local/lib/homelab-deploy/controller.py"
install -m 0644 "$here/homelab-deploy@.service" "$HOME/.config/systemd/user/homelab-deploy@.service"
systemctl --user daemon-reload
echo 'Controller ready. Run a checked main SHA in full mode for the initial baseline.'
+14
View File
@@ -73,6 +73,20 @@ fi
grep -q 'MISSING OR UNREADABLE: app/credentials' "$scratch/secrets.log" grep -q 'MISSING OR UNREADABLE: app/credentials' "$scratch/secrets.log"
# API/rendering errors must not produce an empty reference list and pass. # API/rendering errors must not produce an empty reference list and pass.
kubectl() { return 1; } kubectl() { return 1; }
if ! skip_uninstalled_vmagent_crd "$REPO/prometheus-stack/k8s/vmagent.yaml"; then
echo 'VMAgent preflight did not skip an uninstalled CRD' >&2
exit 1
fi
kubectl() { return 0; }
if skip_uninstalled_vmagent_crd "$REPO/prometheus-stack/k8s/vmagent.yaml"; then
echo 'VMAgent preflight skipped an installed CRD' >&2
exit 1
fi
if skip_uninstalled_vmagent_crd "$REPO/prometheus-stack/k8s/victoria.yaml"; then
echo 'VMAgent preflight skipped an unrelated manifest' >&2
exit 1
fi
kubectl() { return 1; }
if check_referenced_secrets >"$scratch/secrets.log"; then if check_referenced_secrets >"$scratch/secrets.log"; then
echo 'Secret check accepted a failed manifest render' >&2 echo 'Secret check accepted a failed manifest render' >&2
exit 1 exit 1
+356 -386
View File
@@ -1,38 +1,25 @@
name: ci name: ci
"on":
on:
push: push:
branches: branches:
- "**" - main
pull_request: pull_request: null
workflow_dispatch: workflow_dispatch: null
# Every job here is checkout plus local tools. The token needs to read the tree
# and nothing else, and saying so keeps a future step that reaches for the API
# from quietly holding a token that can write to the repository.
permissions: permissions:
contents: read contents: read
actions: read
concurrency: concurrency:
group: ci-${{ github.ref }} group: ci-${{ github.ref }}
cancel-in-progress: ${{ github.ref != 'refs/heads/main' }} cancel-in-progress: ${{ github.ref != 'refs/heads/main' }}
env:
REGISTRY: gcr.forust.xyz
jobs: jobs:
lint-compose: compose:
runs-on: [self-hosted, linux, arch, homelab] name: Compose
timeout-minutes: 10 runs-on: ${{ github.ref == 'refs/heads/main' && 'homelab' || 'homelab-pr' }}
timeout-minutes: 15
steps: steps:
- name: Checkout repository - name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4 uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262
id: source
# Structure check for every committed Compose file, active or not.
# Interpolation, env-file and bind-mount resolution are all switched off,
# because inactive stacks have no .env here and would only fail on their
# ${VAR:?} guards. Active stacks get the full check with interpolation in
# the deploy workflow, where the real .env files live.
- name: Validate Compose files - name: Validate Compose files
shell: bash shell: bash
run: | run: |
@@ -61,35 +48,77 @@ jobs:
exit 1 exit 1
fi fi
echo "checked ${#files[@]} Compose file(s)" echo "checked ${#files[@]} Compose file(s)"
id: check
lint-actionlint: - name: Write the job result
runs-on: [self-hosted, linux, arch, homelab] if: always()
timeout-minutes: 10 env:
SUMMARY_CHECK: Compose
SUMMARY_RESULT: ${{ job.status }}
SUMMARY_FAILED_STEP:
${{ steps.check.conclusion == 'failure' && 'Check or image build' || steps.source.conclusion == 'failure'
&& 'Source checkout' || '' }}
shell: bash
run: |
if [ -f .gitea/workflows/release.py ]; then
python3 .gitea/workflows/release.py check-summary
elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
printf '## %s\n\n- Result: **%s**\n- Failed step: %s\n' "$SUMMARY_CHECK" "$SUMMARY_RESULT" "$SUMMARY_FAILED_STEP" >>"$GITHUB_STEP_SUMMARY" || true
fi
workflows:
name: Workflows
runs-on: ${{ github.ref == 'refs/heads/main' && 'homelab' || 'homelab-pr' }}
timeout-minutes: 15
steps: steps:
- name: Checkout repository - name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4 uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262
id: source
- name: Prepare pinned tools
shell: bash
run: |
set -euo pipefail
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh actionlint shellcheck)"
echo "$tools_dir" >> "$GITHUB_PATH"
id: tools
- name: Lint Gitea Actions workflows with actionlint - name: Lint Gitea Actions workflows with actionlint
shell: bash shell: bash
run: | run: |
set -euo pipefail set -euo pipefail
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh actionlint)"
export PATH="$tools_dir:$PATH"
actionlint -config-file .gitea/actionlint.yaml -color .gitea/workflows/*.yaml actionlint -config-file .gitea/actionlint.yaml -color .gitea/workflows/*.yaml
id: check
lint-shellcheck: - name: Write the job result
runs-on: [self-hosted, linux, arch, homelab] if: always()
timeout-minutes: 10 env:
SUMMARY_CHECK: Workflows
SUMMARY_RESULT: ${{ job.status }}
SUMMARY_FAILED_STEP:
${{ steps.check.conclusion == 'failure' && 'Check or image build' || steps.tools.conclusion == 'failure'
&& 'Tool setup' || steps.source.conclusion == 'failure' && 'Source checkout' || '' }}
shell: bash
run: |
if [ -f .gitea/workflows/release.py ]; then
python3 .gitea/workflows/release.py check-summary
elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
printf '## %s\n\n- Result: **%s**\n- Failed step: %s\n' "$SUMMARY_CHECK" "$SUMMARY_RESULT" "$SUMMARY_FAILED_STEP" >>"$GITHUB_STEP_SUMMARY" || true
fi
shell:
name: Shell
runs-on: ${{ github.ref == 'refs/heads/main' && 'homelab' || 'homelab-pr' }}
timeout-minutes: 15
steps: steps:
- name: Checkout repository - name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4 uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262
id: source
- name: Lint shell scripts with ShellCheck - name: Prepare pinned tools
shell: bash shell: bash
run: | run: |
set -euo pipefail set -euo pipefail
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh shellcheck jq)" tools_dir="$(bash .gitea/workflows/install-ci-tools.sh shellcheck jq)"
export PATH="$tools_dir:$PATH" echo "$tools_dir" >> "$GITHUB_PATH"
id: tools
- name: Lint shell scripts with ShellCheck
shell: bash
run: |
set -euo pipefail
mapfile -t scripts < <( mapfile -t scripts < <(
git ls-files '*.sh' ':(glob)**/*.bash' git ls-files '*.sh' ':(glob)**/*.bash'
) )
@@ -99,20 +128,41 @@ jobs:
fi fi
shellcheck --external-sources --source-path=SCRIPTDIR --severity=style "${scripts[@]}" shellcheck --external-sources --source-path=SCRIPTDIR --severity=style "${scripts[@]}"
bash .gitea/tests/deploy-validation.sh bash .gitea/tests/deploy-validation.sh
id: check
lint-prettier: - name: Write the job result
runs-on: [self-hosted, linux, arch, homelab] if: always()
timeout-minutes: 10 env:
SUMMARY_CHECK: Shell
SUMMARY_RESULT: ${{ job.status }}
SUMMARY_FAILED_STEP:
${{ steps.check.conclusion == 'failure' && 'Check or image build' || steps.tools.conclusion == 'failure'
&& 'Tool setup' || steps.source.conclusion == 'failure' && 'Source checkout' || '' }}
shell: bash
run: |
if [ -f .gitea/workflows/release.py ]; then
python3 .gitea/workflows/release.py check-summary
elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
printf '## %s\n\n- Result: **%s**\n- Failed step: %s\n' "$SUMMARY_CHECK" "$SUMMARY_RESULT" "$SUMMARY_FAILED_STEP" >>"$GITHUB_STEP_SUMMARY" || true
fi
formatting:
name: Formatting
runs-on: ${{ github.ref == 'refs/heads/main' && 'homelab' || 'homelab-pr' }}
timeout-minutes: 15
steps: steps:
- name: Checkout repository - name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4 uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262
id: source
- name: Check formatting with Prettier - name: Prepare pinned tools
shell: bash shell: bash
run: | run: |
set -euo pipefail set -euo pipefail
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh prettier)" tools_dir="$(bash .gitea/workflows/install-ci-tools.sh prettier)"
export PATH="$tools_dir:$PATH" echo "$tools_dir" >> "$GITHUB_PATH"
id: tools
- name: Check formatting with Prettier
shell: bash
run: |
set -euo pipefail
mapfile -t prettier_files < <( mapfile -t prettier_files < <(
git ls-files \ git ls-files \
@@ -126,36 +176,79 @@ jobs:
fi fi
prettier --check --ignore-unknown "${prettier_files[@]}" prettier --check --ignore-unknown "${prettier_files[@]}"
id: check
lint-ruff: - name: Write the job result
runs-on: [self-hosted, linux, arch, homelab] if: always()
timeout-minutes: 10 env:
SUMMARY_CHECK: Formatting
SUMMARY_RESULT: ${{ job.status }}
SUMMARY_FAILED_STEP:
${{ steps.check.conclusion == 'failure' && 'Check or image build' || steps.tools.conclusion == 'failure'
&& 'Tool setup' || steps.source.conclusion == 'failure' && 'Source checkout' || '' }}
shell: bash
run: |
if [ -f .gitea/workflows/release.py ]; then
python3 .gitea/workflows/release.py check-summary
elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
printf '## %s\n\n- Result: **%s**\n- Failed step: %s\n' "$SUMMARY_CHECK" "$SUMMARY_RESULT" "$SUMMARY_FAILED_STEP" >>"$GITHUB_STEP_SUMMARY" || true
fi
python:
name: Python and tests
runs-on: ${{ github.ref == 'refs/heads/main' && 'homelab' || 'homelab-pr' }}
timeout-minutes: 15
steps: steps:
- name: Checkout repository - name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4 uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262
id: source
- name: Prepare pinned tools
shell: bash
run: |
set -euo pipefail
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh ruff jq)"
echo "$tools_dir" >> "$GITHUB_PATH"
id: tools
- name: Lint and format-check Python with Ruff - name: Lint and format-check Python with Ruff
shell: bash shell: bash
run: | run: |
set -euo pipefail set -euo pipefail
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh ruff)" ruff check . .gitea/workflows
export PATH="$tools_dir:$PATH" ruff format --check . .gitea/workflows
ruff check . python3 -m unittest discover -s tests -v
ruff format --check . id: check
- name: Write the job result
lint-yaml: if: always()
runs-on: [self-hosted, linux, arch, homelab] env:
timeout-minutes: 10 SUMMARY_CHECK: Python and tests
SUMMARY_RESULT: ${{ job.status }}
SUMMARY_FAILED_STEP:
${{ steps.check.conclusion == 'failure' && 'Check or image build' || steps.tools.conclusion == 'failure'
&& 'Tool setup' || steps.source.conclusion == 'failure' && 'Source checkout' || '' }}
shell: bash
run: |
if [ -f .gitea/workflows/release.py ]; then
python3 .gitea/workflows/release.py check-summary
elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
printf '## %s\n\n- Result: **%s**\n- Failed step: %s\n' "$SUMMARY_CHECK" "$SUMMARY_RESULT" "$SUMMARY_FAILED_STEP" >>"$GITHUB_STEP_SUMMARY" || true
fi
yaml:
name: YAML
runs-on: ${{ github.ref == 'refs/heads/main' && 'homelab' || 'homelab-pr' }}
timeout-minutes: 15
steps: steps:
- name: Checkout repository - name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4 uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262
id: source
- name: Lint YAML syntax - name: Prepare pinned tools
shell: bash shell: bash
run: | run: |
set -euo pipefail set -euo pipefail
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh yamllint)" tools_dir="$(bash .gitea/workflows/install-ci-tools.sh yamllint)"
export PATH="$tools_dir:$PATH" echo "$tools_dir" >> "$GITHUB_PATH"
id: tools
- name: Lint YAML syntax
shell: bash
run: |
set -euo pipefail
mapfile -t yaml_files < <( mapfile -t yaml_files < <(
git ls-files '*.yaml' '*.yml' \ git ls-files '*.yaml' '*.yml' \
@@ -169,20 +262,41 @@ jobs:
fi fi
yamllint -c .yamllint "${yaml_files[@]}" yamllint -c .yamllint "${yaml_files[@]}"
id: check
lint-dockerfiles: - name: Write the job result
runs-on: [self-hosted, linux, arch, homelab] if: always()
timeout-minutes: 10 env:
SUMMARY_CHECK: YAML
SUMMARY_RESULT: ${{ job.status }}
SUMMARY_FAILED_STEP:
${{ steps.check.conclusion == 'failure' && 'Check or image build' || steps.tools.conclusion == 'failure'
&& 'Tool setup' || steps.source.conclusion == 'failure' && 'Source checkout' || '' }}
shell: bash
run: |
if [ -f .gitea/workflows/release.py ]; then
python3 .gitea/workflows/release.py check-summary
elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
printf '## %s\n\n- Result: **%s**\n- Failed step: %s\n' "$SUMMARY_CHECK" "$SUMMARY_RESULT" "$SUMMARY_FAILED_STEP" >>"$GITHUB_STEP_SUMMARY" || true
fi
dockerfiles:
name: Dockerfiles
runs-on: ${{ github.ref == 'refs/heads/main' && 'homelab' || 'homelab-pr' }}
timeout-minutes: 15
steps: steps:
- name: Checkout repository - name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4 uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262
id: source
- name: Lint Dockerfiles - name: Prepare pinned tools
shell: bash shell: bash
run: | run: |
set -euo pipefail set -euo pipefail
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh hadolint)" tools_dir="$(bash .gitea/workflows/install-ci-tools.sh hadolint)"
export PATH="$tools_dir:$PATH" echo "$tools_dir" >> "$GITHUB_PATH"
id: tools
- name: Lint Dockerfiles
shell: bash
run: |
set -euo pipefail
mapfile -t dockerfiles < <( mapfile -t dockerfiles < <(
git ls-files ':(glob)**/Dockerfile' ':(glob)**/Dockerfile.*' git ls-files ':(glob)**/Dockerfile' ':(glob)**/Dockerfile.*'
@@ -194,20 +308,41 @@ jobs:
fi fi
hadolint -c .hadolint.yaml "${dockerfiles[@]}" hadolint -c .hadolint.yaml "${dockerfiles[@]}"
id: check
validate: - name: Write the job result
runs-on: [self-hosted, linux, arch, homelab] if: always()
timeout-minutes: 20 env:
SUMMARY_CHECK: Dockerfiles
SUMMARY_RESULT: ${{ job.status }}
SUMMARY_FAILED_STEP:
${{ steps.check.conclusion == 'failure' && 'Check or image build' || steps.tools.conclusion == 'failure'
&& 'Tool setup' || steps.source.conclusion == 'failure' && 'Source checkout' || '' }}
shell: bash
run: |
if [ -f .gitea/workflows/release.py ]; then
python3 .gitea/workflows/release.py check-summary
elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
printf '## %s\n\n- Result: **%s**\n- Failed step: %s\n' "$SUMMARY_CHECK" "$SUMMARY_RESULT" "$SUMMARY_FAILED_STEP" >>"$GITHUB_STEP_SUMMARY" || true
fi
kubernetes:
name: Kubernetes
runs-on: ${{ github.ref == 'refs/heads/main' && 'homelab' || 'homelab-pr' }}
timeout-minutes: 15
steps: steps:
- name: Checkout repository - name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4 uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262
id: source
- name: Validate Kubernetes manifests against JSON schemas - name: Prepare pinned tools
shell: bash shell: bash
run: | run: |
set -euo pipefail set -euo pipefail
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh kubeconform)" tools_dir="$(bash .gitea/workflows/install-ci-tools.sh kubeconform)"
export PATH="$tools_dir:$PATH" echo "$tools_dir" >> "$GITHUB_PATH"
id: tools
- name: Validate Kubernetes manifests against JSON schemas
shell: bash
run: |
set -euo pipefail
mapfile -t manifests < <( mapfile -t manifests < <(
git ls-files ':(glob)**/k8s/**/*.yaml' ':(glob)**/k8s/**/*.yml' \ git ls-files ':(glob)**/k8s/**/*.yaml' ':(glob)**/k8s/**/*.yml' \
@@ -224,327 +359,162 @@ jobs:
-ignore-missing-schemas \ -ignore-missing-schemas \
-summary \ -summary \
"${manifests[@]}" "${manifests[@]}"
id: check
# kubeconform has no schemas for CRDs, so every IngressRoute, Certificate, - name: Write the job result
# PrometheusRule, Middleware, ServersTransport and ServiceMonitor is silently if: always()
# skipped above. The live API server knows the real CRD schemas (and runs the env:
# cert-manager / Traefik admission webhooks), so validate there too. SUMMARY_CHECK: Kubernetes
# SUMMARY_RESULT: ${{ job.status }}
# Only services marked with a k8s/active marker are checked: server-side SUMMARY_FAILED_STEP:
# dry-run needs the target namespace to exist, and inactive services are not ${{ steps.check.conclusion == 'failure' && 'Check or image build' || steps.tools.conclusion == 'failure'
# deployed. Services being enabled for the first time are still covered by && 'Tool setup' || steps.source.conclusion == 'failure' && 'Source checkout' || '' }}
# the JSON-schema pass above.
#
# Main pushes only. `--dry-run=server` persists nothing, but it does execute
# the admission webhooks of the production API server, so anyone able to open
# a pull request would be able to run arbitrary manifest content through
# cert-manager and Traefik. A pull request has nothing to gain from it either:
# only main is ever deployed, and this job runs to completion before the
# deploy workflow is allowed to start, so a bad CRD is still caught before
# anything reaches the cluster -- just on the push rather than on the PR.
- name: Note the server-side check is not running here
if: github.event_name == 'pull_request' || github.ref != 'refs/heads/main'
shell: bash shell: bash
run: | run: |
echo "::notice::Skipping the server-side dry-run. It executes the cert-manager and" \ if [ -f .gitea/workflows/release.py ]; then
"Traefik admission webhooks against the production API server, so it is limited" \ python3 .gitea/workflows/release.py check-summary
"to pushes to main. CRDs are still schema-checked by kubeconform above, and the" \ elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
"server-side pass still runs on main before the deploy." printf '## %s\n\n- Result: **%s**\n- Failed step: %s\n' "$SUMMARY_CHECK" "$SUMMARY_RESULT" "$SUMMARY_FAILED_STEP" >>"$GITHUB_STEP_SUMMARY" || true
- name: Validate active manifests against the live API server
if: github.event_name != 'pull_request' && github.ref == 'refs/heads/main'
shell: bash
run: |
set -euo pipefail
if ! kubectl get --raw='/readyz' --request-timeout=10s >/dev/null 2>&1; then
echo "::warning::Cluster unreachable — skipped server-side validation of CRDs (IngressRoute, Certificate, PrometheusRule). Review manifest changes manually."
exit 0
fi fi
image-plan:
mapfile -t k8s_dirs < <( needs: [compose, workflows, shell, formatting, python, yaml, dockerfiles, kubernetes]
git ls-files '*.yaml' '*.yml' \ if: github.event_name != 'pull_request' && github.ref == 'refs/heads/main'
| grep -E '(^|/)k8s/' \ runs-on: homelab
| sed -E 's#((^|.*/)k8s)/.*#\1#' \ timeout-minutes: 10
| sort -u
)
manifests=()
kustomize_apps=()
for dir in "${k8s_dirs[@]}"; do
if [ ! -f "${dir}/active" ]; then
echo "skip (no k8s/active): ${dir}"
continue
fi
if [ -f "${dir}/overlays/prod/kustomization.yaml" ]; then
kustomize_apps+=("${dir}/overlays/prod")
elif [ -f "${dir}/base/kustomization.yaml" ]; then
kustomize_apps+=("${dir}/base")
else
while IFS= read -r f; do
[ -n "$f" ] && manifests+=("$f")
done < <(
git ls-files "${dir}/*.yaml" "${dir}/*.yml" \
| grep -Ev '(^|/)(kustomization\.ya?ml|.*\.example\.ya?ml|.*values\.ya?ml|patch-.*\.ya?ml)$'
)
fi
done
echo "server-side dry-run: ${#manifests[@]} manifests, ${#kustomize_apps[@]} kustomize apps"
failed=0
for m in ${manifests[@]+"${manifests[@]}"}; do
if ! out="$(kubectl apply --dry-run=server -f "$m" 2>&1)"; then
failed=1
echo "::error file=${m}::$(printf '%s' "$out" | head -1)"
fi
done
for k in ${kustomize_apps[@]+"${kustomize_apps[@]}"}; do
if ! out="$(kubectl apply -k "$k" --dry-run=server 2>&1)"; then
failed=1
echo "::error file=${k}::$(printf '%s' "$out" | head -1)"
fi
done
if [ "$failed" -ne 0 ]; then
echo "Server-side validation failed. The API server (or an admission webhook) rejected these manifests."
exit 1
fi
echo "server-side dry-run: all active manifests accepted by the API server"
build:
needs:
# The panel's scan-deps/test-backend/test-frontend jobs gated here until
# userbot moved to its own repo; upstream's code is upstream's gate now.
# The rule is unchanged: publishing and passing the checks are the same
# gate, so a commit that fails any of these still cannot move :prod.
[lint-actionlint, lint-shellcheck, lint-compose, lint-prettier, lint-ruff, lint-yaml, lint-dockerfiles, validate]
if: github.event_name != 'pull_request' && (github.ref_name == 'main' || github.ref_name == 'dev') && !startsWith(github.ref_name, 'renovate/')
runs-on: [self-hosted, linux, arch, homelab]
timeout-minutes: 60
outputs: outputs:
services: ${{ steps.services.outputs.services }} matrix: ${{ steps.plan.outputs.matrix }}
steps: steps:
- name: Checkout repository - name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4 id: source
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262
with: with:
fetch-depth: 0 fetch-depth: 0
- name: Detect build inputs against successful CI
id: plan
env:
GITEA_TOKEN: ${{ github.token }}
run: python3 .gitea/workflows/release.py prepare --output build-plan.json
- name: Store the image plan
id: artifact
uses: actions/upload-artifact@ea165f8d65b6e75b540449e92b4886f43607fa02 # v4.6.2
with:
name: build-plan
path: build-plan.json
if-no-files-found: error
retention-days: 30
- name: Detect changed docker-built services - name: Write the plan result
id: services if: always()
env:
SUMMARY_CHECK: Image plan
SUMMARY_RESULT: ${{ job.status }}
SUMMARY_FAILED_STEP: >-
${{ steps.plan.conclusion == 'failure' && 'Build input detection' ||
steps.artifact.conclusion == 'failure' && 'Plan upload' ||
steps.source.conclusion == 'failure' && 'Source checkout' || '' }}
shell: bash shell: bash
run: | run: |
set -euo pipefail if [ -f .gitea/workflows/release.py ]; then
base="${{ github.event.before }}" python3 .gitea/workflows/release.py check-summary
if [ -z "$base" ] || [ "$base" = "0000000000000000000000000000000000000000" ]; then elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
base="$(git rev-list --max-parents=0 HEAD)" printf '## Image plan\n\nResult: %s\n' "$SUMMARY_RESULT" >>"$GITHUB_STEP_SUMMARY" || true
fi fi
# A failed diff used to leave changed_files empty, which reads exactly images:
# like "nothing to build": the job went green having built nothing and name: Image (${{ matrix.name }})
# the tag never moved. The status is checked, not assumed. needs: [image-plan]
if ! changed="$(git diff --name-only "$base" "${GITHUB_SHA}")"; then if: needs.image-plan.result == 'success'
echo "::error::cannot diff ${base}..${GITHUB_SHA}" runs-on: homelab
exit 1 timeout-minutes: 60
fi strategy:
mapfile -t changed_files <<<"$changed" max-parallel: 1
fail-fast: false
services=() matrix: ${{ fromJSON(needs.image-plan.outputs.matrix || '{"include":[{"name":"inactive"}]}') }}
steps:
add_service() { - name: Checkout repository
local name="$1" id: source
local seen=0 uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262
for existing in "${services[@]}"; do - name: Download the checked image plan
if [ "$existing" = "$name" ]; then id: inputs
seen=1 uses: actions/download-artifact@d3f86a106a0bac45b974a628896c90dbdf5c8093 # v4.3.0
break with:
fi name: build-plan
done - name: Build or reuse this image
if [ "$seen" -eq 0 ]; then id: check
services+=("$name") env:
fi IMAGE_NAME: ${{ matrix.name }}
} REGISTRY_USERNAME: ${{ secrets.REGISTRY_USERNAME }}
REGISTRY_PASSWORD: ${{ secrets.REGISTRY_PASSWORD }}
for file in "${changed_files[@]}"; do run: python3 .gitea/workflows/release.py image --image "$IMAGE_NAME" --output image.json
case "$file" in - name: Store the image result
errorpages/*) id: artifact
add_service errorpages uses: actions/upload-artifact@ea165f8d65b6e75b540449e92b4886f43607fa02 # v4.6.2
;; with:
homepages/*) name: image-${{ matrix.name }}
add_service homepages path: image.json
;; if-no-files-found: error
edu_master/phpsessid-bot/*|edu_master/webinar-checker/*|edu_master/compose.yaml) retention-days: 30
add_service edu_master - name: Write the job result
;; if: always()
esac env:
done SUMMARY_CHECK: Image (${{ matrix.name }})
SUMMARY_RESULT: ${{ job.status }}
if [ "${#services[@]}" -eq 0 ]; then SUMMARY_FAILED_STEP: >-
echo "No docker-built services changed." ${{ steps.check.conclusion == 'failure' && 'Build or tag images' ||
echo "services=" >> "$GITHUB_OUTPUT" steps.artifact.conclusion == 'failure' && 'Artifact upload' ||
exit 0 steps.inputs.conclusion == 'failure' && 'Artifact download' ||
fi steps.source.conclusion == 'failure' && 'Source checkout' || '' }}
printf '%s\n' "${services[@]}" | tee /tmp/services.txt
echo "services=$(paste -sd, /tmp/services.txt)" >> "$GITHUB_OUTPUT"
- name: Log in to registry
# The pin step below also writes (manifest PUTs), and it runs on every
# main push — including manifest-only ones where services is empty. A
# stale persistent login on the old runner used to mask this; a clean
# runner pushes anonymously and gets 401.
if: steps.services.outputs.services != '' || github.ref_name == 'main'
shell: bash shell: bash
# Through env, not by substitution into the script. A secret written run: |
# into a run: block is pasted into the shell source before bash parses if [ -f .gitea/workflows/release.py ]; then
# it, so a password containing a quote, a backtick or $(...) becomes python3 .gitea/workflows/release.py check-summary
# code that runs. Masking the value in the log does not prevent that. elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
printf '## %s\n\n- Result: **%s**\n- Failed step: %s\n' "$SUMMARY_CHECK" "$SUMMARY_RESULT" "$SUMMARY_FAILED_STEP" >>"$GITHUB_STEP_SUMMARY" || true
fi
# Retain the build job name required by the immutable release deployment gate.
build:
needs: [image-plan, images]
runs-on: homelab
timeout-minutes: 15
steps:
- name: Checkout repository
id: source
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262
- name: Download all image results
id: inputs
uses: actions/download-artifact@d3f86a106a0bac45b974a628896c90dbdf5c8093 # v4.3.0
with:
path: artifacts
- name: Pin SHA tags and write the complete release
id: check
env: env:
REGISTRY_USERNAME: ${{ secrets.REGISTRY_USERNAME }} REGISTRY_USERNAME: ${{ secrets.REGISTRY_USERNAME }}
REGISTRY_PASSWORD: ${{ secrets.REGISTRY_PASSWORD }} REGISTRY_PASSWORD: ${{ secrets.REGISTRY_PASSWORD }}
run: | run: >-
set -euo pipefail python3 .gitea/workflows/release.py finalize
printf '%s' "$REGISTRY_PASSWORD" | docker login "${REGISTRY}" \ --plan artifacts/build-plan/build-plan.json
-u "$REGISTRY_USERNAME" \ - name: Store commit release
--password-stdin id: artifact
uses: actions/upload-artifact@ea165f8d65b6e75b540449e92b4886f43607fa02 # v4.6.2
- name: Build and push changed images with:
if: steps.services.outputs.services != '' name: release-${{ github.sha }}
path: release.json
if-no-files-found: error
retention-days: 30
- name: Write the job result
if: always()
env:
SUMMARY_CHECK: Image release and SHA tags
SUMMARY_RESULT: ${{ job.status }}
SUMMARY_FAILED_STEP: >-
${{ steps.check.conclusion == 'failure' && 'Build or tag images' ||
steps.artifact.conclusion == 'failure' && 'Artifact upload' ||
steps.inputs.conclusion == 'failure' && 'Artifact download' ||
steps.source.conclusion == 'failure' && 'Source checkout' || '' }}
shell: bash shell: bash
run: | run: |
# This step was the one run: block in the workflow without it, and it if [ -f .gitea/workflows/release.py ]; then
# is the one that cannot afford it: a docker push that failed partway python3 .gitea/workflows/release.py check-summary
# through the loop used to be followed by more pushes, the loop's exit elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
# status came from the last one, and the job went green with half the printf '## %s\n\n- Result: **%s**\n- Failed step: %s\n' "$SUMMARY_CHECK" "$SUMMARY_RESULT" "$SUMMARY_FAILED_STEP" >>"$GITHUB_STEP_SUMMARY" || true
# images missing from the registry.
set -euo pipefail
IFS=, read -r -a services <<< "${{ steps.services.outputs.services }}"
# Tags for this push. The commit-pinned name is the point of this
# step: the deploy resolves it in preference to :prod, so a deploy
# that sat in the queue behind a later push still gets the build of
# the commit CI validated, instead of whatever :prod points at by the
# time it runs. See render_pinned in deploy-lib.sh.
commit_tag=""
if [ "${GITHUB_REF_NAME}" = "main" ]; then
commit_tag="sha-${GITHUB_SHA:0:12}"
fi fi
set_tags() {
tags=()
case "${GITHUB_REF_NAME}" in
main) tags+=("main" "prod") ;;
dev) tags+=("dev") ;;
esac
if [ -n "$commit_tag" ]; then
tags+=("$commit_tag")
fi
}
for service in "${services[@]}"; do
case "$service" in
errorpages)
image="${REGISTRY}/forust/error-pages"
set_tags
build_args=()
for tag in "${tags[@]}"; do
build_args+=(-t "${image}:${tag}")
done
docker build \
--cache-from "type=registry,ref=${image}:buildcache" \
--cache-to "type=registry,ref=${image}:buildcache,mode=max" \
"${build_args[@]}" errorpages
for tag in "${tags[@]}"; do
docker push "${image}:${tag}"
done
;;
homepages)
for variant in forust xdfnx; do
case "$variant" in
forust)
image="${REGISTRY}/forust/forust-homepage"
;;
xdfnx)
image="${REGISTRY}/forust/xdfnx-homepage"
;;
esac
set_tags
build_args=()
for tag in "${tags[@]}"; do
build_args+=(-t "${image}:${tag}")
done
docker build \
--cache-from "type=registry,ref=${image}:buildcache" \
--cache-to "type=registry,ref=${image}:buildcache,mode=max" \
"${build_args[@]}" -f "homepages/Dockerfile.${variant}" homepages
for tag in "${tags[@]}"; do
docker push "${image}:${tag}"
done
done
;;
edu_master)
for variant in session-keeper webinar-checker; do
case "$variant" in
session-keeper)
context="edu_master/phpsessid-bot"
image="${REGISTRY}/forust/session-keeper"
;;
webinar-checker)
context="edu_master/webinar-checker"
image="${REGISTRY}/forust/webinar-checker"
;;
esac
set_tags
build_args=()
for tag in "${tags[@]}"; do
build_args+=(-t "${image}:${tag}")
done
docker build \
--cache-from "type=registry,ref=${image}:buildcache" \
--cache-to "type=registry,ref=${image}:buildcache,mode=max" \
"${build_args[@]}" "$context"
for tag in "${tags[@]}"; do
docker push "${image}:${tag}"
done
done
;;
esac
done
# Every image the tree names has to carry the commit-pinned name, not only
# the ones this push rebuilt. A push that touches nothing but manifests
# builds nothing, and its deploy would then find no commit-pinned tag to
# resolve and quietly fall back to the moving :prod - which is the whole
# failure the commit-pinned name exists to remove.
#
# Re-tagging copies the manifest list and transfers no layers, so pinning
# six images that already exist costs six registry writes.
#
# The list is derived from the tree rather than written out here, so an
# image added to a manifest is covered without a second place to update.
- name: Pin the commit name on the images this push did not rebuild
if: github.ref_name == 'main'
shell: bash
run: |
set -euo pipefail
commit_tag="sha-${GITHUB_SHA:0:12}"
mapfile -t repos < <(
git grep -hoE 'gcr\.forust\.xyz/forust/[A-Za-z0-9._-]+' -- '*.yaml' '*.yml' \
| sort -u
)
if [ "${#repos[@]}" -eq 0 ]; then
echo "No own images referenced by the tree."
exit 0
fi
echo "pinning ${#repos[@]} image(s) to $commit_tag"
for repo in "${repos[@]}"; do
if docker buildx imagetools inspect "$repo:$commit_tag" >/dev/null 2>&1; then
echo " already built by this push: ${repo##*/}"
continue
fi
if ! docker buildx imagetools inspect "$repo:prod" >/dev/null 2>&1; then
echo " WARNING: ${repo##*/} has no :prod to pin and no build produced it"
continue
fi
docker buildx imagetools create --tag "$repo:$commit_tag" "$repo:prod"
echo " pinned ${repo##*/}"
done
+103
View File
@@ -0,0 +1,103 @@
#!/usr/bin/env python3
"""Resolve Compose images without changing project names or local bind paths."""
import json
import os
import re
import subprocess
import sys
from pathlib import Path
def output(*args, **kwargs):
return subprocess.check_output(args, text=True, **kwargs).strip() # noqa: S603
def resolve(reference):
if '@sha256:' in reference:
return reference
descriptor = json.loads(
output('docker', 'buildx', 'imagetools', 'inspect', reference, '--format', '{{json .Manifest}}')
)
digest = descriptor['digest']
if not re.fullmatch(r'sha256:[0-9a-f]{64}', digest):
raise ValueError(f'Invalid registry digest for {reference}')
# Strip tag only from the final path segment (registry ports are preserved).
repository = reference.rsplit('/', 1)
repository[-1] = repository[-1].split(':')[0]
return '/'.join(repository) + '@' + digest
def prepare(source_file):
config_repo = Path(os.environ['CONFIG_REPO'])
source_repo = Path(os.environ['REPO'])
directory = Path(os.environ['RUN_DIR'])
relative = source_file.relative_to(source_repo)
project_dir = config_repo / relative.parent
base = ['docker', 'compose', '--project-directory', str(project_dir), '-f', str(source_file)]
config = json.loads(output(*base, 'config', '--format', 'json', cwd=config_repo))
project = config['name']
previous_file = directory / 'previous.json'
previous = json.loads(previous_file.read_text()) if previous_file.exists() else {}
images_file = directory / 'compose-images.json'
locks = json.loads(images_file.read_text()) if images_file.exists() else previous.get('compose-images', {})
release = json.loads((directory / 'release.json').read_text())
before = json.loads(json.dumps(config))
for service, settings in config['services'].items():
reference = settings.get('image')
nextcloud_aio_master = project == 'nextcloud' and service == 'nextcloud-aio-mastercontainer'
if not reference or settings.get('build'):
raise ValueError(f'{project}/{service}: Compose deploy requires a published image')
image_repo = reference.split('@')[0].rsplit('/', 1)
image_repo[-1] = image_repo[-1].split(':')[0]
image_repo = '/'.join(image_repo)
# Nextcloud AIO validates the mastercontainer image reference and rejects
# a digest. Keep its configured tag so AIO can start and manage its stack.
if nextcloud_aio_master:
pinned = reference
elif image_repo in release['images']:
pinned = image_repo + '@' + release['images'][image_repo]
elif os.environ.get('REFRESH_IMAGES') != 'true' and reference in locks:
pinned = locks[reference]
else:
pinned = resolve(reference)
settings['image'] = pinned
locks[reference] = pinned
# Capture what is running, not the current value of its mutable tag.
ids = output(
'docker',
'ps',
'-aq',
'--filter',
f'label=com.docker.compose.project={project}',
'--filter',
f'label=com.docker.compose.service={service}',
).splitlines()
actual = set()
for container in ids:
image_id = output('docker', 'inspect', container, '--format', '{{.Image}}')
digests = json.loads(output('docker', 'image', 'inspect', image_id, '--format', '{{json .RepoDigests}}'))
actual.add(next((d for d in digests or [] if d.split('@')[0] == image_repo), image_id))
if len(actual) > 1:
raise ValueError(f'{project}/{service}: mixed running images, cannot capture one recovery config')
# AIO also rejects a digest in its recovery config. Preserve its tag in
# both deploy and recovery files.
if nextcloud_aio_master:
before['services'][service]['image'] = reference
else:
before['services'][service]['image'] = next(iter(actual)) if actual else reference
for name, data in (('compose', config), ('compose-before', before)):
folder = directory / name
folder.mkdir(mode=0o700, exist_ok=True)
destination = folder / f'{relative.parent.name}.json'
destination.write_text(json.dumps(data, indent=2) + '\n')
destination.chmod(0o600)
images_file.write_text(json.dumps(locks, indent=2) + '\n')
print(f'Compose {project}: images pinned; local paths preserved')
print(
f'Recovery: docker compose --project-directory {project_dir} -p {project} -f {directory}/compose-before/{relative.parent.name}.json up -d --pull never'
)
if __name__ == '__main__':
prepare(Path(sys.argv[1]))
+423
View File
@@ -0,0 +1,423 @@
#!/usr/bin/env python3
"""Durable workstation deployment controller. Install with setup-workstation.sh."""
import argparse
import contextlib
import fcntl
import importlib.util
import json
import math
import os
import re
import shutil
import subprocess
import sys
import time
from pathlib import Path
STATE = Path(os.environ.get('HOMELAB_STATE', Path.home() / '.local/state/homelab-deploy'))
CONFIG_REPO = Path(os.environ.get('HOMELAB_REPO', '/srv/homelab'))
RUN_ID = re.compile(r'[0-9]+-[0-9]+')
def command(*args, **kwargs):
return subprocess.check_output(args, text=True, **kwargs).strip() # noqa: S603, S607
def atomic_json(path, data):
temporary = path.with_suffix('.tmp')
temporary.write_text(json.dumps(data, indent=2) + '\n')
temporary.chmod(0o600)
temporary.replace(path)
@contextlib.contextmanager
def lock(name):
STATE.mkdir(mode=0o700, parents=True, exist_ok=True)
with (STATE / name).open('a') as stream:
fcntl.flock(stream, fcntl.LOCK_EX)
yield
def load_module(name, path):
spec = importlib.util.spec_from_file_location(name, path)
module = importlib.util.module_from_spec(spec)
spec.loader.exec_module(module)
return module
def run_directory(run_id):
if not RUN_ID.fullmatch(run_id):
raise ValueError('Run ID must be numeric workflow-id and attempt')
return STATE / 'runs' / run_id
def start(run_id):
payload = sys.stdin.buffer.read(256 * 1024 + 1)
if len(payload) > 256 * 1024:
raise ValueError('Deploy request exceeds 256 KiB')
request = json.loads(payload)
sha = request['release']['sha']
if not re.fullmatch(r'[0-9a-f]{40}', sha) or request['mode'] not in ('changed', 'full', 'plan'):
raise ValueError('Invalid deploy SHA or mode')
if not isinstance(request['refresh_images'], bool):
raise ValueError('refresh_images must be boolean')
directory = run_directory(run_id)
with lock('prepare.lock'):
if (directory / 'request.json').exists():
if json.loads((directory / 'request.json').read_text()) != request:
raise ValueError('Run ID already belongs to a different request')
else:
directory.mkdir(mode=0o700, parents=True, exist_ok=True)
command('git', '-C', str(CONFIG_REPO), 'fetch', '--quiet', 'origin', 'main')
command('git', '-C', str(CONFIG_REPO), 'merge-base', '--is-ancestor', sha, 'origin/main')
if not (directory / 'source').exists():
command('git', '-C', str(CONFIG_REPO), 'worktree', 'add', '--detach', str(directory / 'source'), sha)
if command('git', '-C', str(directory / 'source'), 'rev-parse', 'HEAD') != sha:
raise ValueError('Prepared source does not match deploy SHA')
release_module = load_module('release', directory / 'source/.gitea/workflows/release.py')
release_module.validate_release(request['release'], sha)
atomic_json(directory / 'release.json', request['release'])
atomic_json(directory / 'request.json', request)
if not (directory / 'status.json').exists():
atomic_json(directory / 'status.json', {'state': 'queued', 'stages': {}})
# Starting an existing active or finished ID is idempotent; never re-apply it.
if json.loads((directory / 'status.json').read_text())['state'] == 'queued':
command('systemctl', '--user', 'start', '--no-block', f'homelab-deploy@{run_id}.service')
print(f'Accepted deploy {run_id} ({sha})')
def environment(directory):
request = json.loads((directory / 'request.json').read_text())
return {
**os.environ,
'REPO': str(directory / 'source'),
'CONFIG_REPO': str(CONFIG_REPO),
'RUN_DIR': str(directory),
'DEPLOY_SHA': request['release']['sha'],
'RELEASE_FILE': str(directory / 'release.json'),
'DEPLOY_PLAN': str(directory / 'plan.json'),
'DEPLOY_SNAPSHOT_DIR': str(directory / 'snapshot'),
'REFRESH_IMAGES': str(request['refresh_images']).lower(),
'ROLLOUT_PARALLELISM': '4',
}
def stage(directory, name, budget):
status = json.loads((directory / 'status.json').read_text())
if name in status['stages'] and status['stages'][name].get('result') in ('success', 'failure'):
return status['stages'][name]['result'] == 'success'
started = time.time()
status['stages'][name] = {'result': 'running', 'started': started}
atomic_json(directory / 'status.json', status)
script = directory / 'source/.gitea/workflows/deploy-stage.sh'
with (directory / f'{name}.log').open('a') as log:
# timeout kills the whole stage process group, including children, before recovery.
result = subprocess.run( # noqa: S603, S607
[
shutil.which('timeout') or '/usr/bin/timeout',
'--signal=TERM',
'--kill-after=30s',
str(budget),
'bash',
str(script),
name,
],
env=environment(directory),
stdout=log,
stderr=subprocess.STDOUT,
check=False,
).returncode
status = json.loads((directory / 'status.json').read_text())
status['stages'][name].update(
result='success' if result == 0 else 'failure', exit_code=result, seconds=round(time.time() - started)
)
atomic_json(directory / 'status.json', status)
return result == 0
def make_plan(directory):
source = directory / 'source'
planner = load_module('deploy_plan', source / '.gitea/workflows/deploy-plan.py')
request = json.loads((directory / 'request.json').read_text())
previous = json.loads((STATE / 'last-success.json').read_text()) if (STATE / 'last-success.json').exists() else None
# Helm 4 lists every release status by default and removed the --all flag.
helm = json.loads(command('helm', 'list', '-A', '-o', 'json'))
plan = planner.make_plan(source, CONFIG_REPO, request['release'], previous, request['mode'], helm)
if request['refresh_images']:
plan['selected']['compose'] = plan['active']['compose']
atomic_json(directory / 'plan.json', plan)
if previous:
atomic_json(directory / 'previous.json', previous)
# Local config is deliberately separate from the immutable Git source.
return plan
def finish_success(directory, plan):
# Repeating finalization after a crash is safe while holding deploy.lock.
plan['run_id'] = directory.name
path = directory / 'compose-images.json'
previous = directory / 'previous.json'
plan['compose-images'] = (
json.loads(path.read_text())
if path.exists()
else json.loads(previous.read_text()).get('compose-images', {})
if previous.exists()
else {}
)
atomic_json(STATE / 'last-success.json', plan)
status = json.loads((directory / 'status.json').read_text())
status['state'] = 'success'
atomic_json(directory / 'status.json', status)
try:
retain_completed(directory)
except (OSError, subprocess.CalledProcessError) as error:
print(f'Retention deferred: {error}', flush=True)
def recover(directory, retry=False):
status = json.loads((directory / 'status.json').read_text())
if status['state'] in ('success', 'planned'):
return
completed = ('doctor', 'validate', 'apply-k8s', 'apply-compose', 'verify-k8s', 'smoke')
if all(status['stages'].get(name, {}).get('result') == 'success' for name in completed):
finish_success(directory, json.loads((directory / 'plan.json').read_text()))
return
if retry:
for name in ('verify-k8s', 'smoke'):
if status['stages'].get(name, {}).get('result') == 'failure':
del status['stages'][name]
atomic_json(directory / 'status.json', status)
snapshot = directory / 'snapshot/current'
if snapshot.exists():
stage(directory, 'verify-k8s', 7200)
stage(directory, 'smoke', 600)
status = json.loads((directory / 'status.json').read_text())
status['state'] = 'failure'
atomic_json(directory / 'status.json', status)
def execute(run_id):
directory = run_directory(run_id)
with lock('deploy.lock'):
status = json.loads((directory / 'status.json').read_text())
if status['state'] != 'queued':
return
# A crashed predecessor must be recovered before another apply begins.
for other in (STATE / 'runs').iterdir():
if (
other != directory
and (other / 'status.json').exists()
and json.loads((other / 'status.json').read_text())['state'] == 'running'
):
raise ValueError(f'Interrupted deploy {other.name}; run recover first')
status['state'] = 'running'
atomic_json(directory / 'status.json', status)
phase = 'plan'
try:
plan = make_plan(directory)
print(
json.dumps({'selected': plan['selected'], 'helm': plan['helm'], 'manual_removals': plan['removed']}),
flush=True,
)
phase = 'doctor'
if not stage(directory, 'doctor', 600):
raise RuntimeError('Preflight failed')
phase = 'validate'
if not stage(directory, 'validate', 1200):
raise RuntimeError('Validation failed')
if json.loads((directory / 'request.json').read_text())['mode'] == 'plan':
status = json.loads((directory / 'status.json').read_text())
status['state'] = 'planned'
atomic_json(directory / 'status.json', status)
return
# Budget includes both rollout checks and rollback waves, plus API overhead.
phase = 'Recovery budget'
count = int(
command(
'bash',
str(directory / 'source/.gitea/workflows/deploy-stage.sh'),
'workload-count',
env=environment(directory),
)
)
verify_budget = max(600, 2 * math.ceil(count / 4) * 300 + 120)
if verify_budget > 7200:
raise ValueError('More than two hours of recovery required; split this deploy')
phase = 'apply-k8s'
k8s_ok = stage(directory, 'apply-k8s', 2700)
phase = 'apply-compose'
compose_ok = stage(directory, 'apply-compose', 1800) if k8s_ok else False
phase = 'verify-k8s'
verify_ok = stage(directory, 'verify-k8s', verify_budget)
phase = 'smoke'
smoke_ok = stage(directory, 'smoke', 600)
if not all((k8s_ok, compose_ok, verify_ok, smoke_ok)):
raise RuntimeError('Deploy failed; inspect stage logs and recovery report')
phase = 'Save the successful baseline'
finish_success(directory, plan)
except Exception as error:
status = json.loads((directory / 'status.json').read_text())
status['failure_stage'] = next(
(name for name, result in status['stages'].items() if result.get('result') == 'failure'), phase
)
atomic_json(directory / 'status.json', status)
with (directory / 'controller.log').open('a') as stream:
stream.write(f'{error}\n')
recover(directory)
raise
def retain_completed(current):
finished = []
for directory in (STATE / 'runs').iterdir():
status_file = directory / 'status.json'
if status_file.exists() and json.loads(status_file.read_text())['state'] in ('success', 'planned'):
finished.append(directory)
for directory in sorted(finished, key=lambda p: p.stat().st_mtime, reverse=True)[20:]:
if directory == current:
continue
command('git', '-C', str(CONFIG_REPO), 'worktree', 'remove', '--force', str(directory / 'source'))
shutil.rmtree(directory)
def follow(run_id, phase):
directory = run_directory(run_id)
groups = {
'apply': ('doctor', 'validate', 'apply-k8s', 'apply-compose'),
'verify': ('verify-k8s',),
'smoke': ('smoke',),
}
names = groups[phase]
offsets = {}
while True:
status = json.loads((directory / 'status.json').read_text())
for name in (*names, 'controller'):
path = directory / f'{name}.log'
if path.exists():
with path.open() as stream:
stream.seek(offsets.get(name, 0))
content = stream.read()
if content:
print(content, end='', flush=True)
offsets[name] = stream.tell()
stages = status['stages']
if all(stages.get(name, {}).get('result') in ('success', 'failure') for name in names):
return all(stages[name]['result'] == 'success' for name in names)
if status['state'] in ('success', 'failure', 'planned'):
return status['state'] in ('success', 'planned')
time.sleep(3)
def summary(run_id):
directory = run_directory(run_id)
request = json.loads((directory / 'request.json').read_text())
release = request['release']
plan_file = directory / 'plan.json'
lines = [
f'## Deploy `{release["sha"]}`',
'',
f'- Mode: `{request["mode"]}`',
f'- Refresh third-party images: `{request["refresh_images"]}`',
]
status = json.loads((directory / 'status.json').read_text())
if status.get('failure_stage'):
lines.append(f'- Failed stage: **{status["failure_stage"]}**')
lines.extend(
[
'',
f'- Observed run state: **{status["state"]}**',
'',
'### Stage results',
'| Stage | Result | Exit code |',
'| --- | --- | --- |',
]
)
for name in ('doctor', 'validate', 'apply-k8s', 'apply-compose', 'verify-k8s', 'smoke'):
stage_result = status['stages'].get(name, {})
lines.append(f'| {name} | {stage_result.get("result", "not started")} | {stage_result.get("exit_code", "—")} |')
lines.extend(['', '### Apply and Helm recovery results'])
events_file = directory / 'apply-events.jsonl'
events = []
if events_file.exists():
for line in events_file.read_text().splitlines():
try:
events.append(json.loads(line))
except json.JSONDecodeError:
lines.append('- An operation record is incomplete. Check the stage log.')
latest = {(event['action'], event['target']): event['result'] for event in events}
lines.extend(f'- `{action}` `{target}`: **{result}**' for (action, target), result in latest.items())
if not latest:
lines.append('- No apply results were recorded.')
lines.append('- A completed apply does not confirm health. See verification and smoke results.')
lines.extend(['', '### Kubernetes recovery'])
pointer = directory / 'snapshot/current'
failed = Path(pointer.read_text().strip()) / 'failed-workloads' if pointer.exists() else None
if failed and failed.exists():
contents = failed.read_text()
counts = dict(re.findall(r'^(ROLLED_BACK|UNRECOVERED)=([0-9]+)$', contents, re.MULTILINE))
if not contents.strip():
lines.append('- No failed workloads were recorded. See the verification result above.')
elif counts:
lines.append(f'- Workloads restored: **{counts.get("ROLLED_BACK", "unknown")}**')
lines.append(f'- Workloads that need manual recovery: **{counts.get("UNRECOVERED", "unknown")}**')
else:
lines.append('- Rollback has no recorded result yet. Check the verification log.')
else:
lines.append('- No workload rollback was recorded. This does not confirm health.')
lines.append('- Compose requires manual recovery. Use the saved command in the apply log.')
if not plan_file.exists():
lines.extend(['', 'Plan was not created. Check the controller log.'])
print('\n'.join(lines))
return
plan = json.loads(plan_file.read_text())
lines.extend(['', '### Selected services'])
count = 0
for kind, services in plan['selected'].items():
for service in services:
lines.append(f'- `{kind}`: `{service}`')
count += 1
if not count:
lines.append('- None')
lines.extend(['', '### Selected Helm releases'])
lines.extend(f'- `{release}`' for release in plan.get('helm', []))
if not plan.get('helm'):
lines.append('- None')
lines.extend(['', '### Images pinned in the checked release'])
lines.extend(f'- `{image}@{digest}`' for image, digest in sorted(release['images'].items()))
lines.extend(['', '### Removed resources requiring manual review'])
lines.extend(f'- `{item}`' for item in plan.get('removed', []))
if not plan.get('removed'):
lines.append('- None')
print('\n'.join(lines))
def main():
os.umask(0o077)
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument('action', choices=('start', 'execute', 'recover', 'status', 'follow', 'summary'))
parser.add_argument('run_id')
parser.add_argument('phase', nargs='?', choices=('apply', 'verify', 'smoke'))
parser.add_argument('--retry', action='store_true', help='Retry failed recovery checks; never repeat apply')
args = parser.parse_args()
directory = run_directory(args.run_id)
if args.action == 'start':
start(args.run_id)
elif args.action == 'execute':
execute(args.run_id)
elif args.action == 'recover':
with lock('deploy.lock'):
recover(directory, retry=args.retry)
elif args.action == 'status':
print((directory / 'status.json').read_text())
if (directory / 'plan.json').exists():
plan = json.loads((directory / 'plan.json').read_text())
print(json.dumps({k: plan[k] for k in ('sha', 'selected', 'helm', 'removed')}, indent=2))
elif args.action == 'summary':
summary(args.run_id)
elif not follow(args.run_id, args.phase):
sys.exit(1)
if __name__ == '__main__':
main()
+286 -409
View File
@@ -1,5 +1,5 @@
#!/usr/bin/env bash #!/usr/bin/env bash
# Shared stages for the deploy workflow. Runs on the workstation, invoked as: # Workstation deploy stages; invoked by the durable controller against pinned source.
# REPO=/srv/homelab APPLY_PRUNE=false bash -se <<'EOF' # REPO=/srv/homelab APPLY_PRUNE=false bash -se <<'EOF'
# source "$REPO/.gitea/workflows/deploy-lib.sh" # source "$REPO/.gitea/workflows/deploy-lib.sh"
# run_stage "$STAGE" # run_stage "$STAGE"
@@ -8,8 +8,8 @@ set -euo pipefail
: "${REPO:?REPO must be set}" : "${REPO:?REPO must be set}"
APPLY_PRUNE="${APPLY_PRUNE:-false}" APPLY_PRUNE="${APPLY_PRUNE:-false}"
# Commit CI validated. Empty for a manual workflow_dispatch, which falls back to CONFIG_REPO="${CONFIG_REPO:-$REPO}"
# the current origin/main. # Exact SHA accepted by the CI gate for both manual and automatic deploys.
DEPLOY_SHA="${DEPLOY_SHA:-}" DEPLOY_SHA="${DEPLOY_SHA:-}"
# Handoff point between the apply stage (writes) and the verify stage (reads). # Handoff point between the apply stage (writes) and the verify stage (reads).
# Under the deploy user's own XDG state directory rather than /var/backups: the # Under the deploy user's own XDG state directory rather than /var/backups: the
@@ -20,17 +20,36 @@ DEPLOY_SNAPSHOT_DIR="${DEPLOY_SNAPSHOT_DIR:-${XDG_STATE_HOME:-$HOME/.local/state
# Per-workload rollout budget and how many workloads to watch at once. The whole # Per-workload rollout budget and how many workloads to watch at once. The whole
# apply job has its own timeout-minutes as a backstop. # apply job has its own timeout-minutes as a backstop.
ROLLOUT_TIMEOUT="${ROLLOUT_TIMEOUT:-300}" ROLLOUT_TIMEOUT="${ROLLOUT_TIMEOUT:-300}"
ROLLOUT_PARALLELISM="${ROLLOUT_PARALLELISM:-8}" ROLLOUT_PARALLELISM="${ROLLOUT_PARALLELISM:-4}"
WORKLOAD_KINDS="deployments.apps,statefulsets.apps,daemonsets.apps" WORKLOAD_KINDS="deployments.apps,statefulsets.apps,daemonsets.apps"
log() { log() {
echo "== $* ==" echo "== $* =="
} }
# Store operation results without command output or local configuration values.
record_apply() {
[ -n "${RUN_DIR:-}" ] || return 0
jq -cn --arg action "$1" --arg target "$2" --arg result "$3" \
'{action: $action, target: $target, result: $result}' >>"$RUN_DIR/apply-events.jsonl" \
|| echo 'WARNING: cannot record an apply result' >&2
return 0
}
warn() { warn() {
echo "WARNING: $*" >&2 echo "WARNING: $*" >&2
} }
# Prune needs the complete desired set in one invocation. Per-file pruning
# treats resources from the other files as absent and can delete them.
check_prune_mode() {
if [ "$APPLY_PRUNE" = "true" ]; then
echo "ERROR: APPLY_PRUNE=true is unsupported by the per-file deploy loop." >&2
echo "Disable it; remove obsolete resources explicitly after review." >&2
return 1
fi
}
collect_k8s() { collect_k8s() {
git -C "$REPO" ls-files -- "$1" \ git -C "$REPO" ls-files -- "$1" \
| grep -E '\.ya?ml$' \ | grep -E '\.ya?ml$' \
@@ -50,6 +69,25 @@ kustomize_overlay() {
fi fi
} }
selected_service() {
local kind="$1" service="$2" section=selected
[ -n "${DEPLOY_PLAN:-}" ] || return 0
if [ "${DEPLOY_SMOKE_ALL:-false}" = true ]; then section=active; fi
jq -e --arg kind "$kind" --arg service "$service" --arg section "$section" \
'.[$section][$kind] | index($service) != null' "$DEPLOY_PLAN" >/dev/null
}
# Resolve .env and relative binds on the persistent workstation tree. Locked
# JSON configs keep the same Compose project name and volume names.
compose() {
local cf="$1" locked project_dir
project_dir="$CONFIG_REPO/$(basename "$(dirname "$cf")")"
shift
locked="${RUN_DIR:-/nonexistent}/compose/$(basename "$(dirname "$cf")").json"
if [ -f "$locked" ]; then cf="$locked"; fi
(cd "$CONFIG_REPO" && docker compose --project-directory "$project_dir" -f "$cf" "$@")
}
select_manifests() { select_manifests() {
K8S_MANIFESTS=() K8S_MANIFESTS=()
KUSTOMIZE_APPS=() KUSTOMIZE_APPS=()
@@ -57,6 +95,7 @@ select_manifests() {
local kd_rel kd overlay cf_rel cf f local kd_rel kd overlay cf_rel cf f
while IFS= read -r kd_rel; do while IFS= read -r kd_rel; do
kd="$REPO/$kd_rel" kd="$REPO/$kd_rel"
selected_service k8s "${kd_rel%/k8s}" || continue
if [ ! -f "$kd/active" ]; then if [ ! -f "$kd/active" ]; then
echo "skip (no k8s/active): $kd_rel" echo "skip (no k8s/active): $kd_rel"
continue continue
@@ -78,6 +117,7 @@ select_manifests() {
) )
while IFS= read -r cf_rel; do while IFS= read -r cf_rel; do
cf="$REPO/$cf_rel" cf="$REPO/$cf_rel"
selected_service compose "$(dirname "$cf_rel")" || continue
if [ -f "$(dirname "$cf")/active" ]; then if [ -f "$(dirname "$cf")/active" ]; then
echo "compose: $cf_rel" echo "compose: $cf_rel"
COMPOSE_STACKS+=("$cf") COMPOSE_STACKS+=("$cf")
@@ -94,16 +134,10 @@ select_manifests() {
# on failure roll them back to the revision that was running before, so a bad # on failure roll them back to the revision that was running before, so a bad
# push to main cannot leave a service crash-looping. # push to main cannot leave a service crash-looping.
# #
# Verification lives in its own workflow job, not at the end of the apply stage. # The workstation controller runs apply and verification as separate durable
# Inside a single process it is worthless exactly when it is needed most: a job # stages. Runner jobs only follow their logs. ExecStopPost recovers interrupted
# killed by timeout-minutes or cancelled mid-apply never reaches the rollback # runs using the per-run snapshot, even when the SSH connection has gone away.
# code, and leaves a half-applied cluster behind. Split out, the apply job can
# die in any way and the verify job still runs.
# #
# That split needs a handoff point on the workstation, because the two stages are
# separate processes on separate runner jobs: DEPLOY_SNAPSHOT_DIR/current, written
# before anything is applied, read by the verify stage afterwards.
# Creates this run's snapshot directory and publishes it as the handoff point for # Creates this run's snapshot directory and publishes it as the handoff point for
# the verify stage. Fails hard by design: a deploy that cannot record what it is # the verify stage. Fails hard by design: a deploy that cannot record what it is
# about to change must not start, because then nothing can be rolled back for it # about to change must not start, because then nothing can be rolled back for it
@@ -133,18 +167,36 @@ snapshot_dir() {
} }
save_snapshot() { save_snapshot() {
local dir="$1" local dir="$1" releases revision status
log "Saving pre-apply snapshot to $dir" log "Saving pre-apply snapshot to $dir"
workload_generations >"$dir/generations.before" 2>/dev/null \ workload_generations >"$dir/generations.before" || return 1
|| warn "could not snapshot workload generations" kubectl get "$WORKLOAD_KINDS" -A -o json >"$dir/workloads.json" || return 1
kubectl get "$WORKLOAD_KINDS" -A -o yaml >"$dir/workloads.yaml" 2>/dev/null \ kubectl get controllerrevisions.apps -A -o json >"$dir/controller-revisions.json" || return 1
|| warn "could not snapshot workloads" jq --slurpfile revisions "$dir/controller-revisions.json" '
for release in prometheus-stack loki alloy; do [.items[] | . as $w | {
if helm status "$release" -n prometheus >/dev/null 2>&1; then kind: (.kind | ascii_downcase), namespace: .metadata.namespace, name: .metadata.name, uid: .metadata.uid,
{ revision: (if .kind == "Deployment" then (.metadata.annotations["deployment.kubernetes.io/revision"] // "0" | tonumber)
echo "revision: $(helm history "$release" -n prometheus -o json 2>/dev/null)" else ([$revisions[0].items[] | select(.metadata.namespace == $w.metadata.namespace)
helm get values "$release" -n prometheus --all 2>/dev/null | select(any(.metadata.ownerReferences[]?; .uid == $w.metadata.uid))
} >"$dir/helm-$release.txt" | select($w.kind != "StatefulSet" or .metadata.name == $w.status.currentRevision) | .revision] | max // 0) end)
}]' "$dir/workloads.json" >"$dir/revisions.json" || return 1
# Helm 4 lists every release status by default and removed the --all flag.
releases="$(helm list -A -o json)" || return 1
for entry in "${HELM_RELEASES[@]}"; do
IFS='|' read -r release _ namespace _ _ _ <<<"$entry"
if ! jq -e --arg r "$release" --arg n "$namespace" \
'any(.[]; .name == $r and .namespace == $n)' <<<"$releases" >/dev/null; then
continue
fi
helm status "$release" -n "$namespace" -o json >"$dir/helm-$release.json" || return 1
status="$(jq -r '.info.status' "$dir/helm-$release.json")"
if [ "$status" != deployed ]; then
# Never capture a pending/failed revision as the recovery target.
helm history "$release" -n "$namespace" -o json >"$dir/helm-$release.history.json" || return 1
revision="$(jq '[.[] | select(.status == "deployed" or .status == "superseded") | .revision] | max // 0' \
"$dir/helm-$release.history.json")"
jq --argjson revision "$revision" '.version = $revision' "$dir/helm-$release.json" >"$dir/helm-$release.tmp"
mv "$dir/helm-$release.tmp" "$dir/helm-$release.json"
fi fi
done done
# The verify stage compares this against the commit it is deploying, to refuse # The verify stage compares this against the commit it is deploying, to refuse
@@ -168,294 +220,24 @@ workload_generations() {
# moved since the snapshot, i.e. the ones this apply actually touched. # moved since the snapshot, i.e. the ones this apply actually touched.
changed_workloads() { changed_workloads() {
local before="$1" local before="$1"
local ns name kind gen old local ns name kind gen old current
current="$(workload_generations)" || return 1
while read -r ns name kind gen; do while read -r ns name kind gen; do
[ -n "${gen:-}" ] || continue [ -n "${gen:-}" ] || continue
old="$(awk -v want_ns="$ns" -v want_name="$name" \ old="$(awk -v want_ns="$ns" -v want_name="$name" -v want_kind="$kind" \
'$1 == want_ns && $2 == want_name { print $4; exit }' "$before" 2>/dev/null || true)" '$1 == want_ns && $2 == want_name && $3 == want_kind { print $4; exit }' "$before" 2>/dev/null || true)"
if [ -n "${RUN_DIR:-}" ] && ! grep -qxF "$kind $ns $name" "$RUN_DIR/workload-refs"; then
continue
fi
if [ "$old" != "$gen" ]; then if [ "$old" != "$gen" ]; then
printf '%s %s %s\n' "$kind" "$ns" "$name" printf '%s %s %s\n' "$kind" "$ns" "$name"
fi fi
done < <(workload_generations) done <<<"$current"
} }
# Prints "<ns> <kind>/<name> <image>" for every workload this repository owns that # Resolve owned image references exclusively from the checked CI artifact.
# runs an image from our own registry.
#
# The repository is the scope, deliberately. The cluster also holds workloads on
# our registry that no manifest here declares (they are applied out of band), and
# those are somebody else's to deploy. Walking the manifests rather than the
# cluster means those can never be restarted by this pipeline, now or later.
owned_registry_workloads() {
local kd_rel f
while IFS= read -r kd_rel; do
[ -f "$REPO/$kd_rel/active" ] || continue
while IFS= read -r f; do
[ -n "$f" ] || continue
# A file that does not mention the registry cannot declare a workload on it,
# and parsing costs ~2.5s per file against a millisecond for the grep. The
# filter keeps this at a handful of parses instead of one per manifest.
grep -q 'gcr\.forust\.xyz/forust/' "$REPO/$f" 2>/dev/null || continue
# kubectl prints a bare object for a single-document file and a List for a
# multi-document one, so normalise both shapes before filtering.
kubectl apply --dry-run=client -f "$REPO/$f" -o json 2>/dev/null \
| jq -r '
(if .items then .items[] else . end)
| select(.kind | test("^(Deployment|StatefulSet|DaemonSet)$"))
| select(any((.spec.template.spec.containers // [])[]?;
(.image // "") | test("^gcr\\.forust\\.xyz/forust/")))
| (.metadata.namespace // "default") as $ns
| ([.spec.template.spec.containers[].image
| select(test("^gcr\\.forust\\.xyz/forust/"))][0]) as $img
| "\($ns) \(.kind | ascii_downcase)/\(.metadata.name) \($img)"
' 2>/dev/null || true
done < <(collect_k8s "$kd_rel" || true)
done < <(
git -C "$REPO" ls-files '*.yaml' '*.yml' \
| grep -E '(^|/)k8s/' \
| sed -E 's#((^|.*/)k8s)/.*#\1#' \
| sort -u
)
}
# Prints the digest an image tag resolves to for this cluster's architecture, or
# nothing when it cannot be resolved.
#
# Only the manifest entry matching the node architecture counts. A multi-arch tag
# also carries `unknown/unknown` entries for the build attestation, and a pod's
# imageID is always the per-platform digest, so comparing the wrong entry would
# mark every workload stale forever and restart the whole cluster on every deploy.
registry_digest() {
local arch
arch="$(kubectl get nodes -o jsonpath='{.items[0].status.nodeInfo.architecture}' 2>/dev/null || true)"
[ -n "$arch" ] || arch=amd64
# The || true is load-bearing. Every caller runs under set -euo pipefail, and
# pipefail reports the rightmost non-zero stage, so a ref the registry does not
# have would abort the caller at the assignment instead of yielding an empty
# string. The callers check for empty themselves and report it by name.
#
# Retried with a hard timeout because the registry has a known hang mode (and
# a known blink mode: a single failed lookup aborts the whole apply file in
# render_pinned). A short sleep between attempts lets a restarting registry
# come back instead of failing the deploy on one bad second.
local attempt=0 digest=""
while [ "$attempt" -lt 3 ]; do
digest="$(timeout 25s docker manifest inspect "$1" 2>/dev/null \
| jq -r --arg arch "$arch" '
.manifests[]?
| select(.platform.os == "linux" and .platform.architecture == $arch)
| .digest
' 2>/dev/null \
| head -1 || true)"
[ -n "$digest" ] && break
attempt=$((attempt + 1))
if [ "$attempt" -lt 3 ]; then
echo "WARNING: registry lookup of $1 failed (attempt $attempt/3), retrying in 5s" >&2
sleep 5
fi
done
printf '%s' "$digest"
}
# The commit this deploy is for: what CI validated, or - on a manual dispatch,
# whatever stage_preflight just checked out.
deploy_commit() {
local c="${DEPLOY_SHA:-}"
[ -n "$c" ] || c="$(git -C "$REPO" rev-parse HEAD 2>/dev/null || true)"
printf '%.12s' "${c:-}"
}
# Resolves one of our image refs to the digest THIS commit's build produced.
#
# A manifest naming `:prod` names a pointer, not a version, and the deploy
# resolves it when the apply runs - which is not when CI ran it. Deploy runs are
# queued rather than cancelled (see deploy.yaml), so two pushes in a row leave
# the first deploy resolving the second push's build: the right manifests with
# the wrong code, and nothing anywhere reports it. ci therefore publishes every
# image it ships under `sha-<commit12>`, a name that cannot move, and that is
# the name resolved here.
#
# The fallback to the plain tag is for an image this pipeline never built. It
# reports itself, because a fallback nobody sees is the failure this removes.
pinned_digest() {
local ref="$1" commit pinned
commit="$(deploy_commit)"
if [ -n "$commit" ]; then
pinned="$(registry_digest "${ref%:*}:sha-$commit")"
if [ -n "$pinned" ]; then
printf '%s' "$pinned"
return 0
fi
fi
pinned="$(registry_digest "$ref")"
if [ -n "$pinned" ]; then
echo "WARNING: ${ref} carries no sha-${commit:-<unknown>} tag; resolved the moving tag instead" >&2
fi
printf '%s' "$pinned"
}
# Rewrites our own images to immutable digests on the way into the cluster.
# Reads a manifest stream on stdin, writes the pinned stream to stdout.
#
# A digest is not knowable when a manifest is written, so it is never committed:
# git keeps a readable `:prod` tag and the exact bytes are chosen here, at apply
# time, from the tag ci published for the commit being deployed. That is what
# makes rollback mean something. `kubectl rollout undo` restores the previous
# ReplicaSet's pod template verbatim, and a template naming a digest restores the
# exact bytes that were serving before. A template naming a moving tag does not —
# the tag has already moved by the time the rollback runs, so the "rollback"
# re-pulls the very image that just failed and the cluster stays broken.
#
# imagePullPolicy is deliberately left alone. The manifests no longer set it, and a
# reference that is not `:latest` defaults to IfNotPresent, which is what the
# Kubernetes docs ask for alongside a digest: the bytes under a digest cannot
# change, so pulling again buys nothing.
#
# An image that cannot be resolved is fatal. Carrying on would quietly apply a
# mutable tag again, which is the exact failure this function exists to remove.
render_pinned() { render_pinned() {
local src refs map ref digest missing=0 python3 "$REPO/.gitea/workflows/release.py" render
src="$(mktemp)"
refs="$(mktemp)"
map="$(mktemp)"
cat >"$src"
grep -oE 'gcr\.forust\.xyz/forust/[A-Za-z0-9._-]+:[A-Za-z0-9._-]+' "$src" | sort -u >"$refs" || true
while read -r ref; do
[ -n "$ref" ] || continue
digest="$(pinned_digest "$ref")"
if [ -z "$digest" ]; then
echo "ERROR: cannot resolve ${ref} in the registry; applying nothing." >&2
echo " The build job has to push that tag before the deploy resolves it." >&2
missing=$((missing + 1))
continue
fi
printf '%s\t%s\n' "$ref" "$digest" >>"$map"
done <"$refs"
if [ "$missing" -gt 0 ]; then
rm -f "$src" "$refs" "$map"
return 1
fi
awk -v mapfile="$map" '
BEGIN {
while ((getline line < mapfile) > 0) {
i = index(line, "\t")
d[substr(line, 1, i - 1)] = substr(line, i + 1)
}
}
{
if (match($0, /^[[:space:]]*image:[[:space:]]*gcr\.forust\.xyz\/forust\/[A-Za-z0-9._-]+:[A-Za-z0-9._-]+[[:space:]]*$/)) {
name = $0
sub(/^[[:space:]]*image:[[:space:]]*/, "", name)
sub(/[[:space:]]*$/, "", name)
if (name in d) {
pad = $0
sub(/image:.*/, "", pad)
# Drop the tag: the canonical form used in the docs is repo@sha256:...,
# and leaving :prod next to the digest reads like it still matters.
repo = name
sub(/:[A-Za-z0-9._-]+$/, "", repo)
print pad "image: " repo "@" d[name]
next
}
}
print
}
' "$src"
rm -f "$src" "$refs" "$map"
}
# Restarts every owned workload whose running image is not the one its tag
# resolves to now.
#
# This used to be how a rebuild reached the cluster at all: the manifests pinned
# `:latest`, so a rebuild left the pod template byte-identical, `kubectl apply`
# decided there was nothing to do, and the cluster served the previous build
# indefinitely. The apply now pins digests via render_pinned, so a rebuild moves
# the pod template and rolls out on its own.
#
# What is left is the drift check: a hand-run `kubectl set image`, or anything
# else that edits a live workload behind the deploy's back, is the only way to end
# up serving a digest the tag has moved past. It stays idempotent, so a redeploy
# that changed no image still does not bounce healthy services.
#
# The container is matched on its repository rather than on the exact reference:
# once render_pinned has run, a pod's status reports `repo@sha256:...` while this
# still reads the repository's `:prod` tag out of the manifest.
restart_stale_images() {
local ns target image want selector running entry one
local unchecked=0
local -A digests=()
local -a stale=()
while read -r ns target image; do
[ -n "${target:-}" ] || continue
if [ -z "${digests[$image]:-}" ]; then
digests[$image]="$(pinned_digest "$image")"
fi
want="${digests[$image]}"
if [ -z "$want" ]; then
warn "cannot resolve ${image##*/} in the registry, leaving $target alone"
unchecked=$((unchecked + 1))
continue
fi
selector="$(kubectl get "$target" -n "$ns" -o jsonpath='{.spec.selector.matchLabels}' 2>/dev/null \
| jq -r 'to_entries | map("\(.key)=\(.value)") | join(",")' 2>/dev/null)"
if [ -z "$selector" ]; then
warn "cannot read the pod selector of $target, skipping"
unchecked=$((unchecked + 1))
continue
fi
running="$(kubectl get pods -n "$ns" -l "$selector" -o json 2>/dev/null \
| jq -r --arg repo "${image%%:*}" '
.items[] | .status.containerStatuses[]?
| select(.image == $repo
or (.image | startswith($repo + ":"))
or (.image | startswith($repo + "@")))
| .imageID
' 2>/dev/null)"
if [ -z "$running" ]; then
# Scaled to zero. Nothing is serving stale code, and imagePullPolicy
# resolves the tag when it is scaled back up.
continue
fi
entry=""
while IFS= read -r one; do
[ -n "$one" ] || continue
entry="${one##*@}"
if [ "$entry" != "$want" ]; then
stale+=("$ns $target")
break
fi
done <<<"$running"
done < <(owned_registry_workloads)
if [ "${#stale[@]}" -eq 0 ]; then
if [ "$unchecked" -gt 0 ]; then
# Say so plainly. Reporting "everything is current" after checking nothing
# would tell the operator the deploy is fine when it may not be.
warn "No workload needed a restart, but $unchecked could not be checked"
else
log "All owned workloads already run the image their tag points at"
fi
return 0
fi
log "Restarting ${#stale[@]} workload(s) running an image their tag has moved past"
for ref in "${stale[@]}"; do
log " $ref"
done
local failed=()
for ref in "${stale[@]}"; do
ns="${ref%% *}"
target="${ref#* }"
if ! kubectl rollout restart "$target" -n "$ns" >/dev/null 2>&1; then
failed+=("$ref")
fi
done
if [ "${#failed[@]}" -gt 0 ]; then
warn "could not restart: ${failed[*]}"
return 1
fi
} }
# verify_workloads <failed-file> <kind> <ns> <name> ... # verify_workloads <failed-file> <kind> <ns> <name> ...
@@ -499,31 +281,44 @@ verify_workloads() {
# settle. Prints a report and returns non-zero if any workload is still unhealthy, # settle. Prints a report and returns non-zero if any workload is still unhealthy,
# so the operator knows manual recovery is required. # so the operator knows manual recovery is required.
rollback_workloads() { rollback_workloads() {
local failed_file="$1" local failed_file="$1" snapshot kind ns name index=0 running=0 pid revision uid
local kind ns name unrecovered=() local -a pids=()
local -a recovered=() snapshot="$(cat "$DEPLOY_SNAPSHOT_DIR/current")"
while read -r kind ns name; do while read -r kind ns name; do
[ -n "${kind:-}" ] || continue [[ "$kind" =~ ^(deployment|statefulset|daemonset)$ ]] || continue
# Helm-owned workloads are already rolled back by the release's --rollback-on-failure index=$((index + 1))
# upgrade. `rollout undo` here would step back to the revision Helm just (
# escaped (the failed one), so leave them for the operator instead. if kubectl get "$kind/$name" -n "$ns" -o jsonpath='{.metadata.annotations}' | grep -q 'meta.helm.sh/release-name'; then
if kubectl get "${kind}/${name}" -n "$ns" -o jsonpath='{.metadata.annotations}' 2>/dev/null | grep -q 'meta.helm.sh/release-name'; then echo " skip (Helm recovery owns this workload): $kind/$ns/$name"
echo " skip (helm-managed, needs manual check): ${kind}/${ns}/${name}" exit 1
unrecovered+=("${kind}/${ns}/${name} (helm-managed)") fi
continue revision="$(jq -r --arg ns "$ns" --arg name "$name" --arg kind "$kind" \
fi '.[] | select(.namespace == $ns and .name == $name and .kind == $kind) | .revision' "$snapshot/revisions.json")"
if kubectl rollout undo "${kind}/${name}" -n "$ns" >/dev/null 2>&1 \ uid="$(jq -r --arg ns "$ns" --arg name "$name" --arg kind "$kind" \
&& kubectl rollout status "${kind}/${name}" -n "$ns" --timeout="${ROLLOUT_TIMEOUT}s" >/dev/null 2>&1; then '.[] | select(.namespace == $ns and .name == $name and .kind == $kind) | .uid' "$snapshot/revisions.json")"
echo " rolled back: ${kind}/${ns}/${name}" if [[ ! "$revision" =~ ^[1-9][0-9]*$ ]] || [ "$uid" != "$(kubectl get "$kind/$name" -n "$ns" -o jsonpath='{.metadata.uid}')" ]; then
recovered+=("${kind}/${ns}/${name}") echo " no safe previous revision: $kind/$ns/$name (new or replaced workload)"
else exit 1
echo " NOT RECOVERED: ${kind}/${ns}/${name}" fi
unrecovered+=("${kind}/${ns}/${name}") kubectl rollout undo "$kind/$name" -n "$ns" --to-revision="$revision" \
&& kubectl rollout status "$kind/$name" -n "$ns" --timeout="${ROLLOUT_TIMEOUT}s"
) >"$snapshot/rollback-$index.log" 2>&1 &
pids+=($!)
running=$((running + 1))
if [ "$running" -ge "$ROLLOUT_PARALLELISM" ]; then
wait -n 2>/dev/null || true
running=$((running - 1))
fi fi
done <"$failed_file" done <"$failed_file"
echo "ROLLED_BACK=${#recovered[@]}" >>"$failed_file" local recovered=0 unrecovered=0 i=0
echo "UNRECOVERED=${#unrecovered[@]}" >>"$failed_file" for pid in "${pids[@]}"; do
[ "${#unrecovered[@]}" -eq 0 ] i=$((i + 1))
if wait "$pid"; then recovered=$((recovered + 1)); else unrecovered=$((unrecovered + 1)); fi
cat "$snapshot/rollback-$i.log"
done
echo "ROLLED_BACK=$recovered" >>"$failed_file"
echo "UNRECOVERED=$unrecovered" >>"$failed_file"
[ "$unrecovered" -eq 0 ]
} }
# Helm releases owned by this stage, one line each: # Helm releases owned by this stage, one line each:
@@ -536,6 +331,7 @@ rollback_workloads() {
# have to be declared as custom.regex managers in renovate/renovate.json. # have to be declared as custom.regex managers in renovate/renovate.json.
HELM_RELEASES=( HELM_RELEASES=(
"prometheus-stack|prometheus-community/kube-prometheus-stack|prometheus|86.2.3|prometheus-stack/k8s/grafana-values.yaml|prometheus-stack/k8s/active" "prometheus-stack|prometheus-community/kube-prometheus-stack|prometheus|86.2.3|prometheus-stack/k8s/grafana-values.yaml|prometheus-stack/k8s/active"
"victoria-operator|victoriametrics/victoria-metrics-operator|prometheus|0.68.1|prometheus-stack/k8s/victoria-operator-values.yaml|prometheus-stack/k8s/active"
"loki|grafana/loki|prometheus|7.3.0|loki/k8s/loki-values.yaml|loki/k8s/active" "loki|grafana/loki|prometheus|7.3.0|loki/k8s/loki-values.yaml|loki/k8s/active"
"alloy|grafana/alloy|prometheus|1.12.1|loki/k8s/alloy-values.yaml|loki/k8s/active" "alloy|grafana/alloy|prometheus|1.12.1|loki/k8s/alloy-values.yaml|loki/k8s/active"
"reloader|stakater/reloader|reloader|2.2.17|reloader/k8s/reloader-values.yaml|reloader/k8s/active" "reloader|stakater/reloader|reloader|2.2.17|reloader/k8s/reloader-values.yaml|reloader/k8s/active"
@@ -547,6 +343,7 @@ helm_repo_for() {
prometheus-community/*) echo "prometheus-community https://prometheus-community.github.io/helm-charts" ;; prometheus-community/*) echo "prometheus-community https://prometheus-community.github.io/helm-charts" ;;
grafana/*) echo "grafana https://grafana.github.io/helm-charts" ;; grafana/*) echo "grafana https://grafana.github.io/helm-charts" ;;
stakater/*) echo "stakater https://stakater.github.io/stakater-charts" ;; stakater/*) echo "stakater https://stakater.github.io/stakater-charts" ;;
victoriametrics/*) echo "victoriametrics https://victoriametrics.github.io/helm-charts" ;;
esac esac
} }
@@ -556,8 +353,12 @@ helm_repo_for() {
helm_release_status() { helm_release_status() {
local out local out
if ! out="$(helm status "$1" -n "$2" 2>&1)"; then if ! out="$(helm status "$1" -n "$2" 2>&1)"; then
echo "not-found" if [[ "$out" == *"release: not found"* ]]; then
return 0 echo "not-found"
return 0
fi
printf 'ERROR: cannot read Helm status: %s\n' "$out" >&2
return 1
fi fi
awk '/^STATUS:/{print $2}' <<<"$out" | tr '[:upper:]' '[:lower:]' awk '/^STATUS:/{print $2}' <<<"$out" | tr '[:upper:]' '[:lower:]'
} }
@@ -569,20 +370,35 @@ helm_release_status() {
# pending (deployed, failed, not-found). Returns non-zero when the release is # pending (deployed, failed, not-found). Returns non-zero when the release is
# still not recoverable, so the pipeline fails loud instead of wedging. # still not recoverable, so the pipeline fails loud instead of wedging.
recover_pending_release() { recover_pending_release() {
local release="$1" namespace="$2" status local release="$1" namespace="$2" status revision snapshot
status="$(helm_release_status "$release" "$namespace")" status="$(helm_release_status "$release" "$namespace")" || return 1
case "$status" in case "$status" in
pending-upgrade|pending-rollback|pending-install) pending-upgrade|pending-rollback|pending-install)
log "Release $release is $status, rolling back to the last deployed revision" log "Release $release is $status, rolling back to the last deployed revision"
if ! helm rollback "$release" -n "$namespace" --wait --timeout 10m >/dev/null 2>&1; then revision=""
if [ -s "$DEPLOY_SNAPSHOT_DIR/current" ]; then
snapshot="$(cat "$DEPLOY_SNAPSHOT_DIR/current")"
if [ -s "$snapshot/helm-$release.json" ]; then
revision="$(jq -r '.version' "$snapshot/helm-$release.json")"
fi
fi
if [[ ! "$revision" =~ ^[1-9][0-9]*$ ]]; then
echo "ERROR: no captured Helm revision for $release; manual recovery required"
return 1
fi
record_apply helm-rollback "$namespace/$release" started
if ! helm rollback "$release" "$revision" -n "$namespace" --wait --timeout 10m; then
record_apply helm-rollback "$namespace/$release" failure
echo "WARN: helm rollback of $release did not complete" echo "WARN: helm rollback of $release did not complete"
return 1 return 1
fi fi
status="$(helm_release_status "$release" "$namespace")" status="$(helm_release_status "$release" "$namespace")" || return 1
if [ "$status" != "deployed" ]; then if [ "$status" != "deployed" ]; then
record_apply helm-rollback "$namespace/$release" failure
echo "WARN: $release is $status after rollback" echo "WARN: $release is $status after rollback"
return 1 return 1
fi fi
record_apply helm-rollback "$namespace/$release" success
;; ;;
esac esac
return 0 return 0
@@ -614,11 +430,16 @@ upgrade_helm_releases() {
local entry release chart namespace version values marker repo local entry release chart namespace version values marker repo
for entry in ${HELM_RELEASES[@]+"${HELM_RELEASES[@]}"}; do for entry in ${HELM_RELEASES[@]+"${HELM_RELEASES[@]}"}; do
IFS='|' read -r release chart namespace version values marker <<<"$entry" IFS='|' read -r release chart namespace version values marker <<<"$entry"
if [ -n "${DEPLOY_PLAN:-}" ] && ! jq -e --arg name "$release" '.helm | index($name) != null' "$DEPLOY_PLAN" >/dev/null; then
echo "skip (unchanged Helm release): $release"
continue
fi
if [ ! -f "$REPO/$values" ] && [ -f "$CONFIG_REPO/$values" ]; then values="$CONFIG_REPO/$values"; else values="$REPO/$values"; fi
if [ ! -f "$REPO/$marker" ]; then if [ ! -f "$REPO/$marker" ]; then
echo "skip (no $marker): $release" echo "skip (no $marker): $release"
continue continue
fi fi
if [ ! -f "$REPO/$values" ]; then if [ ! -f "$values" ]; then
echo "ERROR: $values is gitignored but missing on the workstation, restore it first." echo "ERROR: $values is gitignored but missing on the workstation, restore it first."
return 1 return 1
fi fi
@@ -627,8 +448,8 @@ upgrade_helm_releases() {
echo "ERROR: no Helm repository configured for chart $chart" echo "ERROR: no Helm repository configured for chart $chart"
return 1 return 1
fi fi
helm repo add "${repo%% *}" "${repo#* }" >/dev/null 2>&1 || true helm repo add "${repo%% *}" "${repo#* }" >/dev/null
helm repo update "${repo%% *}" >/dev/null 2>&1 || true helm repo update "${repo%% *}" >/dev/null
log "Upgrading $release ($chart $version)" log "Upgrading $release ($chart $version)"
wait_for_calm "helm $release" wait_for_calm "helm $release"
# A previous run with --rollback-on-failure whose own rollback never finished leaves the # A previous run with --rollback-on-failure whose own rollback never finished leaves the
@@ -641,11 +462,13 @@ upgrade_helm_releases() {
# --rollback-on-failure (+ --wait) rolls the release back when the upgrade # --rollback-on-failure (+ --wait) rolls the release back when the upgrade
# times out or the workloads it touches never become ready, so a bad chart # times out or the workloads it touches never become ready, so a bad chart
# bump is not left half applied. (--atomic was this combo; deprecated.) # bump is not left half applied. (--atomic was this combo; deprecated.)
record_apply helm-upgrade "$namespace/$release" started
if ! helm upgrade --install "$release" "$chart" \ if ! helm upgrade --install "$release" "$chart" \
--namespace "$namespace" \ --namespace "$namespace" \
--version "$version" \ --version "$version" \
--values "$REPO/$values" \ --values "$values" \
--wait --rollback-on-failure --cleanup-on-fail --timeout 10m; then --wait --rollback-on-failure --cleanup-on-fail --timeout 10m; then
record_apply helm-upgrade "$namespace/$release" failure
echo "WARN: upgrade of $release failed, checking release state" echo "WARN: upgrade of $release failed, checking release state"
# --rollback-on-failure already attempted its own rollback; finish the job when that # --rollback-on-failure already attempted its own rollback; finish the job when that
# rollback never completed, otherwise the release stays pending-* and # rollback never completed, otherwise the release stays pending-* and
@@ -655,35 +478,49 @@ upgrade_helm_releases() {
else else
echo "ERROR: upgrade of $release failed (release is back on its previous revision)." echo "ERROR: upgrade of $release failed (release is back on its previous revision)."
fi fi
record_apply helm-recovery-state "$namespace/$release" "$(helm_release_status "$release" "$namespace" || echo unknown)"
return 1 return 1
fi fi
record_apply helm-upgrade "$namespace/$release" success
done done
} }
stage_preflight() { stage_doctor() {
if [ ! -d "$REPO/.git" ]; then local tool entry release chart namespace version values marker
echo "Repository not found at $REPO" for tool in git docker kubectl helm jq curl timeout flock python3; do
exit 1 command -v "$tool" >/dev/null || { echo "Missing workstation tool: $tool"; return 1; }
fi done
if [ -n "$DEPLOY_SHA" ]; then docker compose version >/dev/null
log "Checking out the commit CI validated ($DEPLOY_SHA)" docker buildx version >/dev/null
git -C "$REPO" fetch origin --quiet "$DEPLOY_SHA" 2>/dev/null \ [ "$(kubectl config current-context)" = "${KUBE_CONTEXT:?configure KUBE_CONTEXT}" ] || { echo "Unexpected Kubernetes context"; return 1; }
|| git -C "$REPO" fetch origin main [ "$(kubectl get namespace kube-system -o jsonpath='{.metadata.uid}')" = "${EXPECTED_CLUSTER_UID:?configure EXPECTED_CLUSTER_UID}" ] || { echo "Unexpected Kubernetes cluster"; return 1; }
else kubectl get --raw=/readyz --request-timeout=10s >/dev/null
git -C "$REPO" fetch origin main [ "$(git -C "$REPO" rev-parse HEAD)" = "$DEPLOY_SHA" ] || return 1
fi select_manifests
target="${DEPLOY_SHA:-origin/main}" for entry in "${HELM_RELEASES[@]}"; do
log "Workstation state" IFS='|' read -r release chart namespace version values marker <<<"$entry"
echo " local: $(git -C "$REPO" rev-parse --short HEAD)" [ -f "$REPO/$marker" ] || continue
echo " target: $(git -C "$REPO" rev-parse --short "$target")" [ -f "$REPO/$values" ] || [ -f "$CONFIG_REPO/$values" ] || { echo "Missing values: $values"; return 1; }
if [ -n "$(git -C "$REPO" status --porcelain --untracked-files=no)" ]; then done
echo "ERROR: workstation has local tracked modifications, refusing reset:" jq '{sha, selected, helm, removed}' "$DEPLOY_PLAN"
git -C "$REPO" status --porcelain --untracked-files=no local cf
git -C "$REPO" diff --stat for cf in "${COMPOSE_STACKS[@]}"; do
echo "Fix it on the workstation (commit, or 'git restore .'), then re-run the deploy." compose "$cf" config --quiet
exit 1 while IFS= read -r network; do
fi docker network inspect "$network" >/dev/null || return 1
git -C "$REPO" reset --hard "$target" done < <(compose "$cf" config --format json | jq -r '.networks // {} | to_entries[] | select(.value.external == true) | .value.name')
python3 "$REPO/.gitea/workflows/compose-release.py" "$cf"
done
local image refs m k
refs="$(
for m in "${K8S_MANIFESTS[@]}"; do render_pinned <"$m" || return 1; done
for k in "${KUSTOMIZE_APPS[@]}"; do kubectl kustomize "$k" | render_pinned || return 1; done
)" || return 1
refs="$(grep -oE 'gcr\.forust\.xyz/forust/[A-Za-z0-9._-]+@sha256:[0-9a-f]{64}' <<<"$refs" | sort -u || true)"
while IFS= read -r image; do
[ -n "$image" ] || continue
timeout 60s docker buildx imagetools inspect "$image" >/dev/null
done <<<"$refs"
} }
# Required pod Secrets, scoped to the resource namespace. TLS route Secrets are # Required pod Secrets, scoped to the resource namespace. TLS route Secrets are
@@ -693,6 +530,9 @@ check_referenced_secrets() {
local missing=() local missing=()
refs="" refs=""
for m in ${K8S_MANIFESTS[@]+"${K8S_MANIFESTS[@]}"}; do for m in ${K8S_MANIFESTS[@]+"${K8S_MANIFESTS[@]}"}; do
if skip_uninstalled_vmagent_crd "$m"; then
continue
fi
objects="$(kubectl create --dry-run=client --validate=false -f "$m" -o json)" || return 1 objects="$(kubectl create --dry-run=client --validate=false -f "$m" -o json)" || return 1
extracted="$(printf '%s' "$objects" | jq -r -f "$REPO/.gitea/workflows/secret-references.jq")" || return 1 extracted="$(printf '%s' "$objects" | jq -r -f "$REPO/.gitea/workflows/secret-references.jq")" || return 1
refs+="$extracted"$'\n' refs+="$extracted"$'\n'
@@ -719,7 +559,20 @@ check_referenced_secrets() {
fi fi
} }
# The VMAgent CRD is installed by the VictoriaMetrics Operator Helm release in
# stage_apply_k8s, after this preflight stage. Skip only its dry-run until then.
skip_uninstalled_vmagent_crd() {
local manifest="$1"
if [[ "$manifest" == "$REPO/prometheus-stack/k8s/vmagent.yaml" ]] \
&& ! kubectl get crd vmagents.operator.victoriametrics.com >/dev/null 2>&1; then
echo " skip: VMAgent CRD is installed by Helm during apply: ${manifest#"$REPO"/}"
return 0
fi
return 1
}
stage_validate() { stage_validate() {
check_prune_mode || return 1
cd "$REPO" cd "$REPO"
select_manifests select_manifests
local m k cf local m k cf
@@ -731,10 +584,13 @@ stage_validate() {
log "Validate compose stacks" log "Validate compose stacks"
for cf in ${COMPOSE_STACKS[@]+"${COMPOSE_STACKS[@]}"}; do for cf in ${COMPOSE_STACKS[@]+"${COMPOSE_STACKS[@]}"}; do
echo " config: $cf" echo " config: $cf"
validate_compose_file "$cf" compose "$cf" config --quiet
done done
log "Validate k8s manifests (kubectl dry-run=client)" log "Validate k8s manifests (kubectl dry-run=client)"
for m in ${K8S_MANIFESTS[@]+"${K8S_MANIFESTS[@]}"}; do for m in ${K8S_MANIFESTS[@]+"${K8S_MANIFESTS[@]}"}; do
if skip_uninstalled_vmagent_crd "$m"; then
continue
fi
kubectl apply --dry-run=client -f "$m" >/dev/null kubectl apply --dry-run=client -f "$m" >/dev/null
done done
for k in ${KUSTOMIZE_APPS[@]+"${KUSTOMIZE_APPS[@]}"}; do for k in ${KUSTOMIZE_APPS[@]+"${KUSTOMIZE_APPS[@]}"}; do
@@ -742,6 +598,9 @@ stage_validate() {
done done
log "Validate k8s manifests (kubectl dry-run=server)" log "Validate k8s manifests (kubectl dry-run=server)"
for m in ${K8S_MANIFESTS[@]+"${K8S_MANIFESTS[@]}"}; do for m in ${K8S_MANIFESTS[@]+"${K8S_MANIFESTS[@]}"}; do
if skip_uninstalled_vmagent_crd "$m"; then
continue
fi
kubectl apply --dry-run=server -f "$m" >/dev/null kubectl apply --dry-run=server -f "$m" >/dev/null
done done
for k in ${KUSTOMIZE_APPS[@]+"${KUSTOMIZE_APPS[@]}"}; do for k in ${KUSTOMIZE_APPS[@]+"${KUSTOMIZE_APPS[@]}"}; do
@@ -752,34 +611,54 @@ stage_validate() {
check_referenced_secrets check_referenced_secrets
} }
selected_workload_refs() {
local m k
for m in "${K8S_MANIFESTS[@]}"; do
if skip_uninstalled_vmagent_crd "$m" >/dev/null; then continue; fi
kubectl create --dry-run=client --validate=false -f "$m" -o json | jq -r '
(if .kind == "List" then .items[] else . end) | select(.kind | test("^(Deployment|StatefulSet|DaemonSet)$"))
| "\(.kind | ascii_downcase) \(.metadata.namespace // "default") \(.metadata.name)"'
done
for k in "${KUSTOMIZE_APPS[@]}"; do
kubectl kustomize "$k" | kubectl create --dry-run=client --validate=false -f - -o json | jq -r '
(if .kind == "List" then .items[] else . end) | select(.kind | test("^(Deployment|StatefulSet|DaemonSet)$"))
| "\(.kind | ascii_downcase) \(.metadata.namespace // "default") \(.metadata.name)"'
done
}
stage_apply_k8s() { stage_apply_k8s() {
check_prune_mode || return 1
cd "$REPO" cd "$REPO"
select_manifests >/dev/null select_manifests >/dev/null
local ns_files=() other_files=() m k prune_opts=() local ns_files=() other_files=() m k
for m in ${K8S_MANIFESTS[@]+"${K8S_MANIFESTS[@]}"}; do for m in ${K8S_MANIFESTS[@]+"${K8S_MANIFESTS[@]}"}; do
case "$m" in case "$m" in
*/namespace.y?ml) ns_files+=("$m") ;; */namespace.yaml|*/namespace.yml) ns_files+=("$m") ;;
*) other_files+=("$m") ;; *) other_files+=("$m") ;;
esac esac
done done
if [ "$APPLY_PRUNE" = "true" ]; then
prune_opts=(--prune -l app.kubernetes.io/managed-by=homelab-deploy)
fi
# Record what is about to change, and publish it for the verify job, before # Record what is about to change, and publish it for the verify job, before
# the first apply. Both are fatal on failure: see snapshot_dir. # the first apply. Both are fatal on failure: see snapshot_dir.
selected_workload_refs >"$RUN_DIR/workload-refs"
local snapshot local snapshot
snapshot="$(snapshot_dir)" || return 1 snapshot="$(snapshot_dir)" || return 1
save_snapshot "$snapshot" || return 1 save_snapshot "$snapshot" || return 1
touch "$snapshot/ready"
if [ "${#ns_files[@]}" -gt 0 ]; then if [ "${#ns_files[@]}" -gt 0 ]; then
log "Applying namespaces (${#ns_files[@]} files)" log "Applying namespaces (${#ns_files[@]} files)"
for m in "${ns_files[@]}"; do for m in "${ns_files[@]}"; do
kubectl apply -f "$m" record_apply kubectl "${m#"$REPO"/}" started
if ! kubectl apply -f "$m"; then
record_apply kubectl "${m#"$REPO"/}" failure
return 1
fi
record_apply kubectl "${m#"$REPO"/}" success
done done
fi fi
if [ -f "$REPO/prometheus-stack/k8s/active" ]; then if selected_service k8s prometheus-stack && [ -f "$REPO/prometheus-stack/k8s/active" ]; then
if [ ! -f "$REPO/prometheus-stack/k8s/grafana-values.yaml" ]; then if [ ! -f "$CONFIG_REPO/prometheus-stack/k8s/grafana-values.yaml" ]; then
echo "ERROR: prometheus-stack/k8s/grafana-values.yaml (gitignored) missing on workstation, restore it first." echo "ERROR: prometheus-stack/k8s/grafana-values.yaml (gitignored) missing on workstation, restore it first."
exit 1 exit 1
fi fi
@@ -789,20 +668,26 @@ stage_apply_k8s() {
if [ "${#other_files[@]}" -gt 0 ]; then if [ "${#other_files[@]}" -gt 0 ]; then
log "Applying resources (${#other_files[@]} files, our images pinned to digests)" log "Applying resources (${#other_files[@]} files, our images pinned to digests)"
for m in "${other_files[@]}"; do for m in "${other_files[@]}"; do
if ! render_pinned <"$m" | kubectl apply "${prune_opts[@]}" -f -; then log "Applying ${m#"$REPO"/}"
record_apply kubectl "${m#"$REPO"/}" started
if ! render_pinned <"$m" | kubectl apply -f -; then
record_apply kubectl "${m#"$REPO"/}" failure
echo "ERROR: apply failed for ${m#"$REPO"/}" >&2 echo "ERROR: apply failed for ${m#"$REPO"/}" >&2
exit 1 exit 1
fi fi
record_apply kubectl "${m#"$REPO"/}" success
done done
fi fi
for k in ${KUSTOMIZE_APPS[@]+"${KUSTOMIZE_APPS[@]}"}; do for k in ${KUSTOMIZE_APPS[@]+"${KUSTOMIZE_APPS[@]}"}; do
log "Applying kustomize app: ${k#"$REPO"/} (our images pinned to digests)" log "Applying kustomize app: ${k#"$REPO"/} (our images pinned to digests)"
record_apply kustomize "${k#"$REPO"/}" started
if ! kubectl kustomize "$k" | render_pinned | kubectl apply -f -; then if ! kubectl kustomize "$k" | render_pinned | kubectl apply -f -; then
record_apply kustomize "${k#"$REPO"/}" failure
echo "ERROR: apply failed for kustomize app ${k#"$REPO"/}" >&2 echo "ERROR: apply failed for kustomize app ${k#"$REPO"/}" >&2
exit 1 exit 1
fi fi
record_apply kustomize "${k#"$REPO"/}" success
done done
restart_stale_images
# No verification here on purpose. This stage may be killed at any point by # No verification here on purpose. This stage may be killed at any point by
# timeout-minutes, by the runner cancelling the job, or by a dropped SSH # timeout-minutes, by the runner cancelling the job, or by a dropped SSH
@@ -817,7 +702,7 @@ stage_apply_k8s() {
# and rolls back the ones that never became healthy. # and rolls back the ones that never became healthy.
stage_verify_k8s() { stage_verify_k8s() {
local pointer="$DEPLOY_SNAPSHOT_DIR/current" local pointer="$DEPLOY_SNAPSHOT_DIR/current"
local snapshot want have generations local snapshot want have generations changed
local -a touched=() local -a touched=()
if [ ! -s "$pointer" ]; then if [ ! -s "$pointer" ]; then
@@ -828,7 +713,7 @@ stage_verify_k8s() {
return 1 return 1
fi fi
snapshot="$(head -1 "$pointer")" snapshot="$(head -1 "$pointer")"
if [ ! -d "$snapshot" ]; then if [ ! -d "$snapshot" ] || [ ! -f "$snapshot/ready" ]; then
echo "ERROR: snapshot pointer refers to a missing directory: $snapshot" echo "ERROR: snapshot pointer refers to a missing directory: $snapshot"
return 1 return 1
fi fi
@@ -851,6 +736,13 @@ stage_verify_k8s() {
fi fi
echo " snapshot: $snapshot (commit ${have:0:12})" echo " snapshot: $snapshot (commit ${have:0:12})"
local entry release chart namespace version values marker
for entry in "${HELM_RELEASES[@]}"; do
IFS='|' read -r release chart namespace version values marker <<<"$entry"
jq -e --arg name "$release" '.helm | index($name) != null' "$DEPLOY_PLAN" >/dev/null || continue
recover_pending_release "$release" "$namespace" || return 1
done
generations="$snapshot/generations.before" generations="$snapshot/generations.before"
if [ ! -s "$generations" ]; then if [ ! -s "$generations" ]; then
# Without a baseline we cannot tell which workloads the apply touched, so # Without a baseline we cannot tell which workloads the apply touched, so
@@ -859,9 +751,10 @@ stage_verify_k8s() {
: >"$generations" : >"$generations"
fi fi
changed="$(changed_workloads "$generations")" || return 1
while read -r kind ns name; do while read -r kind ns name; do
[ -n "${kind:-}" ] && touched+=("$kind $ns $name") [ -n "${kind:-}" ] && touched+=("$kind $ns $name")
done < <(changed_workloads "$generations") done <<<"$changed"
log "Verifying ${#touched[@]} changed workload(s) (timeout ${ROLLOUT_TIMEOUT}s each)" log "Verifying ${#touched[@]} changed workload(s) (timeout ${ROLLOUT_TIMEOUT}s each)"
if [ "${#touched[@]}" -eq 0 ]; then if [ "${#touched[@]}" -eq 0 ]; then
@@ -900,21 +793,19 @@ stage_verify_k8s() {
verify_compose_stack() { verify_compose_stack() {
local cf="$1" local cf="$1"
local expected running missing=() local expected running missing=()
expected="$(docker compose -f "$cf" config --services 2>/dev/null | sort || true)" expected="$(compose "$cf" config --format json | jq -r ' .services | to_entries[] | select(.value.restart != "no") | .key' | sort)" || return 1
running="$(docker compose -f "$cf" ps --status running --services 2>/dev/null | sort || true)" running="$(compose "$cf" ps --status running --services | sort)" || return 1
[ -n "$expected" ] || return 0 [ -n "$expected" ] || return 0
while IFS= read -r svc; do while IFS= read -r svc; do
[ -n "$svc" ] || continue [ -n "$svc" ] || continue
# restart:"no" services are allowed to have exited. # restart:"no" services are allowed to have exited.
if ! printf '%s\n' "$running" | grep -qx "$svc" \ if ! printf '%s\n' "$running" | grep -qx "$svc"; then
&& ! docker compose -f "$cf" config 2>/dev/null \
| grep -A5 "^ ${svc}:" | grep -qE 'restart:\s*"?no"?'; then
missing+=("$svc") missing+=("$svc")
fi fi
done <<<"$expected" done <<<"$expected"
if [ "${#missing[@]}" -gt 0 ]; then if [ "${#missing[@]}" -gt 0 ]; then
echo " NOT RUNNING: ${missing[*]}" echo " NOT RUNNING: ${missing[*]}"
docker compose -f "$cf" ps --all 2>/dev/null | sed 's/^/ /' || true compose "$cf" ps --all 2>/dev/null | sed 's/^/ /' || true
return 1 return 1
fi fi
echo " all ${#expected} service(s) running" echo " all ${#expected} service(s) running"
@@ -998,10 +889,11 @@ traefik_routed_hosts() {
# cases, so ask Traefik which routes it built and fail on the difference. # cases, so ask Traefik which routes it built and fail on the difference.
stage_smoke() { stage_smoke() {
cd "$REPO" cd "$REPO"
if [ -n "${DEPLOY_PLAN:-}" ] && jq -e '.full_smoke' "$DEPLOY_PLAN" >/dev/null; then
DEPLOY_SMOKE_ALL=true
fi
select_manifests >/dev/null select_manifests >/dev/null
local -a hosts=() local -a hosts=()
# Not named failed: an array of that name already exists in restart_stale_images
# above, and a scalar shadowing an array is a trap rather than a shadow.
local h code rc bad=0 local h code rc bad=0
while IFS= read -r h; do while IFS= read -r h; do
[ -n "$h" ] && hosts+=("$h") [ -n "$h" ] && hosts+=("$h")
@@ -1009,8 +901,8 @@ stage_smoke() {
if [ "${#hosts[@]}" -eq 0 ]; then if [ "${#hosts[@]}" -eq 0 ]; then
# Nothing to probe means the extraction broke, not that the cluster is empty. # Nothing to probe means the extraction broke, not that the cluster is empty.
echo "ERROR: no public hostnames found in active manifests, refusing to report success" echo "No public routes in the selected components"
return 1 return 0
fi fi
log "Probing ${#hosts[@]} public route(s)" log "Probing ${#hosts[@]} public route(s)"
@@ -1094,37 +986,22 @@ stage_apply_compose() {
cd "$REPO" cd "$REPO"
select_manifests >/dev/null select_manifests >/dev/null
local cf local cf
log "Redeploying docker compose stacks (${#COMPOSE_STACKS[@]} stacks)" for cf in "${COMPOSE_STACKS[@]}"; do
for cf in ${COMPOSE_STACKS[@]+"${COMPOSE_STACKS[@]}"}; do log "Applying Compose ${cf#"$REPO"/}"
echo " compose: $cf" record_apply compose "${cf#"$REPO"/}" started
if grep -Eq '^\s+pull_policy:\s*build\b' "$cf"; then if ! compose "$cf" up -d --wait --wait-timeout 180 --pull missing --remove-orphans; then
docker compose -f "$cf" build record_apply compose "${cf#"$REPO"/}" failure
docker compose -f "$cf" push return 1
fi fi
docker compose -f "$cf" up -d --pull always --remove-orphans record_apply compose "${cf#"$REPO"/}" success
verify_compose_stack "$cf"
done done
echo "Compose recovery files: $RUN_DIR/compose-before (manual recovery only)"
local -a broken=()
for cf in ${COMPOSE_STACKS[@]+"${COMPOSE_STACKS[@]}"}; do
echo " verifying: $cf"
if ! verify_compose_stack "$cf"; then
broken+=("$cf")
fi
done
if [ "${#broken[@]}" -gt 0 ]; then
echo
echo "ERROR: ${#broken[@]} compose stack(s) did not come up:"
printf ' - %s\n' "${broken[@]}"
echo "Compose stacks are not rolled back automatically: their images use mutable"
echo "':latest' tags, so there is no previous version to return to. Check the logs"
echo "above, then re-run the deploy once the cause is fixed."
return 1
fi
} }
run_stage() { run_stage() {
case "${1:?stage required}" in case "${1:?stage required}" in
preflight) stage_preflight ;; doctor) stage_doctor ;;
validate) stage_validate ;; validate) stage_validate ;;
apply-k8s) stage_apply_k8s ;; apply-k8s) stage_apply_k8s ;;
verify-k8s) stage_verify_k8s ;; verify-k8s) stage_verify_k8s ;;
+124
View File
@@ -0,0 +1,124 @@
#!/usr/bin/env python3
"""Calculate selected components against the last fully successful deploy."""
import hashlib
import json
import re
import subprocess
from pathlib import Path
def output(*args, **kwargs):
return subprocess.check_output(args, text=True, **kwargs).strip() # noqa: S603
def tracked(repo):
return output('git', '-C', str(repo), 'ls-files').splitlines()
def helm_releases(repo):
text = (repo / '.gitea/workflows/deploy-lib.sh').read_text()
return [line.split('|') for line in re.findall(r'^ "([^"\n]+\|[^"\n]+)"$', text, re.MULTILINE)]
def inventory(repo):
files = tracked(repo)
k8s = sorted(
{f.split('/k8s/')[0] for f in files if '/k8s/' in f and (repo / f.split('/k8s/')[0] / 'k8s/active').is_file()}
)
compose = sorted(
{
str(Path(f).parent)
for f in files
if Path(f).name in ('compose.yaml', 'compose.yml') and (repo / Path(f).parent / 'active').is_file()
}
)
return {'k8s': k8s, 'compose': compose}
def file_hash(path):
return hashlib.sha256(path.read_bytes()).hexdigest() if path.is_file() else 'missing'
def make_plan(repo, config_repo, release, previous, mode, live_helm):
active = inventory(repo)
all_services = set(active['k8s'] + active['compose'])
helm_inputs = {}
helm_selected = []
for name, chart, namespace, version, values, marker in helm_releases(repo):
if not (repo / marker).is_file():
continue
value_path = repo / values if (repo / values).is_file() else config_repo / values
if not value_path.is_file():
raise ValueError(f'Missing Helm values: {values}')
stamp = hashlib.sha256(f'{chart}|{version}|{file_hash(value_path)}'.encode()).hexdigest()
helm_inputs[name] = stamp
live = next((h for h in live_helm if h['name'] == name and h['namespace'] == namespace), None)
if (
mode == 'full'
or previous is None
or previous.get('helm_inputs', {}).get(name) != stamp
or live is None
or live.get('status') != 'deployed'
or live.get('chart') != f'{chart.split("/")[-1]}-{version}'
):
helm_selected.append(name)
local_inputs = {}
for service in all_services:
candidates = [config_repo / service / '.env']
if service in active['compose']:
candidates.append(config_repo / '.env')
cfg = config_repo / service / 'config'
if cfg.is_dir():
candidates.extend(
p for p in cfg.rglob('*') if p.is_file() and p.suffix in ('.yaml', '.yml', '.json', '.conf')
)
local_inputs[service] = hashlib.sha256(
'\n'.join(f'{p.relative_to(config_repo)}:{file_hash(p)}' for p in sorted(candidates)).encode()
).hexdigest()
if previous is None:
if mode == 'changed':
raise ValueError('No successful baseline; run deploy in full mode first')
changed = set(all_services)
removed = []
else:
paths = output('git', '-C', str(repo), 'diff', '--name-only', previous['sha'], release['sha']).splitlines()
changed = {path.split('/')[0] for path in paths}
if any(path.startswith('.gitea/') for path in paths):
changed |= all_services
changed |= {s for s in all_services if previous.get('local_inputs', {}).get(s) != local_inputs[s]}
for file in tracked(repo):
service = file.split('/')[0]
if service not in all_services or not file.endswith(('.yaml', '.yml')):
continue
text = (repo / file).read_text()
if any(
image in text and previous.get('images', {}).get(image) != digest
for image, digest in release['images'].items()
):
changed.add(service)
removed = sorted(
set(previous.get('active', {}).get('k8s', []) + previous.get('active', {}).get('compose', []))
- all_services
)
removed += [path for path in paths if '/k8s/' in path and not (repo / path).exists()]
if mode == 'full':
changed = set(all_services)
dependencies = json.loads((repo / '.gitea/deploy-dependencies.json').read_text())
while True:
expanded = changed | {dependent for service in changed for dependent in dependencies.get(service, [])}
if expanded == changed:
break
changed = expanded
return {
'version': 1,
'sha': release['sha'],
'images': release['images'],
'active': active,
'selected': {kind: sorted(set(services) & changed) for kind, services in active.items()},
'helm': helm_selected,
'helm_inputs': helm_inputs,
'local_inputs': local_inputs,
'removed': sorted(set(removed)),
'full_smoke': mode == 'full' or 'traefik' in changed,
}
+10
View File
@@ -0,0 +1,10 @@
#!/usr/bin/env bash
set -euo pipefail
source "${REPO:?}/.gitea/workflows/deploy-lib.sh"
case "${1:?stage required}" in
workload-count)
select_manifests >/dev/null
selected_workload_refs | sort -u | wc -l
;;
*) run_stage "$1" ;;
esac
+100 -157
View File
@@ -1,197 +1,140 @@
name: deploy name: deploy
on: on:
# Deploy only what CI already validated. workflow_run is used instead of
# workflow_dispatch so a red lint/validate run can never reach the cluster.
workflow_run: workflow_run:
workflows: [ci] workflows: [ci]
branches: [main]
types: [completed] types: [completed]
workflow_dispatch: workflow_dispatch:
inputs:
deploy_ref:
description: "Commit already checked by successful main CI (main or SHA)"
default: main
required: true
deploy_mode:
description: "First deploy requires full; plan changes no production resources"
type: choice
options: [changed, full, plan]
default: changed
refresh_images:
description: "Explicitly refresh mutable third-party Compose tags"
type: boolean
default: false
# The deploy jobs read the tree, then reach the cluster over SSH with the
# deploy key. The Actions token itself is not part of that path, so it gets
# read-only contents and no more.
permissions: permissions:
contents: read contents: read
actions: read
concurrency: concurrency:
group: deploy-main group: deploy-main
# Queue instead of cancelling. Cancelling a run kills the apply job mid-loop and
# takes the verify job down with it, so a superseded deploy would leave the
# cluster half-applied and unchecked — the exact failure the verify job exists
# to catch. kubectl apply and docker compose up are both idempotent, so letting
# the older run finish and then deploying the newer commit costs little.
cancel-in-progress: false cancel-in-progress: false
env: env:
DEPLOY_HOST: ${{ secrets.DEPLOY_HOST }} DEPLOY_HOST: ${{ vars.DEPLOY_HOST || secrets.DEPLOY_HOST }}
DEPLOY_PORT: ${{ secrets.DEPLOY_PORT }} DEPLOY_PORT: ${{ vars.DEPLOY_PORT || secrets.DEPLOY_PORT }}
DEPLOY_USER: ${{ secrets.DEPLOY_USER }} DEPLOY_USER: ${{ vars.DEPLOY_USER || secrets.DEPLOY_USER }}
DEPLOY_PATH: ${{ secrets.DEPLOY_PATH }}
DEPLOY_KEY: ${{ secrets.DEPLOY_SSH_KEY }} DEPLOY_KEY: ${{ secrets.DEPLOY_SSH_KEY }}
APPLY_PRUNE: ${{ vars.APPLY_PRUNE }} DEPLOY_KNOWN_HOSTS: ${{ vars.DEPLOY_KNOWN_HOSTS }}
# workflow_run's own GITHUB_SHA points at the branch head, not at the commit the DEPLOY_RUN_ID: ${{ github.run_id }}-${{ github.run_attempt || 1 }}
# finished ci run checked. Pin the exact validated commit instead, so a push DEPLOY_MODE: ${{ inputs.deploy_mode || 'changed' }}
# landing mid-deploy cannot make the workstation deploy something else. Also REFRESH_IMAGES: ${{ inputs.refresh_images && 'true' || 'false' }}
# what the verify job checks the snapshot against. Empty for workflow_dispatch,
# which falls back to the current origin/main.
DEPLOY_SHA: ${{ github.event.workflow_run.head_sha }}
jobs: jobs:
preflight: gate:
# Autodeploy defaults to OFF: pushes deploy only when the AUTODEPLOY repo
# variable is set to 'true' (Settings -> Actions -> Variables). A manual
# Run workflow always bypasses the switch: dispatching it is the explicit
# intent to deploy.
if: >- if: >-
github.ref == 'refs/heads/main' &&
(vars.AUTODEPLOY == 'true' || github.event_name == 'workflow_dispatch') && (vars.AUTODEPLOY == 'true' || github.event_name == 'workflow_dispatch') &&
(github.event_name != 'workflow_run' || (github.event_name != 'workflow_run' ||
(github.event.workflow_run.conclusion == 'success' && (github.event.workflow_run.conclusion == 'success' && github.event.workflow_run.head_branch == 'main'))
github.event.workflow_run.head_branch == 'main')) runs-on: homelab
runs-on: [self-hosted, linux, arch, homelab, prod]
timeout-minutes: 10 timeout-minutes: 10
outputs:
sha: ${{ steps.release.outputs.sha }}
steps: steps:
- name: Checkout repository - name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4 uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
with:
- name: Fetch and reset workstation fetch-depth: 0
shell: bash - name: Check successful CI and download the exact commit release
id: release
env:
GITEA_TOKEN: ${{ github.token }}
DEPLOY_REF: ${{ inputs.deploy_ref || 'main' }}
EVENT_SHA: ${{ github.event.workflow_run.head_sha }}
run: python3 .gitea/workflows/release.py gate --ref "$DEPLOY_REF" --event-sha "$EVENT_SHA"
- name: Submit durable deploy to workstation
run: bash .gitea/workflows/ssh-run.sh start
- name: Write the request result
if: always()
env:
REQUEST_RESULT: ${{ job.status }}
CHECKED_SHA: ${{ steps.release.outputs.sha }}
run: | run: |
set -euo pipefail if [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
./.gitea/workflows/ssh-run.sh preflight printf '## Deploy request\n\n- Result: **%s**\n- Checked commit: %s\n- Mode: %s\n' "$REQUEST_RESULT" "${CHECKED_SHA:-not checked}" "$DEPLOY_MODE" >>"$GITHUB_STEP_SUMMARY"
if [ "$REQUEST_RESULT" != success ]; then
echo 'Open the failed step log. If SSH submission failed, check the remote controller state.' >>"$GITHUB_STEP_SUMMARY"
fi
fi
validate: apply:
needs: [preflight] needs: [gate]
runs-on: [self-hosted, linux, arch, homelab, prod] runs-on: homelab
timeout-minutes: 20 timeout-minutes: 120
steps: steps:
- name: Checkout repository - name: Checkout checked commit
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4 uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
with:
- name: Dry-run manifests and check Secrets ref: ${{ needs.gate.outputs.sha }}
shell: bash - name: Follow validation and sequential Kubernetes / Compose apply
run: bash .gitea/workflows/ssh-run.sh apply
- name: Write the deploy result
if: always()
run: | run: |
set -euo pipefail if [ -f .gitea/workflows/ssh-run.sh ]; then
./.gitea/workflows/ssh-run.sh validate bash .gitea/workflows/ssh-run.sh summary
elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
echo 'Source checkout failed. The remote deploy state is unknown. Check the job log.' >>"$GITHUB_STEP_SUMMARY"
fi
apply-k8s: verify:
needs: [validate] needs: [gate, apply]
runs-on: [self-hosted, linux, arch, homelab, prod] if: always() && needs.gate.result == 'success'
# Apply only, no verification, so this is just the work itself: snapshot, runs-on: homelab
# then sequential `helm upgrade --install --wait --rollback-on-failure --timeout 10m`, then the apply loop. timeout-minutes: 130
# Verification has its own job and its own budget.
#
# 45 is roughly four times the measured cost of the stage, which is
# deliberately not raised on a theory:
#
# helm, healthy 3 no-op upgrades ~3-5 min
# helm, one release bad rollback-on-failure spends its 10m, ~10-15 min
# then rolls that one back
# apply loop ~40 manifests, 4 of which ~1 min
# resolve an image digest
# restart_stale_images 7.6s to find 8 workloads, ~0.5 min
# 9.8s to resolve their digests
#
# The helm figure is one release, not three: `set -e` aborts
# upgrade_helm_releases on the first failure, so a broken release costs
# 10m and the other two are never attempted. Multiplying 10m by three
# overstates the worst case by 20 minutes.
#
# The 45 minutes this was last raised to 45 were still not enough, and the
# job logs for those runs no longer exist, so what actually consumed the
# budget is not known - the two measurable candidates above account for
# ~15 of it. The unbounded `docker manifest inspect` against the registry's
# known hang mode is now bounded inside registry_digest (25s timeout, 3
# attempts): a dead registry fails each owned image after ~85s instead of
# hanging the stage, and a blinking one is retried instead of failing the
# whole apply file. Still open: make the stage announce which manifest it
# is working on, so a killed run leaves a diagnosable last line.
timeout-minutes: 45
steps: steps:
- name: Checkout repository - name: Checkout checked commit
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4 uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
with:
- name: Apply Kubernetes manifests ref: ${{ needs.gate.outputs.sha }}
shell: bash - name: Follow workload verification and recovery
run: bash .gitea/workflows/ssh-run.sh verify
- name: Write the deploy result
if: always()
run: | run: |
set -euo pipefail if [ -f .gitea/workflows/ssh-run.sh ]; then
./.gitea/workflows/ssh-run.sh apply-k8s bash .gitea/workflows/ssh-run.sh summary
elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
echo 'Source checkout failed. The remote deploy state is unknown. Check the job log.' >>"$GITHUB_STEP_SUMMARY"
fi
apply-compose:
needs: [validate]
runs-on: [self-hosted, linux, arch, homelab, prod]
timeout-minutes: 30
steps:
- name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
- name: Redeploy docker compose stacks
shell: bash
run: |
set -euo pipefail
./.gitea/workflows/ssh-run.sh apply-compose
# Watches the workloads this deploy changed and rolls back the ones that never
# became healthy. Runs even when the apply jobs failed, timed out or were
# cancelled — that is the whole point of splitting it out. `always()` is what
# lets it start after a failed dependency; the needs on apply-compose are a
# barrier, so verification begins only once both applies are done.
verify-k8s:
needs: [apply-k8s, apply-compose]
if: >-
always() &&
needs.apply-k8s.result != 'skipped' &&
needs.apply-compose.result != 'skipped'
runs-on: [self-hosted, linux, arch, homelab, prod]
# Not raised, because the arithmetic does not close.
#
# 32 workloads are under management and the wave width is 8, so the verify
# itself is 4 waves of ROLLOUT_TIMEOUT (300s) = 20 minutes worst case, when
# every rollout times out rather than converging. That is already 20 of 30.
#
# The other 10 would have to absorb rollback, and rollback_workloads is a
# serial `while read` loop at 300s per failed workload. 10 minutes buys two.
# Any larger number is buying a bigger multiple of an unbounded term rather
# than covering a known cost: 60 minutes buys eight, and 60 minutes is
# therefore not a bound, it is a guess with two digits.
#
# The number becomes derivable the moment rollback uses the same wave width
# as the verify: 32 failures then cost 4 waves = 20 minutes instead of 160,
# and 45 covers verify plus rollback at full width. That change is to the
# recovery path and is not folded into a timeout edit.
timeout-minutes: 30
steps:
- name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
- name: Verify workloads and roll back on failure
shell: bash
run: |
set -euo pipefail
./.gitea/workflows/ssh-run.sh verify-k8s
# Asks the public route of every active service whether it is actually
# serving, which the rollout check above structurally cannot: a pod can
# converge and still be crash-looping, or be listening on a port no Service
# points at, or answer 500.
#
# `always()` for the same reason verify-k8s has it, and it runs after that job
# specifically because a rollback is when a route most needs re-checking. The
# needs is a barrier, not a filter: whether verify-k8s passed, failed or was
# cancelled, the probes are what say whether the cluster is serving, and
# suppressing them on a rollback would hide the one run where the answer
# matters most.
smoke: smoke:
needs: [verify-k8s] needs: [gate, verify]
if: always() && needs.verify-k8s.result != 'skipped' if: always() && needs.gate.result == 'success'
runs-on: [self-hosted, linux, arch, homelab, prod] runs-on: homelab
timeout-minutes: 10 timeout-minutes: 15
steps: steps:
- name: Checkout repository - name: Checkout checked commit
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4 uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
with:
- name: Probe the public route of every active service ref: ${{ needs.gate.outputs.sha }}
shell: bash - name: Follow public route checks
run: bash .gitea/workflows/ssh-run.sh smoke
- name: Write the deploy result
if: always()
run: | run: |
set -euo pipefail if [ -f .gitea/workflows/ssh-run.sh ]; then
./.gitea/workflows/ssh-run.sh smoke bash .gitea/workflows/ssh-run.sh summary
elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
echo 'Source checkout failed. The remote deploy state is unknown. Check the job log.' >>"$GITHUB_STEP_SUMMARY"
fi
+51 -30
View File
@@ -13,9 +13,14 @@ here="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
# shellcheck source=tool-versions.env # shellcheck source=tool-versions.env
. "$here/tool-versions.env" . "$here/tool-versions.env"
TOOLS_DIR="${TOOLS_DIR:-${RUNNER_TEMP:-/tmp}/homelab-tools}" TOOLS_DIR="${TOOLS_DIR:-${XDG_CACHE_HOME:-$HOME/.cache}/homelab-ci}"
BIN_DIR="$TOOLS_DIR/bin" BIN_DIR="$TOOLS_DIR/bin"
mkdir -p "$BIN_DIR" mkdir -p "$BIN_DIR"
# A runner may accept overlapping workflows even though each workflow is sequential.
exec 9>"$TOOLS_DIR/install.lock"
flock -w 300 9
export UV_TOOL_DIR="$TOOLS_DIR/uv-tools"
export UV_CACHE_DIR="$TOOLS_DIR/uv-cache"
# The just-installed tools must resolve inside this script too: callers only # The just-installed tools must resolve inside this script too: callers only
# prepend BIN_DIR to PATH after the script exits, so a bare `uv` below would # prepend BIN_DIR to PATH after the script exits, so a bare `uv` below would
# miss the binary install_uv just placed (exit 127 on a clean runner). # miss the binary install_uv just placed (exit 127 on a clean runner).
@@ -48,7 +53,7 @@ esac
fetch() { fetch() {
# fetch <url> <dest> # fetch <url> <dest>
if command -v curl >/dev/null 2>&1; then if command -v curl >/dev/null 2>&1; then
curl -sSLf --retry 3 -o "$2" "$1" curl -sSLf --connect-timeout 15 --max-time 120 --retry 3 -o "$2" "$1"
elif command -v wget >/dev/null 2>&1; then elif command -v wget >/dev/null 2>&1; then
wget -q -O "$2" "$1" wget -q -O "$2" "$1"
else else
@@ -88,10 +93,13 @@ installed_version() {
# at_version <command> <expected> # at_version <command> <expected>
at_version() { at_version() {
case "$(installed_version "$1")" in local version expected="${2#v}"
*"$2"*) return 0 ;; version="$(installed_version "$1")"
*) return 1 ;; if [[ "$version" =~ (^|[^0-9.])v?([0-9]+(\.[0-9]+)+) ]]; then
esac [ "${BASH_REMATCH[2]}" = "$expected" ]
else
return 1
fi
} }
install_kubeconform() { install_kubeconform() {
@@ -176,6 +184,7 @@ install_pip_audit() {
} }
install_prettier() { install_prettier() {
install_node
if at_version prettier "${PRETTIER_VERSION}"; then if at_version prettier "${PRETTIER_VERSION}"; then
return 0 return 0
fi fi
@@ -236,29 +245,41 @@ install_actionlint() {
rm -rf "$tmp" rm -rf "$tmp"
} }
wanted=("$@") main() {
if [ "${#wanted[@]}" -eq 0 ]; then wanted=("$@")
wanted=(kubeconform shellcheck actionlint prettier ruff yamllint hadolint) if [ "${#wanted[@]}" -eq 0 ]; then
fi wanted=(node jq kubeconform shellcheck actionlint prettier ruff yamllint hadolint)
fi
for tool in "${wanted[@]}"; do for tool in "${wanted[@]}"; do
case "$tool" in case "$tool" in
kubeconform) install_kubeconform ;; kubeconform) install_kubeconform ;;
shellcheck) install_shellcheck ;; shellcheck) install_shellcheck ;;
jq) install_jq ;; jq) install_jq ;;
actionlint) install_actionlint ;; actionlint) install_actionlint ;;
prettier) install_prettier ;; prettier) install_prettier ;;
ruff) install_ruff ;; ruff) install_ruff ;;
yamllint) install_yamllint ;; yamllint) install_yamllint ;;
pip-audit) install_pip_audit ;; pip-audit) install_pip_audit ;;
hadolint) install_hadolint ;; hadolint) install_hadolint ;;
node) install_node ;; node) install_node ;;
uv) install_uv ;; uv) install_uv ;;
*) *)
echo "install-ci-tools: unknown tool: $tool" >&2 echo "install-ci-tools: unknown tool: $tool" >&2
exit 1 exit 1
;; ;;
esac esac
done done
printf '%s\n' "$BIN_DIR" for old in "$BIN_DIR"/node-* "$BIN_DIR"/prettier-*; do
[ -d "$old" ] || continue
case "$(basename "$old")" in
"node-$NODE_VERSION"|"prettier-$PRETTIER_VERSION") ;;
*) rm -rf "$old" ;;
esac
done
if [ -x "$BIN_DIR/uv" ]; then "$BIN_DIR/uv" cache prune >/dev/null; fi
printf '%s\n' "$BIN_DIR"
}
if [ "${BASH_SOURCE[0]}" = "$0" ]; then main "$@"; fi
+516
View File
@@ -0,0 +1,516 @@
#!/usr/bin/env python3
"""CI release artifacts and the SHA-specific Gitea deployment gate (stdlib only)."""
import argparse
import hashlib
import io
import itertools
import json
import os
import re
import shutil
import subprocess
import sys
import tempfile
import urllib.error
import urllib.parse
import urllib.request
import zipfile
from pathlib import Path
SHA = re.compile(r'[0-9a-f]{40}')
DIGEST = re.compile(r'sha256:[0-9a-f]{64}')
IMAGES = {
'error-pages': ('errorpages', 'errorpages/Dockerfile'),
'forust-homepage': ('homepages', 'homepages/Dockerfile.forust'),
'xdfnx-homepage': ('homepages', 'homepages/Dockerfile.xdfnx'),
}
def command(*args, **kwargs):
"""Arguments are passed directly to the executable, never to a shell."""
return subprocess.check_output(args, text=True, **kwargs).strip() # noqa: S603, S607
def validate_release(data, sha=None):
if data.get('version') != 1 or not SHA.fullmatch(data.get('sha', '')):
raise ValueError('Invalid release version or SHA')
if sha is not None and data['sha'] != sha:
raise ValueError('Release SHA does not match the checked CI commit')
expected = {f'gcr.forust.xyz/forust/{name}' for name in IMAGES}
if set(data.get('images', {})) != expected:
raise ValueError('Release must contain all owned images')
if not all(DIGEST.fullmatch(value) for value in data['images'].values()):
raise ValueError('Release has an invalid image digest')
if set(data.get('inputs', {})) != expected or not all(
re.fullmatch(r'[0-9a-f]{64}', value) for value in data['inputs'].values()
):
raise ValueError('Release has invalid build input fingerprints')
return data
class NoRedirect(urllib.request.HTTPRedirectHandler):
def redirect_request(self, _req, _fp, _code, _msg, _headers, _newurl):
return None
class Gitea:
def __init__(self):
self.origin = os.environ['GITHUB_SERVER_URL'].rstrip('/')
if urllib.parse.urlsplit(self.origin).scheme != 'https':
raise ValueError('Gitea API must use HTTPS')
self.repository = os.environ['GITHUB_REPOSITORY']
if not re.fullmatch(r'[\w.-]+/[\w.-]+', self.repository):
raise ValueError('Invalid Gitea repository')
self.token = os.environ['GITEA_TOKEN']
self.base = f'{self.origin}/api/v1/repos/{self.repository}'
def request(self, url, *, archive=False):
if not url.startswith(self.base + '/'):
raise ValueError('Refusing to send the Actions token to another origin')
req = urllib.request.Request(url, headers={'Authorization': f'token {self.token}'}) # noqa: S310 -- HTTPS origin validated above
opener = urllib.request.build_opener(NoRedirect())
try:
response = opener.open(req, timeout=30) # noqa: S310
except urllib.error.HTTPError as error:
if not archive or error.code not in (301, 302, 303, 307, 308):
raise RuntimeError(f'Gitea API returned HTTP {error.code}') from None
target = urllib.parse.urljoin(url, error.headers['Location'])
if urllib.parse.urlsplit(target).scheme != 'https':
raise ValueError('Artifact redirect must use HTTPS') from None
# Signed storage redirects must never receive the Gitea token.
response = urllib.request.urlopen(target, timeout=30) # noqa: S310
with response:
payload = response.read(8 * 1024 * 1024 + 1)
if len(payload) > 8 * 1024 * 1024:
raise ValueError('Gitea response exceeds 8 MiB')
return payload if archive else json.loads(payload)
def pages(self, path, key, **params):
for page in range(1, 101):
query = urllib.parse.urlencode({**params, 'page': page, 'limit': 50})
data = self.request(f'{self.base}/{path}?{query}')
entries = data[key]
yield from entries
if len(entries) < 50:
return
raise RuntimeError('Gitea pagination limit exceeded')
def successful_runs(self, sha=None):
params = {'branch': 'main', 'status': 'success', 'exclude_pull_requests': 'true'}
if sha:
params['head_sha'] = sha
for run in self.pages('actions/workflows/ci.yaml/runs', 'workflow_runs', **params):
if (
run.get('status') == 'completed'
and run.get('conclusion') == 'success'
and run.get('head_branch') == 'main'
and run.get('event') in ('push', 'workflow_dispatch')
and (run.get('repository') or {}).get('full_name') == self.repository
and (run.get('head_repository') or run.get('repository') or {}).get('full_name') == self.repository
and (sha is None or run.get('head_sha') == sha)
):
yield run
def release(self, run):
sha = run['head_sha']
jobs = list(self.pages(f'actions/runs/{run["id"]}/jobs', 'jobs'))
# A green workflow with a skipped build must not authorize a deploy.
if not any(job.get('name') == 'build' and job.get('conclusion') == 'success' for job in jobs):
raise ValueError('CI build job did not succeed')
artifacts = self.request(f'{self.base}/actions/runs/{run["id"]}/artifacts')['artifacts']
matching = [a for a in artifacts if a['name'] == f'release-{sha}' and not a.get('expired')]
if len(matching) != 1:
raise ValueError('CI release artifact is missing, expired or ambiguous; rerun CI')
blob = self.request(f'{self.base}/actions/artifacts/{matching[0]["id"]}/zip', archive=True)
with zipfile.ZipFile(io.BytesIO(blob)) as archive:
files = [entry for entry in archive.infolist() if not entry.is_dir()]
if len(files) != 1 or files[0].filename != 'release.json' or files[0].file_size > 256 * 1024:
raise ValueError('Unexpected release archive contents')
return validate_release(json.loads(archive.read(files[0])), sha)
def fingerprint(context, dockerfile):
tree = command('git', 'ls-tree', '-r', 'HEAD', '--', context, dockerfile, '.gitea/workflows/release.py')
return hashlib.sha256(tree.encode()).hexdigest()
def gate(output, requested_ref, event_sha):
command('git', 'fetch', '--quiet', 'origin', 'main')
if event_sha:
if not SHA.fullmatch(event_sha):
raise ValueError('Invalid workflow_run SHA')
sha = event_sha
else:
if requested_ref == 'main':
requested_ref = 'origin/main'
sha = command('git', 'rev-parse', '--verify', '--end-of-options', f'{requested_ref}^{{commit}}')
if not SHA.fullmatch(sha):
raise ValueError('Invalid deploy SHA')
command('git', 'merge-base', '--is-ancestor', sha, 'origin/main')
api = Gitea()
runs = list(api.successful_runs(sha))
if not runs:
raise ValueError(f'No successful main CI for {sha}; run CI before deploying')
release = api.release(max(runs, key=lambda run: run['id']))
output.write_text(json.dumps(release, indent=2) + '\n')
if os.environ.get('GITHUB_OUTPUT'):
with Path(os.environ['GITHUB_OUTPUT']).open('a') as stream:
stream.write(f'sha={sha}\n')
print(f'CI gate accepted {sha}')
def prepare_images(output):
sha = command('git', 'rev-parse', 'HEAD')
if sha != os.environ['GITHUB_SHA'] or not SHA.fullmatch(sha):
raise ValueError('Build checkout does not match GITHUB_SHA')
api = Gitea()
previous = None
for run in sorted(itertools.islice(api.successful_runs(), 50), key=lambda item: item['id'], reverse=True):
if str(run['id']) == os.environ.get('GITHUB_RUN_ID'):
continue
try:
previous = api.release(run)
break
except ValueError:
# Expired artifacts only cost a rebuild; mutable tags are never a fallback.
continue
targets = []
for name, (context, dockerfile) in IMAGES.items():
image = f'gcr.forust.xyz/forust/{name}'
inputs = fingerprint(context, dockerfile)
old_digest = (previous or {}).get('images', {}).get(image)
targets.append(
{
'name': name,
'image': image,
'context': context,
'dockerfile': dockerfile,
'inputs': inputs,
'reuse_digest': old_digest if (previous or {}).get('inputs', {}).get(image) == inputs else None,
}
)
output.write_text(json.dumps({'sha': sha, 'targets': targets}, indent=2) + '\n')
if os.environ.get('GITHUB_OUTPUT'):
with Path(os.environ['GITHUB_OUTPUT']).open('a') as stream:
stream.write('matrix=' + json.dumps({'include': targets}, separators=(',', ':')) + '\n')
print(f'Prepared {len(targets)} image jobs; {sum(t["reuse_digest"] is None for t in targets)} require builds')
def checked_plan(path):
data = json.loads(path.read_text())
sha = command('git', 'rev-parse', 'HEAD')
if data.get('sha') != sha or sha != os.environ['GITHUB_SHA'] or not SHA.fullmatch(sha):
raise ValueError('Image plan does not match the checked source commit')
targets = data.get('targets', [])
if sorted(t['name'] for t in targets) != sorted(IMAGES):
raise ValueError('Image plan must contain each owned image once')
for target in targets:
name = target['name']
context, dockerfile = IMAGES[name]
if (target['context'], target['dockerfile'], target['image']) != (
context,
dockerfile,
f'gcr.forust.xyz/forust/{name}',
) or target['inputs'] != fingerprint(context, dockerfile):
raise ValueError('Image plan has invalid build inputs')
if target['reuse_digest'] is not None and not DIGEST.fullmatch(target['reuse_digest']):
raise ValueError('Image plan has an invalid reuse digest')
return data
def build_images(output, report, name, plan):
data = checked_plan(plan)
sha = data['sha']
target = next(t for t in data['targets'] if t['name'] == name)
context, dockerfile = IMAGES[name]
docker_config = tempfile.mkdtemp(prefix='homelab-registry-')
builder_config = Path.home() / '.cache/homelab-ci/buildx'
builder_config.mkdir(parents=True, exist_ok=True)
env = {**os.environ, 'DOCKER_CONFIG': docker_config, 'BUILDX_CONFIG': str(builder_config)}
try:
report['phase'] = 'Registry login'
subprocess.run( # noqa: S603, S607
[
shutil.which('docker') or '/usr/bin/docker',
'login',
'gcr.forust.xyz',
'-u',
os.environ['REGISTRY_USERNAME'],
'--password-stdin',
],
input=os.environ['REGISTRY_PASSWORD'],
text=True,
check=True,
env=env,
)
report['phase'] = 'Prepare the builder'
builder = 'homelab-ci'
versions = dict(
re.findall(r'^([A-Z_]+)="([^"\n]+)"$', Path('.gitea/workflows/tool-versions.env').read_text(), re.MULTILINE)
)
image = versions['BUILDKIT_IMAGE']
signature = builder_config / 'homelab-ci-image'
exists = (
subprocess.run( # noqa: S603
[shutil.which('docker') or '/usr/bin/docker', 'buildx', 'inspect', builder],
capture_output=True,
env=env,
).returncode
== 0
)
if exists and (not signature.exists() or signature.read_text().strip() != image):
command('docker', 'buildx', 'rm', '--keep-state', builder, env=env)
exists = False
if not exists:
command(
'docker',
'buildx',
'create',
'--name',
builder,
'--driver',
'docker-container',
'--driver-opt',
f'image={image}',
'--buildkitd-config',
'.gitea/runner/buildkitd.toml',
env=env,
)
signature.write_text(image + '\n')
release = {'version': 1, 'sha': sha, 'images': {}, 'inputs': {}}
report['images'] = release['images']
report['phase'] = f'Build or reuse {name}'
report['current'] = name
image = f'gcr.forust.xyz/forust/{name}'
inputs = target['inputs']
old_digest = target['reuse_digest']
exists = False
if old_digest:
exists = (
subprocess.run( # noqa: S603, S607
[
shutil.which('docker') or '/usr/bin/docker',
'buildx',
'imagetools',
'inspect',
f'{image}@{old_digest}',
],
capture_output=True,
env=env,
timeout=60,
).returncode
== 0
)
if exists:
print(f'Reuse {name}: inputs unchanged')
digest = old_digest
report['reused'].append(name)
else:
print(f'Build {name}', flush=True)
metadata = Path(docker_config) / 'metadata.json'
command(
'docker',
'buildx',
'build',
'--builder',
builder,
'--platform',
'linux/amd64',
'--provenance=false',
'--cache-from',
f'type=registry,ref={image}:buildcache',
'--cache-to',
f'type=registry,ref={image}:buildcache,mode=max',
'--output',
f'type=image,name={image},push-by-digest=true,name-canonical=true,push=true',
'--metadata-file',
str(metadata),
'--file',
dockerfile,
context,
env=env,
)
digest = json.loads(metadata.read_text())['containerimage.digest']
report['built'].append(name)
release['images'][image] = digest
release['inputs'][image] = inputs
if not DIGEST.fullmatch(digest):
raise ValueError('Image job returned an invalid digest')
output.write_text(json.dumps(release, indent=2) + '\n')
report['current'] = None
report['phase'] = 'Release file saved'
finally:
# Cleanup errors must neither leak credentials nor mask the original build error.
try:
subprocess.run( # noqa: S603
[
shutil.which('docker') or '/usr/bin/docker',
'buildx',
'prune',
'--builder',
'homelab-ci',
'--force',
'--max-used-space',
'1gb',
],
env=env,
timeout=60,
)
except (OSError, subprocess.TimeoutExpired):
print('CI builder cache cleanup deferred', flush=True)
finally:
shutil.rmtree(docker_config)
def write_summary(lines):
path = os.environ.get('GITHUB_STEP_SUMMARY')
if path:
try:
with Path(path).open('a') as stream:
stream.write('\n'.join(lines) + '\n\n')
except OSError:
print('WARNING: cannot write the job summary')
def check_summary():
lines = [
f'## {os.environ["SUMMARY_CHECK"]}',
'',
f'- Commit: `{os.environ.get("GITHUB_SHA", "unknown")}`',
f'- Result: **{os.environ["SUMMARY_RESULT"]}**',
]
if os.environ.get('SUMMARY_FAILED_STEP'):
lines.append(f'- Failed step: {os.environ["SUMMARY_FAILED_STEP"]}')
if os.environ['SUMMARY_RESULT'] != 'success':
lines.append('- Open the failed step log for the error details.')
write_summary(lines)
def build(output, name, plan):
report = {'phase': 'Check the source commit', 'current': None, 'built': [], 'reused': [], 'images': {}}
result = 'failure'
try:
build_images(output, report, name, plan)
result = 'success'
finally:
lines = [
f'## Image release `{os.environ.get("GITHUB_SHA", "unknown")}`',
'',
f'- Result: **{result}**',
f'- Last stage: {report["phase"]}',
]
if result == 'failure':
lines.append('- No release from this build can be deployed. Open the failed step log.')
if report['current']:
lines.append(f'- Image at the failure: `{report["current"]}`')
for title, key in (('Built', 'built'), ('Reused from successful CI', 'reused')):
lines.extend(['', f'### {title}'])
lines.extend(f'- `{name}`' for name in report[key])
if not report[key]:
lines.append('- None')
lines.extend(['', '### Completed image digests'])
lines.extend(f'- `{image}@{digest}`' for image, digest in report['images'].items())
if not report['images']:
lines.append('- None')
write_summary(lines)
def render(stream, destination):
release = validate_release(json.loads(Path(os.environ['RELEASE_FILE']).read_text()), os.environ['DEPLOY_SHA'])
image_line = re.compile(
r"^(\s*(?:-\s*)?image:\s*)(['\"]?)(gcr\.forust\.xyz/forust/[\w.-]+)(?::[\w.-]+|@sha256:[0-9a-f]{64})\2(\s*(?:#.*)?)$"
)
rendered = []
for line in stream:
match = image_line.fullmatch(line.rstrip('\n'))
if match:
prefix, quote, image, tail = match.groups()
if image not in release['images']:
raise ValueError(f'Owned image missing from checked release: {image}')
line = f'{prefix}{quote}{image}@{release["images"][image]}{quote}{tail}\n'
elif re.match(r'\s*(?:-\s*)?image:', line) and 'gcr.forust.xyz/forust/' in line:
raise ValueError('Unsupported owned image syntax; refusing to apply a mutable tag')
rendered.append(line)
destination.writelines(rendered)
def finalize_images(output, fragments, plan):
data = checked_plan(plan)
sha = data['sha']
release = {'version': 1, 'sha': sha, 'images': {}, 'inputs': {}}
for name in IMAGES:
fragment = json.loads((fragments / f'image-{name}' / 'image.json').read_text())
image = f'gcr.forust.xyz/forust/{name}'
if fragment.get('sha') != sha or fragment.get('version') != 1 or set(fragment.get('images', {})) != {image}:
raise ValueError('Image job artifact is missing or belongs to another commit')
target = next(t for t in data['targets'] if t['name'] == name)
if fragment.get('inputs') != {image: target['inputs']}:
raise ValueError('Image artifact does not match the build plan')
release['images'].update(fragment['images'])
release['inputs'].update(fragment['inputs'])
validate_release(release, sha)
# Only a complete set of successful image jobs can publish the release tags.
docker_config = tempfile.mkdtemp(prefix='homelab-registry-')
env = {**os.environ, 'DOCKER_CONFIG': docker_config}
try:
subprocess.run( # noqa: S603, S607
[
shutil.which('docker') or '/usr/bin/docker',
'login',
'gcr.forust.xyz',
'-u',
os.environ['REGISTRY_USERNAME'],
'--password-stdin',
],
input=os.environ['REGISTRY_PASSWORD'],
text=True,
check=True,
env=env,
)
for image, digest in release['images'].items():
command(
'docker',
'buildx',
'imagetools',
'create',
'--prefer-index=false',
'--tag',
f'{image}:sha-{sha}',
f'{image}@{digest}',
env=env,
timeout=90,
)
output.write_text(json.dumps(release, indent=2) + '\n')
finally:
shutil.rmtree(docker_config)
def main():
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument('action', choices=('prepare', 'image', 'finalize', 'gate', 'render', 'check-summary'))
parser.add_argument('--output', type=Path, default=Path('release.json'))
parser.add_argument('--ref', default='main')
parser.add_argument('--event-sha', default='')
parser.add_argument('--image', choices=IMAGES)
parser.add_argument('--plan', type=Path, default=Path('build-plan.json'))
parser.add_argument('--fragments', type=Path, default=Path('artifacts'))
args = parser.parse_args()
if args.action == 'check-summary':
check_summary()
elif args.action == 'render':
render(sys.stdin, sys.stdout)
elif args.action == 'gate':
gate(args.output, args.ref, args.event_sha)
elif args.action == 'prepare':
prepare_images(args.output)
elif args.action == 'image':
if not args.image:
parser.error('--image is required')
build(args.output, args.image, args.plan)
else:
finalize_images(args.output, args.fragments, args.plan)
if __name__ == '__main__':
main()
+41 -17
View File
@@ -1,10 +1,26 @@
name: renovate-ci name: renovate-ci
on: on:
pull_request: # Read the workflow from the trusted base branch. PR code runs only on the
# unprivileged runner selected below.
pull_request_target:
paths:
- "renovate/**"
- ".gitea/workflows/renovate-ci.yaml"
- ".gitea/workflows/sync-renovate-configmap.sh"
- ".gitea/workflows/compose-lint.sh"
- ".gitea/workflows/install-ci-tools.sh"
- ".gitea/workflows/tool-versions.env"
push: push:
branches: branches:
- main - main
paths:
- "renovate/**"
- ".gitea/workflows/renovate-ci.yaml"
- ".gitea/workflows/sync-renovate-configmap.sh"
- ".gitea/workflows/compose-lint.sh"
- ".gitea/workflows/install-ci-tools.sh"
- ".gitea/workflows/tool-versions.env"
workflow_dispatch: workflow_dispatch:
permissions: permissions:
@@ -12,37 +28,47 @@ permissions:
jobs: jobs:
validate-renovate: validate-renovate:
runs-on: [self-hosted, linux, arch, homelab] runs-on: ${{ github.event_name == 'push' && github.ref == 'refs/heads/main' && 'homelab' || 'homelab-pr' }}
timeout-minutes: 20 timeout-minutes: 20
steps: steps:
- name: Checkout repository - name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4 uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
with:
ref: ${{ github.event_name == 'pull_request_target' && github.event.pull_request.head.sha || github.sha }}
# renovate/k8s/cronjob.yaml is the single source of truth for the image tag, # renovate/k8s/cronjob.yaml is the single source of truth for the version.
# so the same version that runs in the cluster is the one validated here. - name: Resolve the deployed Renovate version
- name: Resolve the deployed Renovate image
id: image id: image
shell: bash shell: bash
run: | run: |
set -euo pipefail set -euo pipefail
image="$(sed -n 's|.*image:[[:space:]]*\(renovate/renovate:[^[:space:]]*\).*|\1|p' \ image="$(sed -n 's|.*image:[[:space:]]*\(renovate/renovate:[^[:space:]]*\).*|\1|p' \
renovate/k8s/cronjob.yaml | head -1)" renovate/k8s/cronjob.yaml | head -1)"
if [ -z "$image" ]; then if [[ ! "$image" =~ ^renovate/renovate:([0-9]+\.[0-9]+\.[0-9]+)$ ]]; then
echo "::error::no renovate/renovate image found in renovate/k8s/cronjob.yaml" echo "::error::expected a pinned renovate/renovate semantic version in renovate/k8s/cronjob.yaml"
exit 1 exit 1
fi fi
echo "using $image" version="${BASH_REMATCH[1]}"
echo "image=$image" >> "$GITHUB_OUTPUT" echo "using Renovate $version"
printf 'version=%s\n' "$version" >> "$GITHUB_OUTPUT"
- name: Validate Renovate repository config - name: Prepare pinned validation tools
shell: bash shell: bash
run: | run: |
set -euo pipefail set -euo pipefail
docker run --rm \ tools_dir="$(bash .gitea/workflows/install-ci-tools.sh kubeconform node)"
-v "$PWD/renovate:/opt/renovate:ro" \ echo "$tools_dir" >> "$GITHUB_PATH"
-e RENOVATE_CONFIG_FILE=/opt/renovate/renovate.json \
"${{ steps.image.outputs.image }}" \ - name: Validate Renovate repository config
renovate-config-validator /opt/renovate/renovate.json shell: bash
env:
RENOVATE_VERSION: ${{ steps.image.outputs.version }}
run: |
set -euo pipefail
npm_cache="$(mktemp -d "${RUNNER_TEMP:-/tmp}/renovate-npm-cache.XXXXXXXX")"
trap 'rm -rf "$npm_cache"' EXIT
NPM_CONFIG_CACHE="$npm_cache" RENOVATE_CONFIG_FILE="$PWD/renovate/renovate.json" \
npm exec --yes --package="renovate@${RENOVATE_VERSION}" -- renovate-config-validator
# The CronJob cannot read the repository, so renovate/k8s/configmap.yaml # The CronJob cannot read the repository, so renovate/k8s/configmap.yaml
# carries an inlined copy of the config. Fail if it no longer matches. # carries an inlined copy of the config. Fail if it no longer matches.
@@ -56,8 +82,6 @@ jobs:
shell: bash shell: bash
run: | run: |
set -euo pipefail set -euo pipefail
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh kubeconform)"
export PATH="$tools_dir:$PATH"
kubeconform \ kubeconform \
-strict \ -strict \
-ignore-missing-schemas \ -ignore-missing-schemas \
+12 -6
View File
@@ -32,11 +32,14 @@ concurrency:
jobs: jobs:
run-renovate: run-renovate:
runs-on: [self-hosted, linux, arch, homelab] if: github.ref == 'refs/heads/main'
runs-on: homelab
timeout-minutes: 60 timeout-minutes: 60
steps: steps:
- name: Checkout repository - name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4 uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
with:
ref: refs/heads/main
# renovate/k8s/cronjob.yaml is the single source of truth for the image tag. # renovate/k8s/cronjob.yaml is the single source of truth for the image tag.
# Reading it here means this workflow validates and runs the exact version # Reading it here means this workflow validates and runs the exact version
@@ -48,21 +51,23 @@ jobs:
set -euo pipefail set -euo pipefail
image="$(sed -n 's|.*image:[[:space:]]*\(renovate/renovate:[^[:space:]]*\).*|\1|p' \ image="$(sed -n 's|.*image:[[:space:]]*\(renovate/renovate:[^[:space:]]*\).*|\1|p' \
renovate/k8s/cronjob.yaml | head -1)" renovate/k8s/cronjob.yaml | head -1)"
if [ -z "$image" ]; then if [[ ! "$image" =~ ^renovate/renovate:[0-9]+\.[0-9]+\.[0-9]+$ ]]; then
echo "::error::no renovate/renovate image found in renovate/k8s/cronjob.yaml" echo "::error::expected a pinned renovate/renovate semantic version in renovate/k8s/cronjob.yaml"
exit 1 exit 1
fi fi
echo "using $image" echo "using $image"
echo "image=$image" >> "$GITHUB_OUTPUT" printf 'image=%s\n' "$image" >> "$GITHUB_OUTPUT"
- name: Validate Renovate config - name: Validate Renovate config
shell: bash shell: bash
env:
RENOVATE_IMAGE: ${{ steps.image.outputs.image }}
run: | run: |
set -euo pipefail set -euo pipefail
docker run --rm \ docker run --rm \
-v "$PWD/renovate/renovate.json:/opt/renovate/renovate.json:ro" \ -v "$PWD/renovate/renovate.json:/opt/renovate/renovate.json:ro" \
-e RENOVATE_CONFIG_FILE=/opt/renovate/renovate.json \ -e RENOVATE_CONFIG_FILE=/opt/renovate/renovate.json \
"${{ steps.image.outputs.image }}" \ "$RENOVATE_IMAGE" \
renovate-config-validator renovate-config-validator
- name: Run Renovate - name: Run Renovate
@@ -73,6 +78,7 @@ jobs:
RENOVATE_REPOSITORIES: ${{ inputs.repositories }} RENOVATE_REPOSITORIES: ${{ inputs.repositories }}
RENOVATE_DRY_RUN: ${{ inputs.dry_run && 'full' || '' }} RENOVATE_DRY_RUN: ${{ inputs.dry_run && 'full' || '' }}
LOG_LEVEL: ${{ inputs.log_level }} LOG_LEVEL: ${{ inputs.log_level }}
RENOVATE_IMAGE: ${{ steps.image.outputs.image }}
run: | run: |
set -euo pipefail set -euo pipefail
@@ -89,4 +95,4 @@ jobs:
-e RENOVATE_CONFIG_FILE=/opt/renovate/renovate.json \ -e RENOVATE_CONFIG_FILE=/opt/renovate/renovate.json \
-e RENOVATE_BASE_DIR=/tmp/renovate \ -e RENOVATE_BASE_DIR=/tmp/renovate \
-e LOG_LEVEL="${LOG_LEVEL:-info}" \ -e LOG_LEVEL="${LOG_LEVEL:-info}" \
"${{ steps.image.outputs.image }}" "$RENOVATE_IMAGE"
+62 -63
View File
@@ -1,71 +1,70 @@
#!/usr/bin/env bash #!/usr/bin/env bash
# usage: ssh-run.sh <stage> # The SSH client submits once and follows durable stages on workstation.
# Runs one deploy-lib.sh stage on the workstation over SSH.
set -euo pipefail set -euo pipefail
: "${DEPLOY_HOST:?missing DEPLOY_HOST}" : "${DEPLOY_HOST:?missing DEPLOY_HOST}"
: "${DEPLOY_USER:?missing DEPLOY_USER}" : "${DEPLOY_USER:?missing DEPLOY_USER}"
: "${DEPLOY_KEY:?missing DEPLOY_SSH_KEY}" : "${DEPLOY_KEY:?missing DEPLOY_SSH_KEY}"
: "${DEPLOY_KNOWN_HOSTS:?configure pinned DEPLOY_KNOWN_HOSTS}"
deploy_port="${DEPLOY_PORT:-22}" : "${DEPLOY_RUN_ID:?missing DEPLOY_RUN_ID}"
deploy_path="${DEPLOY_PATH:-/srv/homelab}" [[ "$DEPLOY_USER" =~ ^[A-Za-z_][A-Za-z0-9_.-]*$ ]] || exit 1
deploy_path="$(printf '%s' "$deploy_path" | tr -d '\"' | tr -d '\r' | xargs)" [[ "$DEPLOY_HOST" =~ ^[A-Za-z0-9_.:-]+$ ]] || exit 1
[[ "$DEPLOY_RUN_ID" =~ ^[0-9]+-[0-9]+$ ]] || exit 1
# The private key is written to a per-run directory that is removed on exit, so a [[ "${DEPLOY_PORT:-22}" =~ ^[0-9]+$ ]] || exit 1
# failed or cancelled job cannot leave deploy credentials in the runner's temp
# directory. Do not use a fixed path: apply-k8s and apply-compose run in parallel.
key_dir="$(mktemp -d "${RUNNER_TEMP:-/tmp}/homelab-deploy-key.XXXXXXXX")" key_dir="$(mktemp -d "${RUNNER_TEMP:-/tmp}/homelab-deploy-key.XXXXXXXX")"
trap 'rm -rf "$key_dir"' EXIT INT TERM trap 'rm -rf "$key_dir"' EXIT
chmod 700 "$key_dir"
ssh_key="$key_dir/deploy_key" printf '%s\n' "$DEPLOY_KEY" >"$key_dir/key"
printf '%s\n' "$DEPLOY_KEY" > "$ssh_key" printf '%s\n' "$DEPLOY_KNOWN_HOSTS" >"$key_dir/known_hosts"
chmod 600 "$ssh_key" chmod 600 "$key_dir/key" "$key_dir/known_hosts"
ssh_opts=(-i "$key_dir/key" -p "${DEPLOY_PORT:-22}" -o BatchMode=yes -o StrictHostKeyChecking=yes
# A connection that died silently used to hang until the job timeout, and the -o "UserKnownHostsFile=$key_dir/known_hosts" -o ConnectTimeout=15
# stage was never re-run: one flaky TCP session cost a whole 45-minute apply. -o ServerAliveInterval=15 -o ServerAliveCountMax=4)
# ServerAlive* bounds how long a dead peer goes unnoticed, ConnectTimeout bounds controller=.local/lib/homelab-deploy/controller.py
# setup. Only exit 255 - ssh's own transport failures - is retried. A stage that case "${1:?start, apply, verify, smoke or summary required}" in
# fails on its own merits exits with the remote's status, so a real failure start)
# still surfaces its own log instead of burning three attempts. The stages are python3 - <<'PY' >"$key_dir/request.json"
# declarative applies, so re-running one that had already committed is harmless. import json
ssh_opts=( import os
-i "$ssh_key" -p "$deploy_port" from pathlib import Path
-o BatchMode=yes -o StrictHostKeyChecking=accept-new release = json.loads(Path('release.json').read_text())
-o ConnectTimeout=15 print(json.dumps({'release': release, 'mode': os.environ.get('DEPLOY_MODE', 'changed'),
-o ServerAliveInterval=15 -o ServerAliveCountMax=4 'refresh_images': os.environ.get('REFRESH_IMAGES', 'false') == 'true'}))
) PY
for attempt in 1 2 3; do
rc=0 rc=0
# apply-k8s and apply-compose are separate workflow jobs so the graph stays # shellcheck disable=SC2029 # The run ID and operation are validated local arguments, not remote variables.
# intact for the verify job, but on a single node they must not run at once: ssh "${ssh_opts[@]}" "$DEPLOY_USER@$DEPLOY_HOST" python3 "$controller" start "$DEPLOY_RUN_ID" <"$key_dir/request.json" || rc=$?
# host docker churn on top of cluster churn is what melts the node (load 40+, [ "$rc" -eq 0 ] && exit 0
# netbird/ssh die, helm is left pending-*). Serialize them on the workstation [ "$rc" -eq 255 ] || exit "$rc"
# with a shared lock; whoever arrives second waits. sleep 5
remote_cmd=(bash -se) done
case "$1" in exit "$rc"
apply-k8s | apply-compose)
remote_cmd=(flock -w 5400 /tmp/homelab-apply.lock bash -se)
;; ;;
apply|verify|smoke)
result=0
for attempt in 1 2 3; do
rc=0
# shellcheck disable=SC2029 # The run ID and operation are validated local arguments, not remote variables.
ssh "${ssh_opts[@]}" "$DEPLOY_USER@$DEPLOY_HOST" python3 "$controller" follow "$DEPLOY_RUN_ID" "$1" || rc=$?
[ "$rc" -eq 0 ] && break
[ "$rc" -eq 255 ] || { result="$rc"; break; }
echo "SSH disconnected; reconnecting to the existing deploy ($attempt/3)"
if [ "$attempt" -eq 3 ]; then result=255; break; fi
sleep 5
done
exit "$result"
;;
summary)
if [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
rc=0
# shellcheck disable=SC2029 # The run ID is validated above.
ssh "${ssh_opts[@]}" "$DEPLOY_USER@$DEPLOY_HOST" python3 "$controller" summary "$DEPLOY_RUN_ID" >"$key_dir/deploy-summary.md" || rc=$?
if [ "$rc" -eq 0 ]; then
cat "$key_dir/deploy-summary.md" >>"$GITHUB_STEP_SUMMARY" || echo "WARNING: cannot write the deploy summary"
else
echo 'Deploy summary is unavailable. The SSH connection failed or the controller did not respond. Check the job log.' >>"$GITHUB_STEP_SUMMARY" || true
fi
fi
;;
*) echo "Unknown SSH operation: $1" >&2; exit 1 ;;
esac esac
for attempt in 1 2 3; do
if [ "$attempt" -gt 1 ]; then
echo ":: warning::ssh transport failed, retrying (${attempt}/3)"
sleep $((attempt * 5))
fi
rc=0
# shellcheck disable=SC2029 # remote_cmd/ssh_opts expand on the client on purpose: they select the local ssh invocation, only the heredoc runs remotely.
ssh "${ssh_opts[@]}" "${DEPLOY_USER}@${DEPLOY_HOST}" \
env "REPO=$deploy_path" "APPLY_PRUNE=${APPLY_PRUNE:-false}" \
"DEPLOY_SHA=${DEPLOY_SHA:-}" "DEPLOY_SNAPSHOT_DIR=${DEPLOY_SNAPSHOT_DIR:-}" \
"STAGE=$1" "${remote_cmd[@]}" <<'EOF' || rc=$?
source "$REPO/.gitea/workflows/deploy-lib.sh"
run_stage "$STAGE"
EOF
[ "$rc" -eq 0 ] && break
[ "$rc" -ne 255 ] && break
done
if [ "$rc" -ne 0 ]; then
echo ":: error::stage $1 failed over ssh (exit $rc)"
fi
exit "$rc"
+3
View File
@@ -34,3 +34,6 @@ NODE_VERSION="22.23.3"
# Secret-reference regression tests parse rendered Kubernetes objects. # Secret-reference regression tests parse rendered Kubernetes objects.
JQ_VERSION="1.8.1" JQ_VERSION="1.8.1"
# BuildKit is the only auxiliary CI container; jobs themselves stay on the host.
BUILDKIT_IMAGE="moby/buildkit:v0.33.1"
-1
View File
@@ -94,7 +94,6 @@ replacements.txt
.idea .idea
# Temp files # Temp files
edu_master/temp/
temp/* temp/*
# Local-only tooling scratch space (pinned CI tools, verification scripts) # Local-only tooling scratch space (pinned CI tools, verification scripts)
tmp/ tmp/
+1 -1
View File
@@ -31,7 +31,7 @@ services:
- "traefik.http.routers.adguard-dev.entrypoints=websecure" - "traefik.http.routers.adguard-dev.entrypoints=websecure"
- "traefik.http.routers.adguard-dev.tls=true" - "traefik.http.routers.adguard-dev.tls=true"
# DoH Router # DoH Router
- "traefik.http.routers.dns-over-https.rule=(Host(`dns.forust.xyz` || Host(`adguard.forust.xyz`)) && PathPrefix(`/dns-query`))" - "traefik.http.routers.dns-over-https.rule=(Host(`dns.forust.xyz`) || Host(`adguard.forust.xyz`)) && PathPrefix(`/dns-query`)"
- "traefik.http.routers.dns-over-https.entrypoints=websecure" - "traefik.http.routers.dns-over-https.entrypoints=websecure"
- "traefik.http.routers.dns-over-https.tls.certresolver=letsencrypt" - "traefik.http.routers.dns-over-https.tls.certresolver=letsencrypt"
+1 -1
View File
@@ -20,7 +20,7 @@ spec:
spec: spec:
containers: containers:
- name: cloudflared - name: cloudflared
image: cloudflare/cloudflared:2026.9.3 image: cloudflare/cloudflared:2026.10.0
imagePullPolicy: IfNotPresent imagePullPolicy: IfNotPresent
args: args:
- tunnel - tunnel
-14
View File
@@ -1,14 +0,0 @@
EDU_LOGIN=your_edu_login_here
EDU_PASSWORD=your_edu_password_here
EDU_URL_LOGIN=https://edu.edu.vn.ua/user/login
EDU_URL_VERIFY=https://edu.edu.vn.ua/course/userlist
PHPSESSID_INTERVAL=10
USER_AGENT="Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/142.0.0.0 Safari/537.36"
WEBINAR_URL=https://edu.edu.vn.ua/webinar/useractive
WEBINAR_CHECK_INTERVAL=60
REDIS_HOST=redis
REDIS_PORT=6379
PLAYWRIGHT_WS=ws://playwright-service:3000/ws
TZ=Europe/Kyiv
WEBINAR_TELEGRAM_TOKEN=your_telegram_bot_token_here
WEBINAR_ADMIN_ID=123456789
-1
View File
@@ -1 +0,0 @@
1.56.0
-49
View File
@@ -1,49 +0,0 @@
services:
redis:
image: redis:8.10.2-alpine
restart: unless-stopped
volumes:
- redis-data:/data
healthcheck:
test: ["CMD", "redis-cli", "ping"]
interval: 5s
timeout: 3s
retries: 5
playwright-service:
image: mcr.microsoft.com/playwright:v1.56.0-jammy
restart: unless-stopped
command: npx -y playwright@1.56.0 run-server --port 3000 --path /ws
session-keeper:
build: ./phpsessid-bot
image: gcr.forust.xyz/forust/session-keeper:prod
pull_policy: build
env_file: .env
restart: unless-stopped
depends_on:
redis:
condition: service_healthy
healthcheck:
test: ["CMD-SHELL", "redis-cli -h redis EXISTS EDU_PHPSESSID | grep -q 1"]
interval: 30s
timeout: 5s
retries: 10
start_period: 60s
webinar-checker:
build: ./webinar-checker
image: gcr.forust.xyz/forust/webinar-checker:prod
pull_policy: build
env_file: .env
restart: unless-stopped
depends_on:
redis:
condition: service_healthy
session-keeper:
condition: service_healthy
playwright-service:
condition: service_started
volumes:
redis-data:
View File
Whitespace-only changes.
-96
View File
@@ -1,96 +0,0 @@
apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
name: edu-master-webinar
namespace: edu-master
labels:
release: prometheus-stack
spec:
groups:
- name: edu_master.webinar
rules:
# No successful webinar check for 5m (~2-3 missed 2-min checks).
# Catches: playwright hangs/timeouts, version skew, site changes, hung job.
# The last_success > 0 guard is mandatory: checker.py initialises
# last_success to 0, so without it `time() - 0` equals the current epoch
# and humanizeDuration renders ~20722d on every pod restart. Keep the
# duration expression on the left so $value stays the real gap.
- alert: WebinarCheckerNoSuccessfulCheck
expr: |
((time() - webinar_check_last_success_timestamp_seconds) > 300)
and (webinar_check_last_success_timestamp_seconds > 0)
and (webinar_check_last_run_timestamp_seconds > 0)
for: 2m
labels:
severity: critical
annotations:
summary: "Webinar checker has no successful check for 5m"
description: "edu-master/webinar-checker: last successful webinar check was {{ $value | humanizeDuration }} ago. Checks are failing or hanging (see consecutive failures alert). Notifications about new webinars are NOT being sent."
# Checks are running but none has ever succeeded since pod start.
# Split out from the rule above so a zeroed gauge never feeds
# humanizeDuration.
- alert: WebinarCheckerNeverSucceeded
expr: |
(webinar_check_last_success_timestamp_seconds == 0)
and (webinar_check_last_run_timestamp_seconds > 0)
for: 10m
labels:
severity: critical
annotations:
summary: "Webinar checker has never completed a successful check"
description: 'edu-master/webinar-checker: checks have been running for 10m but not one has ever succeeded since the pod started, so every check is failing. Check pod logs (Loki: {namespace="edu-master", container="webinar-checker"}).'
# Fast path: 3 consecutive failures (~6+ min at 2-min interval).
- alert: WebinarCheckerConsecutiveFailures
expr: |
webinar_check_consecutive_failures >= 3
for: 5m
labels:
severity: critical
annotations:
summary: "Webinar checker failing consecutively"
description: 'edu-master/webinar-checker: {{ $value }} consecutive webinar check failures (timeout / playwright error / page error). Check pod logs (Loki: {namespace="edu-master", container="webinar-checker"}).'
# Metrics endpoint not scraped for 10m: pod down, metrics server dead, or ServiceMonitor broken.
- alert: WebinarCheckerScrapeDown
expr: |
absent(webinar_check_last_run_timestamp_seconds) == 1
for: 10m
labels:
severity: critical
annotations:
summary: "Webinar checker metrics missing"
description: "edu-master/webinar-checker: no metrics series for 10m. Pod may be down, metrics server dead, or ServiceMonitor/Service broken. Webinar checks are unobserved."
# EDU session lost: session-keeper down or credentials expired. Without PHPSESSID every check is skipped.
- alert: EduPhpsessidMissing
expr: |
edu_phpsessid_present == 0
for: 10m
labels:
severity: critical
annotations:
summary: "EDU_PHPSESSID missing"
description: "edu-master: EDU_PHPSESSID absent from redis for 10m. Webinar/diari/schedule checks are all skipped. Check session-keeper logs and EDU credentials."
# Hard deps: checker and playwright deployments unavailable.
- alert: WebinarCheckerDeploymentDown
expr: |
kube_deployment_status_replicas_unavailable{deployment="webinar-checker", namespace="edu-master"} > 0
for: 10m
labels:
severity: critical
annotations:
summary: "Webinar checker deployment unavailable"
description: "edu-master/webinar-checker deployment has {{ $value }} unavailable replica(s) for 10m."
- alert: PlaywrightServiceDown
expr: |
kube_deployment_status_replicas_unavailable{deployment="playwright-service", namespace="edu-master"} > 0
for: 10m
labels:
severity: critical
annotations:
summary: "Playwright service unavailable"
description: "edu-master/playwright-service deployment has {{ $value }} unavailable replica(s) for 10m. All webinar/diari/schedule checks fail without it."
-69
View File
@@ -1,69 +0,0 @@
apiVersion: apps/v1
kind: Deployment
metadata:
name: playwright-service
namespace: edu-master
labels:
app: edu-master-playwright
spec:
replicas: 1
selector:
matchLabels:
app: edu-master-playwright
strategy:
type: Recreate
template:
metadata:
labels:
app: edu-master-playwright
spec:
containers:
- name: playwright
# renovate: datasource=docker depName=mcr.microsoft.com/playwright versioning=docker
image: mcr.microsoft.com/playwright:v1.56.0-jammy
imagePullPolicy: IfNotPresent
# p95 412M, max 478M over 7 days, no limit before. Request is set at p95
# so the pod is not an eviction candidate; the limit stays above 2x the
# request because browser page lifetimes are unpredictable.
resources:
requests:
cpu: "200m"
memory: "416Mi"
limits:
memory: "1Gi"
command:
- npx
- -y
- playwright@1.56.0
- run-server
- --port
- "3000"
- --path
- /ws
ports:
- containerPort: 3000
readinessProbe:
tcpSocket:
port: 3000
initialDelaySeconds: 5
periodSeconds: 10
timeoutSeconds: 3
livenessProbe:
tcpSocket:
port: 3000
initialDelaySeconds: 15
periodSeconds: 20
timeoutSeconds: 3
---
apiVersion: v1
kind: Service
metadata:
name: playwright-service
namespace: edu-master
spec:
selector:
app: edu-master-playwright
ports:
- name: ws
port: 3000
targetPort: 3000
-75
View File
@@ -1,75 +0,0 @@
apiVersion: apps/v1
kind: StatefulSet
metadata:
name: redis
namespace: edu-master
labels:
app: edu-master-redis
spec:
serviceName: redis
replicas: 1
selector:
matchLabels:
app: edu-master-redis
template:
metadata:
labels:
app: edu-master-redis
spec:
containers:
- name: redis
image: redis:8.10.2-alpine
imagePullPolicy: IfNotPresent
ports:
- containerPort: 6379
volumeMounts:
- name: redis-data
mountPath: /data
resources:
requests:
cpu: 25m
memory: 32Mi
limits:
cpu: 250m
memory: 128Mi
readinessProbe:
exec:
command: ["redis-cli", "ping"]
initialDelaySeconds: 5
periodSeconds: 5
timeoutSeconds: 3
livenessProbe:
exec:
command: ["redis-cli", "ping"]
initialDelaySeconds: 10
periodSeconds: 10
timeoutSeconds: 3
volumes:
- name: redis-data
persistentVolumeClaim:
claimName: redis-data-pvc
---
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: redis-data-pvc
namespace: edu-master
spec:
accessModes:
- ReadWriteOnce
resources:
requests:
storage: 1Gi
---
apiVersion: v1
kind: Service
metadata:
name: redis
namespace: edu-master
spec:
selector:
app: edu-master-redis
ports:
- name: redis
port: 6379
targetPort: 6379
@@ -1,50 +0,0 @@
# One-time Job to migrate redis state from docker compose to k8s (maintenance window).
# The .example file is not applied by the deploy pipeline (mask *.example.yaml).
#
# Runbook:
# 1. docker compose -f <repo>/edu_master/compose.yaml stop # SIGTERM -> redis will flush dump.rdb
# 2. docker run --rm -v edu_master_redis-data:/data \
# -v /tmp/edu-master-backup:/backup \
# redis:alpine sh -c "cp /data/dump.rdb /backup/ && ls -la /backup"
# 3. kubectl apply -f edu_master/k8s/namespace.yaml
# 4. kubectl apply -f <only the PVC from redis.yaml> # seed must come BEFORE redis pod starts
# 5. kubectl apply -f edu_master/k8s/restore-seed-job.yaml.example
# kubectl wait --for=condition=complete job/redis-restore-seed -n edu-master --timeout=120s
# 6. kubectl delete job redis-restore-seed -n edu-master
# 7. kubectl apply -f edu_master/k8s/ -R # apply remaining manifests
apiVersion: batch/v1
kind: Job
metadata:
name: redis-restore-seed
namespace: edu-master
spec:
backoffLimit: 2
ttlSecondsAfterFinished: 3600
template:
spec:
restartPolicy: Never
containers:
- name: seed
image: redis:alpine
command:
- /bin/sh
- -ec
- |
ls -la /backup
cp /backup/dump.rdb /data/dump.rdb
chmod 644 /data/dump.rdb
ls -la /data
volumeMounts:
- name: redis-data
mountPath: /data
- name: backup
mountPath: /backup
readOnly: true
volumes:
- name: redis-data
persistentVolumeClaim:
claimName: redis-data-pvc
- name: backup
hostPath:
path: /tmp/edu-master-backup
type: DirectoryOrCreate
-29
View File
@@ -1,29 +0,0 @@
apiVersion: v1
kind: Secret
metadata:
name: edu-master-secrets
namespace: edu-master
type: Opaque
stringData:
# Session keeper credentials
KEEPER_LOGIN: ""
KEEPER_PASSWORD: ""
KEEPER_INTERVAL: "10"
# EDU links
EDU_URL_BASE: "https://edu.edu.vn.ua"
EDU_URL_LOGIN: "/user/login"
EDU_URL_COURSES: "/course/userlist"
EDU_URL_WEBINAR: "/webinar/useractive"
# Playwright
USER_AGENT: ""
PLAYWRIGHT_WS: "ws://playwright-service:3000/ws"
# Webinar-checker
WEBINAR_TELEGRAM_TOKEN: ""
WEBINAR_ADMIN_ID: ""
WEBINAR_CHECK_INTERVAL: "60"
# Prometheus metrics endpoint (scraped via ServiceMonitor, alerts in k8s/alerts.yaml)
METRICS_PORT: "8000"
# Database
REDIS_HOST: "redis"
REDIS_PORT: "6379"
TZ: "Europe/Kyiv"
-15
View File
@@ -1,15 +0,0 @@
apiVersion: v1
kind: Service
metadata:
name: webinar-checker
namespace: edu-master
labels:
app: edu-master-webinar-checker
spec:
selector:
app: edu-master-webinar-checker
ports:
- name: metrics
port: 8000
targetPort: metrics
protocol: TCP
-55
View File
@@ -1,55 +0,0 @@
apiVersion: apps/v1
kind: Deployment
metadata:
annotations:
reloader.stakater.com/auto: "true"
name: session-keeper
namespace: edu-master
labels:
app: edu-master-session-keeper
spec:
replicas: 1
selector:
matchLabels:
app: edu-master-session-keeper
strategy:
type: Recreate
template:
metadata:
labels:
app: edu-master-session-keeper
spec:
initContainers:
- name: wait-redis
image: redis:8.10.2-alpine
command:
- /bin/sh
- -ec
- |
i=0
until redis-cli -h redis ping | grep -q PONG; do
i=$((i+1))
[ "$i" -ge 300 ] && echo "TIMEOUT: redis not ready" && exit 1
sleep 2
done
echo "redis is ready"
containers:
- name: session-keeper
image: gcr.forust.xyz/forust/session-keeper:prod
envFrom:
- secretRef:
name: edu-master-secrets
resources:
requests:
cpu: 25m
memory: 32Mi
limits:
cpu: 250m
memory: 128Mi
readinessProbe:
exec:
command: ["/bin/sh", "-ec", "redis-cli -h redis EXISTS EDU_PHPSESSID | grep -q 1"]
initialDelaySeconds: 15
periodSeconds: 30
timeoutSeconds: 5
failureThreshold: 10
-77
View File
@@ -1,77 +0,0 @@
apiVersion: apps/v1
kind: Deployment
metadata:
annotations:
reloader.stakater.com/auto: "true"
name: webinar-checker
namespace: edu-master
labels:
app: edu-master-webinar-checker
spec:
replicas: 1
selector:
matchLabels:
app: edu-master-webinar-checker
strategy:
type: Recreate
template:
metadata:
labels:
app: edu-master-webinar-checker
spec:
# Enforces dependency order like compose depends_on:
# redis healthy -> session-keeper healthy (EXISTS EDU_PHPSESSID) -> playwright started
initContainers:
- name: wait-deps
image: redis:8.10.2-alpine
command:
- /bin/sh
- -ec
- |
i=0
until redis-cli -h redis ping | grep -q PONG; do
i=$((i+1))
[ "$i" -ge 300 ] && echo "TIMEOUT: redis not ready" && exit 1
sleep 2
done
echo "redis ok"
until [ "$(redis-cli -h redis EXISTS EDU_PHPSESSID)" = "1" ]; do
i=$((i+1))
[ "$i" -ge 300 ] && echo "TIMEOUT: no PHPSESSID (session-keeper down?)" && exit 1
sleep 2
done
echo "PHPSESSID ok"
until nc -z playwright-service 3000; do
i=$((i+1))
[ "$i" -ge 300 ] && echo "TIMEOUT: playwright-service not reachable" && exit 1
sleep 2
done
echo "playwright ok"
containers:
- name: webinar-checker
image: gcr.forust.xyz/forust/webinar-checker:prod
ports:
- name: metrics
containerPort: 8000
protocol: TCP
readinessProbe:
httpGet:
path: /health
port: metrics
periodSeconds: 10
timeoutSeconds: 3
failureThreshold: 12
initialDelaySeconds: 10
envFrom:
- secretRef:
name: edu-master-secrets
env:
- name: TZ
value: "Europe/Kyiv"
resources:
requests:
cpu: "50m"
memory: "192Mi"
limits:
cpu: "600m"
memory: "384Mi"
-15
View File
@@ -1,15 +0,0 @@
FROM python:3.11-slim
WORKDIR /app
# Install system dependencies
RUN apt-get update && apt-get install -y --no-install-recommends redis-tools && rm -rf /var/lib/apt/lists/*
# Install dependencies
RUN pip install --no-cache-dir requests==2.32.3 redis==5.2.1
# Copy application code
COPY . .
# Run the bot
CMD ["python", "bot.py"]
-132
View File
@@ -1,132 +0,0 @@
import logging
import os
import time
from datetime import datetime
import redis
import requests
# Configure logging
logging.basicConfig(level=logging.INFO, format='%(asctime)s - %(levelname)s - %(message)s')
logger = logging.getLogger(__name__)
# Load configuration (adapted to .env keys)
def _env(key, default=None):
v = os.getenv(key, default)
if isinstance(v, str) and len(v) >= 2 and ((v[0] == '"' and v[-1] == '"') or (v[0] == "'" and v[-1] == "'")):
return v[1:-1]
return v
LOGIN = _env('KEEPER_LOGIN')
PASSWORD = _env('KEEPER_PASSWORD')
EDU_BASE = _env('EDU_URL_BASE', 'https://edu.edu.vn.ua')
EDU_LOGIN_PATH = _env('EDU_URL_LOGIN', '/user/login')
EDU_COURSES_PATH = _env('EDU_URL_COURSES', '/course/userlist')
URL_LOGIN = f'{EDU_BASE.rstrip("/")}/{EDU_LOGIN_PATH.lstrip("/")}'
URL_VERIFY = f'{EDU_BASE.rstrip("/")}/{EDU_COURSES_PATH.lstrip("/")}'
INTERVAL = int(_env('KEEPER_INTERVAL', 10))
USER_AGENT = _env(
'USER_AGENT',
'Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/142.0.0.0 Safari/537.36',
)
REDIS_HOST = _env('REDIS_HOST', 'redis')
REDIS_PORT = int(_env('REDIS_PORT', 6379))
SUCCESS_FILE = '/tmp/last_success' # noqa: S108
def touch_success_file():
"""Updates the timestamp of the success file for healthchecks."""
try:
with open(SUCCESS_FILE, 'w') as f:
f.write(str(datetime.now().timestamp()))
except Exception as e:
logger.error(f'Failed to touch success file: {e}')
def main():
logger.info('Starting Session Keeper Bot')
# Connect to Redis
try:
redis_client = redis.Redis(host=REDIS_HOST, port=REDIS_PORT, decode_responses=True)
redis_client.ping()
logger.info(f'Connected to Redis at {REDIS_HOST}:{REDIS_PORT}')
except Exception as e:
logger.error(f'Failed to connect to Redis: {e}')
return
session = requests.Session()
# Set headers
headers = {
'User-Agent': USER_AGENT,
'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,image/apng,*/*;q=0.8,application/signed-exchange;v=b3;q=0.7',
'Accept-Language': 'en-US,en;q=0.9',
'Cache-Control': 'max-age=0',
'Upgrade-Insecure-Requests': '1',
'Sec-Fetch-Site': 'same-origin',
'Sec-Fetch-Mode': 'navigate',
'Sec-Fetch-User': '?1',
'Sec-Fetch-Dest': 'document',
'Sec-Ch-Ua': '"Not_A Brand";v="99", "Chromium";v="142"',
'Sec-Ch-Ua-Mobile': '?0',
'Sec-Ch-Ua-Platform': '"Linux"',
'Accept-Encoding': 'gzip, deflate, br',
'Priority': 'u=0, i',
}
session.headers.update(headers)
while True:
try:
logger.info('Attempting login...')
# Login payload
payload = {'login': LOGIN, 'password': PASSWORD}
# Perform Login
# Note: The user request shows a POST to /user/login with form data
# We need to make sure we handle the PHPSESSID correctly.
# If we already have a PHPSESSID, requests will send it.
login_response = session.post(URL_LOGIN, data=payload, allow_redirects=True)
logger.info(f'Login Response Status: {login_response.status_code}')
logger.info(f'Cookies after login: {session.cookies.get_dict()}')
# Verify Session
logger.info('Verifying session...')
verify_response = session.get(URL_VERIFY, allow_redirects=False)
logger.info(f'Verify Response Status: {verify_response.status_code}')
if verify_response.status_code == 200:
logger.info('Session verification SUCCESS (200 OK).')
touch_success_file()
# Save PHPSESSID to Redis
phpsessid = session.cookies.get('PHPSESSID')
if phpsessid:
try:
redis_client.set('EDU_PHPSESSID', phpsessid)
logger.info(f'Saved PHPSESSID to Redis: {phpsessid}')
except Exception as e:
logger.error(f'Failed to save PHPSESSID to Redis: {e}')
elif verify_response.status_code == 302:
logger.warning('Session verification FAILED (302 Redirect). Session might be invalid.')
else:
logger.warning(f'Session verification returned unexpected status: {verify_response.status_code}')
except Exception as e:
logger.error(f'An error occurred: {e}')
logger.info(f'Sleeping for {INTERVAL} minutes...')
time.sleep(INTERVAL * 60)
if __name__ == '__main__':
main()
-13
View File
@@ -1,13 +0,0 @@
FROM python:3.11-slim
WORKDIR /app
# renovate: datasource=pypi depName=playwright versioning=pep440
ARG PLAYWRIGHT_VERSION=1.56.0
# Install dependencies - PLAYWRIGHT_VERSION is single-source, renovate updates ARG above and all other places via regexManagers
RUN pip install --no-cache-dir pip==25.0.1 && pip install --no-cache-dir playwright==${PLAYWRIGHT_VERSION} redis==5.2.1 requests==2.32.3 "python-telegram-bot[job-queue]==21.10"
COPY checker.py .
CMD ["python", "checker.py"]
File diff suppressed because it is too large. Load diff
+2
View File
@@ -20,6 +20,8 @@ data:
GITEA__mailer__ENABLED: "false" GITEA__mailer__ENABLED: "false"
GITEA__metrics__ENABLED: "true"
# No code/issue search needed: bleve reindexes the whole issue index on # No code/issue search needed: bleve reindexes the whole issue index on
# every pod restart (cron.rebuild_issue_indexer RUN_AT_START) and hammers # every pod restart (cron.rebuild_issue_indexer RUN_AT_START) and hammers
# the rotational disk for an hour. "db" serves issue search from postgres. # the rotational disk for an hour. "db" serves issue search from postgres.
+2
View File
@@ -3,6 +3,8 @@ kind: Service
metadata: metadata:
name: gitea-service name: gitea-service
namespace: gitea namespace: gitea
labels:
app: gitea
spec: spec:
selector: selector:
app: gitea app: gitea
+2 -1
View File
@@ -7,7 +7,8 @@ spec:
entryPoints: entryPoints:
- websecure - websecure
routes: routes:
- match: Host(`gitea.forust.xyz`) || Host(`git.forust.xyz`) # Metrics are scraped directly through the cluster Service.
- match: (Host(`gitea.forust.xyz`) || Host(`git.forust.xyz`)) && !PathPrefix(`/metrics`)
kind: Rule kind: Rule
services: services:
- name: gitea-service - name: gitea-service
+16
View File
@@ -0,0 +1,16 @@
apiVersion: monitoring.coreos.com/v1
kind: ServiceMonitor
metadata:
name: gitea
namespace: gitea
labels:
release: prometheus-stack
spec:
selector:
matchLabels:
app: gitea
endpoints:
- port: http
path: /metrics
interval: 30s
scrapeTimeout: 10s
+1 -1
View File
@@ -71,7 +71,7 @@ spec:
name: glance-config name: glance-config
- name: glance-assets - name: glance-assets
configMap: configMap:
name: glance-config name: glance-assets
- name: docker-socket - name: docker-socket
hostPath: hostPath:
path: /var/run/docker.sock path: /var/run/docker.sock
+1 -1
View File
@@ -7,7 +7,7 @@
# # Dev server_url # # Dev server_url
# server_url: https://hs.dev_internal_domain.internal # server_url: https://hs.dev_internal_domain.internal
listen_addr: 0.0.0.0:8080 listen_addr: 0.0.0.0:8080
metrics_listen_addr: 127.0.0.1:9090 metrics_listen_addr: 0.0.0.0:9090
grpc_listen_addr: 127.0.0.1:50443 grpc_listen_addr: 127.0.0.1:50443
grpc_allow_insecure: false grpc_allow_insecure: false
noise: noise:
@@ -3,6 +3,8 @@ kind: Service
metadata: metadata:
name: headscale-server-external name: headscale-server-external
namespace: headscale namespace: headscale
labels:
app: headscale
spec: spec:
ports: ports:
- port: 8080 - port: 8080
+18
View File
@@ -0,0 +1,18 @@
apiVersion: operator.victoriametrics.com/v1beta1
kind: VMServiceScrape
metadata:
name: headscale
namespace: headscale
labels:
release: prometheus-stack
spec:
# The external Service has a manually managed EndpointSlice, not Endpoints.
discoveryRole: endpointslice
selector:
matchLabels:
app: headscale
endpoints:
- port: metrics
path: /metrics
interval: 30s
scrapeTimeout: 10s
+1 -1
View File
@@ -1,7 +1,7 @@
services: services:
homarr: homarr:
container_name: homarr container_name: homarr
image: ghcr.io/homarr-labs/homarr:v2.1.2 image: ghcr.io/homarr-labs/homarr:v2.2.0
restart: unless-stopped restart: unless-stopped
volumes: volumes:
- ./appdata:/appdata - ./appdata:/appdata
+1 -1
View File
@@ -32,7 +32,7 @@ spec:
serviceAccountName: homarr serviceAccountName: homarr
containers: containers:
- name: homarr - name: homarr
image: ghcr.io/homarr-labs/homarr:v2.1.2 image: ghcr.io/homarr-labs/homarr:v2.2.0
envFrom: envFrom:
- configMapRef: - configMapRef:
name: homarr-config name: homarr-config
+4
View File
@@ -6,6 +6,10 @@ metadata:
data: data:
TZ: "Europe/Bratislava" TZ: "Europe/Bratislava"
IMMICH_TELEMETRY_INCLUDE: "all"
IMMICH_API_METRICS_PORT: "8081"
IMMICH_MICROSERVICES_METRICS_PORT: "8082"
# The database in this namespace, not the shared one in the database # The database in this namespace, not the shared one in the database
# namespace: v3 needs VectorChord, and only the dedicated image carries it. # namespace: v3 needs VectorChord, and only the dedicated image carries it.
DB_HOSTNAME: "immich-postgres" DB_HOSTNAME: "immich-postgres"
+12
View File
@@ -3,6 +3,8 @@ kind: Service
metadata: metadata:
name: immich-service name: immich-service
namespace: immich namespace: immich
labels:
app: immich
spec: spec:
selector: selector:
app: immich app: immich
@@ -10,6 +12,12 @@ spec:
- name: http - name: http
port: 2283 port: 2283
targetPort: 2283 targetPort: 2283
- name: api-metrics
port: 8081
targetPort: api-metrics
- name: worker-metrics
port: 8082
targetPort: worker-metrics
--- ---
apiVersion: apps/v1 apiVersion: apps/v1
kind: Deployment kind: Deployment
@@ -41,6 +49,10 @@ spec:
ports: ports:
- name: http - name: http
containerPort: 2283 containerPort: 2283
- name: api-metrics
containerPort: 8081
- name: worker-metrics
containerPort: 8082
volumeMounts: volumeMounts:
- name: immich-data - name: immich-data
mountPath: /data mountPath: /data
+20
View File
@@ -0,0 +1,20 @@
apiVersion: monitoring.coreos.com/v1
kind: ServiceMonitor
metadata:
name: immich
namespace: immich
labels:
release: prometheus-stack
spec:
selector:
matchLabels:
app: immich
endpoints:
- port: api-metrics
path: /metrics
interval: 30s
scrapeTimeout: 10s
- port: worker-metrics
path: /metrics
interval: 30s
scrapeTimeout: 10s
+1 -1
View File
@@ -1,6 +1,6 @@
services: services:
n8n: n8n:
image: docker.n8n.io/n8nio/n8n:2.42.3 image: docker.n8n.io/n8nio/n8n:2.43.0
container_name: n8n container_name: n8n
restart: unless-stopped restart: unless-stopped
environment: environment:
+1 -1
View File
@@ -31,7 +31,7 @@ spec:
spec: spec:
containers: containers:
- name: n8n - name: n8n
image: docker.n8n.io/n8nio/n8n:2.42.3 image: docker.n8n.io/n8nio/n8n:2.43.0
envFrom: envFrom:
- configMapRef: - configMapRef:
name: n8n-config name: n8n-config
+109
View File
@@ -0,0 +1,109 @@
#!/bin/sh
set -eu
umask 077
TEMPLATE_PATH=/opt/netbird/config.template.yaml
RENDERED_PATH=/run/netbird/config.yaml
RELAY_SECRET_PATH=/run/secrets/relay_auth_secret
ENCRYPTION_KEY_PATH=/run/secrets/datastore_encryption_key
is_valid_proxy_subnet() {
candidate="$1"
case "$candidate" in
0.0.0.0/0)
return 1
;;
*/*)
address="${candidate%%/*}"
prefix="${candidate#*/}"
;;
*)
return 1
;;
esac
case "$prefix" in
0|[1-9]|[1-2][0-9]|3[0-2]) ;;
*)
return 1
;;
esac
old_ifs="$IFS"
IFS=.
# shellcheck disable=SC2086
set -- $address
IFS="$old_ifs"
[ "$#" -eq 4 ] || return 1
for octet do
case "$octet" in
0|[1-9]|[1-9][0-9]|1[0-9][0-9]|2[0-4][0-9]|25[0-5]) ;;
*)
return 1
;;
esac
done
}
read_secret() {
secret_path="$1"
if [ ! -r "$secret_path" ]; then
echo "Required secret is not readable: $secret_path" >&2
exit 1
fi
secret_value="$(cat "$secret_path")"
if [ -z "$secret_value" ]; then
echo "Required secret is empty: $secret_path" >&2
exit 1
fi
printf '%s' "$secret_value"
}
if [ -z "${NETBIRD_DOMAIN:-}" ]; then
echo "NETBIRD_DOMAIN must be set" >&2
exit 1
fi
case "$NETBIRD_DOMAIN" in
*[!A-Za-z0-9.-]*)
echo "NETBIRD_DOMAIN contains unsupported characters" >&2
exit 1
;;
esac
if [ -z "${NETBIRD_PROXY_SUBNET:-}" ] || [ "$NETBIRD_PROXY_SUBNET" = "auto" ]; then
echo "NETBIRD_PROXY_SUBNET must be an explicit IPv4 CIDR; run netbird/setup.sh first" >&2
exit 1
fi
if ! is_valid_proxy_subnet "$NETBIRD_PROXY_SUBNET"; then
echo "NETBIRD_PROXY_SUBNET must be a non-default IPv4 CIDR, for example 172.20.0.0/16" >&2
exit 1
fi
if [ "$#" -ne 2 ] || [ "$1" != "--config" ] || [ "$2" != "$RENDERED_PATH" ]; then
echo "Expected: --config $RENDERED_PATH" >&2
exit 1
fi
relay_secret="$(read_secret "$RELAY_SECRET_PATH")"
encryption_key="$(read_secret "$ENCRYPTION_KEY_PATH")"
mkdir -p "$(dirname "$RENDERED_PATH")"
sed \
-e "s|__NETBIRD_DOMAIN__|${NETBIRD_DOMAIN}|g" \
-e "s|__NETBIRD_AUTH_SECRET__|${relay_secret}|g" \
-e "s|__NETBIRD_ENCRYPTION_KEY__|${encryption_key}|g" \
-e "s|__NETBIRD_PROXY_SUBNET__|${NETBIRD_PROXY_SUBNET}|g" \
"$TEMPLATE_PATH" >"$RENDERED_PATH"
if grep -q '__NETBIRD_' "$RENDERED_PATH"; then
echo "Rendered NetBird configuration still contains unresolved placeholders" >&2
exit 1
fi
exec /go/bin/netbird-server "$@"
+9
View File
@@ -3,6 +3,8 @@ kind: Service
metadata: metadata:
name: netbird-server-service name: netbird-server-service
namespace: netbird namespace: netbird
labels:
app: netbird-server
spec: spec:
selector: selector:
app: netbird-server app: netbird-server
@@ -11,6 +13,10 @@ spec:
name: http name: http
targetPort: 80 targetPort: 80
protocol: TCP protocol: TCP
- port: 9090
name: metrics
targetPort: metrics
protocol: TCP
- port: 3478 - port: 3478
name: stun name: stun
targetPort: 3478 targetPort: 3478
@@ -59,6 +65,9 @@ spec:
- containerPort: 80 - containerPort: 80
name: http name: http
protocol: TCP protocol: TCP
- containerPort: 9090
name: metrics
protocol: TCP
- containerPort: 3478 - containerPort: 3478
name: stun name: stun
protocol: UDP protocol: UDP
@@ -1,14 +1,14 @@
apiVersion: monitoring.coreos.com/v1 apiVersion: monitoring.coreos.com/v1
kind: ServiceMonitor kind: ServiceMonitor
metadata: metadata:
name: webinar-checker name: netbird-server
namespace: edu-master namespace: netbird
labels: labels:
release: prometheus-stack release: prometheus-stack
spec: spec:
selector: selector:
matchLabels: matchLabels:
app: edu-master-webinar-checker app: netbird-server
endpoints: endpoints:
- port: metrics - port: metrics
path: /metrics path: /metrics
+38
View File
@@ -0,0 +1,38 @@
#!/usr/bin/env bash
# Prepare local Compose configuration without replacing existing credentials.
set -euo pipefail
cd "$(dirname "${BASH_SOURCE[0]}")"
umask 077
if [ ! -f .env ]; then
cp .env.example .env
fi
if grep -q '^NETBIRD_PROXY_SUBNET=auto$' .env; then
subnet="$(docker network inspect proxy --format '{{range .IPAM.Config}}{{println .Subnet}}{{end}}' | awk '/^[0-9]+\./ { print; exit }')"
if [ -z "$subnet" ]; then
echo "No IPv4 subnet found on the Docker proxy network. Set NETBIRD_PROXY_SUBNET in .env." >&2
exit 1
fi
# The detected value must be safe to substitute into the env file.
if [[ ! "$subnet" =~ ^[0-9.]+/[0-9]+$ ]]; then
echo "Unexpected Docker network subnet: $subnet" >&2
exit 1
fi
sed -i "s|^NETBIRD_PROXY_SUBNET=auto$|NETBIRD_PROXY_SUBNET=$subnet|" .env
fi
mkdir -p secrets
chmod 700 secrets
for name in relay-auth-secret datastore-encryption-key; do
path="secrets/$name"
if [ -e "$path" ]; then
if [ ! -s "$path" ]; then
echo "Existing secret is empty: $path. Restore it before continuing." >&2
exit 1
fi
else
openssl rand -base64 32 >"$path"
fi
chmod 600 "$path"
done
printf '%s\n' 'Local files are ready. Review .env, then run docker compose config --quiet.'
+1 -1
View File
@@ -1,6 +1,6 @@
services: services:
netronome: netronome:
image: ghcr.io/autobrr/netronome:v0.15.0 image: ghcr.io/autobrr/netronome:v0.16.0
restart: unless-stopped restart: unless-stopped
container_name: netronome container_name: netronome
ports: ports:
+1 -1
View File
@@ -34,7 +34,7 @@ spec:
spec: spec:
containers: containers:
- name: netronome - name: netronome
image: ghcr.io/autobrr/netronome:v0.15.0 image: ghcr.io/autobrr/netronome:v0.16.0
ports: ports:
- name: netronome-port - name: netronome-port
protocol: TCP protocol: TCP
+16
View File
@@ -0,0 +1,16 @@
PAPERLESS_URL=https://papers.forust.xyz
PAPERLESS_ALLOWED_HOSTS=papers.forust.xyz,papers.workstation.internal
PAPERLESS_CSRF_TRUSTED_ORIGINS=https://papers.forust.xyz,https://papers.workstation.internal
PAPERLESS_TIME_ZONE=Europe/Bratislava
PAPERLESS_REDIS=redis://valkey:6379
PAPERLESS_DBENGINE=postgresql
PAPERLESS_DBHOST=homelab-postgres
PAPERLESS_DBNAME=paperless
PAPERLESS_DBUSER=paperless
PAPERLESS_DBPASS=<SET_THE_SAME_VALUE_AS_SHARED_POSTGRES_PAPERLESS_DB_PASSWORD>
PAPERLESS_OCR_LANGUAGE=rus+eng
PAPERLESS_OCR_LANGUAGES=rus
PAPERLESS_TASK_WORKERS=1
PAPERLESS_ADMIN_USER=admin
PAPERLESS_SECRET_KEY=<GENERATE_WITH_python3_-c_import_secrets;_print(secrets.token_urlsafe(64))>
PAPERLESS_ADMIN_PASSWORD=<SET_A_LONG_UNIQUE_PASSWORD>
+79
View File
@@ -0,0 +1,79 @@
# Paperless-ngx
Paperless-ngx runs in the `paperless` namespace. It uses the shared PostgreSQL
service in the `database` namespace and Valkey for its task queue. The document
library, exports, and consume folder are stored on the `local-path-retain`
volume. The PVC size is fixed at 50 GiB because this storage class does not
support volume expansion.
The local route is `https://papers.workstation.internal`; the public route is
`https://papers.forust.xyz`. Both use TLS. Paperless keeps its own login and
password authentication. OCR is configured for Russian and English documents.
Keep `papers.forust.xyz` in the existing Cloudflare DDNS `DOMAINS` setting so
the public record follows the workstation address.
## Compose alternative
`compose.yaml` is an alternative to the active Kubernetes deployment. Do not
run both at the same time: they use the same Paperless database and route
names, but have separate document volumes.
The Compose variant uses the shared Compose PostgreSQL service on the
`homelab-database` Docker network. It does not start a PostgreSQL container.
The shared Compose database must be running and must have the `paperless`
database and role. Set `PAPERLESS_DBPASS` to the same password as
`PAPERLESS_DB_PASSWORD` in the shared PostgreSQL configuration.
To prepare and start the Compose variant:
```sh
cp paperless/.env.example paperless/.env
cd paperless
docker compose -f compose.yaml config --quiet
docker compose -f compose.yaml up -d
```
Create unique values for `PAPERLESS_SECRET_KEY` and
`PAPERLESS_ADMIN_PASSWORD` in `.env`. This directory has no Compose `active`
marker, so the repository deploy workflow does not start this alternative.
Stop the Kubernetes Paperless deployment before switching to Compose. Back up
and migrate the media files as well as the database; the Compose named volumes
are separate from the Kubernetes PVC.
## Prepare the secret
Create `k8s/secrets.yaml` on the workstation from
`k8s/secrets.yaml.example`. Set a unique random `PAPERLESS_SECRET_KEY`, a long
`PAPERLESS_ADMIN_PASSWORD`, and `PAPERLESS_DB_PASSWORD`.
Add the same `PAPERLESS_DB_PASSWORD` value to the local
`postgres/k8s/secrets.yaml` file. Keep both secret files out of Git. The
database bootstrap Job creates the `paperless` role and database from the
shared PostgreSQL secret. The job runs in the `database` namespace and needs
that namespace's existing `postgres-shared-secrets` Secret.
For example, generate a key with:
```sh
python3 -c 'import secrets; print(secrets.token_urlsafe(64))'
```
Then apply the secret before enabling the service:
```sh
kubectl apply -f paperless/k8s/namespace.yaml
kubectl apply -f postgres/k8s/secrets.yaml
kubectl apply -f paperless/k8s/secrets.yaml
```
The normal deploy workflow applies the remaining manifests when
`paperless/k8s/active` is present. Verify the rollout and ingress after deploy:
```sh
kubectl -n paperless rollout status deployment/paperless
kubectl -n paperless get pods,pvc,services
```
Back up the `paperless-data` PVC and the shared PostgreSQL database. The PVC
contains the originals, archived PDFs, and export/consume folders. Valkey has
no persistent volume; queued tasks are recreated after a restart.
+77
View File
@@ -0,0 +1,77 @@
services:
paperless:
image: ghcr.io/paperless-ngx/paperless-ngx:3.2.1
container_name: paperless
restart: unless-stopped
env_file:
- .env
depends_on:
valkey:
condition: service_healthy
deploy:
resources:
limits:
cpus: "2.0"
memory: 2G
reservations:
cpus: "0.10"
memory: 512M
volumes:
- paperless-data:/usr/src/paperless/data
- paperless-media:/usr/src/paperless/media
- paperless-export:/usr/src/paperless/export
- paperless-consume:/usr/src/paperless/consume
networks:
- default
- proxy
- database
labels:
- "traefik.enable=true"
- "traefik.docker.network=proxy"
- "traefik.http.services.paperless-compose.loadbalancer.server.port=8000"
- "traefik.http.routers.paperless-compose.rule=Host(`papers.forust.xyz`)"
- "traefik.http.routers.paperless-compose.entrypoints=websecure"
- "traefik.http.routers.paperless-compose.tls.certresolver=letsencrypt"
- "traefik.http.routers.paperless-compose-local.rule=Host(`papers.workstation.internal`)"
- "traefik.http.routers.paperless-compose-local.entrypoints=websecure"
- "traefik.http.routers.paperless-compose-local.tls=true"
valkey:
image: valkey/valkey:9.0.3-alpine
container_name: paperless-valkey
restart: unless-stopped
command:
- valkey-server
- --save
- ""
- --appendonly
- "no"
healthcheck:
test: ["CMD", "valkey-cli", "ping"]
interval: 10s
timeout: 5s
retries: 5
deploy:
resources:
limits:
cpus: "0.25"
memory: 256M
reservations:
cpus: "0.025"
memory: 64M
networks:
- default
volumes:
paperless-data:
paperless-media:
paperless-export:
paperless-consume:
networks:
default:
proxy:
external: true
database:
name: homelab-database
external: true
+1
View File
@@ -0,0 +1 @@
+28
View File
@@ -0,0 +1,28 @@
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: paperless-prod-tls
namespace: paperless
spec:
secretName: paperless-prod-tls
dnsNames:
- papers.forust.xyz
issuerRef:
name: letsencrypt-prod
kind: ClusterIssuer
---
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: internal-wildcard-tls
namespace: paperless
spec:
secretName: internal-wildcard-tls
dnsNames:
- "*.workstation.internal"
- "*.gigaforust.internal"
- workstation.internal
- gigaforust.internal
issuerRef:
name: internal-ca
kind: ClusterIssuer
+58
View File
@@ -0,0 +1,58 @@
apiVersion: batch/v1
kind: Job
metadata:
name: paperless-database-init
namespace: database
spec:
backoffLimit: 5
template:
metadata:
labels:
app.kubernetes.io/name: paperless-database-init
spec:
restartPolicy: OnFailure
containers:
- name: create-database
image: postgres:17.11-alpine
command:
- /bin/sh
- -ec
- |
PGPASSWORD="$POSTGRES_ADMIN_PASSWORD" psql \
--host postgres \
--username postgres \
--dbname postgres \
--set ON_ERROR_STOP=1 \
--set paperless_password="$PAPERLESS_DB_PASSWORD" <<'SQL'
SELECT format(
'CREATE ROLE paperless LOGIN PASSWORD %L',
:'paperless_password'
)
WHERE NOT EXISTS (
SELECT FROM pg_roles WHERE rolname = 'paperless'
)
\gexec
SELECT format('CREATE DATABASE paperless OWNER paperless')
WHERE NOT EXISTS (
SELECT FROM pg_database WHERE datname = 'paperless'
)
\gexec
SQL
env:
- name: POSTGRES_ADMIN_PASSWORD
valueFrom:
secretKeyRef:
name: postgres-shared-secrets
key: POSTGRES_ADMIN_PASSWORD
- name: PAPERLESS_DB_PASSWORD
valueFrom:
secretKeyRef:
name: postgres-shared-secrets
key: PAPERLESS_DB_PASSWORD
resources:
requests:
cpu: 10m
memory: 32Mi
limits:
cpu: 100m
memory: 128Mi
+33
View File
@@ -0,0 +1,33 @@
apiVersion: traefik.io/v1alpha1
kind: IngressRoute
metadata:
name: paperless-prod
namespace: paperless
spec:
entryPoints:
- websecure
routes:
- match: Host(`papers.forust.xyz`)
kind: Rule
services:
- name: paperless
port: 8000
tls:
secretName: paperless-prod-tls
---
apiVersion: traefik.io/v1alpha1
kind: IngressRoute
metadata:
name: paperless-local
namespace: paperless
spec:
entryPoints:
- websecure
routes:
- match: Host(`papers.workstation.internal`)
kind: Rule
services:
- name: paperless
port: 8000
tls:
secretName: internal-wildcard-tls
@@ -1,4 +1,4 @@
apiVersion: v1 apiVersion: v1
kind: Namespace kind: Namespace
metadata: metadata:
name: edu-master name: paperless
+39
View File
@@ -0,0 +1,39 @@
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: paperless-ingress
namespace: paperless
spec:
podSelector:
matchLabels:
app.kubernetes.io/name: paperless
policyTypes:
- Ingress
ingress:
- from:
- namespaceSelector:
matchLabels:
kubernetes.io/metadata.name: traefik
ports:
- protocol: TCP
port: 8000
---
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: paperless-valkey-ingress
namespace: paperless
spec:
podSelector:
matchLabels:
app.kubernetes.io/name: paperless-valkey
policyTypes:
- Ingress
ingress:
- from:
- podSelector:
matchLabels:
app.kubernetes.io/name: paperless
ports:
- protocol: TCP
port: 6379
+210
View File
@@ -0,0 +1,210 @@
apiVersion: v1
kind: Service
metadata:
name: paperless
namespace: paperless
labels:
app.kubernetes.io/name: paperless
spec:
selector:
app.kubernetes.io/name: paperless
ports:
- name: http
port: 8000
targetPort: http
---
apiVersion: apps/v1
kind: Deployment
metadata:
name: paperless
namespace: paperless
labels:
app.kubernetes.io/name: paperless
spec:
replicas: 1
strategy:
type: Recreate
selector:
matchLabels:
app.kubernetes.io/name: paperless
template:
metadata:
labels:
app.kubernetes.io/name: paperless
spec:
enableServiceLinks: false
containers:
- name: paperless
image: ghcr.io/paperless-ngx/paperless-ngx:3.2.1
ports:
- name: http
containerPort: 8000
env:
- name: PAPERLESS_URL
value: https://papers.forust.xyz
- name: PAPERLESS_ALLOWED_HOSTS
value: papers.forust.xyz,papers.workstation.internal
- name: PAPERLESS_CSRF_TRUSTED_ORIGINS
value: https://papers.forust.xyz,https://papers.workstation.internal
- name: PAPERLESS_TIME_ZONE
value: Europe/Bratislava
- name: PAPERLESS_REDIS
value: redis://paperless-valkey:6379
- name: PAPERLESS_DBENGINE
value: postgresql
- name: PAPERLESS_DBHOST
value: postgres.database.svc.cluster.local
- name: PAPERLESS_DBNAME
value: paperless
- name: PAPERLESS_DBUSER
value: paperless
- name: PAPERLESS_DBPASS
valueFrom:
secretKeyRef:
name: paperless-secrets
key: PAPERLESS_DB_PASSWORD
- name: PAPERLESS_OCR_LANGUAGE
value: rus+eng
- name: PAPERLESS_OCR_LANGUAGES
value: rus
- name: PAPERLESS_TASK_WORKERS
value: "1"
- name: PAPERLESS_ADMIN_USER
value: admin
- name: PAPERLESS_SECRET_KEY
valueFrom:
secretKeyRef:
name: paperless-secrets
key: PAPERLESS_SECRET_KEY
- name: PAPERLESS_ADMIN_PASSWORD
valueFrom:
secretKeyRef:
name: paperless-secrets
key: PAPERLESS_ADMIN_PASSWORD
volumeMounts:
- name: documents
mountPath: /usr/src/paperless/data
subPath: data
- name: documents
mountPath: /usr/src/paperless/media
subPath: media
- name: documents
mountPath: /usr/src/paperless/export
subPath: export
- name: documents
mountPath: /usr/src/paperless/consume
subPath: consume
startupProbe:
httpGet:
path: /
port: http
httpHeaders:
- name: Host
value: papers.workstation.internal
failureThreshold: 60
periodSeconds: 10
timeoutSeconds: 5
readinessProbe:
httpGet:
path: /
port: http
httpHeaders:
- name: Host
value: papers.workstation.internal
periodSeconds: 10
timeoutSeconds: 5
livenessProbe:
httpGet:
path: /
port: http
httpHeaders:
- name: Host
value: papers.workstation.internal
periodSeconds: 30
timeoutSeconds: 5
resources:
requests:
cpu: 100m
memory: 512Mi
limits:
cpu: "2"
memory: 2Gi
volumes:
- name: documents
persistentVolumeClaim:
claimName: paperless-data
---
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: paperless-data
namespace: paperless
spec:
accessModes:
- ReadWriteOnce
storageClassName: local-path-retain
resources:
requests:
storage: 50Gi
---
apiVersion: v1
kind: Service
metadata:
name: paperless-valkey
namespace: paperless
labels:
app.kubernetes.io/name: paperless-valkey
spec:
selector:
app.kubernetes.io/name: paperless-valkey
ports:
- name: redis
port: 6379
targetPort: redis
---
apiVersion: apps/v1
kind: Deployment
metadata:
name: paperless-valkey
namespace: paperless
labels:
app.kubernetes.io/name: paperless-valkey
spec:
replicas: 1
strategy:
type: Recreate
selector:
matchLabels:
app.kubernetes.io/name: paperless-valkey
template:
metadata:
labels:
app.kubernetes.io/name: paperless-valkey
spec:
containers:
- name: valkey
image: valkey/valkey:9.0.3-alpine
args:
- valkey-server
- --save
- ""
- --appendonly
- "no"
ports:
- name: redis
containerPort: 6379
readinessProbe:
exec:
command: ["valkey-cli", "ping"]
periodSeconds: 10
livenessProbe:
exec:
command: ["valkey-cli", "ping"]
periodSeconds: 30
resources:
requests:
cpu: 25m
memory: 64Mi
limits:
cpu: 250m
memory: 256Mi
+10
View File
@@ -0,0 +1,10 @@
apiVersion: v1
kind: Secret
metadata:
name: paperless-secrets
namespace: paperless
type: Opaque
stringData:
PAPERLESS_SECRET_KEY: "<GENERATE_WITH_python3_-c_import_secrets;_print(secrets.token_urlsafe(64))>"
PAPERLESS_ADMIN_PASSWORD: "<SET_A_LONG_UNIQUE_PASSWORD>"
PAPERLESS_DB_PASSWORD: "<SET_THE_SAME_VALUE_AS_database_PAPERLESS_DB_PASSWORD>"
+2
View File
@@ -1,6 +1,8 @@
POSTGRES_ADMIN_PASSWORD= POSTGRES_ADMIN_PASSWORD=
AUTHENTIK_DB_PASSWORD= AUTHENTIK_DB_PASSWORD=
GITEA_DB_PASSWORD= GITEA_DB_PASSWORD=
NETBOX_DB_PASSWORD=
NETRONOME_DB_PASSWORD= NETRONOME_DB_PASSWORD=
PENPOT_DB_PASSWORD= PENPOT_DB_PASSWORD=
PAPERLESS_DB_PASSWORD=
STATUSPAGE_DB_PASSWORD= STATUSPAGE_DB_PASSWORD=
+12 -11
View File
@@ -1,20 +1,21 @@
# Shared PostgreSQL # Shared PostgreSQL
This directory contains the shared PostgreSQL 17 deployment for Authentik, This directory contains the shared PostgreSQL 17 deployment for Authentik,
Gitea, NetBox, Netronome, and Statuspage. It creates one database and one login role Gitea, NetBox, Netronome, Paperless-ngx, and Statuspage. It creates one database
per service. Per-service standalone databases were removed after the and one login role per service. Per-service standalone databases were removed
migration (Sep 2026); Penpot stays on its own compose PostgreSQL (archived, after the migration (Sep 2026); Penpot stays on its own compose PostgreSQL
not part of the shared instance). (archived, not part of the shared instance).
## Compatibility baseline ## Compatibility baseline
| Service | Current application | Shared PostgreSQL 17 | | Service | Current application | Shared PostgreSQL 17 |
| ---------- | ------------------- | -------------------------------------- | | ------------- | ------------------- | -------------------------------------- |
| Authentik | 2025.10.x | Supported (Authentik requires 14+) | | Authentik | 2025.10.x | Supported (Authentik requires 14+) |
| Gitea | 1.27.3 | Supported (Gitea requires 12+) | | Gitea | 1.27.3 | Supported (Gitea requires 12+) |
| NetBox | 4.7.x | Supported (NetBox 4.x requires 13+) | | NetBox | 4.7.x | Supported (NetBox 4.x requires 13+) |
| Netronome | 0.14.0 | Supported (upstream's example uses 17) | | Netronome | 0.14.0 | Supported (upstream's example uses 17) |
| Statuspage | custom | Supported | | Paperless-ngx | 3.2.1 | Supported |
| Statuspage | custom | Supported |
A major-version change must use a logical dump/restore; changing only the A major-version change must use a logical dump/restore; changing only the
image tag while keeping a data directory is not supported. image tag while keeping a data directory is not supported.
+2
View File
@@ -6,6 +6,7 @@ set -euo pipefail
: "${NETBOX_DB_PASSWORD:?NETBOX_DB_PASSWORD is required}" : "${NETBOX_DB_PASSWORD:?NETBOX_DB_PASSWORD is required}"
: "${NETRONOME_DB_PASSWORD:?NETRONOME_DB_PASSWORD is required}" : "${NETRONOME_DB_PASSWORD:?NETRONOME_DB_PASSWORD is required}"
: "${PENPOT_DB_PASSWORD:?PENPOT_DB_PASSWORD is required}" : "${PENPOT_DB_PASSWORD:?PENPOT_DB_PASSWORD is required}"
: "${PAPERLESS_DB_PASSWORD:?PAPERLESS_DB_PASSWORD is required}"
: "${STATUSPAGE_DB_PASSWORD:?STATUSPAGE_DB_PASSWORD is required}" : "${STATUSPAGE_DB_PASSWORD:?STATUSPAGE_DB_PASSWORD is required}"
create_role_and_database() { create_role_and_database() {
@@ -27,4 +28,5 @@ create_role_and_database gitea gitea "$GITEA_DB_PASSWORD"
create_role_and_database netbox netbox "$NETBOX_DB_PASSWORD" create_role_and_database netbox netbox "$NETBOX_DB_PASSWORD"
create_role_and_database netronome netronome "$NETRONOME_DB_PASSWORD" create_role_and_database netronome netronome "$NETRONOME_DB_PASSWORD"
create_role_and_database penpot penpot "$PENPOT_DB_PASSWORD" create_role_and_database penpot penpot "$PENPOT_DB_PASSWORD"
create_role_and_database paperless paperless "$PAPERLESS_DB_PASSWORD"
create_role_and_database statuspage statuspage "$STATUSPAGE_DB_PASSWORD" create_role_and_database statuspage statuspage "$STATUSPAGE_DB_PASSWORD"
+12
View File
@@ -26,6 +26,18 @@ spec:
- namespaceSelector: - namespaceSelector:
matchLabels: matchLabels:
kubernetes.io/metadata.name: penpot kubernetes.io/metadata.name: penpot
- namespaceSelector:
matchLabels:
kubernetes.io/metadata.name: paperless
podSelector:
matchLabels:
app.kubernetes.io/name: paperless
- namespaceSelector:
matchLabels:
kubernetes.io/metadata.name: database
podSelector:
matchLabels:
app.kubernetes.io/name: paperless-database-init
- namespaceSelector: - namespaceSelector:
matchLabels: matchLabels:
kubernetes.io/metadata.name: statuspage kubernetes.io/metadata.name: statuspage
+2
View File
@@ -126,6 +126,7 @@ data:
: "${NETBOX_DB_PASSWORD:?NETBOX_DB_PASSWORD is required}" : "${NETBOX_DB_PASSWORD:?NETBOX_DB_PASSWORD is required}"
: "${NETRONOME_DB_PASSWORD:?NETRONOME_DB_PASSWORD is required}" : "${NETRONOME_DB_PASSWORD:?NETRONOME_DB_PASSWORD is required}"
: "${PENPOT_DB_PASSWORD:?PENPOT_DB_PASSWORD is required}" : "${PENPOT_DB_PASSWORD:?PENPOT_DB_PASSWORD is required}"
: "${PAPERLESS_DB_PASSWORD:?PAPERLESS_DB_PASSWORD is required}"
: "${STATUSPAGE_DB_PASSWORD:?STATUSPAGE_DB_PASSWORD is required}" : "${STATUSPAGE_DB_PASSWORD:?STATUSPAGE_DB_PASSWORD is required}"
create_role_and_database() { create_role_and_database() {
@@ -147,4 +148,5 @@ data:
create_role_and_database netbox netbox "$NETBOX_DB_PASSWORD" create_role_and_database netbox netbox "$NETBOX_DB_PASSWORD"
create_role_and_database netronome netronome "$NETRONOME_DB_PASSWORD" create_role_and_database netronome netronome "$NETRONOME_DB_PASSWORD"
create_role_and_database penpot penpot "$PENPOT_DB_PASSWORD" create_role_and_database penpot penpot "$PENPOT_DB_PASSWORD"
create_role_and_database paperless paperless "$PAPERLESS_DB_PASSWORD"
create_role_and_database statuspage statuspage "$STATUSPAGE_DB_PASSWORD" create_role_and_database statuspage statuspage "$STATUSPAGE_DB_PASSWORD"
+1
View File
@@ -11,4 +11,5 @@ stringData:
NETBOX_DB_PASSWORD: "" NETBOX_DB_PASSWORD: ""
NETRONOME_DB_PASSWORD: "" NETRONOME_DB_PASSWORD: ""
PENPOT_DB_PASSWORD: "" PENPOT_DB_PASSWORD: ""
PAPERLESS_DB_PASSWORD: ""
STATUSPAGE_DB_PASSWORD: "" STATUSPAGE_DB_PASSWORD: ""
+43
View File
@@ -0,0 +1,43 @@
# VictoriaMetrics
The `victoria-operator` Helm release converts Prometheus Operator
`ServiceMonitor` resources into owned `VMServiceScrape` resources. The
`VMAgent` selects converted scrapes labeled `release: prometheus-stack` in all
namespaces and writes them to the existing single-node VictoriaMetrics
instance. Changes to selected `ServiceMonitor` resources are reconciled
automatically; there is no copied Prometheus scrape-config blob to regenerate.
The agent drops targets for the Prometheus server service to avoid duplicating
its self-scrape. `scraper: victoria` identifies the samples ingested by this
VMAgent.
The VictoriaMetrics Operator chart and its CRDs are installed before the
Kubernetes manifests by the normal deploy workflow. On a cluster where the
operator CRDs are not installed yet, CI skips the server-side dry-run of the
`VMAgent` resource; the deploy installs the chart before applying that resource.
## Application metrics
The application ServiceMonitors use a 30s interval and a 10s timeout:
- Headscale: the external Service points to the Compose host on port 19090.
A VMServiceScrape uses EndpointSlice discovery for this manually managed target.
The Compose configuration must bind metrics to `0.0.0.0:9090`.
- NetBird: the combined server exports `/metrics` on port 9090. The existing
`server.metricsPort` setting enables the listener.
- Gitea: `GITEA__metrics__ENABLED` enables `/metrics` on the HTTP port. The public
ingress excludes this path. The monitor uses the internal Service directly.
- Immich: `IMMICH_TELEMETRY_INCLUDE=all` enables API and worker metrics on ports
8081 and 8082. The monitor scrapes both ports on each server replica.
Deploy through the existing CI and deploy workflow. Gitea and Immich reload their
ConfigMap changes through Reloader. Check the VMAgent targets after deployment
and query `up{scraper="victoria",namespace=~"netbird|gitea|immich|headscale"}` in
VictoriaMetrics. All targets should report 1.
For rollback, revert the application metrics changes, run CI, and deploy the
revert. Remove the three application ServiceMonitors and the Headscale VMServiceScrape explicitly: the deployment
workflow applies manifests and does not prune removed resources.
For Headscale rollback, remove its VMServiceScrape and Service label, restore the
previous Compose metrics bind address, and restart only the Headscale service.
+11
View File
@@ -38,6 +38,8 @@ grafana:
# One block covers both the dashboards and datasources sidecars (p95 91M / 80M). # One block covers both the dashboards and datasources sidecars (p95 91M / 80M).
sidecar: sidecar:
datasources:
defaultDatasourceEnabled: false
resources: resources:
requests: requests:
memory: "96Mi" memory: "96Mi"
@@ -50,9 +52,18 @@ grafana:
type: loki type: loki
url: http://loki-gateway.prometheus.svc.cluster.local url: http://loki-gateway.prometheus.svc.cluster.local
access: proxy access: proxy
- name: VictoriaMetrics
type: prometheus
url: http://victoria-metrics.prometheus.svc.cluster.local:8428
access: proxy
isDefault: true
prometheus: prometheus:
prometheusSpec: prometheusSpec:
# VM trial: vmagent scrapes and remote-writes to VictoriaMetrics, so the
# Prometheus server itself stands down. Encoded here (not a kubectl patch)
# so helm keeps owning spec.replicas and upgrades do not conflict on it.
replicas: 0
retention: 60d retention: 60d
retentionSize: 32GB retentionSize: 32GB
storageSpec: storageSpec:
+102
View File
@@ -31,3 +31,105 @@ spec:
port: 80 port: 80
tls: tls:
secretName: internal-wildcard-tls secretName: internal-wildcard-tls
---
apiVersion: traefik.io/v1alpha1
kind: IngressRoute
metadata:
name: prometheus-local
namespace: prometheus
spec:
entryPoints:
- websecure
routes:
- match: Host(`prom.workstation.internal`) || Host(`prom.gigaforust.internal`)
kind: Rule
services:
- name: prometheus-stack-kube-prom-prometheus
port: 9090
tls:
secretName: internal-wildcard-tls
---
apiVersion: traefik.io/v1alpha1
kind: IngressRoute
metadata:
name: alertmanager-local
namespace: prometheus
spec:
entryPoints:
- websecure
routes:
- match: Host(`am.workstation.internal`) || Host(`am.gigaforust.internal`)
kind: Rule
services:
- name: prometheus-stack-kube-prom-alertmanager
port: 9093
tls:
secretName: internal-wildcard-tls
---
apiVersion: traefik.io/v1alpha1
kind: IngressRoute
metadata:
name: loki-local
namespace: prometheus
spec:
entryPoints:
- websecure
routes:
- match: Host(`loki.workstation.internal`) || Host(`loki.gigaforust.internal`)
kind: Rule
services:
- name: loki-gateway
port: 80
tls:
secretName: internal-wildcard-tls
---
apiVersion: traefik.io/v1alpha1
kind: IngressRoute
metadata:
name: alloy-local
namespace: prometheus
spec:
entryPoints:
- websecure
routes:
- match: Host(`alloy.workstation.internal`) || Host(`alloy.gigaforust.internal`)
kind: Rule
services:
- name: alloy
port: 12345
tls:
secretName: internal-wildcard-tls
---
apiVersion: traefik.io/v1alpha1
kind: IngressRoute
metadata:
name: victoria-local
namespace: prometheus
spec:
entryPoints:
- websecure
routes:
- match: Host(`victoria.workstation.internal`) || Host(`victoria.gigaforust.internal`)
kind: Rule
services:
- name: victoria-metrics
port: 8428
tls:
secretName: internal-wildcard-tls
---
apiVersion: traefik.io/v1alpha1
kind: IngressRoute
metadata:
name: vmalert-local
namespace: prometheus
spec:
entryPoints:
- websecure
routes:
- match: Host(`vmalert.workstation.internal`) || Host(`vmalert.gigaforust.internal`)
kind: Rule
services:
- name: vmalert
port: 8880
tls:
secretName: internal-wildcard-tls
@@ -0,0 +1,12 @@
nameOverride: victoria-operator
operator:
enable_converter_ownership: true
resources:
requests:
cpu: 50m
memory: 96Mi
limits:
cpu: 200m
memory: 256Mi
+79
View File
@@ -0,0 +1,79 @@
apiVersion: v1
kind: Service
metadata:
name: victoria-metrics
namespace: prometheus
spec:
selector:
app: victoria-metrics
ports:
- port: 8428
targetPort: 8428
---
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: victoria-pvc
namespace: prometheus
spec:
resources:
requests:
storage: 10Gi
volumeMode: Filesystem
accessModes:
- ReadWriteOnce
---
apiVersion: apps/v1
kind: Deployment
metadata:
name: victoria-deployment
namespace: prometheus
spec:
replicas: 1
selector:
matchLabels:
app: victoria-metrics
strategy:
type: Recreate
template:
metadata:
labels:
app: victoria-metrics
spec:
containers:
- name: victoria
image: victoriametrics/victoria-metrics:v1.153.0-scratch
args:
- -storageDataPath=/vmdata
- -retentionPeriod=30d
- -httpListenAddr=:8428
ports:
- containerPort: 8428
readinessProbe:
httpGet:
path: /health
port: 8428
initialDelaySeconds: 15
periodSeconds: 10
failureThreshold: 6
livenessProbe:
httpGet:
path: /health
port: 8428
initialDelaySeconds: 60
periodSeconds: 30
failureThreshold: 3
volumeMounts:
- name: vmdata
mountPath: /vmdata
resources:
requests:
cpu: "100m"
memory: "256Mi"
limits:
cpu: "1000m"
memory: "1Gi"
volumes:
- name: vmdata
persistentVolumeClaim:
claimName: victoria-pvc
+29
View File
@@ -0,0 +1,29 @@
apiVersion: operator.victoriametrics.com/v1beta1
kind: VMAgent
metadata:
name: vmagent
namespace: prometheus
spec:
image:
tag: v1.153.0
scrapeInterval: 30s
externalLabels:
scraper: victoria
serviceScrapeNamespaceSelector: {}
serviceScrapeSelector:
matchLabels:
release: prometheus-stack
globalScrapeRelabelConfigs:
- action: drop
source_labels:
- __meta_kubernetes_service_name
regex: prometheus-stack-kube-prom-prometheus
remoteWrite:
- url: http://victoria-metrics.prometheus.svc.cluster.local:8428/api/v1/write
resources:
requests:
cpu: 100m
memory: 256Mi
limits:
cpu: "1000m"
memory: 1Gi
+70
View File
@@ -0,0 +1,70 @@
apiVersion: v1
kind: Service
metadata:
name: vmalert
namespace: prometheus
spec:
selector:
app: vmalert
ports:
- port: 8880
targetPort: 8880
---
apiVersion: apps/v1
kind: Deployment
metadata:
name: vmalert-deployment
namespace: prometheus
spec:
replicas: 1
selector:
matchLabels:
app: vmalert
strategy:
type: Recreate
template:
metadata:
labels:
app: vmalert
spec:
containers:
- name: vmalert
image: victoriametrics/vmalert:v1.153.0
args:
- -datasource.url=http://victoria-metrics.prometheus.svc.cluster.local:8428
- -remoteWrite.url=http://victoria-metrics.prometheus.svc.cluster.local:8428
- -notifier.url=http://prometheus-stack-kube-prom-alertmanager.prometheus.svc.cluster.local:9093
- -rule=/etc/vm/rules/*.yaml
- -evaluationInterval=60s
- -httpListenAddr=:8880
ports:
- containerPort: 8880
readinessProbe:
httpGet:
path: /metrics
port: 8880
initialDelaySeconds: 15
periodSeconds: 10
failureThreshold: 6
livenessProbe:
httpGet:
path: /metrics
port: 8880
initialDelaySeconds: 60
periodSeconds: 30
failureThreshold: 3
volumeMounts:
- name: rules
mountPath: /etc/vm/rules
readOnly: true
resources:
requests:
cpu: "50m"
memory: "64Mi"
limits:
cpu: "200m"
memory: "256Mi"
volumes:
- name: rules
configMap:
name: prometheus-prometheus-stack-kube-prom-prometheus-rulefiles-0
+43 -49
View File
@@ -19,6 +19,9 @@ data:
"dependencyDashboard": true, "dependencyDashboard": true,
"prCreation": "immediate", "prCreation": "immediate",
"labels": ["dependencies", "automated"], "labels": ["dependencies", "automated"],
"docker-compose": {
"managerFilePatterns": ["renovate/renovate-compose.yaml"]
},
"helm-values": { "helm-values": {
"managerFilePatterns": ["/k8s/.+values\\.ya?ml$/"] "managerFilePatterns": ["/k8s/.+values\\.ya?ml$/"]
}, },
@@ -26,35 +29,28 @@ data:
"managerFilePatterns": ["/k8s/.+\\.ya?ml$/"] "managerFilePatterns": ["/k8s/.+\\.ya?ml$/"]
}, },
"customManagers": [ "customManagers": [
{
"customType": "regex",
"description": "singlesource: playwright npm version pinned in npx command (k8s + compose)",
"managerFilePatterns": ["^edu_master/k8s/playwright\\.yaml$", "^edu_master/compose\\.yaml$"],
"matchStrings": ["playwright@(?<currentValue>\\d+\\.\\d+\\.\\d+)"],
"datasourceTemplate": "npm",
"depNameTemplate": "playwright"
},
{
"customType": "regex",
"description": "singlesource: PLAYWRIGHT_VERSION file",
"managerFilePatterns": ["^edu_master/PLAYWRIGHT_VERSION$"],
"matchStrings": ["^(?<currentValue>\\d+\\.\\d+\\.\\d+)$"],
"datasourceTemplate": "pypi",
"depNameTemplate": "playwright"
},
{ {
"customType": "regex", "customType": "regex",
"description": "kube-prometheus-stack chart version pinned in the deploy workflow", "description": "kube-prometheus-stack chart version pinned in the deploy workflow",
"managerFilePatterns": ["^\\.gitea/workflows/deploy-lib\\.sh$"], "managerFilePatterns": [".gitea/workflows/deploy-lib.sh"],
"matchStrings": ["\\|prometheus-community/kube-prometheus-stack\\|prometheus\\|(?<currentValue>[0-9.]+)\\|"], "matchStrings": ["\\|prometheus-community/kube-prometheus-stack\\|prometheus\\|(?<currentValue>[0-9.]+)\\|"],
"datasourceTemplate": "helm", "datasourceTemplate": "helm",
"depNameTemplate": "kube-prometheus-stack", "depNameTemplate": "kube-prometheus-stack",
"registryUrlTemplate": "https://prometheus-community.github.io/helm-charts" "registryUrlTemplate": "https://prometheus-community.github.io/helm-charts"
}, },
{
"customType": "regex",
"description": "VictoriaMetrics Operator chart version pinned in the deploy workflow",
"managerFilePatterns": [".gitea/workflows/deploy-lib.sh"],
"matchStrings": ["\\|victoriametrics/victoria-metrics-operator\\|prometheus\\|(?<currentValue>[0-9.]+)\\|"],
"datasourceTemplate": "helm",
"depNameTemplate": "victoria-metrics-operator",
"registryUrlTemplate": "https://victoriametrics.github.io/helm-charts"
},
{ {
"customType": "regex", "customType": "regex",
"description": "grafana/loki chart version pinned in the deploy workflow", "description": "grafana/loki chart version pinned in the deploy workflow",
"managerFilePatterns": ["^\\.gitea/workflows/deploy-lib\\.sh$"], "managerFilePatterns": [".gitea/workflows/deploy-lib.sh"],
"matchStrings": ["\\|grafana/loki\\|prometheus\\|(?<currentValue>[0-9.]+)\\|"], "matchStrings": ["\\|grafana/loki\\|prometheus\\|(?<currentValue>[0-9.]+)\\|"],
"datasourceTemplate": "helm", "datasourceTemplate": "helm",
"depNameTemplate": "loki", "depNameTemplate": "loki",
@@ -63,7 +59,7 @@ data:
{ {
"customType": "regex", "customType": "regex",
"description": "grafana/alloy chart version pinned in the deploy workflow", "description": "grafana/alloy chart version pinned in the deploy workflow",
"managerFilePatterns": ["^\\.gitea/workflows/deploy-lib\\.sh$"], "managerFilePatterns": [".gitea/workflows/deploy-lib.sh"],
"matchStrings": ["\\|grafana/alloy\\|prometheus\\|(?<currentValue>[0-9.]+)\\|"], "matchStrings": ["\\|grafana/alloy\\|prometheus\\|(?<currentValue>[0-9.]+)\\|"],
"datasourceTemplate": "helm", "datasourceTemplate": "helm",
"depNameTemplate": "alloy", "depNameTemplate": "alloy",
@@ -72,7 +68,7 @@ data:
{ {
"customType": "regex", "customType": "regex",
"description": "actionlint version used by the ci workflow", "description": "actionlint version used by the ci workflow",
"managerFilePatterns": ["^\\.gitea/workflows/tool-versions\\.env$"], "managerFilePatterns": [".gitea/workflows/tool-versions.env"],
"matchStrings": ["(?:^|\\n)ACTIONLINT_VERSION=\"(?<currentValue>[0-9.]+)\""], "matchStrings": ["(?:^|\\n)ACTIONLINT_VERSION=\"(?<currentValue>[0-9.]+)\""],
"datasourceTemplate": "github-tags", "datasourceTemplate": "github-tags",
"depNameTemplate": "rhysd/actionlint" "depNameTemplate": "rhysd/actionlint"
@@ -80,7 +76,7 @@ data:
{ {
"customType": "regex", "customType": "regex",
"description": "shellcheck version used by the ci workflow", "description": "shellcheck version used by the ci workflow",
"managerFilePatterns": ["^\\.gitea/workflows/tool-versions\\.env$"], "managerFilePatterns": [".gitea/workflows/tool-versions.env"],
"matchStrings": ["(?:^|\\n)SHELLCHECK_VERSION=\"(?<currentValue>[0-9.]+)\""], "matchStrings": ["(?:^|\\n)SHELLCHECK_VERSION=\"(?<currentValue>[0-9.]+)\""],
"datasourceTemplate": "github-tags", "datasourceTemplate": "github-tags",
"depNameTemplate": "koalaman/shellcheck" "depNameTemplate": "koalaman/shellcheck"
@@ -88,7 +84,7 @@ data:
{ {
"customType": "regex", "customType": "regex",
"description": "kubeconform version used by the ci workflow", "description": "kubeconform version used by the ci workflow",
"managerFilePatterns": ["^\\.gitea/workflows/tool-versions\\.env$"], "managerFilePatterns": [".gitea/workflows/tool-versions.env"],
"matchStrings": ["(?:^|\\n)KUBECONFORM_VERSION=\"(?<currentValue>[0-9.]+)\""], "matchStrings": ["(?:^|\\n)KUBECONFORM_VERSION=\"(?<currentValue>[0-9.]+)\""],
"datasourceTemplate": "github-tags", "datasourceTemplate": "github-tags",
"depNameTemplate": "yannh/kubeconform" "depNameTemplate": "yannh/kubeconform"
@@ -96,7 +92,7 @@ data:
{ {
"customType": "regex", "customType": "regex",
"description": "uv version used to build the pytest venv", "description": "uv version used to build the pytest venv",
"managerFilePatterns": ["^\\.gitea/workflows/tool-versions\\.env$"], "managerFilePatterns": [".gitea/workflows/tool-versions.env"],
"matchStrings": ["(?:^|\\n)UV_VERSION=\"(?<currentValue>[0-9.]+)\""], "matchStrings": ["(?:^|\\n)UV_VERSION=\"(?<currentValue>[0-9.]+)\""],
"datasourceTemplate": "github-tags", "datasourceTemplate": "github-tags",
"depNameTemplate": "astral-sh/uv" "depNameTemplate": "astral-sh/uv"
@@ -104,7 +100,7 @@ data:
{ {
"customType": "regex", "customType": "regex",
"description": "prettier version used by the ci workflow", "description": "prettier version used by the ci workflow",
"managerFilePatterns": ["^\\.gitea/workflows/tool-versions\\.env$"], "managerFilePatterns": [".gitea/workflows/tool-versions.env"],
"matchStrings": ["(?:^|\\n)PRETTIER_VERSION=\"(?<currentValue>[0-9.]+)\""], "matchStrings": ["(?:^|\\n)PRETTIER_VERSION=\"(?<currentValue>[0-9.]+)\""],
"datasourceTemplate": "npm", "datasourceTemplate": "npm",
"depNameTemplate": "prettier" "depNameTemplate": "prettier"
@@ -112,7 +108,7 @@ data:
{ {
"customType": "regex", "customType": "regex",
"description": "ruff version used by the ci workflow", "description": "ruff version used by the ci workflow",
"managerFilePatterns": ["^\\.gitea/workflows/tool-versions\\.env$"], "managerFilePatterns": [".gitea/workflows/tool-versions.env"],
"matchStrings": ["(?:^|\\n)RUFF_VERSION=\"(?<currentValue>[0-9.]+)\""], "matchStrings": ["(?:^|\\n)RUFF_VERSION=\"(?<currentValue>[0-9.]+)\""],
"datasourceTemplate": "pypi", "datasourceTemplate": "pypi",
"depNameTemplate": "ruff" "depNameTemplate": "ruff"
@@ -120,7 +116,7 @@ data:
{ {
"customType": "regex", "customType": "regex",
"description": "pip-audit version used by the ci workflow", "description": "pip-audit version used by the ci workflow",
"managerFilePatterns": ["^\\.gitea/workflows/tool-versions\\.env$"], "managerFilePatterns": [".gitea/workflows/tool-versions.env"],
"matchStrings": ["(?:^|\\n)PIP_AUDIT_VERSION=\"(?<currentValue>[0-9.]+)\""], "matchStrings": ["(?:^|\\n)PIP_AUDIT_VERSION=\"(?<currentValue>[0-9.]+)\""],
"datasourceTemplate": "pypi", "datasourceTemplate": "pypi",
"depNameTemplate": "pip-audit" "depNameTemplate": "pip-audit"
@@ -128,7 +124,7 @@ data:
{ {
"customType": "regex", "customType": "regex",
"description": "yamllint version used by the ci workflow", "description": "yamllint version used by the ci workflow",
"managerFilePatterns": ["^\\.gitea/workflows/tool-versions\\.env$"], "managerFilePatterns": [".gitea/workflows/tool-versions.env"],
"matchStrings": ["(?:^|\\n)YAMLLINT_VERSION=\"(?<currentValue>[0-9.]+)\""], "matchStrings": ["(?:^|\\n)YAMLLINT_VERSION=\"(?<currentValue>[0-9.]+)\""],
"datasourceTemplate": "pypi", "datasourceTemplate": "pypi",
"depNameTemplate": "yamllint" "depNameTemplate": "yamllint"
@@ -136,7 +132,7 @@ data:
{ {
"customType": "regex", "customType": "regex",
"description": "hadolint version used by the ci workflow", "description": "hadolint version used by the ci workflow",
"managerFilePatterns": ["^\\.gitea/workflows/tool-versions\\.env$"], "managerFilePatterns": [".gitea/workflows/tool-versions.env"],
"matchStrings": ["(?:^|\\n)HADOLINT_VERSION=\"(?<currentValue>[0-9.]+)\""], "matchStrings": ["(?:^|\\n)HADOLINT_VERSION=\"(?<currentValue>[0-9.]+)\""],
"datasourceTemplate": "github-tags", "datasourceTemplate": "github-tags",
"depNameTemplate": "hadolint/hadolint" "depNameTemplate": "hadolint/hadolint"
@@ -144,7 +140,7 @@ data:
{ {
"customType": "regex", "customType": "regex",
"description": "node version the ci workflow runs npm with", "description": "node version the ci workflow runs npm with",
"managerFilePatterns": ["^\\.gitea/workflows/tool-versions\\.env$"], "managerFilePatterns": [".gitea/workflows/tool-versions.env"],
"matchStrings": ["(?:^|\\n)NODE_VERSION=\"(?<currentValue>[0-9.]+)\""], "matchStrings": ["(?:^|\\n)NODE_VERSION=\"(?<currentValue>[0-9.]+)\""],
"datasourceTemplate": "node", "datasourceTemplate": "node",
"depNameTemplate": "node" "depNameTemplate": "node"
@@ -152,16 +148,23 @@ data:
{ {
"customType": "regex", "customType": "regex",
"description": "stakater/reloader chart version pinned in the deploy workflow", "description": "stakater/reloader chart version pinned in the deploy workflow",
"managerFilePatterns": ["^\\.gitea/workflows/deploy-lib\\.sh$"], "managerFilePatterns": [".gitea/workflows/deploy-lib.sh"],
"matchStrings": ["\\|stakater/reloader\\|reloader\\|(?<currentValue>[0-9.]+)\\|"], "matchStrings": ["\\|stakater/reloader\\|reloader\\|(?<currentValue>[0-9.]+)\\|"],
"datasourceTemplate": "helm", "datasourceTemplate": "helm",
"depNameTemplate": "reloader", "depNameTemplate": "reloader",
"registryUrlTemplate": "https://stakater.github.io/stakater-charts" "registryUrlTemplate": "https://stakater.github.io/stakater-charts"
},
{
"customType": "regex",
"description": "Pinned CI BuildKit helper image",
"managerFilePatterns": [".gitea/workflows/tool-versions.env"],
"matchStrings": ["BUILDKIT_IMAGE=\"(?<depName>moby/buildkit):(?<currentValue>[^\"\\n]+)\""],
"datasourceTemplate": "docker"
} }
], ],
"packageRules": [ "packageRules": [
{ {
"description": "Automerge digest and patch updates - safe by definition, review adds nothing, keeps the renovate queue and the deploy line short. Specific no-automerge rules below still override this for playwright, helm and majors.", "description": "Automerge ordinary digest and patch updates after successful checks; specific manual-review rules below override this.",
"matchUpdateTypes": ["digest", "patch"], "matchUpdateTypes": ["digest", "patch"],
"automerge": true "automerge": true
}, },
@@ -172,23 +175,18 @@ data:
"groupSlug": "all-minor", "groupSlug": "all-minor",
"automerge": false "automerge": false
}, },
{
"description": "Group ordinary patch updates; the specific groups and manual-review rules below take precedence",
"matchUpdateTypes": ["patch"],
"groupName": "all patch updates",
"groupSlug": "all-patch"
},
{ {
"description": "Keep private homelab images unchanged", "description": "Keep private homelab images unchanged",
"matchDatasources": ["docker"], "matchDatasources": ["docker"],
"matchPackageNames": ["/gcr\\.forust\\.xyz\\/forust\\/.+/"], "matchPackageNames": ["/gcr\\.forust\\.xyz\\/forust\\/.+/"],
"enabled": false "enabled": false
}, },
{
"description": "singlesource playwright - use whichever version is found, keep docker+pypi+npm in sync",
"matchPackageNames": ["playwright", "mcr.microsoft.com/playwright"],
"groupName": "playwright singlesource",
"groupSlug": "playwright"
},
{
"description": "playwright must not automerge - version skew breaks the WS handshake (checker.py:1523 vs playwright.yaml:20)",
"matchPackageNames": ["playwright", "mcr.microsoft.com/playwright"],
"automerge": false
},
{ {
"description": "Renovate updates itself in lockstep across the CronJob and the Compose file", "description": "Renovate updates itself in lockstep across the CronJob and the Compose file",
"matchPackageNames": ["renovate/renovate"], "matchPackageNames": ["renovate/renovate"],
@@ -196,7 +194,7 @@ data:
"automerge": false "automerge": false
}, },
{ {
"description": "CI runs npm on the node the panel image is built from - the NODE_VERSION pin in tool-versions.env and node:22-alpine in the Dockerfile are the same dependency and move as one", "description": "Keep CI Node runtime updates in a separate, manually reviewed group",
"matchPackageNames": ["node"], "matchPackageNames": ["node"],
"groupName": "node runtime", "groupName": "node runtime",
"groupSlug": "node", "groupSlug": "node",
@@ -205,6 +203,8 @@ data:
{ {
"description": "Helm chart bumps change PVC fields and admission behaviour, keep them reviewable", "description": "Helm chart bumps change PVC fields and admission behaviour, keep them reviewable",
"matchDatasources": ["helm"], "matchDatasources": ["helm"],
"groupName": "Helm chart {{depName}}",
"groupSlug": "helm-{{depName}}",
"automerge": false "automerge": false
}, },
{ {
@@ -213,12 +213,6 @@ data:
"dependencyDashboardApproval": true, "dependencyDashboardApproval": true,
"automerge": false "automerge": false
}, },
{
"description": "Group patch updates from all sources - automerge still applies via the digest/patch rule above (helm/playwright stay manual via their own rules)",
"matchUpdateTypes": ["patch"],
"groupName": "all patch updates",
"groupSlug": "all-patch"
},
{ {
"description": "Python Y-bumps break compat (3.11->3.12->3.13->3.14) - keep the base image out of the shared minor/patch groups, review every bump separately. Placed last so its groupName wins.", "description": "Python Y-bumps break compat (3.11->3.12->3.13->3.14) - keep the base image out of the shared minor/patch groups, review every bump separately. Placed last so its groupName wins.",
"matchDatasources": ["docker"], "matchDatasources": ["docker"],
+1 -1
View File
@@ -19,7 +19,7 @@ spec:
restartPolicy: Never restartPolicy: Never
containers: containers:
- name: renovate - name: renovate
image: renovate/renovate:44.136.0 image: renovate/renovate:44.140.0
env: env:
- name: RENOVATE_PLATFORM - name: RENOVATE_PLATFORM
value: gitea value: gitea
+1 -1
View File
@@ -2,7 +2,7 @@ services:
renovate: renovate:
# Kept in step with renovate/k8s/cronjob.yaml by the "renovate self-update" # Kept in step with renovate/k8s/cronjob.yaml by the "renovate self-update"
# package rule in renovate/renovate.json. # package rule in renovate/renovate.json.
image: renovate/renovate:44.115.9 image: renovate/renovate:44.136.0
container_name: renovate container_name: renovate
restart: "no" restart: "no"
env_file: env_file:
+43 -49
View File
@@ -8,6 +8,9 @@
"dependencyDashboard": true, "dependencyDashboard": true,
"prCreation": "immediate", "prCreation": "immediate",
"labels": ["dependencies", "automated"], "labels": ["dependencies", "automated"],
"docker-compose": {
"managerFilePatterns": ["renovate/renovate-compose.yaml"]
},
"helm-values": { "helm-values": {
"managerFilePatterns": ["/k8s/.+values\\.ya?ml$/"] "managerFilePatterns": ["/k8s/.+values\\.ya?ml$/"]
}, },
@@ -15,35 +18,28 @@
"managerFilePatterns": ["/k8s/.+\\.ya?ml$/"] "managerFilePatterns": ["/k8s/.+\\.ya?ml$/"]
}, },
"customManagers": [ "customManagers": [
{
"customType": "regex",
"description": "singlesource: playwright npm version pinned in npx command (k8s + compose)",
"managerFilePatterns": ["^edu_master/k8s/playwright\\.yaml$", "^edu_master/compose\\.yaml$"],
"matchStrings": ["playwright@(?<currentValue>\\d+\\.\\d+\\.\\d+)"],
"datasourceTemplate": "npm",
"depNameTemplate": "playwright"
},
{
"customType": "regex",
"description": "singlesource: PLAYWRIGHT_VERSION file",
"managerFilePatterns": ["^edu_master/PLAYWRIGHT_VERSION$"],
"matchStrings": ["^(?<currentValue>\\d+\\.\\d+\\.\\d+)$"],
"datasourceTemplate": "pypi",
"depNameTemplate": "playwright"
},
{ {
"customType": "regex", "customType": "regex",
"description": "kube-prometheus-stack chart version pinned in the deploy workflow", "description": "kube-prometheus-stack chart version pinned in the deploy workflow",
"managerFilePatterns": ["^\\.gitea/workflows/deploy-lib\\.sh$"], "managerFilePatterns": [".gitea/workflows/deploy-lib.sh"],
"matchStrings": ["\\|prometheus-community/kube-prometheus-stack\\|prometheus\\|(?<currentValue>[0-9.]+)\\|"], "matchStrings": ["\\|prometheus-community/kube-prometheus-stack\\|prometheus\\|(?<currentValue>[0-9.]+)\\|"],
"datasourceTemplate": "helm", "datasourceTemplate": "helm",
"depNameTemplate": "kube-prometheus-stack", "depNameTemplate": "kube-prometheus-stack",
"registryUrlTemplate": "https://prometheus-community.github.io/helm-charts" "registryUrlTemplate": "https://prometheus-community.github.io/helm-charts"
}, },
{
"customType": "regex",
"description": "VictoriaMetrics Operator chart version pinned in the deploy workflow",
"managerFilePatterns": [".gitea/workflows/deploy-lib.sh"],
"matchStrings": ["\\|victoriametrics/victoria-metrics-operator\\|prometheus\\|(?<currentValue>[0-9.]+)\\|"],
"datasourceTemplate": "helm",
"depNameTemplate": "victoria-metrics-operator",
"registryUrlTemplate": "https://victoriametrics.github.io/helm-charts"
},
{ {
"customType": "regex", "customType": "regex",
"description": "grafana/loki chart version pinned in the deploy workflow", "description": "grafana/loki chart version pinned in the deploy workflow",
"managerFilePatterns": ["^\\.gitea/workflows/deploy-lib\\.sh$"], "managerFilePatterns": [".gitea/workflows/deploy-lib.sh"],
"matchStrings": ["\\|grafana/loki\\|prometheus\\|(?<currentValue>[0-9.]+)\\|"], "matchStrings": ["\\|grafana/loki\\|prometheus\\|(?<currentValue>[0-9.]+)\\|"],
"datasourceTemplate": "helm", "datasourceTemplate": "helm",
"depNameTemplate": "loki", "depNameTemplate": "loki",
@@ -52,7 +48,7 @@
{ {
"customType": "regex", "customType": "regex",
"description": "grafana/alloy chart version pinned in the deploy workflow", "description": "grafana/alloy chart version pinned in the deploy workflow",
"managerFilePatterns": ["^\\.gitea/workflows/deploy-lib\\.sh$"], "managerFilePatterns": [".gitea/workflows/deploy-lib.sh"],
"matchStrings": ["\\|grafana/alloy\\|prometheus\\|(?<currentValue>[0-9.]+)\\|"], "matchStrings": ["\\|grafana/alloy\\|prometheus\\|(?<currentValue>[0-9.]+)\\|"],
"datasourceTemplate": "helm", "datasourceTemplate": "helm",
"depNameTemplate": "alloy", "depNameTemplate": "alloy",
@@ -61,7 +57,7 @@
{ {
"customType": "regex", "customType": "regex",
"description": "actionlint version used by the ci workflow", "description": "actionlint version used by the ci workflow",
"managerFilePatterns": ["^\\.gitea/workflows/tool-versions\\.env$"], "managerFilePatterns": [".gitea/workflows/tool-versions.env"],
"matchStrings": ["(?:^|\\n)ACTIONLINT_VERSION=\"(?<currentValue>[0-9.]+)\""], "matchStrings": ["(?:^|\\n)ACTIONLINT_VERSION=\"(?<currentValue>[0-9.]+)\""],
"datasourceTemplate": "github-tags", "datasourceTemplate": "github-tags",
"depNameTemplate": "rhysd/actionlint" "depNameTemplate": "rhysd/actionlint"
@@ -69,7 +65,7 @@
{ {
"customType": "regex", "customType": "regex",
"description": "shellcheck version used by the ci workflow", "description": "shellcheck version used by the ci workflow",
"managerFilePatterns": ["^\\.gitea/workflows/tool-versions\\.env$"], "managerFilePatterns": [".gitea/workflows/tool-versions.env"],
"matchStrings": ["(?:^|\\n)SHELLCHECK_VERSION=\"(?<currentValue>[0-9.]+)\""], "matchStrings": ["(?:^|\\n)SHELLCHECK_VERSION=\"(?<currentValue>[0-9.]+)\""],
"datasourceTemplate": "github-tags", "datasourceTemplate": "github-tags",
"depNameTemplate": "koalaman/shellcheck" "depNameTemplate": "koalaman/shellcheck"
@@ -77,7 +73,7 @@
{ {
"customType": "regex", "customType": "regex",
"description": "kubeconform version used by the ci workflow", "description": "kubeconform version used by the ci workflow",
"managerFilePatterns": ["^\\.gitea/workflows/tool-versions\\.env$"], "managerFilePatterns": [".gitea/workflows/tool-versions.env"],
"matchStrings": ["(?:^|\\n)KUBECONFORM_VERSION=\"(?<currentValue>[0-9.]+)\""], "matchStrings": ["(?:^|\\n)KUBECONFORM_VERSION=\"(?<currentValue>[0-9.]+)\""],
"datasourceTemplate": "github-tags", "datasourceTemplate": "github-tags",
"depNameTemplate": "yannh/kubeconform" "depNameTemplate": "yannh/kubeconform"
@@ -85,7 +81,7 @@
{ {
"customType": "regex", "customType": "regex",
"description": "uv version used to build the pytest venv", "description": "uv version used to build the pytest venv",
"managerFilePatterns": ["^\\.gitea/workflows/tool-versions\\.env$"], "managerFilePatterns": [".gitea/workflows/tool-versions.env"],
"matchStrings": ["(?:^|\\n)UV_VERSION=\"(?<currentValue>[0-9.]+)\""], "matchStrings": ["(?:^|\\n)UV_VERSION=\"(?<currentValue>[0-9.]+)\""],
"datasourceTemplate": "github-tags", "datasourceTemplate": "github-tags",
"depNameTemplate": "astral-sh/uv" "depNameTemplate": "astral-sh/uv"
@@ -93,7 +89,7 @@
{ {
"customType": "regex", "customType": "regex",
"description": "prettier version used by the ci workflow", "description": "prettier version used by the ci workflow",
"managerFilePatterns": ["^\\.gitea/workflows/tool-versions\\.env$"], "managerFilePatterns": [".gitea/workflows/tool-versions.env"],
"matchStrings": ["(?:^|\\n)PRETTIER_VERSION=\"(?<currentValue>[0-9.]+)\""], "matchStrings": ["(?:^|\\n)PRETTIER_VERSION=\"(?<currentValue>[0-9.]+)\""],
"datasourceTemplate": "npm", "datasourceTemplate": "npm",
"depNameTemplate": "prettier" "depNameTemplate": "prettier"
@@ -101,7 +97,7 @@
{ {
"customType": "regex", "customType": "regex",
"description": "ruff version used by the ci workflow", "description": "ruff version used by the ci workflow",
"managerFilePatterns": ["^\\.gitea/workflows/tool-versions\\.env$"], "managerFilePatterns": [".gitea/workflows/tool-versions.env"],
"matchStrings": ["(?:^|\\n)RUFF_VERSION=\"(?<currentValue>[0-9.]+)\""], "matchStrings": ["(?:^|\\n)RUFF_VERSION=\"(?<currentValue>[0-9.]+)\""],
"datasourceTemplate": "pypi", "datasourceTemplate": "pypi",
"depNameTemplate": "ruff" "depNameTemplate": "ruff"
@@ -109,7 +105,7 @@
{ {
"customType": "regex", "customType": "regex",
"description": "pip-audit version used by the ci workflow", "description": "pip-audit version used by the ci workflow",
"managerFilePatterns": ["^\\.gitea/workflows/tool-versions\\.env$"], "managerFilePatterns": [".gitea/workflows/tool-versions.env"],
"matchStrings": ["(?:^|\\n)PIP_AUDIT_VERSION=\"(?<currentValue>[0-9.]+)\""], "matchStrings": ["(?:^|\\n)PIP_AUDIT_VERSION=\"(?<currentValue>[0-9.]+)\""],
"datasourceTemplate": "pypi", "datasourceTemplate": "pypi",
"depNameTemplate": "pip-audit" "depNameTemplate": "pip-audit"
@@ -117,7 +113,7 @@
{ {
"customType": "regex", "customType": "regex",
"description": "yamllint version used by the ci workflow", "description": "yamllint version used by the ci workflow",
"managerFilePatterns": ["^\\.gitea/workflows/tool-versions\\.env$"], "managerFilePatterns": [".gitea/workflows/tool-versions.env"],
"matchStrings": ["(?:^|\\n)YAMLLINT_VERSION=\"(?<currentValue>[0-9.]+)\""], "matchStrings": ["(?:^|\\n)YAMLLINT_VERSION=\"(?<currentValue>[0-9.]+)\""],
"datasourceTemplate": "pypi", "datasourceTemplate": "pypi",
"depNameTemplate": "yamllint" "depNameTemplate": "yamllint"
@@ -125,7 +121,7 @@
{ {
"customType": "regex", "customType": "regex",
"description": "hadolint version used by the ci workflow", "description": "hadolint version used by the ci workflow",
"managerFilePatterns": ["^\\.gitea/workflows/tool-versions\\.env$"], "managerFilePatterns": [".gitea/workflows/tool-versions.env"],
"matchStrings": ["(?:^|\\n)HADOLINT_VERSION=\"(?<currentValue>[0-9.]+)\""], "matchStrings": ["(?:^|\\n)HADOLINT_VERSION=\"(?<currentValue>[0-9.]+)\""],
"datasourceTemplate": "github-tags", "datasourceTemplate": "github-tags",
"depNameTemplate": "hadolint/hadolint" "depNameTemplate": "hadolint/hadolint"
@@ -133,7 +129,7 @@
{ {
"customType": "regex", "customType": "regex",
"description": "node version the ci workflow runs npm with", "description": "node version the ci workflow runs npm with",
"managerFilePatterns": ["^\\.gitea/workflows/tool-versions\\.env$"], "managerFilePatterns": [".gitea/workflows/tool-versions.env"],
"matchStrings": ["(?:^|\\n)NODE_VERSION=\"(?<currentValue>[0-9.]+)\""], "matchStrings": ["(?:^|\\n)NODE_VERSION=\"(?<currentValue>[0-9.]+)\""],
"datasourceTemplate": "node", "datasourceTemplate": "node",
"depNameTemplate": "node" "depNameTemplate": "node"
@@ -141,16 +137,23 @@
{ {
"customType": "regex", "customType": "regex",
"description": "stakater/reloader chart version pinned in the deploy workflow", "description": "stakater/reloader chart version pinned in the deploy workflow",
"managerFilePatterns": ["^\\.gitea/workflows/deploy-lib\\.sh$"], "managerFilePatterns": [".gitea/workflows/deploy-lib.sh"],
"matchStrings": ["\\|stakater/reloader\\|reloader\\|(?<currentValue>[0-9.]+)\\|"], "matchStrings": ["\\|stakater/reloader\\|reloader\\|(?<currentValue>[0-9.]+)\\|"],
"datasourceTemplate": "helm", "datasourceTemplate": "helm",
"depNameTemplate": "reloader", "depNameTemplate": "reloader",
"registryUrlTemplate": "https://stakater.github.io/stakater-charts" "registryUrlTemplate": "https://stakater.github.io/stakater-charts"
},
{
"customType": "regex",
"description": "Pinned CI BuildKit helper image",
"managerFilePatterns": [".gitea/workflows/tool-versions.env"],
"matchStrings": ["BUILDKIT_IMAGE=\"(?<depName>moby/buildkit):(?<currentValue>[^\"\\n]+)\""],
"datasourceTemplate": "docker"
} }
], ],
"packageRules": [ "packageRules": [
{ {
"description": "Automerge digest and patch updates - safe by definition, review adds nothing, keeps the renovate queue and the deploy line short. Specific no-automerge rules below still override this for playwright, helm and majors.", "description": "Automerge ordinary digest and patch updates after successful checks; specific manual-review rules below override this.",
"matchUpdateTypes": ["digest", "patch"], "matchUpdateTypes": ["digest", "patch"],
"automerge": true "automerge": true
}, },
@@ -161,23 +164,18 @@
"groupSlug": "all-minor", "groupSlug": "all-minor",
"automerge": false "automerge": false
}, },
{
"description": "Group ordinary patch updates; the specific groups and manual-review rules below take precedence",
"matchUpdateTypes": ["patch"],
"groupName": "all patch updates",
"groupSlug": "all-patch"
},
{ {
"description": "Keep private homelab images unchanged", "description": "Keep private homelab images unchanged",
"matchDatasources": ["docker"], "matchDatasources": ["docker"],
"matchPackageNames": ["/gcr\\.forust\\.xyz\\/forust\\/.+/"], "matchPackageNames": ["/gcr\\.forust\\.xyz\\/forust\\/.+/"],
"enabled": false "enabled": false
}, },
{
"description": "singlesource playwright - use whichever version is found, keep docker+pypi+npm in sync",
"matchPackageNames": ["playwright", "mcr.microsoft.com/playwright"],
"groupName": "playwright singlesource",
"groupSlug": "playwright"
},
{
"description": "playwright must not automerge - version skew breaks the WS handshake (checker.py:1523 vs playwright.yaml:20)",
"matchPackageNames": ["playwright", "mcr.microsoft.com/playwright"],
"automerge": false
},
{ {
"description": "Renovate updates itself in lockstep across the CronJob and the Compose file", "description": "Renovate updates itself in lockstep across the CronJob and the Compose file",
"matchPackageNames": ["renovate/renovate"], "matchPackageNames": ["renovate/renovate"],
@@ -185,7 +183,7 @@
"automerge": false "automerge": false
}, },
{ {
"description": "CI runs npm on the node the panel image is built from - the NODE_VERSION pin in tool-versions.env and node:22-alpine in the Dockerfile are the same dependency and move as one", "description": "Keep CI Node runtime updates in a separate, manually reviewed group",
"matchPackageNames": ["node"], "matchPackageNames": ["node"],
"groupName": "node runtime", "groupName": "node runtime",
"groupSlug": "node", "groupSlug": "node",
@@ -194,6 +192,8 @@
{ {
"description": "Helm chart bumps change PVC fields and admission behaviour, keep them reviewable", "description": "Helm chart bumps change PVC fields and admission behaviour, keep them reviewable",
"matchDatasources": ["helm"], "matchDatasources": ["helm"],
"groupName": "Helm chart {{depName}}",
"groupSlug": "helm-{{depName}}",
"automerge": false "automerge": false
}, },
{ {
@@ -202,12 +202,6 @@
"dependencyDashboardApproval": true, "dependencyDashboardApproval": true,
"automerge": false "automerge": false
}, },
{
"description": "Group patch updates from all sources - automerge still applies via the digest/patch rule above (helm/playwright stay manual via their own rules)",
"matchUpdateTypes": ["patch"],
"groupName": "all patch updates",
"groupSlug": "all-patch"
},
{ {
"description": "Python Y-bumps break compat (3.11->3.12->3.13->3.14) - keep the base image out of the shared minor/patch groups, review every bump separately. Placed last so its groupName wins.", "description": "Python Y-bumps break compat (3.11->3.12->3.13->3.14) - keep the base image out of the shared minor/patch groups, review every bump separately. Placed last so its groupName wins.",
"matchDatasources": ["docker"], "matchDatasources": ["docker"],
+2 -2
View File
@@ -12,11 +12,11 @@ services:
- "traefik.enable=true" - "traefik.enable=true"
- "traefik.http.services.searxng.loadbalancer.server.port=8080" - "traefik.http.services.searxng.loadbalancer.server.port=8080"
# Prod Router # Prod Router
- "traefik.http.routers.searxng.rule=Host(`s.forust.xyz` || `search.forust.xyz`)" - "traefik.http.routers.searxng.rule=Host(`s.forust.xyz`) || Host(`search.forust.xyz`)"
- "traefik.http.routers.searxng.entrypoints=websecure" - "traefik.http.routers.searxng.entrypoints=websecure"
- "traefik.http.routers.searxng.tls.certresolver=letsencrypt" - "traefik.http.routers.searxng.tls.certresolver=letsencrypt"
# Local Router # Local Router
- "traefik.http.routers.searxng-local.rule=Host(`s.workstation.internal` || `searxng.workstation.internal`)" - "traefik.http.routers.searxng-local.rule=Host(`s.workstation.internal`) || Host(`searxng.workstation.internal`)"
- "traefik.http.routers.searxng-local.entrypoints=websecure" - "traefik.http.routers.searxng-local.entrypoints=websecure"
- "traefik.http.routers.searxng-local.tls=true" - "traefik.http.routers.searxng-local.tls=true"
# Dev Router # Dev Router
View File
Whitespace-only changes.
+1 -1
View File
@@ -19,7 +19,7 @@ services:
- streaming - streaming
qbittorrent: qbittorrent:
image: lscr.io/linuxserver/qbittorrent:5.2.4 image: lscr.io/linuxserver/qbittorrent:20.04.1
container_name: qbittorrent container_name: qbittorrent
restart: unless-stopped restart: unless-stopped
environment: environment:
View File
Whitespace-only changes.
+1 -1
View File
@@ -1,6 +1,6 @@
services: services:
termix: termix:
image: ghcr.io/lukegus/termix:2.9.1 image: ghcr.io/lukegus/termix:2.9.2
container_name: termix container_name: termix
restart: unless-stopped restart: unless-stopped
# ports: # ports:
Loaded 100 of 108 files, more files were not shown because too many files have changed in this diff. Show more