diff --git a/.gitea/EDU_HANDOFF.md b/.gitea/EDU_HANDOFF.md new file mode 100644 index 0000000..a0b8dd8 --- /dev/null +++ b/.gitea/EDU_HANDOFF.md @@ -0,0 +1,50 @@ +# EDU ownership handoff + +## Status + +The EDU ownership handoff is complete. The homelab repository no longer owns +EDU workloads, images, routes, alerts, or deployment selection. The EDU +repository is the only deployment owner: [forust/edu-master](https://git.forust.xyz/forust/edu-master). + +Homelab PRs #99 and #105 are merged. PR #105 removed the EDU subtree and its +build, deploy, rollback, verification, route-probe, and registry references. +It also added the serial image build matrix for the homelab services. This +handoff record is the only remaining EDU-specific file in homelab Git. + +The dedicated workstation checkout is `/srv/edu-master`, at release +`4f2b2a0e37dc11ac2c75441a15076c178e219d37`. It contains `k8s/active`; root +`active` is absent. The old untracked `/srv/homelab/edu_master` checkout was +moved outside the homelab repository to +`/srv/edu-master-legacy-archive-20261007/edu_master`. Its private files remain +mode `0600` inside an archive directory with mode `0700`. The homelab deploy +checkout has no EDU marker or tracked EDU application/deployment files. +`AUTODEPLOY=false` remains in place for homelab deployment. + +## Release evidence + +EDU PR #4 merged after its review and CI checks. Main-push CI run 1652 passed +all validation and both image builds. Deploy run 1653 passed for the exact main +SHA above. + +The workstation rollout completed for both Deployments. The deployment +verified `/health` and `/live` with HTTP 200, Redis AUTH, session TTL of 1058 +seconds, a delivery backlog of zero, and all nine EDU vmalert rules with +matching expressions and healthy evaluation. + +The images now run by digest: + +- Session keeper: `sha256:998dea51aa3015fd9cabefb0f53b030157a650c3bef72e02fe84f17d5762613d` +- Webinar checker: `sha256:92f3c1fa2bb7f9b4680a9fc76a5b33dfbea8ef3dd9c6490ebc45876fd4c54461` + +Redis StatefulSet was unchanged. PVC `redis-data-pvc` remains bound to PV +`pvc-a4f2a79a-363a-4c12-ae91-92cdfc2a0d2e` with capacity 1 GiB. The existing +runtime Secret and Fernet key were preserved during the handoff. Notification +delivery was verified before closeout, as confirmed by the operator. The +deployment did not record downtime. + +The release rollback snapshot is +`/home/forust/.local/state/edu-master-deploy/20261007T180541Z-4f2b2a0e37dc11ac2c75441a15076c178e219d37`. +The handoff data snapshot remains at +`/home/forust/.local/state/edu-master-deploy/handoff-20261007T080838Z`. +Both snapshots are outside Git. Do not restore old Redis data unless recovery +requires it. Never delete or recreate the Redis PVC. diff --git a/.gitea/README.md b/.gitea/README.md index 856cd9f..4d906b8 100644 --- a/.gitea/README.md +++ b/.gitea/README.md @@ -1,88 +1,64 @@ -# Build and deployment workflows +# CI and deployment -Gitea Actions checks this repository, builds its custom images, and deploys -selected services to the workstation. Workflows use the self-hosted runner labels -`linux`, `arch`, and `homelab`; deployment jobs also require `prod`. +Gitea Actions validates changes, builds the repository's custom images, and can +deploy selected services to the workstation. CI and production deployment use +separate workflows. See the [runner and recovery guide](runner/README.md) for +installation, configuration, and operator commands. -## Checks +## CI -`ci.yaml` runs Compose validation, actionlint, ShellCheck, Prettier, Ruff, -yamllint, hadolint, and kubeconform. Tool versions are pinned in -`workflows/tool-versions.env` and installed by `install-ci-tools.sh`. +`workflows/ci.yaml` runs Compose, workflow, shell, formatting, Python and unit +test, YAML, Dockerfile, and Kubernetes checks. Pull requests and non-main refs +use the unprivileged `homelab-pr` runner. Main-branch CI uses `homelab`. Tool +versions are pinned in `workflows/tool-versions.env`. -Compose CI checks structure without resolving local environment files or paths. -On the reviewed main commit it only discovers standard filenames; the -`fix/deploy-validation` branch adds the manual Compose entry points too. +Compose CI checks every committed Compose file without requiring ignored `.env` +files. Kubernetes checks validate known schemas; unknown CRDs are skipped. -Kubeconform validates known resource schemas. Unknown CRDs are skipped. On main, -CI also attempts server-side dry-runs for marked services; these require an -existing namespace and contact the cluster's admission webhooks. A cluster that -is unreachable produces a warning and skips that CI pass. Deploy validation has -its own dry-run stage. +On main, CI plans builds for the three owned images: `error-pages`, +`forust-homepage`, and `xdfnx-homepage`. It builds changed inputs or reuses a +digest from a successful earlier main run. The successful build job publishes a +release artifact for the exact commit SHA. Pull requests do not publish images. -`renovate-ci.yaml` validates Renovate settings and checks that its generated -ConfigMap matches `renovate/renovate.json`. +## Deployment gate -## Image builds +`workflows/deploy.yaml` starts a deployment after successful main CI when the +`AUTODEPLOY` Actions variable is `true`. Manual dispatch uses the same gate: the +requested `main` ref or commit must have successful main CI and its matching +release artifact. A manual dispatch does not bypass validation. -CI builds changed custom images for `errorpages`, both `homepages` variants, and -the two `edu_master` Python services. Main builds publish `main`, `prod`, and a -commit tag. Dev builds publish `dev`. Build jobs wait for the lint and manifest -checks. +The workflow supports these modes: -Kubernetes deployment resolves the lab's own registry images to digests, preferring -commit-specific tags. Third-party image versions remain declared in the manifests. +- `changed`: select active services changed since the last successful deploy. +- `full`: select all active services; use this for the first baseline. +- `plan`: validate and show the selection without applying production resources. -## Deploy selection +`refresh_images=true` explicitly refreshes mutable third-party Compose tags. -`workflows/deploy-lib.sh` owns the stage logic; `ssh-run.sh` invokes it on the -workstation through SSH. Kubernetes selection uses `k8s/active`; Compose selection -uses an `active` file beside a standard `compose.yaml` or `compose.yml`. -Kustomize overlays are supported, although the current tree primarily contains -plain manifests. +## Selection and rollout -Secret files, examples, Helm values, and patch files are excluded from plain -manifest selection. Create local Kubernetes Secrets separately in their target -namespaces. The Helm table lists Prometheus, Loki, Alloy, and Reloader, with each -release controlled by its configured marker. Other charts need separate setup. +The active markers define automatic deployment. `/active` selects a +standard Compose file; `/k8s/active` selects Kubernetes resources. Helm +releases have their own markers in `workflows/deploy-lib.sh`. Service +dependencies are declared in `deploy-dependencies.json`. Removed resources are +reported for manual review; the workflow does not prune them automatically. -## Trigger and required settings +The workstation controller runs the checked source in a per-SHA worktree. It +validates configuration, applies Kubernetes and Compose changes in sequence, +verifies changed Kubernetes workloads, and checks public routes. A durable +systemd service continues the rollout if the Actions SSH client disconnects. +The workflow checks the exact CI release before it submits a deployment. -Automatic deployment follows a successful main CI run when the repository Actions -variable `AUTODEPLOY` is `true`. The manual deploy workflow bypasses that switch -and targets the fetched main branch when no validated commit SHA is provided. -A manual dispatch does not prove that this commit passed CI. +Kubernetes recovery uses captured workload revisions. It does not restore +ConfigMaps, Secrets, database schemas, or persistent data. Compose recovery is +manual and does not restore volume data or reverse migrations. Keep backups for +stateful services. The runner guide documents status, retry, logs, and recovery +commands. -Configure the Actions secrets `DEPLOY_HOST`, `DEPLOY_USER`, `DEPLOY_SSH_KEY`, and, -where needed, `DEPLOY_PORT` and `DEPLOY_PATH`. Registry publishing uses -`REGISTRY_USERNAME` and `REGISTRY_PASSWORD`. The remote user needs access to Git, -Docker, kubectl, Helm, jq, and the state directory used for snapshots. +## Settings -Keep `APPLY_PRUNE` false on the reviewed implementation: its per-file prune loop -is unsafe. `fix/deploy-prune-guard` rejects that option before changes are applied. - -Preflight fetches and resets the remote checkout. It refuses when tracked files -have local changes; ignored local env and Secret files stay in place. Do not use -a development checkout with uncommitted tracked changes as the deployment target. - -## Stages and recovery - -1. Preflight fetches the target commit and checks the remote working tree. -2. Validate selects services, parses Compose, performs Kubernetes dry-runs, and - checks referenced Secrets. -3. Apply Kubernetes records a workload snapshot, upgrades selected Helm releases, - applies resources, and refreshes owned custom images. -4. Apply Compose recreates marked stacks and checks container state. -5. Verify Kubernetes checks changed workloads and attempts rollback for failures. -6. Smoke probes public routes after verification. - -The two apply jobs share a remote lock. Workflow concurrency queues deployments -rather than interrupting an older apply. Snapshots live under -`$XDG_STATE_HOME/homelab-deploy`, or `~/.local/state/homelab-deploy` by default. -They contain the pre-apply workload data and commit identifier. - -Rollback uses workload revisions. It does not restore ConfigMaps, Secrets, -database schemas, or data. Helm-owned workloads are handled through the Helm -upgrade's rollback path; the generic rollback skips them. Compose has no automatic -rollback. See the [review](../docs/repository-review.md) for remaining recovery -limitations, including SSH retries and serial rollback timing. +Configure `DEPLOY_HOST`, `DEPLOY_USER`, `DEPLOY_PORT`, and the verified +`DEPLOY_KNOWN_HOSTS` entry as Actions variables. Keep `DEPLOY_SSH_KEY`, +`REGISTRY_USERNAME`, and `REGISTRY_PASSWORD` in Actions secrets. The workstation +also needs its existing registry authentication. Set `AUTODEPLOY=false` until +automatic production deploys are intended. diff --git a/.gitea/actionlint.yaml b/.gitea/actionlint.yaml index b7c699a..b5a9cc3 100644 --- a/.gitea/actionlint.yaml +++ b/.gitea/actionlint.yaml @@ -7,4 +7,5 @@ self-hosted-runner: labels: - arch - homelab + - homelab-pr - prod diff --git a/.gitea/deploy-dependencies.json b/.gitea/deploy-dependencies.json new file mode 100644 index 0000000..7c5f3f8 --- /dev/null +++ b/.gitea/deploy-dependencies.json @@ -0,0 +1,3 @@ +{ + "postgres": ["authentik", "gitea", "immich", "n8n", "netbox", "netronome"] +} diff --git a/.gitea/runner/README.md b/.gitea/runner/README.md new file mode 100644 index 0000000..98f2a24 --- /dev/null +++ b/.gitea/runner/README.md @@ -0,0 +1,181 @@ +# Homelab CI/CD + +The native Gitea runners run on **vps**; production runs on **workstation**. +Main-branch checks and image builds use `homelab:host`. Pull request and +non-main checks use `homelab-pr:host` under a separate account without Docker +access. The `homelab-pr` runner is registered at User scope for `forust`, so +any repository under that account can schedule jobs that request this label. +Each runner accepts one job at a time; the build waits for every check to pass. +CI and deploy runs also show a summary with +the release SHA, image build or reuse results, deploy mode, selected services, +and image digests. Failed runs keep a summary of completed image builds, stage +results, apply results, and recorded Kubernetes recovery. The final deploy +summary is in the smoke job; earlier jobs show the state observed at that time. +Apply success is separate from health and recovery. Update the installed +workstation controller with `setup-workstation.sh` when no deploy is running. +No job images or Kubernetes credentials are needed on the VPS. Builds use one +pinned BuildKit helper container. CI and deploy are separate workflows. + +## Runner installation + +Install Docker Engine with Compose and Buildx, Git, Python 3.11+, Bash, curl, +GNU tar/xz, flock and systemd using the host's package manager. Keep the existing +Gitea runner 3.0.2 binary at `/usr/local/bin/gitea-runner`. + +From this checkout on the VPS: + +```sh +sudo bash .gitea/runner/setup-runner.sh +``` + +The installer reuses `/var/lib/gitea-runner/.runner` and the existing service. +For a new host, install the same runner binary and register as `gitea-runner` +using the registration token interactively, label `homelab:host`, and working +directory `/var/lib/gitea-runner`; then rerun the installer. Tokens never belong +in this repository or command-line examples. + +Pinned tools live in the runner user's `~/.cache/homelab-ci`; CI repairs version +drift there. Installations are locked. Buildx uses only the `homelab-ci` builder, +pushes directly to the registry, and caps retained local cache at 1 GiB with a +2 GiB free-space target. This is not a hard limit on peak build disk usage. +Nothing runs `docker system prune`, removes unrelated images, or deletes volumes. + +### Pull request runner + +Install the unprivileged host runner on the VPS: + +```sh +sudo bash .gitea/runner/setup-pr-runner.sh +``` + +Get a registration token from the user Actions runner settings. Run the +installer in a terminal. It asks for the token without echoing it, registers the +runner as `homelab-pr` with label `homelab-pr:host`, then enables the service. +The work directory is `/var/lib/gitea-pr-runner`. Confirm that Gitea lists the +runner as User scope before merging the workflow change. An unmatched label can +fall back to the default job image. + +Renovate PR validation uses `pull_request_target`, which reads the workflow from +the base branch. It checks out the PR head only after runner selection and runs +that code on `homelab-pr`. Keep this workflow read-only and do not add secrets. + +The PR runner has a separate home and tool cache. Do not add it to the `docker` +group or give it access to `/var/run/docker.sock`. It runs repository code from +pull requests, so keep its registration and permissions separate from the +trusted `homelab` runner. This separates users and host permissions, but both +runners still share the VPS kernel and network. Use a disposable VM if PRs from +untrusted external authors must be fully isolated. + +## Workstation setup + +As the existing SSH deploy user on workstation: + +```sh +sudo loginctl enable-linger forust +bash .gitea/runner/setup-workstation.sh +``` + +The controller uses `/srv/homelab` as the persistent configuration tree and makes +a detached source worktree for each SHA. It never resets `/srv/homelab`, moves +local configuration, renames Compose projects, or changes volume names. +The installer records the current Kubernetes context and cluster UID in +`~/.config/homelab-deploy/environment`. Check these before installing. + +Configure Gitea Actions Variables: + +- `DEPLOY_HOST`, `DEPLOY_USER`, `DEPLOY_PORT`: the existing VPS-to-workstation SSH endpoint. +- `DEPLOY_KNOWN_HOSTS`: workstation's verified SSH host key entry for that endpoint. +- `AUTODEPLOY`: `false` initially; `true` enables deployment after successful main CI. + +Keep `DEPLOY_SSH_KEY`, `REGISTRY_USERNAME` and `REGISTRY_PASSWORD` in Actions +Secrets. Legacy endpoint secrets remain accepted during migration. The Actions +token must have repository read and Actions read access for release downloads. +The deploy user's existing Docker registry authentication remains necessary. + +## Releases and deployment + +CI publishes `release-` as a Gitea artifact with all three owned image +digests and build input fingerprints. Unchanged images are reused only from a +successful main CI artifact, never from `:prod`. Expired artifacts cause CI to +rebuild images; they block deployment until CI is rerun. + +Run deploy from main with `deploy_ref=main` or a checked SHA: + +- `full`: required for the first baseline; reconcile all active components. +- `changed`: compare with the last fully successful production deploy. +- `plan`: validate configuration and show selection without changing production resources. +- `refresh_images=true`: explicitly refresh mutable third-party Compose tags. + +The manual and automatic paths both require successful CI, a successful build +job and the exact SHA's release artifact. PRs cannot publish images or deploy. +Removed resources are reported and require explicit removal; no automatic prune. +Service dependencies are listed in `.gitea/deploy-dependencies.json`. + +A workstation user systemd service holds the deploy lock across validation, +sequential apply, verification and smoke checks. SSH clients only submit/follow: +disconnecting or cancelling the Actions client does not kill production apply. +Retrying the same run ID does not start another apply. `ExecStopPost` recovers +interrupted runs before the unit finishes. Kubernetes rolls back to captured +revisions; configuration and persistent data are not reverted. + +## Status and recovery + +`--retry` repeats failed verification and smoke checks, never apply. Recovery +keeps a failed deploy out of the successful baseline, even after rollback. + +On workstation (replace the numeric ID with Actions run ID and attempt): + +```sh +python3 ~/.local/lib/homelab-deploy/controller.py status 123-1 +python3 ~/.local/lib/homelab-deploy/controller.py recover 123-1 --retry +journalctl --user -u homelab-deploy@123-1 +``` + +Runs live in `~/.local/state/homelab-deploy/runs`. Compose stores resolved configs +with restricted permissions; these may contain credentials and must never be +uploaded as CI artifacts. Stage logs print the exact manual recovery command +using `compose-before/.json`, the original project directory and project +name. Compose does not automatically roll back, and Nextcloud AIO's child +containers remain managed by AIO. Preserve its own backups for data recovery. + +The controller retains twenty successful/planned runs and preserves failures. +Update the workstation dispatcher only when no deploy is running. + +## Validation and migration rollback + +```sh +python3 -m unittest discover -s tests -v +bash .gitea/tests/deploy-validation.sh +``` + +Test on a separate namespace before the initial production `full` run. Check a +failed rollout, interrupted SSH and repeated run ID, and verify that an isolated +service change does not upgrade unrelated Helm releases or Compose stacks. + +To roll back the migration, disable autodeploy and finish or recover the remote +run first. Restore the runner config/unit from `.before-` backups, +reload systemd and restart the runner. Restore the prior workflows from Git. +Production data and persistent volumes stay where they were. Do not remove run +state or Compose recovery files until recovery is confirmed. + +### Compose configuration recovery + +Successful deploys save the complete resolved Compose configuration in +`~/.local/state/homelab-deploy/compose-configs/`. These files can contain secrets. +Keep them private and do not commit or upload them. +The next deploy uses this configuration for its recovery file, including old +commands, environment, mounts, ports, and removed services. The recovery command +uses `--remove-orphans` to remove services added by the failed deploy. It does +not restore volume data or reverse database migrations. + +On the first run after this update, the controller can use the Compose file +from the previous successful run. If that file is absent, it reads the persistent +checkout and checks its service configuration hashes against existing containers. +A mismatch stops preflight. Restore the previous configuration before retrying. +Update the installed controller with `bash .gitea/runner/setup-workstation.sh` +from the reviewed checkout before using this change. + +New namespaces are checked during preflight. Server validation of their resources +runs after namespace creation and before application resources are applied. +Plan mode does not create namespaces. A failed deferred check can leave an empty +namespace; inspect it before removing it. diff --git a/.gitea/runner/buildkitd.toml b/.gitea/runner/buildkitd.toml new file mode 100644 index 0000000..9b23273 --- /dev/null +++ b/.gitea/runner/buildkitd.toml @@ -0,0 +1,11 @@ +[worker.oci] + gc = true + reservedSpace = "256MB" + maxUsedSpace = "1GB" + minFreeSpace = "2GB" + +[[worker.oci.gcpolicy]] + reservedSpace = "256MB" + maxUsedSpace = "1GB" + minFreeSpace = "2GB" + all = true diff --git a/.gitea/runner/config.yaml b/.gitea/runner/config.yaml new file mode 100644 index 0000000..77028e5 --- /dev/null +++ b/.gitea/runner/config.yaml @@ -0,0 +1,10 @@ +runner: + file: /var/lib/gitea-runner/.runner + capacity: 1 + timeout: 5h + labels: + - homelab:host +cache: + enabled: false +container: + docker_host: unix:///var/run/docker.sock diff --git a/.gitea/runner/gitea-runner.service b/.gitea/runner/gitea-runner.service new file mode 100644 index 0000000..8665bfc --- /dev/null +++ b/.gitea/runner/gitea-runner.service @@ -0,0 +1,18 @@ +[Unit] +Description=Gitea Actions runner +After=network-online.target docker.service +Wants=network-online.target + +[Service] +User=gitea-runner +Group=gitea-runner +SupplementaryGroups=docker +WorkingDirectory=/var/lib/gitea-runner +Environment=PATH=/var/lib/gitea-runner/.cache/homelab-ci/bin:/usr/local/bin:/usr/bin:/bin +ExecStart=/usr/local/bin/gitea-runner daemon --config /etc/gitea-runner/config.yaml +Restart=on-failure +RestartSec=5 +UMask=0077 + +[Install] +WantedBy=multi-user.target diff --git a/.gitea/runner/homelab-deploy@.service b/.gitea/runner/homelab-deploy@.service new file mode 100644 index 0000000..1e9ceab --- /dev/null +++ b/.gitea/runner/homelab-deploy@.service @@ -0,0 +1,12 @@ +[Unit] +Description=Homelab deploy %i + +[Service] +Type=exec +EnvironmentFile=%h/.config/homelab-deploy/environment +ExecStart=/usr/bin/python3 %h/.local/lib/homelab-deploy/controller.py execute %i +ExecStopPost=/usr/bin/python3 %h/.local/lib/homelab-deploy/controller.py recover %i +RuntimeMaxSec=5h +TimeoutStopSec=135min +KillMode=control-group +UMask=0077 diff --git a/.gitea/runner/pr-config.yaml b/.gitea/runner/pr-config.yaml new file mode 100644 index 0000000..2b6ee49 --- /dev/null +++ b/.gitea/runner/pr-config.yaml @@ -0,0 +1,8 @@ +runner: + file: /var/lib/gitea-pr-runner/.runner + capacity: 1 + timeout: 5h + labels: + - homelab-pr:host +cache: + enabled: false diff --git a/.gitea/runner/pr-runner.service b/.gitea/runner/pr-runner.service new file mode 100644 index 0000000..9c6fd5a --- /dev/null +++ b/.gitea/runner/pr-runner.service @@ -0,0 +1,27 @@ +[Unit] +Description=Gitea Actions untrusted pull request runner +After=network-online.target +Wants=network-online.target + +[Service] +User=gitea-pr-runner +Group=gitea-pr-runner +WorkingDirectory=/var/lib/gitea-pr-runner +Environment=HOME=/var/lib/gitea-pr-runner +Environment=PATH=/var/lib/gitea-pr-runner/.cache/homelab-ci/bin:/usr/local/bin:/usr/bin:/bin +ExecStart=/usr/local/bin/gitea-runner daemon --config /etc/gitea-pr-runner/config.yaml +Restart=on-failure +RestartSec=5 +NoNewPrivileges=yes +PrivateTmp=yes +ProtectSystem=full +ProtectHome=yes +ProtectKernelTunables=yes +ProtectKernelModules=yes +ProtectControlGroups=yes +RestrictSUIDSGID=yes +LockPersonality=yes +UMask=0077 + +[Install] +WantedBy=multi-user.target diff --git a/.gitea/runner/setup-pr-runner.sh b/.gitea/runner/setup-pr-runner.sh new file mode 100755 index 0000000..7c71679 --- /dev/null +++ b/.gitea/runner/setup-pr-runner.sh @@ -0,0 +1,56 @@ +#!/usr/bin/env bash +# Install a native runner for untrusted PR jobs without Docker access. +set -euo pipefail +here="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" +[ "$(id -u)" -eq 0 ] || { echo 'Run with sudo on the runner host' >&2; exit 1; } +for tool in cp cut date getent id install runuser systemctl useradd; do + command -v "$tool" >/dev/null || { echo "Install missing prerequisite: $tool" >&2; exit 1; } +done +command -v /usr/local/bin/gitea-runner >/dev/null || { + echo 'Install gitea-runner 3.0.2 at /usr/local/bin/gitea-runner first' >&2 + exit 1 +} + +id gitea-pr-runner >/dev/null 2>&1 || \ + useradd --system --create-home --home-dir /var/lib/gitea-pr-runner --shell /usr/bin/bash gitea-pr-runner +runner_home="$(getent passwd gitea-pr-runner | cut -d: -f6)" +[ "$runner_home" = /var/lib/gitea-pr-runner ] || { + echo 'Unexpected PR runner home; inspect the existing service first' >&2 + exit 1 +} +case " $(id -nG gitea-pr-runner) " in + *' docker '*) + echo 'The PR runner account must not belong to the docker group' >&2 + exit 1 + ;; +esac + +install -d -m 0755 /etc/gitea-pr-runner +stamp="$(date -u +%Y%m%dT%H%M%SZ)" +for existing in /etc/gitea-pr-runner/config.yaml /etc/systemd/system/gitea-pr-runner.service; do + [ ! -f "$existing" ] || cp -p "$existing" "$existing.before-$stamp" +done +install -m 0644 "$here/pr-config.yaml" /etc/gitea-pr-runner/config.yaml +install -m 0644 "$here/pr-runner.service" /etc/systemd/system/gitea-pr-runner.service + +if [ ! -f /var/lib/gitea-pr-runner/.runner ]; then + read -r -s -p 'Enter the Gitea repository runner registration token: ' runner_token + printf '\n' + [ -n "$runner_token" ] || { echo 'Runner token is required' >&2; exit 1; } + export GITEA_RUNNER_REGISTRATION_TOKEN="$runner_token" + unset runner_token + runuser --preserve-environment -u gitea-pr-runner -- \ + /usr/local/bin/gitea-runner register \ + --config /etc/gitea-pr-runner/config.yaml \ + --instance https://gitea.forust.xyz \ + --name homelab-pr \ + --labels homelab-pr:host \ + --no-interactive + unset GITEA_RUNNER_REGISTRATION_TOKEN +fi +chmod 0600 /var/lib/gitea-pr-runner/.runner + +systemctl daemon-reload +systemctl enable --now gitea-pr-runner.service +systemctl restart gitea-pr-runner.service +echo "PR runner ready. Configuration backups: *.before-$stamp" diff --git a/.gitea/runner/setup-runner.sh b/.gitea/runner/setup-runner.sh new file mode 100755 index 0000000..bacd7e1 --- /dev/null +++ b/.gitea/runner/setup-runner.sh @@ -0,0 +1,48 @@ +#!/usr/bin/env bash +# Native host runner, with pinned user-space tools and no extra CI images. +set -euo pipefail +here="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" +[ "$(id -u)" -eq 0 ] || { echo 'Run with sudo on the runner host' >&2; exit 1; } +for tool in docker curl python3 git tar xz flock runuser systemctl; do + command -v "$tool" >/dev/null || { echo "Install missing prerequisite: $tool" >&2; exit 1; } +done +docker info >/dev/null +docker compose version >/dev/null +docker buildx version >/dev/null +id gitea-runner >/dev/null 2>&1 || useradd --system --create-home --home-dir /var/lib/gitea-runner --shell /usr/bin/bash gitea-runner +# Reuse the established service account and runner registration. +runner_home="$(getent passwd gitea-runner | cut -d: -f6)" +[ "$runner_home" = /var/lib/gitea-runner ] || { echo 'Unexpected runner home; inspect the existing service first' >&2; exit 1; } +runuser -u gitea-runner -- docker info >/dev/null || { echo "The runner user needs access to Docker before setup" >&2; exit 1; } +command -v gitea-runner >/dev/null || { echo 'Install gitea-runner 3.0.2 at /usr/local/bin/gitea-runner first' >&2; exit 1; } +mkdir -p /etc/gitea-runner +stamp="$(date -u +%Y%m%dT%H%M%SZ)" +for existing in /etc/gitea-runner/config.yaml /etc/systemd/system/gitea-runner.service; do + [ ! -f "$existing" ] || cp -p "$existing" "$existing.before-$stamp" +done +scratch="$(mktemp -d)" +trap 'rm -rf "$scratch"' EXIT +chmod 755 "$scratch" +install -m 0644 "$here/../workflows/install-ci-tools.sh" "$here/../workflows/tool-versions.env" "$scratch/" +runuser -u gitea-runner -- bash "$scratch/install-ci-tools.sh" +install -m 0644 "$here/config.yaml" /etc/gitea-runner/config.yaml +python3 - <<'PYLABELS' +import json +from pathlib import Path +registration = Path('/var/lib/gitea-runner/.runner') +if registration.exists(): + labels = json.loads(registration.read_text()).get('labels', []) + labels = [label for label in labels if isinstance(label, str) and label.split(':')[0] != 'homelab'] + labels.append('homelab:host') + config = Path('/etc/gitea-runner/config.yaml') + config.write_text(config.read_text().replace(' - homelab:host', '\n'.join(' - ' + json.dumps(label) for label in labels))) +PYLABELS +install -m 0644 "$here/gitea-runner.service" /etc/systemd/system/gitea-runner.service +if [ ! -f /var/lib/gitea-runner/.runner ]; then + echo 'Register once as gitea-runner with homelab:host before starting the service.' + exit 0 +fi +systemctl daemon-reload +systemctl enable --now gitea-runner.service +systemctl restart gitea-runner.service +echo "Runner ready. Configuration backups: *.before-$stamp" diff --git a/.gitea/runner/setup-workstation.sh b/.gitea/runner/setup-workstation.sh new file mode 100755 index 0000000..a6d5f9b --- /dev/null +++ b/.gitea/runner/setup-workstation.sh @@ -0,0 +1,32 @@ +#!/usr/bin/env bash +# Run as the existing deploy user on workstation. Never resets the working tree. +set -euo pipefail +here="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" +repo="${HOMELAB_REPO:-/srv/homelab}" +for tool in python3 git kubectl helm docker flock timeout; do + command -v "$tool" >/dev/null || { echo "Install missing dependency: $tool" >&2; exit 1; } +done +[ -d "$repo/.git" ] || { echo "Missing deploy checkout: $repo" >&2; exit 1; } +[[ "$repo" =~ ^/[A-Za-z0-9_./-]+$ ]] || { echo 'Deploy path must be absolute and contain no whitespace' >&2; exit 1; } +if [ "$(loginctl show-user "$USER" -p Linger --value)" != yes ]; then + echo "Run once: sudo loginctl enable-linger $USER" >&2 + exit 1 +fi +config="${XDG_CONFIG_HOME:-$HOME/.config}/homelab-deploy" +mkdir -p "$config" "$HOME/.local/lib/homelab-deploy" "$HOME/.config/systemd/user" +chmod 700 "$config" +if [ ! -f "$config/environment" ]; then + context="$(kubectl config current-context)" + cluster_uid="$(kubectl get namespace kube-system -o jsonpath='{.metadata.uid}')" + printf 'HOMELAB_REPO=%s\nKUBE_CONTEXT=%s\nEXPECTED_CLUSTER_UID=%s\n' "$repo" "$context" "$cluster_uid" >"$config/environment" + chmod 600 "$config/environment" +fi +# Do not replace a dispatcher while an existing deploy uses it. +if systemctl --user list-units 'homelab-deploy@*' --state=running --no-legend | grep -q .; then + echo 'An existing deploy is running; wait before updating the controller' >&2 + exit 1 +fi +install -m 0755 "$here/../workflows/deploy-controller.py" "$HOME/.local/lib/homelab-deploy/controller.py" +install -m 0644 "$here/homelab-deploy@.service" "$HOME/.config/systemd/user/homelab-deploy@.service" +systemctl --user daemon-reload +echo 'Controller ready. Run a checked main SHA in full mode for the initial baseline.' diff --git a/.gitea/tests/deploy-validation.sh b/.gitea/tests/deploy-validation.sh new file mode 100755 index 0000000..b93ea3b --- /dev/null +++ b/.gitea/tests/deploy-validation.sh @@ -0,0 +1,156 @@ +#!/usr/bin/env bash +# Local regressions only: kubectl is mocked and Docker is used for config parsing. +set -euo pipefail +repo="$(cd "$(dirname "${BASH_SOURCE[0]}")/../.." && pwd)" +scratch="$(mktemp -d)" +trap 'rm -rf "$scratch"' EXIT + +mkdir -p "$scratch/repo/app" "$scratch/repo/postgres" "$scratch/repo/netbird" "$scratch/repo/renovate" +git -C "$scratch/repo" init -q +for file in app/compose.yaml postgres/shared-compose.yaml netbird/client.compose.yaml renovate/renovate-compose.yaml; do + touch "$scratch/repo/$file" +done +git -C "$scratch/repo" add . +# shellcheck source=../workflows/compose-lint.sh +source "$repo/.gitea/workflows/compose-lint.sh" +actual="$(cd "$scratch/repo" && compose_files)" +expected=$'app/compose.yaml\nnetbird/client.compose.yaml\npostgres/shared-compose.yaml\nrenovate/renovate-compose.yaml' +[ "$actual" = "$expected" ] || { echo 'Compose discovery missed a file' >&2; exit 1; } + +cat >"$scratch/compose.yaml" <<'YAML' +services: + example: + image: busybox:1.37.0 + environment: + REQUIRED: ${HOMELAB_TEST_REQUIRED:?required for this regression} +YAML +unset HOMELAB_TEST_REQUIRED +if validate_compose_file "$scratch/compose.yaml" >"$scratch/config.log" 2>&1; then + echo 'Full Compose validation accepted a missing variable' >&2 + exit 1 +fi +grep -q 'required for this regression' "$scratch/config.log" +HOMELAB_TEST_REQUIRED=present validate_compose_file "$scratch/compose.yaml" + +cat >"$scratch/resources.json" <<'JSON' +{"kind":"List","items":[ + {"kind":"Deployment","metadata":{"namespace":"app"},"spec":{"template":{"spec":{ + "containers":[{"envFrom":[{"secretRef":{"name":"credentials"}},{"secretRef":{"name":"optional","optional":true}}],"env":[{"valueFrom":{"secretKeyRef":{"name":"credentials","key":"password"}}}]}], + "initContainers":[{"envFrom":[{"secretRef":{"name":"init"}}]}], + "imagePullSecrets":[{"name":"registry"}], + "volumes":[{"secret":{"secretName":"mounted"}},{"projected":{"sources":[{"secret":{"name":"projected"}},{"secret":{"name":"optional-projected","optional":true}}]}}] + }}}}, + {"kind":"CronJob","metadata":{},"spec":{"jobTemplate":{"spec":{"template":{"spec":{"containers":[{"envFrom":[{"secretRef":{"name":"cron"}}]}]}}}}}}, + {"kind":"IngressRoute","metadata":{"namespace":"app"},"spec":{"tls":{"secretName":"controller-issued-tls"}}} +]} +JSON +actual="$(jq -r -f "$repo/.gitea/workflows/secret-references.jq" "$scratch/resources.json" | sort)" +expected=$'app credentials\napp init\napp mounted\napp projected\napp registry\ndefault cron' +[ "$actual" = "$expected" ] || { echo "Unexpected Secret references: $actual" >&2; exit 1; } + +REPO="$repo" +# shellcheck source=../workflows/deploy-lib.sh +source "$repo/.gitea/workflows/deploy-lib.sh" +K8S_MANIFESTS=("$scratch/resources.json") +KUSTOMIZE_APPS=() +# No live cluster access. Reject credentials in app even if they exist elsewhere. +kubectl() { + case "$1" in + create) cat "$scratch/resources.json" ;; + get) + if [ "$3" = credentials ] && [ "$5" = app ]; then + return 1 + fi + return 0 + ;; + *) echo "Unexpected kubectl invocation: $*" >&2; return 1 ;; + esac +} +if check_referenced_secrets >"$scratch/secrets.log"; then + echo 'Namespace-scoped Secret check accepted a missing Secret' >&2 + exit 1 +fi +grep -q 'MISSING OR UNREADABLE: app/credentials' "$scratch/secrets.log" +# API/rendering errors must not produce an empty reference list and pass. +kubectl() { return 1; } +if ! skip_uninstalled_vmagent_crd "$REPO/prometheus-stack/k8s/vmagent.yaml"; then + echo 'VMAgent preflight did not skip an uninstalled CRD' >&2 + exit 1 +fi +kubectl() { return 0; } +if skip_uninstalled_vmagent_crd "$REPO/prometheus-stack/k8s/vmagent.yaml"; then + echo 'VMAgent preflight skipped an installed CRD' >&2 + exit 1 +fi +if skip_uninstalled_vmagent_crd "$REPO/prometheus-stack/k8s/victoria.yaml"; then + echo 'VMAgent preflight skipped an unrelated manifest' >&2 + exit 1 +fi +kubectl() { return 1; } +if check_referenced_secrets >"$scratch/secrets.log"; then + echo 'Secret check accepted a failed manifest render' >&2 + exit 1 +fi + +# New declared namespaces defer only their own resources during preflight. +render_selected_resources() { + cat <<'JSON' +{"apiVersion":"v1","kind":"List","items":[ + {"apiVersion":"v1","kind":"Namespace","metadata":{"name":"new"}}, + {"apiVersion":"v1","kind":"ConfigMap","metadata":{"name":"new-config","namespace":"new"}}, + {"apiVersion":"v1","kind":"ConfigMap","metadata":{"name":"existing-config","namespace":"default"}} +]} +JSON +} +kubectl() { + case "$1" in + get) printf '%s\n' '{"items":[{"metadata":{"name":"default"}}]}' ;; + apply) cat >"$scratch/server-input.json" ;; + *) return 1 ;; + esac +} +validate_server_resources true +jq -e '.items | length == 2 and all(.metadata.name != "new-config")' "$scratch/server-input.json" >/dev/null +if validate_server_resources false 2>"$scratch/deferred.log"; then + echo 'Post-namespace validation accepted a missing namespace' >&2 + exit 1 +fi +kubectl() { + case "$1" in + get) printf '%s\n' '{"items":[{"metadata":{"name":"default"}},{"metadata":{"name":"new"}}]}' ;; + apply) cat >"$scratch/server-input.json" ;; + *) return 1 ;; + esac +} +validate_server_resources false +jq -e '.items | length == 3' "$scratch/server-input.json" >/dev/null +render_selected_resources() { + printf '%s\n' '{"items":[{"kind":"ConfigMap","metadata":{"name":"bad","namespace":"undeclared"}}]}' +} +if validate_server_resources true 2>"$scratch/undeclared.log"; then + echo 'Preflight accepted an undeclared missing namespace' >&2 + exit 1 +fi +# Count services, not characters in the newline-separated service names. +compose() { + case "$*" in + *'config --format json') printf '%s\n' '{"services":{"headscale":{},"headplane":{},"web":{},"init":{"restart":"no"}}}' ;; + *'ps --status running --services') printf '%s\n' headscale headplane web ;; + *) return 1 ;; + esac +} +verify_compose_stack example.yaml >"$scratch/compose-count.log" +grep -qF 'all 3 service(s) running' "$scratch/compose-count.log" +compose() { + case "$*" in + *'config --format json') printf '%s\n' '{"services":{"headscale":{},"headplane":{},"web":{}}}' ;; + *'ps --status running --services') printf '%s\n' headscale headplane ;; + *) return 0 ;; + esac +} +if verify_compose_stack example.yaml >"$scratch/compose-missing.log"; then + echo 'Compose verification accepted a missing service' >&2 + exit 1 +fi +grep -qF 'NOT RUNNING: web' "$scratch/compose-missing.log" +printf '%s\n' 'Deploy validation regressions passed.' diff --git a/.gitea/workflows/ci.yaml b/.gitea/workflows/ci.yaml index 2480c13..75ddd15 100644 --- a/.gitea/workflows/ci.yaml +++ b/.gitea/workflows/ci.yaml @@ -1,38 +1,25 @@ name: ci - -on: +"on": push: branches: - - "**" - pull_request: - workflow_dispatch: - -# Every job here is checkout plus local tools. The token needs to read the tree -# and nothing else, and saying so keeps a future step that reaches for the API -# from quietly holding a token that can write to the repository. + - main + pull_request: null + workflow_dispatch: null permissions: contents: read - + actions: read concurrency: group: ci-${{ github.ref }} cancel-in-progress: ${{ github.ref != 'refs/heads/main' }} - -env: - REGISTRY: gcr.forust.xyz - jobs: - lint-compose: - runs-on: [self-hosted, linux, arch, homelab] - timeout-minutes: 10 + compose: + name: Compose + runs-on: ${{ github.ref == 'refs/heads/main' && 'homelab' || 'homelab-pr' }} + timeout-minutes: 15 steps: - name: Checkout repository - uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4 - - # Structure check for every committed Compose file, active or not. - # Interpolation, env-file and bind-mount resolution are all switched off, - # because inactive stacks have no .env here and would only fail on their - # ${VAR:?} guards. Active stacks get the full check with interpolation in - # the deploy workflow, where the real .env files live. + uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 + id: source - name: Validate Compose files shell: bash run: | @@ -61,35 +48,77 @@ jobs: exit 1 fi echo "checked ${#files[@]} Compose file(s)" - - lint-actionlint: - runs-on: [self-hosted, linux, arch, homelab] - timeout-minutes: 10 + id: check + - name: Write the job result + if: always() + env: + SUMMARY_CHECK: Compose + SUMMARY_RESULT: ${{ job.status }} + SUMMARY_FAILED_STEP: + ${{ steps.check.conclusion == 'failure' && 'Check or image build' || steps.source.conclusion == 'failure' + && 'Source checkout' || '' }} + shell: bash + run: | + if [ -f .gitea/workflows/release.py ]; then + python3 .gitea/workflows/release.py check-summary + elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then + printf '## %s\n\n- Result: **%s**\n- Failed step: %s\n' "$SUMMARY_CHECK" "$SUMMARY_RESULT" "$SUMMARY_FAILED_STEP" >>"$GITHUB_STEP_SUMMARY" || true + fi + workflows: + name: Workflows + runs-on: ${{ github.ref == 'refs/heads/main' && 'homelab' || 'homelab-pr' }} + timeout-minutes: 15 steps: - name: Checkout repository - uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4 - + uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 + id: source + - name: Prepare pinned tools + shell: bash + run: | + set -euo pipefail + tools_dir="$(bash .gitea/workflows/install-ci-tools.sh actionlint shellcheck)" + echo "$tools_dir" >> "$GITHUB_PATH" + id: tools - name: Lint Gitea Actions workflows with actionlint shell: bash run: | set -euo pipefail - tools_dir="$(bash .gitea/workflows/install-ci-tools.sh actionlint)" - export PATH="$tools_dir:$PATH" actionlint -config-file .gitea/actionlint.yaml -color .gitea/workflows/*.yaml - - lint-shellcheck: - runs-on: [self-hosted, linux, arch, homelab] - timeout-minutes: 10 + id: check + - name: Write the job result + if: always() + env: + SUMMARY_CHECK: Workflows + SUMMARY_RESULT: ${{ job.status }} + SUMMARY_FAILED_STEP: + ${{ steps.check.conclusion == 'failure' && 'Check or image build' || steps.tools.conclusion == 'failure' + && 'Tool setup' || steps.source.conclusion == 'failure' && 'Source checkout' || '' }} + shell: bash + run: | + if [ -f .gitea/workflows/release.py ]; then + python3 .gitea/workflows/release.py check-summary + elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then + printf '## %s\n\n- Result: **%s**\n- Failed step: %s\n' "$SUMMARY_CHECK" "$SUMMARY_RESULT" "$SUMMARY_FAILED_STEP" >>"$GITHUB_STEP_SUMMARY" || true + fi + shell: + name: Shell + runs-on: ${{ github.ref == 'refs/heads/main' && 'homelab' || 'homelab-pr' }} + timeout-minutes: 15 steps: - name: Checkout repository - uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4 - + uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 + id: source + - name: Prepare pinned tools + shell: bash + run: | + set -euo pipefail + tools_dir="$(bash .gitea/workflows/install-ci-tools.sh shellcheck jq)" + echo "$tools_dir" >> "$GITHUB_PATH" + id: tools - name: Lint shell scripts with ShellCheck shell: bash run: | set -euo pipefail - tools_dir="$(bash .gitea/workflows/install-ci-tools.sh shellcheck)" - export PATH="$tools_dir:$PATH" mapfile -t scripts < <( git ls-files '*.sh' ':(glob)**/*.bash' ) @@ -98,20 +127,42 @@ jobs: exit 0 fi shellcheck --external-sources --source-path=SCRIPTDIR --severity=style "${scripts[@]}" - - lint-prettier: - runs-on: [self-hosted, linux, arch, homelab] - timeout-minutes: 10 + bash .gitea/tests/deploy-validation.sh + id: check + - name: Write the job result + if: always() + env: + SUMMARY_CHECK: Shell + SUMMARY_RESULT: ${{ job.status }} + SUMMARY_FAILED_STEP: + ${{ steps.check.conclusion == 'failure' && 'Check or image build' || steps.tools.conclusion == 'failure' + && 'Tool setup' || steps.source.conclusion == 'failure' && 'Source checkout' || '' }} + shell: bash + run: | + if [ -f .gitea/workflows/release.py ]; then + python3 .gitea/workflows/release.py check-summary + elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then + printf '## %s\n\n- Result: **%s**\n- Failed step: %s\n' "$SUMMARY_CHECK" "$SUMMARY_RESULT" "$SUMMARY_FAILED_STEP" >>"$GITHUB_STEP_SUMMARY" || true + fi + formatting: + name: Formatting + runs-on: ${{ github.ref == 'refs/heads/main' && 'homelab' || 'homelab-pr' }} + timeout-minutes: 15 steps: - name: Checkout repository - uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4 - - - name: Check formatting with Prettier + uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 + id: source + - name: Prepare pinned tools shell: bash run: | set -euo pipefail tools_dir="$(bash .gitea/workflows/install-ci-tools.sh prettier)" - export PATH="$tools_dir:$PATH" + echo "$tools_dir" >> "$GITHUB_PATH" + id: tools + - name: Check formatting with Prettier + shell: bash + run: | + set -euo pipefail mapfile -t prettier_files < <( git ls-files \ @@ -125,36 +176,79 @@ jobs: fi prettier --check --ignore-unknown "${prettier_files[@]}" - - lint-ruff: - runs-on: [self-hosted, linux, arch, homelab] - timeout-minutes: 10 + id: check + - name: Write the job result + if: always() + env: + SUMMARY_CHECK: Formatting + SUMMARY_RESULT: ${{ job.status }} + SUMMARY_FAILED_STEP: + ${{ steps.check.conclusion == 'failure' && 'Check or image build' || steps.tools.conclusion == 'failure' + && 'Tool setup' || steps.source.conclusion == 'failure' && 'Source checkout' || '' }} + shell: bash + run: | + if [ -f .gitea/workflows/release.py ]; then + python3 .gitea/workflows/release.py check-summary + elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then + printf '## %s\n\n- Result: **%s**\n- Failed step: %s\n' "$SUMMARY_CHECK" "$SUMMARY_RESULT" "$SUMMARY_FAILED_STEP" >>"$GITHUB_STEP_SUMMARY" || true + fi + python: + name: Python and tests + runs-on: ${{ github.ref == 'refs/heads/main' && 'homelab' || 'homelab-pr' }} + timeout-minutes: 15 steps: - name: Checkout repository - uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4 - + uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 + id: source + - name: Prepare pinned tools + shell: bash + run: | + set -euo pipefail + tools_dir="$(bash .gitea/workflows/install-ci-tools.sh ruff jq)" + echo "$tools_dir" >> "$GITHUB_PATH" + id: tools - name: Lint and format-check Python with Ruff shell: bash run: | set -euo pipefail - tools_dir="$(bash .gitea/workflows/install-ci-tools.sh ruff)" - export PATH="$tools_dir:$PATH" - ruff check . - ruff format --check . - - lint-yaml: - runs-on: [self-hosted, linux, arch, homelab] - timeout-minutes: 10 + ruff check . .gitea/workflows + ruff format --check . .gitea/workflows + python3 -m unittest discover -s tests -v + id: check + - name: Write the job result + if: always() + env: + SUMMARY_CHECK: Python and tests + SUMMARY_RESULT: ${{ job.status }} + SUMMARY_FAILED_STEP: + ${{ steps.check.conclusion == 'failure' && 'Check or image build' || steps.tools.conclusion == 'failure' + && 'Tool setup' || steps.source.conclusion == 'failure' && 'Source checkout' || '' }} + shell: bash + run: | + if [ -f .gitea/workflows/release.py ]; then + python3 .gitea/workflows/release.py check-summary + elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then + printf '## %s\n\n- Result: **%s**\n- Failed step: %s\n' "$SUMMARY_CHECK" "$SUMMARY_RESULT" "$SUMMARY_FAILED_STEP" >>"$GITHUB_STEP_SUMMARY" || true + fi + yaml: + name: YAML + runs-on: ${{ github.ref == 'refs/heads/main' && 'homelab' || 'homelab-pr' }} + timeout-minutes: 15 steps: - name: Checkout repository - uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4 - - - name: Lint YAML syntax + uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 + id: source + - name: Prepare pinned tools shell: bash run: | set -euo pipefail tools_dir="$(bash .gitea/workflows/install-ci-tools.sh yamllint)" - export PATH="$tools_dir:$PATH" + echo "$tools_dir" >> "$GITHUB_PATH" + id: tools + - name: Lint YAML syntax + shell: bash + run: | + set -euo pipefail mapfile -t yaml_files < <( git ls-files '*.yaml' '*.yml' \ @@ -168,20 +262,41 @@ jobs: fi yamllint -c .yamllint "${yaml_files[@]}" - - lint-dockerfiles: - runs-on: [self-hosted, linux, arch, homelab] - timeout-minutes: 10 + id: check + - name: Write the job result + if: always() + env: + SUMMARY_CHECK: YAML + SUMMARY_RESULT: ${{ job.status }} + SUMMARY_FAILED_STEP: + ${{ steps.check.conclusion == 'failure' && 'Check or image build' || steps.tools.conclusion == 'failure' + && 'Tool setup' || steps.source.conclusion == 'failure' && 'Source checkout' || '' }} + shell: bash + run: | + if [ -f .gitea/workflows/release.py ]; then + python3 .gitea/workflows/release.py check-summary + elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then + printf '## %s\n\n- Result: **%s**\n- Failed step: %s\n' "$SUMMARY_CHECK" "$SUMMARY_RESULT" "$SUMMARY_FAILED_STEP" >>"$GITHUB_STEP_SUMMARY" || true + fi + dockerfiles: + name: Dockerfiles + runs-on: ${{ github.ref == 'refs/heads/main' && 'homelab' || 'homelab-pr' }} + timeout-minutes: 15 steps: - name: Checkout repository - uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4 - - - name: Lint Dockerfiles + uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 + id: source + - name: Prepare pinned tools shell: bash run: | set -euo pipefail tools_dir="$(bash .gitea/workflows/install-ci-tools.sh hadolint)" - export PATH="$tools_dir:$PATH" + echo "$tools_dir" >> "$GITHUB_PATH" + id: tools + - name: Lint Dockerfiles + shell: bash + run: | + set -euo pipefail mapfile -t dockerfiles < <( git ls-files ':(glob)**/Dockerfile' ':(glob)**/Dockerfile.*' @@ -193,20 +308,41 @@ jobs: fi hadolint -c .hadolint.yaml "${dockerfiles[@]}" - - validate: - runs-on: [self-hosted, linux, arch, homelab] - timeout-minutes: 20 + id: check + - name: Write the job result + if: always() + env: + SUMMARY_CHECK: Dockerfiles + SUMMARY_RESULT: ${{ job.status }} + SUMMARY_FAILED_STEP: + ${{ steps.check.conclusion == 'failure' && 'Check or image build' || steps.tools.conclusion == 'failure' + && 'Tool setup' || steps.source.conclusion == 'failure' && 'Source checkout' || '' }} + shell: bash + run: | + if [ -f .gitea/workflows/release.py ]; then + python3 .gitea/workflows/release.py check-summary + elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then + printf '## %s\n\n- Result: **%s**\n- Failed step: %s\n' "$SUMMARY_CHECK" "$SUMMARY_RESULT" "$SUMMARY_FAILED_STEP" >>"$GITHUB_STEP_SUMMARY" || true + fi + kubernetes: + name: Kubernetes + runs-on: ${{ github.ref == 'refs/heads/main' && 'homelab' || 'homelab-pr' }} + timeout-minutes: 15 steps: - name: Checkout repository - uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4 - - - name: Validate Kubernetes manifests against JSON schemas + uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 + id: source + - name: Prepare pinned tools shell: bash run: | set -euo pipefail tools_dir="$(bash .gitea/workflows/install-ci-tools.sh kubeconform)" - export PATH="$tools_dir:$PATH" + echo "$tools_dir" >> "$GITHUB_PATH" + id: tools + - name: Validate Kubernetes manifests against JSON schemas + shell: bash + run: | + set -euo pipefail mapfile -t manifests < <( git ls-files ':(glob)**/k8s/**/*.yaml' ':(glob)**/k8s/**/*.yml' \ @@ -223,327 +359,162 @@ jobs: -ignore-missing-schemas \ -summary \ "${manifests[@]}" - - # kubeconform has no schemas for CRDs, so every IngressRoute, Certificate, - # PrometheusRule, Middleware, ServersTransport and ServiceMonitor is silently - # skipped above. The live API server knows the real CRD schemas (and runs the - # cert-manager / Traefik admission webhooks), so validate there too. - # - # Only services marked with a k8s/active marker are checked: server-side - # dry-run needs the target namespace to exist, and inactive services are not - # deployed. Services being enabled for the first time are still covered by - # the JSON-schema pass above. - # - # Main pushes only. `--dry-run=server` persists nothing, but it does execute - # the admission webhooks of the production API server, so anyone able to open - # a pull request would be able to run arbitrary manifest content through - # cert-manager and Traefik. A pull request has nothing to gain from it either: - # only main is ever deployed, and this job runs to completion before the - # deploy workflow is allowed to start, so a bad CRD is still caught before - # anything reaches the cluster -- just on the push rather than on the PR. - - name: Note the server-side check is not running here - if: github.event_name == 'pull_request' || github.ref != 'refs/heads/main' + id: check + - name: Write the job result + if: always() + env: + SUMMARY_CHECK: Kubernetes + SUMMARY_RESULT: ${{ job.status }} + SUMMARY_FAILED_STEP: + ${{ steps.check.conclusion == 'failure' && 'Check or image build' || steps.tools.conclusion == 'failure' + && 'Tool setup' || steps.source.conclusion == 'failure' && 'Source checkout' || '' }} shell: bash run: | - echo "::notice::Skipping the server-side dry-run. It executes the cert-manager and" \ - "Traefik admission webhooks against the production API server, so it is limited" \ - "to pushes to main. CRDs are still schema-checked by kubeconform above, and the" \ - "server-side pass still runs on main before the deploy." - - - name: Validate active manifests against the live API server - if: github.event_name != 'pull_request' && github.ref == 'refs/heads/main' - shell: bash - run: | - set -euo pipefail - - if ! kubectl get --raw='/readyz' --request-timeout=10s >/dev/null 2>&1; then - echo "::warning::Cluster unreachable — skipped server-side validation of CRDs (IngressRoute, Certificate, PrometheusRule). Review manifest changes manually." - exit 0 + if [ -f .gitea/workflows/release.py ]; then + python3 .gitea/workflows/release.py check-summary + elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then + printf '## %s\n\n- Result: **%s**\n- Failed step: %s\n' "$SUMMARY_CHECK" "$SUMMARY_RESULT" "$SUMMARY_FAILED_STEP" >>"$GITHUB_STEP_SUMMARY" || true fi - - mapfile -t k8s_dirs < <( - git ls-files '*.yaml' '*.yml' \ - | grep -E '(^|/)k8s/' \ - | sed -E 's#((^|.*/)k8s)/.*#\1#' \ - | sort -u - ) - - manifests=() - kustomize_apps=() - for dir in "${k8s_dirs[@]}"; do - if [ ! -f "${dir}/active" ]; then - echo "skip (no k8s/active): ${dir}" - continue - fi - if [ -f "${dir}/overlays/prod/kustomization.yaml" ]; then - kustomize_apps+=("${dir}/overlays/prod") - elif [ -f "${dir}/base/kustomization.yaml" ]; then - kustomize_apps+=("${dir}/base") - else - while IFS= read -r f; do - [ -n "$f" ] && manifests+=("$f") - done < <( - git ls-files "${dir}/*.yaml" "${dir}/*.yml" \ - | grep -Ev '(^|/)(kustomization\.ya?ml|.*\.example\.ya?ml|.*values\.ya?ml|patch-.*\.ya?ml)$' - ) - fi - done - - echo "server-side dry-run: ${#manifests[@]} manifests, ${#kustomize_apps[@]} kustomize apps" - failed=0 - for m in ${manifests[@]+"${manifests[@]}"}; do - if ! out="$(kubectl apply --dry-run=server -f "$m" 2>&1)"; then - failed=1 - echo "::error file=${m}::$(printf '%s' "$out" | head -1)" - fi - done - for k in ${kustomize_apps[@]+"${kustomize_apps[@]}"}; do - if ! out="$(kubectl apply -k "$k" --dry-run=server 2>&1)"; then - failed=1 - echo "::error file=${k}::$(printf '%s' "$out" | head -1)" - fi - done - - if [ "$failed" -ne 0 ]; then - echo "Server-side validation failed. The API server (or an admission webhook) rejected these manifests." - exit 1 - fi - echo "server-side dry-run: all active manifests accepted by the API server" - - build: - needs: - # The panel's scan-deps/test-backend/test-frontend jobs gated here until - # userbot moved to its own repo; upstream's code is upstream's gate now. - # The rule is unchanged: publishing and passing the checks are the same - # gate, so a commit that fails any of these still cannot move :prod. - [lint-actionlint, lint-shellcheck, lint-compose, lint-prettier, lint-ruff, lint-yaml, lint-dockerfiles, validate] - if: github.event_name != 'pull_request' && (github.ref_name == 'main' || github.ref_name == 'dev') && !startsWith(github.ref_name, 'renovate/') - runs-on: [self-hosted, linux, arch, homelab] - timeout-minutes: 60 + image-plan: + needs: [compose, workflows, shell, formatting, python, yaml, dockerfiles, kubernetes] + if: github.event_name != 'pull_request' && github.ref == 'refs/heads/main' + runs-on: homelab + timeout-minutes: 10 outputs: - services: ${{ steps.services.outputs.services }} + matrix: ${{ steps.plan.outputs.matrix }} steps: - name: Checkout repository - uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4 + id: source + uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 with: fetch-depth: 0 + - name: Detect build inputs against successful CI + id: plan + env: + GITEA_TOKEN: ${{ github.token }} + run: python3 .gitea/workflows/release.py prepare --output build-plan.json + - name: Store the image plan + id: artifact + uses: actions/upload-artifact@ea165f8d65b6e75b540449e92b4886f43607fa02 # v4.6.2 + with: + name: build-plan + path: build-plan.json + if-no-files-found: error + retention-days: 30 - - name: Detect changed docker-built services - id: services + - name: Write the plan result + if: always() + env: + SUMMARY_CHECK: Image plan + SUMMARY_RESULT: ${{ job.status }} + SUMMARY_FAILED_STEP: >- + ${{ steps.plan.conclusion == 'failure' && 'Build input detection' || + steps.artifact.conclusion == 'failure' && 'Plan upload' || + steps.source.conclusion == 'failure' && 'Source checkout' || '' }} shell: bash run: | - set -euo pipefail - base="${{ github.event.before }}" - if [ -z "$base" ] || [ "$base" = "0000000000000000000000000000000000000000" ]; then - base="$(git rev-list --max-parents=0 HEAD)" + if [ -f .gitea/workflows/release.py ]; then + python3 .gitea/workflows/release.py check-summary + elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then + printf '## Image plan\n\nResult: %s\n' "$SUMMARY_RESULT" >>"$GITHUB_STEP_SUMMARY" || true fi - # A failed diff used to leave changed_files empty, which reads exactly - # like "nothing to build": the job went green having built nothing and - # the tag never moved. The status is checked, not assumed. - if ! changed="$(git diff --name-only "$base" "${GITHUB_SHA}")"; then - echo "::error::cannot diff ${base}..${GITHUB_SHA}" - exit 1 - fi - mapfile -t changed_files <<<"$changed" - - services=() - - add_service() { - local name="$1" - local seen=0 - for existing in "${services[@]}"; do - if [ "$existing" = "$name" ]; then - seen=1 - break - fi - done - if [ "$seen" -eq 0 ]; then - services+=("$name") - fi - } - - for file in "${changed_files[@]}"; do - case "$file" in - errorpages/*) - add_service errorpages - ;; - homepages/*) - add_service homepages - ;; - edu_master/phpsessid-bot/*|edu_master/webinar-checker/*|edu_master/compose.yaml) - add_service edu_master - ;; - esac - done - - if [ "${#services[@]}" -eq 0 ]; then - echo "No docker-built services changed." - echo "services=" >> "$GITHUB_OUTPUT" - exit 0 - fi - - printf '%s\n' "${services[@]}" | tee /tmp/services.txt - echo "services=$(paste -sd, /tmp/services.txt)" >> "$GITHUB_OUTPUT" - - - name: Log in to registry - # The pin step below also writes (manifest PUTs), and it runs on every - # main push — including manifest-only ones where services is empty. A - # stale persistent login on the old runner used to mask this; a clean - # runner pushes anonymously and gets 401. - if: steps.services.outputs.services != '' || github.ref_name == 'main' + images: + name: Image (${{ matrix.name }}) + needs: [image-plan] + if: needs.image-plan.result == 'success' + runs-on: homelab + timeout-minutes: 60 + strategy: + max-parallel: 1 + fail-fast: false + matrix: ${{ fromJSON(needs.image-plan.outputs.matrix || '{"include":[{"name":"inactive"}]}') }} + steps: + - name: Checkout repository + id: source + uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 + - name: Download the checked image plan + id: inputs + uses: actions/download-artifact@d3f86a106a0bac45b974a628896c90dbdf5c8093 # v4.3.0 + with: + name: build-plan + - name: Build or reuse this image + id: check + env: + IMAGE_NAME: ${{ matrix.name }} + REGISTRY_USERNAME: ${{ secrets.REGISTRY_USERNAME }} + REGISTRY_PASSWORD: ${{ secrets.REGISTRY_PASSWORD }} + run: python3 .gitea/workflows/release.py image --image "$IMAGE_NAME" --output image.json + - name: Store the image result + id: artifact + uses: actions/upload-artifact@ea165f8d65b6e75b540449e92b4886f43607fa02 # v4.6.2 + with: + name: image-${{ matrix.name }} + path: image.json + if-no-files-found: error + retention-days: 30 + - name: Write the job result + if: always() + env: + SUMMARY_CHECK: Image (${{ matrix.name }}) + SUMMARY_RESULT: ${{ job.status }} + SUMMARY_FAILED_STEP: >- + ${{ steps.check.conclusion == 'failure' && 'Build or tag images' || + steps.artifact.conclusion == 'failure' && 'Artifact upload' || + steps.inputs.conclusion == 'failure' && 'Artifact download' || + steps.source.conclusion == 'failure' && 'Source checkout' || '' }} shell: bash - # Through env, not by substitution into the script. A secret written - # into a run: block is pasted into the shell source before bash parses - # it, so a password containing a quote, a backtick or $(...) becomes - # code that runs. Masking the value in the log does not prevent that. + run: | + if [ -f .gitea/workflows/release.py ]; then + python3 .gitea/workflows/release.py check-summary + elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then + printf '## %s\n\n- Result: **%s**\n- Failed step: %s\n' "$SUMMARY_CHECK" "$SUMMARY_RESULT" "$SUMMARY_FAILED_STEP" >>"$GITHUB_STEP_SUMMARY" || true + fi + + # Retain the build job name required by the immutable release deployment gate. + build: + needs: [image-plan, images] + runs-on: homelab + timeout-minutes: 15 + steps: + - name: Checkout repository + id: source + uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 + - name: Download all image results + id: inputs + uses: actions/download-artifact@d3f86a106a0bac45b974a628896c90dbdf5c8093 # v4.3.0 + with: + path: artifacts + - name: Pin SHA tags and write the complete release + id: check env: REGISTRY_USERNAME: ${{ secrets.REGISTRY_USERNAME }} REGISTRY_PASSWORD: ${{ secrets.REGISTRY_PASSWORD }} - run: | - set -euo pipefail - printf '%s' "$REGISTRY_PASSWORD" | docker login "${REGISTRY}" \ - -u "$REGISTRY_USERNAME" \ - --password-stdin - - - name: Build and push changed images - if: steps.services.outputs.services != '' + run: >- + python3 .gitea/workflows/release.py finalize + --plan artifacts/build-plan/build-plan.json + - name: Store commit release + id: artifact + uses: actions/upload-artifact@ea165f8d65b6e75b540449e92b4886f43607fa02 # v4.6.2 + with: + name: release-${{ github.sha }} + path: release.json + if-no-files-found: error + retention-days: 30 + - name: Write the job result + if: always() + env: + SUMMARY_CHECK: Image release and SHA tags + SUMMARY_RESULT: ${{ job.status }} + SUMMARY_FAILED_STEP: >- + ${{ steps.check.conclusion == 'failure' && 'Build or tag images' || + steps.artifact.conclusion == 'failure' && 'Artifact upload' || + steps.inputs.conclusion == 'failure' && 'Artifact download' || + steps.source.conclusion == 'failure' && 'Source checkout' || '' }} shell: bash run: | - # This step was the one run: block in the workflow without it, and it - # is the one that cannot afford it: a docker push that failed partway - # through the loop used to be followed by more pushes, the loop's exit - # status came from the last one, and the job went green with half the - # images missing from the registry. - set -euo pipefail - IFS=, read -r -a services <<< "${{ steps.services.outputs.services }}" - - # Tags for this push. The commit-pinned name is the point of this - # step: the deploy resolves it in preference to :prod, so a deploy - # that sat in the queue behind a later push still gets the build of - # the commit CI validated, instead of whatever :prod points at by the - # time it runs. See render_pinned in deploy-lib.sh. - commit_tag="" - if [ "${GITHUB_REF_NAME}" = "main" ]; then - commit_tag="sha-${GITHUB_SHA:0:12}" + if [ -f .gitea/workflows/release.py ]; then + python3 .gitea/workflows/release.py check-summary + elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then + printf '## %s\n\n- Result: **%s**\n- Failed step: %s\n' "$SUMMARY_CHECK" "$SUMMARY_RESULT" "$SUMMARY_FAILED_STEP" >>"$GITHUB_STEP_SUMMARY" || true fi - - set_tags() { - tags=() - case "${GITHUB_REF_NAME}" in - main) tags+=("main" "prod") ;; - dev) tags+=("dev") ;; - esac - if [ -n "$commit_tag" ]; then - tags+=("$commit_tag") - fi - } - - for service in "${services[@]}"; do - case "$service" in - errorpages) - image="${REGISTRY}/forust/error-pages" - set_tags - build_args=() - for tag in "${tags[@]}"; do - build_args+=(-t "${image}:${tag}") - done - docker build \ - --cache-from "type=registry,ref=${image}:buildcache" \ - --cache-to "type=registry,ref=${image}:buildcache,mode=max" \ - "${build_args[@]}" errorpages - for tag in "${tags[@]}"; do - docker push "${image}:${tag}" - done - ;; - homepages) - for variant in forust xdfnx; do - case "$variant" in - forust) - image="${REGISTRY}/forust/forust-homepage" - ;; - xdfnx) - image="${REGISTRY}/forust/xdfnx-homepage" - ;; - esac - set_tags - build_args=() - for tag in "${tags[@]}"; do - build_args+=(-t "${image}:${tag}") - done - docker build \ - --cache-from "type=registry,ref=${image}:buildcache" \ - --cache-to "type=registry,ref=${image}:buildcache,mode=max" \ - "${build_args[@]}" -f "homepages/Dockerfile.${variant}" homepages - for tag in "${tags[@]}"; do - docker push "${image}:${tag}" - done - done - ;; - edu_master) - for variant in session-keeper webinar-checker; do - case "$variant" in - session-keeper) - context="edu_master/phpsessid-bot" - image="${REGISTRY}/forust/session-keeper" - ;; - webinar-checker) - context="edu_master/webinar-checker" - image="${REGISTRY}/forust/webinar-checker" - ;; - esac - set_tags - build_args=() - for tag in "${tags[@]}"; do - build_args+=(-t "${image}:${tag}") - done - docker build \ - --cache-from "type=registry,ref=${image}:buildcache" \ - --cache-to "type=registry,ref=${image}:buildcache,mode=max" \ - "${build_args[@]}" "$context" - for tag in "${tags[@]}"; do - docker push "${image}:${tag}" - done - done - ;; - esac - done - - # Every image the tree names has to carry the commit-pinned name, not only - # the ones this push rebuilt. A push that touches nothing but manifests - # builds nothing, and its deploy would then find no commit-pinned tag to - # resolve and quietly fall back to the moving :prod - which is the whole - # failure the commit-pinned name exists to remove. - # - # Re-tagging copies the manifest list and transfers no layers, so pinning - # six images that already exist costs six registry writes. - # - # The list is derived from the tree rather than written out here, so an - # image added to a manifest is covered without a second place to update. - - name: Pin the commit name on the images this push did not rebuild - if: github.ref_name == 'main' - shell: bash - run: | - set -euo pipefail - commit_tag="sha-${GITHUB_SHA:0:12}" - mapfile -t repos < <( - git grep -hoE 'gcr\.forust\.xyz/forust/[A-Za-z0-9._-]+' -- '*.yaml' '*.yml' \ - | sort -u - ) - if [ "${#repos[@]}" -eq 0 ]; then - echo "No own images referenced by the tree." - exit 0 - fi - echo "pinning ${#repos[@]} image(s) to $commit_tag" - for repo in "${repos[@]}"; do - if docker buildx imagetools inspect "$repo:$commit_tag" >/dev/null 2>&1; then - echo " already built by this push: ${repo##*/}" - continue - fi - if ! docker buildx imagetools inspect "$repo:prod" >/dev/null 2>&1; then - echo " WARNING: ${repo##*/} has no :prod to pin and no build produced it" - continue - fi - docker buildx imagetools create --tag "$repo:$commit_tag" "$repo:prod" - echo " pinned ${repo##*/}" - done diff --git a/.gitea/workflows/compose-lint.sh b/.gitea/workflows/compose-lint.sh index c27e29f..30e79bb 100644 --- a/.gitea/workflows/compose-lint.sh +++ b/.gitea/workflows/compose-lint.sh @@ -21,8 +21,7 @@ # All committed Compose files, including the ones deploy never starts. compose_files() { git ls-files \ - '*/compose.yaml' '*/compose.yml' 'compose.yaml' 'compose.yml' \ - '*/docker-compose.yaml' '*/docker-compose.yml' + '*compose.yaml' '*compose.yml' } # Prints the flags that turn `docker compose config` into the general check. diff --git a/.gitea/workflows/compose-release.py b/.gitea/workflows/compose-release.py new file mode 100644 index 0000000..a12b778 --- /dev/null +++ b/.gitea/workflows/compose-release.py @@ -0,0 +1,164 @@ +#!/usr/bin/env python3 +"""Resolve Compose images without changing project names or local bind paths.""" + +import json +import os +import re +import subprocess +import sys +from pathlib import Path + + +def output(*args, **kwargs): + return subprocess.check_output(args, text=True, **kwargs).strip() # noqa: S603 + + +def resolve(reference): + if '@sha256:' in reference: + return reference + descriptor = json.loads( + output('docker', 'buildx', 'imagetools', 'inspect', reference, '--format', '{{json .Manifest}}') + ) + digest = descriptor['digest'] + if not re.fullmatch(r'sha256:[0-9a-f]{64}', digest): + raise ValueError(f'Invalid registry digest for {reference}') + # Strip tag only from the final path segment (registry ports are preserved). + repository = reference.rsplit('/', 1) + repository[-1] = repository[-1].split(':')[0] + return '/'.join(repository) + '@' + digest + + +def prepare(source_file): + config_repo = Path(os.environ['CONFIG_REPO']) + source_repo = Path(os.environ['REPO']) + directory = Path(os.environ['RUN_DIR']) + relative = source_file.relative_to(source_repo) + project_dir = config_repo / relative.parent + base = ['docker', 'compose', '--project-directory', str(project_dir), '-f', str(source_file)] + config = json.loads(output(*base, 'config', '--format', 'json', cwd=config_repo)) + project = config['name'] + previous_file = directory / 'previous.json' + previous = json.loads(previous_file.read_text()) if previous_file.exists() else {} + images_file = directory / 'compose-images.json' + locks = json.loads(images_file.read_text()) if images_file.exists() else previous.get('compose-images', {}) + release = json.loads((directory / 'release.json').read_text()) + state = Path(os.environ.get('HOMELAB_STATE', Path.home() / '.local/state/homelab-deploy')) + baseline = state / 'compose-configs' / f'{relative.parent.name}.json' + if not baseline.exists() and re.fullmatch(r'[0-9]+-[0-9]+', previous.get('run_id', '')): + baseline = state / 'runs' / previous['run_id'] / 'compose' / baseline.name + bootstrap = not baseline.exists() + if not bootstrap: + before = json.loads(baseline.read_text()) + else: + # Bootstrap from the persistent configuration, never from the new source. + persistent_file = config_repo / relative + if persistent_file.exists(): + before = json.loads( + output( + 'docker', + 'compose', + '--project-directory', + str(project_dir), + '-f', + str(persistent_file), + 'config', + '--format', + 'json', + cwd=config_repo, + ) + ) + elif output('docker', 'ps', '-aq', '--filter', f'label=com.docker.compose.project={project}'): + raise ValueError(f'{project}: no previous Compose configuration; restore it before deploy') + else: + before = {'name': project, 'services': {}} + if before['name'] != project: + raise ValueError('Compose project name changed; manual migration is required') + for service, settings in config['services'].items(): + reference = settings.get('image') + nextcloud_aio_master = project == 'nextcloud' and service == 'nextcloud-aio-mastercontainer' + if not reference or settings.get('build'): + raise ValueError(f'{project}/{service}: Compose deploy requires a published image') + image_repo = reference.split('@')[0].rsplit('/', 1) + image_repo[-1] = image_repo[-1].split(':')[0] + image_repo = '/'.join(image_repo) + # Nextcloud AIO validates the mastercontainer image reference and rejects + # a digest. Keep its configured tag so AIO can start and manage its stack. + if nextcloud_aio_master: + pinned = reference + elif image_repo in release['images']: + pinned = image_repo + '@' + release['images'][image_repo] + elif os.environ.get('REFRESH_IMAGES') != 'true' and reference in locks: + pinned = locks[reference] + else: + pinned = resolve(reference) + settings['image'] = pinned + locks[reference] = pinned + for service, settings in before['services'].items(): + reference = settings['image'] + image_repo = reference.split('@')[0].rsplit('/', 1) + image_repo[-1] = image_repo[-1].split(':')[0] + image_repo = '/'.join(image_repo) + nextcloud_aio_master = project == 'nextcloud' and service == 'nextcloud-aio-mastercontainer' + # Capture what is running, not the current value of its mutable tag. + ids = output( + 'docker', + 'ps', + '-aq', + '--filter', + f'label=com.docker.compose.project={project}', + '--filter', + f'label=com.docker.compose.service={service}', + ).splitlines() + actual = set() + if bootstrap and ids: + expected_hash = output( + 'docker', + 'compose', + '--project-directory', + str(project_dir), + '-f', + str(persistent_file), + 'config', + '--hash', + service, + cwd=config_repo, + ).split()[-1] + for container in ids: + running_hash = output( + 'docker', + 'inspect', + container, + '--format', + '{{ index .Config.Labels "com.docker.compose.config-hash" }}', + ) + if running_hash != expected_hash: + raise ValueError( + f'{project}/{service}: persistent config differs from running config; restore the previous config' + ) + for container in ids: + image_id = output('docker', 'inspect', container, '--format', '{{.Image}}') + digests = json.loads(output('docker', 'image', 'inspect', image_id, '--format', '{{json .RepoDigests}}')) + actual.add(next((d for d in digests or [] if d.split('@')[0] == image_repo), image_id)) + if len(actual) > 1: + raise ValueError(f'{project}/{service}: mixed running images, cannot capture one recovery config') + # AIO also rejects a digest in its recovery config. Preserve its tag in + # both deploy and recovery files. + if nextcloud_aio_master: + before['services'][service]['image'] = reference + else: + before['services'][service]['image'] = next(iter(actual)) if actual else reference + for name, data in (('compose', config), ('compose-before', before)): + folder = directory / name + folder.mkdir(mode=0o700, exist_ok=True) + destination = folder / f'{relative.parent.name}.json' + destination.write_text(json.dumps(data, indent=2) + '\n') + destination.chmod(0o600) + images_file.write_text(json.dumps(locks, indent=2) + '\n') + print(f'Compose {project}: images pinned; local paths preserved') + print( + f'Recovery: docker compose --project-directory {project_dir} -p {project} -f {directory}/compose-before/{relative.parent.name}.json up -d --pull never --remove-orphans' + ) + + +if __name__ == '__main__': + prepare(Path(sys.argv[1])) diff --git a/.gitea/workflows/deploy-controller.py b/.gitea/workflows/deploy-controller.py new file mode 100644 index 0000000..d56469a --- /dev/null +++ b/.gitea/workflows/deploy-controller.py @@ -0,0 +1,427 @@ +#!/usr/bin/env python3 +"""Durable workstation deployment controller. Install with setup-workstation.sh.""" + +import argparse +import contextlib +import fcntl +import importlib.util +import json +import math +import os +import re +import shutil +import subprocess +import sys +import time +from pathlib import Path + +STATE = Path(os.environ.get('HOMELAB_STATE', Path.home() / '.local/state/homelab-deploy')) +CONFIG_REPO = Path(os.environ.get('HOMELAB_REPO', '/srv/homelab')) +RUN_ID = re.compile(r'[0-9]+-[0-9]+') + + +def command(*args, **kwargs): + return subprocess.check_output(args, text=True, **kwargs).strip() # noqa: S603, S607 + + +def atomic_json(path, data): + temporary = path.with_suffix('.tmp') + temporary.write_text(json.dumps(data, indent=2) + '\n') + temporary.chmod(0o600) + temporary.replace(path) + + +@contextlib.contextmanager +def lock(name): + STATE.mkdir(mode=0o700, parents=True, exist_ok=True) + with (STATE / name).open('a') as stream: + fcntl.flock(stream, fcntl.LOCK_EX) + yield + + +def load_module(name, path): + spec = importlib.util.spec_from_file_location(name, path) + module = importlib.util.module_from_spec(spec) + spec.loader.exec_module(module) + return module + + +def run_directory(run_id): + if not RUN_ID.fullmatch(run_id): + raise ValueError('Run ID must be numeric workflow-id and attempt') + return STATE / 'runs' / run_id + + +def start(run_id): + payload = sys.stdin.buffer.read(256 * 1024 + 1) + if len(payload) > 256 * 1024: + raise ValueError('Deploy request exceeds 256 KiB') + request = json.loads(payload) + sha = request['release']['sha'] + if not re.fullmatch(r'[0-9a-f]{40}', sha) or request['mode'] not in ('changed', 'full', 'plan'): + raise ValueError('Invalid deploy SHA or mode') + if not isinstance(request['refresh_images'], bool): + raise ValueError('refresh_images must be boolean') + directory = run_directory(run_id) + with lock('prepare.lock'): + if (directory / 'request.json').exists(): + if json.loads((directory / 'request.json').read_text()) != request: + raise ValueError('Run ID already belongs to a different request') + else: + directory.mkdir(mode=0o700, parents=True, exist_ok=True) + command('git', '-C', str(CONFIG_REPO), 'fetch', '--quiet', 'origin', 'main') + command('git', '-C', str(CONFIG_REPO), 'merge-base', '--is-ancestor', sha, 'origin/main') + if not (directory / 'source').exists(): + command('git', '-C', str(CONFIG_REPO), 'worktree', 'add', '--detach', str(directory / 'source'), sha) + if command('git', '-C', str(directory / 'source'), 'rev-parse', 'HEAD') != sha: + raise ValueError('Prepared source does not match deploy SHA') + release_module = load_module('release', directory / 'source/.gitea/workflows/release.py') + release_module.validate_release(request['release'], sha) + atomic_json(directory / 'release.json', request['release']) + atomic_json(directory / 'request.json', request) + if not (directory / 'status.json').exists(): + atomic_json(directory / 'status.json', {'state': 'queued', 'stages': {}}) + # Starting an existing active or finished ID is idempotent; never re-apply it. + if json.loads((directory / 'status.json').read_text())['state'] == 'queued': + command('systemctl', '--user', 'start', '--no-block', f'homelab-deploy@{run_id}.service') + print(f'Accepted deploy {run_id} ({sha})') + + +def environment(directory): + request = json.loads((directory / 'request.json').read_text()) + return { + **os.environ, + 'REPO': str(directory / 'source'), + 'CONFIG_REPO': str(CONFIG_REPO), + 'RUN_DIR': str(directory), + 'DEPLOY_SHA': request['release']['sha'], + 'RELEASE_FILE': str(directory / 'release.json'), + 'DEPLOY_PLAN': str(directory / 'plan.json'), + 'DEPLOY_SNAPSHOT_DIR': str(directory / 'snapshot'), + 'REFRESH_IMAGES': str(request['refresh_images']).lower(), + 'ROLLOUT_PARALLELISM': '4', + } + + +def stage(directory, name, budget): + status = json.loads((directory / 'status.json').read_text()) + if name in status['stages'] and status['stages'][name].get('result') in ('success', 'failure'): + return status['stages'][name]['result'] == 'success' + started = time.time() + status['stages'][name] = {'result': 'running', 'started': started} + atomic_json(directory / 'status.json', status) + script = directory / 'source/.gitea/workflows/deploy-stage.sh' + with (directory / f'{name}.log').open('a') as log: + # timeout kills the whole stage process group, including children, before recovery. + result = subprocess.run( # noqa: S603, S607 + [ + shutil.which('timeout') or '/usr/bin/timeout', + '--signal=TERM', + '--kill-after=30s', + str(budget), + 'bash', + str(script), + name, + ], + env=environment(directory), + stdout=log, + stderr=subprocess.STDOUT, + check=False, + ).returncode + status = json.loads((directory / 'status.json').read_text()) + status['stages'][name].update( + result='success' if result == 0 else 'failure', exit_code=result, seconds=round(time.time() - started) + ) + atomic_json(directory / 'status.json', status) + return result == 0 + + +def make_plan(directory): + source = directory / 'source' + planner = load_module('deploy_plan', source / '.gitea/workflows/deploy-plan.py') + request = json.loads((directory / 'request.json').read_text()) + previous = json.loads((STATE / 'last-success.json').read_text()) if (STATE / 'last-success.json').exists() else None + # Helm 4 lists every release status by default and removed the --all flag. + helm = json.loads(command('helm', 'list', '-A', '-o', 'json')) + plan = planner.make_plan(source, CONFIG_REPO, request['release'], previous, request['mode'], helm) + if request['refresh_images']: + plan['selected']['compose'] = plan['active']['compose'] + atomic_json(directory / 'plan.json', plan) + if previous: + atomic_json(directory / 'previous.json', previous) + # Local config is deliberately separate from the immutable Git source. + return plan + + +def finish_success(directory, plan): + # Repeating finalization after a crash is safe while holding deploy.lock. + plan['run_id'] = directory.name + path = directory / 'compose-images.json' + previous = directory / 'previous.json' + plan['compose-images'] = ( + json.loads(path.read_text()) + if path.exists() + else json.loads(previous.read_text()).get('compose-images', {}) + if previous.exists() + else {} + ) + configs = STATE / 'compose-configs' + configs.mkdir(mode=0o700, exist_ok=True) + for config in (directory / 'compose').glob('*.json'): + atomic_json(configs / config.name, json.loads(config.read_text())) + atomic_json(STATE / 'last-success.json', plan) + status = json.loads((directory / 'status.json').read_text()) + status['state'] = 'success' + atomic_json(directory / 'status.json', status) + try: + retain_completed(directory) + except (OSError, subprocess.CalledProcessError) as error: + print(f'Retention deferred: {error}', flush=True) + + +def recover(directory, retry=False): + status = json.loads((directory / 'status.json').read_text()) + if status['state'] in ('success', 'planned'): + return + completed = ('doctor', 'validate', 'apply-k8s', 'apply-compose', 'verify-k8s', 'smoke') + if all(status['stages'].get(name, {}).get('result') == 'success' for name in completed): + finish_success(directory, json.loads((directory / 'plan.json').read_text())) + return + if retry: + for name in ('verify-k8s', 'smoke'): + if status['stages'].get(name, {}).get('result') == 'failure': + del status['stages'][name] + atomic_json(directory / 'status.json', status) + snapshot = directory / 'snapshot/current' + if snapshot.exists(): + stage(directory, 'verify-k8s', 7200) + stage(directory, 'smoke', 600) + status = json.loads((directory / 'status.json').read_text()) + status['state'] = 'failure' + atomic_json(directory / 'status.json', status) + + +def execute(run_id): + directory = run_directory(run_id) + with lock('deploy.lock'): + status = json.loads((directory / 'status.json').read_text()) + if status['state'] != 'queued': + return + # A crashed predecessor must be recovered before another apply begins. + for other in (STATE / 'runs').iterdir(): + if ( + other != directory + and (other / 'status.json').exists() + and json.loads((other / 'status.json').read_text())['state'] == 'running' + ): + raise ValueError(f'Interrupted deploy {other.name}; run recover first') + status['state'] = 'running' + atomic_json(directory / 'status.json', status) + phase = 'plan' + try: + plan = make_plan(directory) + print( + json.dumps({'selected': plan['selected'], 'helm': plan['helm'], 'manual_removals': plan['removed']}), + flush=True, + ) + phase = 'doctor' + if not stage(directory, 'doctor', 600): + raise RuntimeError('Preflight failed') + phase = 'validate' + if not stage(directory, 'validate', 1200): + raise RuntimeError('Validation failed') + if json.loads((directory / 'request.json').read_text())['mode'] == 'plan': + status = json.loads((directory / 'status.json').read_text()) + status['state'] = 'planned' + atomic_json(directory / 'status.json', status) + return + # Budget includes both rollout checks and rollback waves, plus API overhead. + phase = 'Recovery budget' + count = int( + command( + 'bash', + str(directory / 'source/.gitea/workflows/deploy-stage.sh'), + 'workload-count', + env=environment(directory), + ) + ) + verify_budget = max(600, 2 * math.ceil(count / 4) * 300 + 120) + if verify_budget > 7200: + raise ValueError('More than two hours of recovery required; split this deploy') + phase = 'apply-k8s' + k8s_ok = stage(directory, 'apply-k8s', 2700) + phase = 'apply-compose' + compose_ok = stage(directory, 'apply-compose', 1800) if k8s_ok else False + phase = 'verify-k8s' + verify_ok = stage(directory, 'verify-k8s', verify_budget) + phase = 'smoke' + smoke_ok = stage(directory, 'smoke', 600) + if not all((k8s_ok, compose_ok, verify_ok, smoke_ok)): + raise RuntimeError('Deploy failed; inspect stage logs and recovery report') + phase = 'Save the successful baseline' + finish_success(directory, plan) + except Exception as error: + status = json.loads((directory / 'status.json').read_text()) + status['failure_stage'] = next( + (name for name, result in status['stages'].items() if result.get('result') == 'failure'), phase + ) + atomic_json(directory / 'status.json', status) + with (directory / 'controller.log').open('a') as stream: + stream.write(f'{error}\n') + recover(directory) + raise + + +def retain_completed(current): + finished = [] + for directory in (STATE / 'runs').iterdir(): + status_file = directory / 'status.json' + if status_file.exists() and json.loads(status_file.read_text())['state'] in ('success', 'planned'): + finished.append(directory) + for directory in sorted(finished, key=lambda p: p.stat().st_mtime, reverse=True)[20:]: + if directory == current: + continue + command('git', '-C', str(CONFIG_REPO), 'worktree', 'remove', '--force', str(directory / 'source')) + shutil.rmtree(directory) + + +def follow(run_id, phase): + directory = run_directory(run_id) + groups = { + 'apply': ('doctor', 'validate', 'apply-k8s', 'apply-compose'), + 'verify': ('verify-k8s',), + 'smoke': ('smoke',), + } + names = groups[phase] + offsets = {} + while True: + status = json.loads((directory / 'status.json').read_text()) + for name in (*names, 'controller'): + path = directory / f'{name}.log' + if path.exists(): + with path.open() as stream: + stream.seek(offsets.get(name, 0)) + content = stream.read() + if content: + print(content, end='', flush=True) + offsets[name] = stream.tell() + stages = status['stages'] + if all(stages.get(name, {}).get('result') in ('success', 'failure') for name in names): + return all(stages[name]['result'] == 'success' for name in names) + if status['state'] in ('success', 'failure', 'planned'): + return status['state'] in ('success', 'planned') + time.sleep(3) + + +def summary(run_id): + directory = run_directory(run_id) + request = json.loads((directory / 'request.json').read_text()) + release = request['release'] + plan_file = directory / 'plan.json' + lines = [ + f'## Deploy `{release["sha"]}`', + '', + f'- Mode: `{request["mode"]}`', + f'- Refresh third-party images: `{request["refresh_images"]}`', + ] + status = json.loads((directory / 'status.json').read_text()) + if status.get('failure_stage'): + lines.append(f'- Failed stage: **{status["failure_stage"]}**') + lines.extend( + [ + '', + f'- Observed run state: **{status["state"]}**', + '', + '### Stage results', + '| Stage | Result | Exit code |', + '| --- | --- | --- |', + ] + ) + for name in ('doctor', 'validate', 'apply-k8s', 'apply-compose', 'verify-k8s', 'smoke'): + stage_result = status['stages'].get(name, {}) + lines.append(f'| {name} | {stage_result.get("result", "not started")} | {stage_result.get("exit_code", "—")} |') + lines.extend(['', '### Apply and Helm recovery results']) + events_file = directory / 'apply-events.jsonl' + events = [] + if events_file.exists(): + for line in events_file.read_text().splitlines(): + try: + events.append(json.loads(line)) + except json.JSONDecodeError: + lines.append('- An operation record is incomplete. Check the stage log.') + latest = {(event['action'], event['target']): event['result'] for event in events} + lines.extend(f'- `{action}` `{target}`: **{result}**' for (action, target), result in latest.items()) + if not latest: + lines.append('- No apply results were recorded.') + lines.append('- A completed apply does not confirm health. See verification and smoke results.') + lines.extend(['', '### Kubernetes recovery']) + pointer = directory / 'snapshot/current' + failed = Path(pointer.read_text().strip()) / 'failed-workloads' if pointer.exists() else None + if failed and failed.exists(): + contents = failed.read_text() + counts = dict(re.findall(r'^(ROLLED_BACK|UNRECOVERED)=([0-9]+)$', contents, re.MULTILINE)) + if not contents.strip(): + lines.append('- No failed workloads were recorded. See the verification result above.') + elif counts: + lines.append(f'- Workloads restored: **{counts.get("ROLLED_BACK", "unknown")}**') + lines.append(f'- Workloads that need manual recovery: **{counts.get("UNRECOVERED", "unknown")}**') + else: + lines.append('- Rollback has no recorded result yet. Check the verification log.') + else: + lines.append('- No workload rollback was recorded. This does not confirm health.') + lines.append('- Compose requires manual recovery. Use the saved command in the apply log.') + if not plan_file.exists(): + lines.extend(['', 'Plan was not created. Check the controller log.']) + print('\n'.join(lines)) + return + plan = json.loads(plan_file.read_text()) + lines.extend(['', '### Selected services']) + count = 0 + for kind, services in plan['selected'].items(): + for service in services: + lines.append(f'- `{kind}`: `{service}`') + count += 1 + if not count: + lines.append('- None') + lines.extend(['', '### Selected Helm releases']) + lines.extend(f'- `{release}`' for release in plan.get('helm', [])) + if not plan.get('helm'): + lines.append('- None') + lines.extend(['', '### Images pinned in the checked release']) + lines.extend(f'- `{image}@{digest}`' for image, digest in sorted(release['images'].items())) + lines.extend(['', '### Removed resources requiring manual review']) + lines.extend(f'- `{item}`' for item in plan.get('removed', [])) + if not plan.get('removed'): + lines.append('- None') + print('\n'.join(lines)) + + +def main(): + os.umask(0o077) + parser = argparse.ArgumentParser(description=__doc__) + parser.add_argument('action', choices=('start', 'execute', 'recover', 'status', 'follow', 'summary')) + parser.add_argument('run_id') + parser.add_argument('phase', nargs='?', choices=('apply', 'verify', 'smoke')) + parser.add_argument('--retry', action='store_true', help='Retry failed recovery checks; never repeat apply') + args = parser.parse_args() + directory = run_directory(args.run_id) + if args.action == 'start': + start(args.run_id) + elif args.action == 'execute': + execute(args.run_id) + elif args.action == 'recover': + with lock('deploy.lock'): + recover(directory, retry=args.retry) + elif args.action == 'status': + print((directory / 'status.json').read_text()) + if (directory / 'plan.json').exists(): + plan = json.loads((directory / 'plan.json').read_text()) + print(json.dumps({k: plan[k] for k in ('sha', 'selected', 'helm', 'removed')}, indent=2)) + elif args.action == 'summary': + summary(args.run_id) + elif not follow(args.run_id, args.phase): + sys.exit(1) + + +if __name__ == '__main__': + main() diff --git a/.gitea/workflows/deploy-lib.sh b/.gitea/workflows/deploy-lib.sh index fa7452e..ab010df 100644 --- a/.gitea/workflows/deploy-lib.sh +++ b/.gitea/workflows/deploy-lib.sh @@ -1,5 +1,5 @@ #!/usr/bin/env bash -# Shared stages for the deploy workflow. Runs on the workstation, invoked as: +# Workstation deploy stages; invoked by the durable controller against pinned source. # REPO=/srv/homelab APPLY_PRUNE=false bash -se <<'EOF' # source "$REPO/.gitea/workflows/deploy-lib.sh" # run_stage "$STAGE" @@ -8,8 +8,8 @@ set -euo pipefail : "${REPO:?REPO must be set}" APPLY_PRUNE="${APPLY_PRUNE:-false}" -# Commit CI validated. Empty for a manual workflow_dispatch, which falls back to -# the current origin/main. +CONFIG_REPO="${CONFIG_REPO:-$REPO}" +# Exact SHA accepted by the CI gate for both manual and automatic deploys. DEPLOY_SHA="${DEPLOY_SHA:-}" # Handoff point between the apply stage (writes) and the verify stage (reads). # Under the deploy user's own XDG state directory rather than /var/backups: the @@ -20,17 +20,36 @@ DEPLOY_SNAPSHOT_DIR="${DEPLOY_SNAPSHOT_DIR:-${XDG_STATE_HOME:-$HOME/.local/state # Per-workload rollout budget and how many workloads to watch at once. The whole # apply job has its own timeout-minutes as a backstop. ROLLOUT_TIMEOUT="${ROLLOUT_TIMEOUT:-300}" -ROLLOUT_PARALLELISM="${ROLLOUT_PARALLELISM:-8}" +ROLLOUT_PARALLELISM="${ROLLOUT_PARALLELISM:-4}" WORKLOAD_KINDS="deployments.apps,statefulsets.apps,daemonsets.apps" log() { echo "== $* ==" } +# Store operation results without command output or local configuration values. +record_apply() { + [ -n "${RUN_DIR:-}" ] || return 0 + jq -cn --arg action "$1" --arg target "$2" --arg result "$3" \ + '{action: $action, target: $target, result: $result}' >>"$RUN_DIR/apply-events.jsonl" \ + || echo 'WARNING: cannot record an apply result' >&2 + return 0 +} + warn() { echo "WARNING: $*" >&2 } +# Prune needs the complete desired set in one invocation. Per-file pruning +# treats resources from the other files as absent and can delete them. +check_prune_mode() { + if [ "$APPLY_PRUNE" = "true" ]; then + echo "ERROR: APPLY_PRUNE=true is unsupported by the per-file deploy loop." >&2 + echo "Disable it; remove obsolete resources explicitly after review." >&2 + return 1 + fi +} + collect_k8s() { git -C "$REPO" ls-files -- "$1" \ | grep -E '\.ya?ml$' \ @@ -50,6 +69,25 @@ kustomize_overlay() { fi } +selected_service() { + local kind="$1" service="$2" section=selected + [ -n "${DEPLOY_PLAN:-}" ] || return 0 + if [ "${DEPLOY_SMOKE_ALL:-false}" = true ]; then section=active; fi + jq -e --arg kind "$kind" --arg service "$service" --arg section "$section" \ + '.[$section][$kind] | index($service) != null' "$DEPLOY_PLAN" >/dev/null +} + +# Resolve .env and relative binds on the persistent workstation tree. Locked +# JSON configs keep the same Compose project name and volume names. +compose() { + local cf="$1" locked project_dir + project_dir="$CONFIG_REPO/$(basename "$(dirname "$cf")")" + shift + locked="${RUN_DIR:-/nonexistent}/compose/$(basename "$(dirname "$cf")").json" + if [ -f "$locked" ]; then cf="$locked"; fi + (cd "$CONFIG_REPO" && docker compose --project-directory "$project_dir" -f "$cf" "$@") +} + select_manifests() { K8S_MANIFESTS=() KUSTOMIZE_APPS=() @@ -57,6 +95,7 @@ select_manifests() { local kd_rel kd overlay cf_rel cf f while IFS= read -r kd_rel; do kd="$REPO/$kd_rel" + selected_service k8s "${kd_rel%/k8s}" || continue if [ ! -f "$kd/active" ]; then echo "skip (no k8s/active): $kd_rel" continue @@ -78,6 +117,7 @@ select_manifests() { ) while IFS= read -r cf_rel; do cf="$REPO/$cf_rel" + selected_service compose "$(dirname "$cf_rel")" || continue if [ -f "$(dirname "$cf")/active" ]; then echo "compose: $cf_rel" COMPOSE_STACKS+=("$cf") @@ -94,16 +134,10 @@ select_manifests() { # on failure roll them back to the revision that was running before, so a bad # push to main cannot leave a service crash-looping. # -# Verification lives in its own workflow job, not at the end of the apply stage. -# Inside a single process it is worthless exactly when it is needed most: a job -# killed by timeout-minutes or cancelled mid-apply never reaches the rollback -# code, and leaves a half-applied cluster behind. Split out, the apply job can -# die in any way and the verify job still runs. +# The workstation controller runs apply and verification as separate durable +# stages. Runner jobs only follow their logs. ExecStopPost recovers interrupted +# runs using the per-run snapshot, even when the SSH connection has gone away. # -# That split needs a handoff point on the workstation, because the two stages are -# separate processes on separate runner jobs: DEPLOY_SNAPSHOT_DIR/current, written -# before anything is applied, read by the verify stage afterwards. - # Creates this run's snapshot directory and publishes it as the handoff point for # the verify stage. Fails hard by design: a deploy that cannot record what it is # about to change must not start, because then nothing can be rolled back for it @@ -133,18 +167,36 @@ snapshot_dir() { } save_snapshot() { - local dir="$1" + local dir="$1" releases revision status log "Saving pre-apply snapshot to $dir" - workload_generations >"$dir/generations.before" 2>/dev/null \ - || warn "could not snapshot workload generations" - kubectl get "$WORKLOAD_KINDS" -A -o yaml >"$dir/workloads.yaml" 2>/dev/null \ - || warn "could not snapshot workloads" - for release in prometheus-stack loki alloy; do - if helm status "$release" -n prometheus >/dev/null 2>&1; then - { - echo "revision: $(helm history "$release" -n prometheus -o json 2>/dev/null)" - helm get values "$release" -n prometheus --all 2>/dev/null - } >"$dir/helm-$release.txt" + workload_generations >"$dir/generations.before" || return 1 + kubectl get "$WORKLOAD_KINDS" -A -o json >"$dir/workloads.json" || return 1 + kubectl get controllerrevisions.apps -A -o json >"$dir/controller-revisions.json" || return 1 + jq --slurpfile revisions "$dir/controller-revisions.json" ' + [.items[] | . as $w | { + kind: (.kind | ascii_downcase), namespace: .metadata.namespace, name: .metadata.name, uid: .metadata.uid, + revision: (if .kind == "Deployment" then (.metadata.annotations["deployment.kubernetes.io/revision"] // "0" | tonumber) + else ([$revisions[0].items[] | select(.metadata.namespace == $w.metadata.namespace) + | select(any(.metadata.ownerReferences[]?; .uid == $w.metadata.uid)) + | select($w.kind != "StatefulSet" or .metadata.name == $w.status.currentRevision) | .revision] | max // 0) end) + }]' "$dir/workloads.json" >"$dir/revisions.json" || return 1 + # Helm 4 lists every release status by default and removed the --all flag. + releases="$(helm list -A -o json)" || return 1 + for entry in "${HELM_RELEASES[@]}"; do + IFS='|' read -r release _ namespace _ _ _ <<<"$entry" + if ! jq -e --arg r "$release" --arg n "$namespace" \ + 'any(.[]; .name == $r and .namespace == $n)' <<<"$releases" >/dev/null; then + continue + fi + helm status "$release" -n "$namespace" -o json >"$dir/helm-$release.json" || return 1 + status="$(jq -r '.info.status' "$dir/helm-$release.json")" + if [ "$status" != deployed ]; then + # Never capture a pending/failed revision as the recovery target. + helm history "$release" -n "$namespace" -o json >"$dir/helm-$release.history.json" || return 1 + revision="$(jq '[.[] | select(.status == "deployed" or .status == "superseded") | .revision] | max // 0' \ + "$dir/helm-$release.history.json")" + jq --argjson revision "$revision" '.version = $revision' "$dir/helm-$release.json" >"$dir/helm-$release.tmp" + mv "$dir/helm-$release.tmp" "$dir/helm-$release.json" fi done # The verify stage compares this against the commit it is deploying, to refuse @@ -168,294 +220,24 @@ workload_generations() { # moved since the snapshot, i.e. the ones this apply actually touched. changed_workloads() { local before="$1" - local ns name kind gen old + local ns name kind gen old current + current="$(workload_generations)" || return 1 while read -r ns name kind gen; do [ -n "${gen:-}" ] || continue - old="$(awk -v want_ns="$ns" -v want_name="$name" \ - '$1 == want_ns && $2 == want_name { print $4; exit }' "$before" 2>/dev/null || true)" + old="$(awk -v want_ns="$ns" -v want_name="$name" -v want_kind="$kind" \ + '$1 == want_ns && $2 == want_name && $3 == want_kind { print $4; exit }' "$before" 2>/dev/null || true)" + if [ -n "${RUN_DIR:-}" ] && ! grep -qxF "$kind $ns $name" "$RUN_DIR/workload-refs"; then + continue + fi if [ "$old" != "$gen" ]; then printf '%s %s %s\n' "$kind" "$ns" "$name" fi - done < <(workload_generations) + done <<<"$current" } -# Prints " / " for every workload this repository owns that -# runs an image from our own registry. -# -# The repository is the scope, deliberately. The cluster also holds workloads on -# our registry that no manifest here declares (they are applied out of band), and -# those are somebody else's to deploy. Walking the manifests rather than the -# cluster means those can never be restarted by this pipeline, now or later. -owned_registry_workloads() { - local kd_rel f - while IFS= read -r kd_rel; do - [ -f "$REPO/$kd_rel/active" ] || continue - while IFS= read -r f; do - [ -n "$f" ] || continue - # A file that does not mention the registry cannot declare a workload on it, - # and parsing costs ~2.5s per file against a millisecond for the grep. The - # filter keeps this at a handful of parses instead of one per manifest. - grep -q 'gcr\.forust\.xyz/forust/' "$REPO/$f" 2>/dev/null || continue - # kubectl prints a bare object for a single-document file and a List for a - # multi-document one, so normalise both shapes before filtering. - kubectl apply --dry-run=client -f "$REPO/$f" -o json 2>/dev/null \ - | jq -r ' - (if .items then .items[] else . end) - | select(.kind | test("^(Deployment|StatefulSet|DaemonSet)$")) - | select(any((.spec.template.spec.containers // [])[]?; - (.image // "") | test("^gcr\\.forust\\.xyz/forust/"))) - | (.metadata.namespace // "default") as $ns - | ([.spec.template.spec.containers[].image - | select(test("^gcr\\.forust\\.xyz/forust/"))][0]) as $img - | "\($ns) \(.kind | ascii_downcase)/\(.metadata.name) \($img)" - ' 2>/dev/null || true - done < <(collect_k8s "$kd_rel" || true) - done < <( - git -C "$REPO" ls-files '*.yaml' '*.yml' \ - | grep -E '(^|/)k8s/' \ - | sed -E 's#((^|.*/)k8s)/.*#\1#' \ - | sort -u - ) -} - -# Prints the digest an image tag resolves to for this cluster's architecture, or -# nothing when it cannot be resolved. -# -# Only the manifest entry matching the node architecture counts. A multi-arch tag -# also carries `unknown/unknown` entries for the build attestation, and a pod's -# imageID is always the per-platform digest, so comparing the wrong entry would -# mark every workload stale forever and restart the whole cluster on every deploy. -registry_digest() { - local arch - arch="$(kubectl get nodes -o jsonpath='{.items[0].status.nodeInfo.architecture}' 2>/dev/null || true)" - [ -n "$arch" ] || arch=amd64 - # The || true is load-bearing. Every caller runs under set -euo pipefail, and - # pipefail reports the rightmost non-zero stage, so a ref the registry does not - # have would abort the caller at the assignment instead of yielding an empty - # string. The callers check for empty themselves and report it by name. - # - # Retried with a hard timeout because the registry has a known hang mode (and - # a known blink mode: a single failed lookup aborts the whole apply file in - # render_pinned). A short sleep between attempts lets a restarting registry - # come back instead of failing the deploy on one bad second. - local attempt=0 digest="" - while [ "$attempt" -lt 3 ]; do - digest="$(timeout 25s docker manifest inspect "$1" 2>/dev/null \ - | jq -r --arg arch "$arch" ' - .manifests[]? - | select(.platform.os == "linux" and .platform.architecture == $arch) - | .digest - ' 2>/dev/null \ - | head -1 || true)" - [ -n "$digest" ] && break - attempt=$((attempt + 1)) - if [ "$attempt" -lt 3 ]; then - echo "WARNING: registry lookup of $1 failed (attempt $attempt/3), retrying in 5s" >&2 - sleep 5 - fi - done - printf '%s' "$digest" -} - -# The commit this deploy is for: what CI validated, or - on a manual dispatch, -# whatever stage_preflight just checked out. -deploy_commit() { - local c="${DEPLOY_SHA:-}" - [ -n "$c" ] || c="$(git -C "$REPO" rev-parse HEAD 2>/dev/null || true)" - printf '%.12s' "${c:-}" -} - -# Resolves one of our image refs to the digest THIS commit's build produced. -# -# A manifest naming `:prod` names a pointer, not a version, and the deploy -# resolves it when the apply runs - which is not when CI ran it. Deploy runs are -# queued rather than cancelled (see deploy.yaml), so two pushes in a row leave -# the first deploy resolving the second push's build: the right manifests with -# the wrong code, and nothing anywhere reports it. ci therefore publishes every -# image it ships under `sha-`, a name that cannot move, and that is -# the name resolved here. -# -# The fallback to the plain tag is for an image this pipeline never built. It -# reports itself, because a fallback nobody sees is the failure this removes. -pinned_digest() { - local ref="$1" commit pinned - commit="$(deploy_commit)" - if [ -n "$commit" ]; then - pinned="$(registry_digest "${ref%:*}:sha-$commit")" - if [ -n "$pinned" ]; then - printf '%s' "$pinned" - return 0 - fi - fi - pinned="$(registry_digest "$ref")" - if [ -n "$pinned" ]; then - echo "WARNING: ${ref} carries no sha-${commit:-} tag; resolved the moving tag instead" >&2 - fi - printf '%s' "$pinned" -} - -# Rewrites our own images to immutable digests on the way into the cluster. -# Reads a manifest stream on stdin, writes the pinned stream to stdout. -# -# A digest is not knowable when a manifest is written, so it is never committed: -# git keeps a readable `:prod` tag and the exact bytes are chosen here, at apply -# time, from the tag ci published for the commit being deployed. That is what -# makes rollback mean something. `kubectl rollout undo` restores the previous -# ReplicaSet's pod template verbatim, and a template naming a digest restores the -# exact bytes that were serving before. A template naming a moving tag does not — -# the tag has already moved by the time the rollback runs, so the "rollback" -# re-pulls the very image that just failed and the cluster stays broken. -# -# imagePullPolicy is deliberately left alone. The manifests no longer set it, and a -# reference that is not `:latest` defaults to IfNotPresent, which is what the -# Kubernetes docs ask for alongside a digest: the bytes under a digest cannot -# change, so pulling again buys nothing. -# -# An image that cannot be resolved is fatal. Carrying on would quietly apply a -# mutable tag again, which is the exact failure this function exists to remove. +# Resolve owned image references exclusively from the checked CI artifact. render_pinned() { - local src refs map ref digest missing=0 - src="$(mktemp)" - refs="$(mktemp)" - map="$(mktemp)" - - cat >"$src" - grep -oE 'gcr\.forust\.xyz/forust/[A-Za-z0-9._-]+:[A-Za-z0-9._-]+' "$src" | sort -u >"$refs" || true - - while read -r ref; do - [ -n "$ref" ] || continue - digest="$(pinned_digest "$ref")" - if [ -z "$digest" ]; then - echo "ERROR: cannot resolve ${ref} in the registry; applying nothing." >&2 - echo " The build job has to push that tag before the deploy resolves it." >&2 - missing=$((missing + 1)) - continue - fi - printf '%s\t%s\n' "$ref" "$digest" >>"$map" - done <"$refs" - if [ "$missing" -gt 0 ]; then - rm -f "$src" "$refs" "$map" - return 1 - fi - - awk -v mapfile="$map" ' - BEGIN { - while ((getline line < mapfile) > 0) { - i = index(line, "\t") - d[substr(line, 1, i - 1)] = substr(line, i + 1) - } - } - { - if (match($0, /^[[:space:]]*image:[[:space:]]*gcr\.forust\.xyz\/forust\/[A-Za-z0-9._-]+:[A-Za-z0-9._-]+[[:space:]]*$/)) { - name = $0 - sub(/^[[:space:]]*image:[[:space:]]*/, "", name) - sub(/[[:space:]]*$/, "", name) - if (name in d) { - pad = $0 - sub(/image:.*/, "", pad) - # Drop the tag: the canonical form used in the docs is repo@sha256:..., - # and leaving :prod next to the digest reads like it still matters. - repo = name - sub(/:[A-Za-z0-9._-]+$/, "", repo) - print pad "image: " repo "@" d[name] - next - } - } - print - } - ' "$src" - rm -f "$src" "$refs" "$map" -} - -# Restarts every owned workload whose running image is not the one its tag -# resolves to now. -# -# This used to be how a rebuild reached the cluster at all: the manifests pinned -# `:latest`, so a rebuild left the pod template byte-identical, `kubectl apply` -# decided there was nothing to do, and the cluster served the previous build -# indefinitely. The apply now pins digests via render_pinned, so a rebuild moves -# the pod template and rolls out on its own. -# -# What is left is the drift check: a hand-run `kubectl set image`, or anything -# else that edits a live workload behind the deploy's back, is the only way to end -# up serving a digest the tag has moved past. It stays idempotent, so a redeploy -# that changed no image still does not bounce healthy services. -# -# The container is matched on its repository rather than on the exact reference: -# once render_pinned has run, a pod's status reports `repo@sha256:...` while this -# still reads the repository's `:prod` tag out of the manifest. -restart_stale_images() { - local ns target image want selector running entry one - local unchecked=0 - local -A digests=() - local -a stale=() - while read -r ns target image; do - [ -n "${target:-}" ] || continue - if [ -z "${digests[$image]:-}" ]; then - digests[$image]="$(pinned_digest "$image")" - fi - want="${digests[$image]}" - if [ -z "$want" ]; then - warn "cannot resolve ${image##*/} in the registry, leaving $target alone" - unchecked=$((unchecked + 1)) - continue - fi - selector="$(kubectl get "$target" -n "$ns" -o jsonpath='{.spec.selector.matchLabels}' 2>/dev/null \ - | jq -r 'to_entries | map("\(.key)=\(.value)") | join(",")' 2>/dev/null)" - if [ -z "$selector" ]; then - warn "cannot read the pod selector of $target, skipping" - unchecked=$((unchecked + 1)) - continue - fi - running="$(kubectl get pods -n "$ns" -l "$selector" -o json 2>/dev/null \ - | jq -r --arg repo "${image%%:*}" ' - .items[] | .status.containerStatuses[]? - | select(.image == $repo - or (.image | startswith($repo + ":")) - or (.image | startswith($repo + "@"))) - | .imageID - ' 2>/dev/null)" - if [ -z "$running" ]; then - # Scaled to zero. Nothing is serving stale code, and imagePullPolicy - # resolves the tag when it is scaled back up. - continue - fi - entry="" - while IFS= read -r one; do - [ -n "$one" ] || continue - entry="${one##*@}" - if [ "$entry" != "$want" ]; then - stale+=("$ns $target") - break - fi - done <<<"$running" - done < <(owned_registry_workloads) - if [ "${#stale[@]}" -eq 0 ]; then - if [ "$unchecked" -gt 0 ]; then - # Say so plainly. Reporting "everything is current" after checking nothing - # would tell the operator the deploy is fine when it may not be. - warn "No workload needed a restart, but $unchecked could not be checked" - else - log "All owned workloads already run the image their tag points at" - fi - return 0 - fi - log "Restarting ${#stale[@]} workload(s) running an image their tag has moved past" - for ref in "${stale[@]}"; do - log " $ref" - done - local failed=() - for ref in "${stale[@]}"; do - ns="${ref%% *}" - target="${ref#* }" - if ! kubectl rollout restart "$target" -n "$ns" >/dev/null 2>&1; then - failed+=("$ref") - fi - done - if [ "${#failed[@]}" -gt 0 ]; then - warn "could not restart: ${failed[*]}" - return 1 - fi + python3 "$REPO/.gitea/workflows/release.py" render } # verify_workloads ... @@ -499,31 +281,44 @@ verify_workloads() { # settle. Prints a report and returns non-zero if any workload is still unhealthy, # so the operator knows manual recovery is required. rollback_workloads() { - local failed_file="$1" - local kind ns name unrecovered=() - local -a recovered=() + local failed_file="$1" snapshot kind ns name index=0 running=0 pid revision uid + local -a pids=() + snapshot="$(cat "$DEPLOY_SNAPSHOT_DIR/current")" while read -r kind ns name; do - [ -n "${kind:-}" ] || continue - # Helm-owned workloads are already rolled back by the release's --rollback-on-failure - # upgrade. `rollout undo` here would step back to the revision Helm just - # escaped (the failed one), so leave them for the operator instead. - if kubectl get "${kind}/${name}" -n "$ns" -o jsonpath='{.metadata.annotations}' 2>/dev/null | grep -q 'meta.helm.sh/release-name'; then - echo " skip (helm-managed, needs manual check): ${kind}/${ns}/${name}" - unrecovered+=("${kind}/${ns}/${name} (helm-managed)") - continue - fi - if kubectl rollout undo "${kind}/${name}" -n "$ns" >/dev/null 2>&1 \ - && kubectl rollout status "${kind}/${name}" -n "$ns" --timeout="${ROLLOUT_TIMEOUT}s" >/dev/null 2>&1; then - echo " rolled back: ${kind}/${ns}/${name}" - recovered+=("${kind}/${ns}/${name}") - else - echo " NOT RECOVERED: ${kind}/${ns}/${name}" - unrecovered+=("${kind}/${ns}/${name}") + [[ "$kind" =~ ^(deployment|statefulset|daemonset)$ ]] || continue + index=$((index + 1)) + ( + if kubectl get "$kind/$name" -n "$ns" -o jsonpath='{.metadata.annotations}' | grep -q 'meta.helm.sh/release-name'; then + echo " skip (Helm recovery owns this workload): $kind/$ns/$name" + exit 1 + fi + revision="$(jq -r --arg ns "$ns" --arg name "$name" --arg kind "$kind" \ + '.[] | select(.namespace == $ns and .name == $name and .kind == $kind) | .revision' "$snapshot/revisions.json")" + uid="$(jq -r --arg ns "$ns" --arg name "$name" --arg kind "$kind" \ + '.[] | select(.namespace == $ns and .name == $name and .kind == $kind) | .uid' "$snapshot/revisions.json")" + if [[ ! "$revision" =~ ^[1-9][0-9]*$ ]] || [ "$uid" != "$(kubectl get "$kind/$name" -n "$ns" -o jsonpath='{.metadata.uid}')" ]; then + echo " no safe previous revision: $kind/$ns/$name (new or replaced workload)" + exit 1 + fi + kubectl rollout undo "$kind/$name" -n "$ns" --to-revision="$revision" \ + && kubectl rollout status "$kind/$name" -n "$ns" --timeout="${ROLLOUT_TIMEOUT}s" + ) >"$snapshot/rollback-$index.log" 2>&1 & + pids+=($!) + running=$((running + 1)) + if [ "$running" -ge "$ROLLOUT_PARALLELISM" ]; then + wait -n 2>/dev/null || true + running=$((running - 1)) fi done <"$failed_file" - echo "ROLLED_BACK=${#recovered[@]}" >>"$failed_file" - echo "UNRECOVERED=${#unrecovered[@]}" >>"$failed_file" - [ "${#unrecovered[@]}" -eq 0 ] + local recovered=0 unrecovered=0 i=0 + for pid in "${pids[@]}"; do + i=$((i + 1)) + if wait "$pid"; then recovered=$((recovered + 1)); else unrecovered=$((unrecovered + 1)); fi + cat "$snapshot/rollback-$i.log" + done + echo "ROLLED_BACK=$recovered" >>"$failed_file" + echo "UNRECOVERED=$unrecovered" >>"$failed_file" + [ "$unrecovered" -eq 0 ] } # Helm releases owned by this stage, one line each: @@ -535,10 +330,11 @@ rollback_workloads() { # written straight into a `helm upgrade` command would never be updated: these # have to be declared as custom.regex managers in renovate/renovate.json. HELM_RELEASES=( - "prometheus-stack|prometheus-community/kube-prometheus-stack|prometheus|86.2.3|prometheus-stack/k8s/grafana-values.yaml|prometheus-stack/k8s/active" + "prometheus-stack|prometheus-community/kube-prometheus-stack|prometheus|86.3.2|prometheus-stack/k8s/grafana-values.yaml|prometheus-stack/k8s/active" + "victoria-operator|victoriametrics/victoria-metrics-operator|prometheus|0.68.1|prometheus-stack/k8s/victoria-operator-values.yaml|prometheus-stack/k8s/active" "loki|grafana/loki|prometheus|7.3.0|loki/k8s/loki-values.yaml|loki/k8s/active" "alloy|grafana/alloy|prometheus|1.12.1|loki/k8s/alloy-values.yaml|loki/k8s/active" - "reloader|stakater/reloader|reloader|2.2.17|reloader/k8s/reloader-values.yaml|reloader/k8s/active" + "reloader|stakater/reloader|reloader|2.2.18|reloader/k8s/reloader-values.yaml|reloader/k8s/active" ) # "name url" for the Helm repository hosting a chart, empty if unknown. @@ -547,6 +343,7 @@ helm_repo_for() { prometheus-community/*) echo "prometheus-community https://prometheus-community.github.io/helm-charts" ;; grafana/*) echo "grafana https://grafana.github.io/helm-charts" ;; stakater/*) echo "stakater https://stakater.github.io/stakater-charts" ;; + victoriametrics/*) echo "victoriametrics https://victoriametrics.github.io/helm-charts" ;; esac } @@ -556,8 +353,12 @@ helm_repo_for() { helm_release_status() { local out if ! out="$(helm status "$1" -n "$2" 2>&1)"; then - echo "not-found" - return 0 + if [[ "$out" == *"release: not found"* ]]; then + echo "not-found" + return 0 + fi + printf 'ERROR: cannot read Helm status: %s\n' "$out" >&2 + return 1 fi awk '/^STATUS:/{print $2}' <<<"$out" | tr '[:upper:]' '[:lower:]' } @@ -569,20 +370,35 @@ helm_release_status() { # pending (deployed, failed, not-found). Returns non-zero when the release is # still not recoverable, so the pipeline fails loud instead of wedging. recover_pending_release() { - local release="$1" namespace="$2" status - status="$(helm_release_status "$release" "$namespace")" + local release="$1" namespace="$2" status revision snapshot + status="$(helm_release_status "$release" "$namespace")" || return 1 case "$status" in pending-upgrade|pending-rollback|pending-install) log "Release $release is $status, rolling back to the last deployed revision" - if ! helm rollback "$release" -n "$namespace" --wait --timeout 10m >/dev/null 2>&1; then + revision="" + if [ -s "$DEPLOY_SNAPSHOT_DIR/current" ]; then + snapshot="$(cat "$DEPLOY_SNAPSHOT_DIR/current")" + if [ -s "$snapshot/helm-$release.json" ]; then + revision="$(jq -r '.version' "$snapshot/helm-$release.json")" + fi + fi + if [[ ! "$revision" =~ ^[1-9][0-9]*$ ]]; then + echo "ERROR: no captured Helm revision for $release; manual recovery required" + return 1 + fi + record_apply helm-rollback "$namespace/$release" started + if ! helm rollback "$release" "$revision" -n "$namespace" --wait --timeout 10m; then + record_apply helm-rollback "$namespace/$release" failure echo "WARN: helm rollback of $release did not complete" return 1 fi - status="$(helm_release_status "$release" "$namespace")" + status="$(helm_release_status "$release" "$namespace")" || return 1 if [ "$status" != "deployed" ]; then + record_apply helm-rollback "$namespace/$release" failure echo "WARN: $release is $status after rollback" return 1 fi + record_apply helm-rollback "$namespace/$release" success ;; esac return 0 @@ -614,11 +430,16 @@ upgrade_helm_releases() { local entry release chart namespace version values marker repo for entry in ${HELM_RELEASES[@]+"${HELM_RELEASES[@]}"}; do IFS='|' read -r release chart namespace version values marker <<<"$entry" + if [ -n "${DEPLOY_PLAN:-}" ] && ! jq -e --arg name "$release" '.helm | index($name) != null' "$DEPLOY_PLAN" >/dev/null; then + echo "skip (unchanged Helm release): $release" + continue + fi + if [ ! -f "$REPO/$values" ] && [ -f "$CONFIG_REPO/$values" ]; then values="$CONFIG_REPO/$values"; else values="$REPO/$values"; fi if [ ! -f "$REPO/$marker" ]; then echo "skip (no $marker): $release" continue fi - if [ ! -f "$REPO/$values" ]; then + if [ ! -f "$values" ]; then echo "ERROR: $values is gitignored but missing on the workstation, restore it first." return 1 fi @@ -627,8 +448,8 @@ upgrade_helm_releases() { echo "ERROR: no Helm repository configured for chart $chart" return 1 fi - helm repo add "${repo%% *}" "${repo#* }" >/dev/null 2>&1 || true - helm repo update "${repo%% *}" >/dev/null 2>&1 || true + helm repo add "${repo%% *}" "${repo#* }" >/dev/null + helm repo update "${repo%% *}" >/dev/null log "Upgrading $release ($chart $version)" wait_for_calm "helm $release" # A previous run with --rollback-on-failure whose own rollback never finished leaves the @@ -641,11 +462,13 @@ upgrade_helm_releases() { # --rollback-on-failure (+ --wait) rolls the release back when the upgrade # times out or the workloads it touches never become ready, so a bad chart # bump is not left half applied. (--atomic was this combo; deprecated.) + record_apply helm-upgrade "$namespace/$release" started if ! helm upgrade --install "$release" "$chart" \ --namespace "$namespace" \ --version "$version" \ - --values "$REPO/$values" \ + --values "$values" \ --wait --rollback-on-failure --cleanup-on-fail --timeout 10m; then + record_apply helm-upgrade "$namespace/$release" failure echo "WARN: upgrade of $release failed, checking release state" # --rollback-on-failure already attempted its own rollback; finish the job when that # rollback never completed, otherwise the release stays pending-* and @@ -655,132 +478,222 @@ upgrade_helm_releases() { else echo "ERROR: upgrade of $release failed (release is back on its previous revision)." fi + record_apply helm-recovery-state "$namespace/$release" "$(helm_release_status "$release" "$namespace" || echo unknown)" return 1 fi + record_apply helm-upgrade "$namespace/$release" success done } -stage_preflight() { - if [ ! -d "$REPO/.git" ]; then - echo "Repository not found at $REPO" - exit 1 +stage_doctor() { + local tool entry release chart namespace version values marker + for tool in git docker kubectl helm jq curl timeout flock python3; do + command -v "$tool" >/dev/null || { echo "Missing workstation tool: $tool"; return 1; } + done + docker compose version >/dev/null + docker buildx version >/dev/null + [ "$(kubectl config current-context)" = "${KUBE_CONTEXT:?configure KUBE_CONTEXT}" ] || { echo "Unexpected Kubernetes context"; return 1; } + [ "$(kubectl get namespace kube-system -o jsonpath='{.metadata.uid}')" = "${EXPECTED_CLUSTER_UID:?configure EXPECTED_CLUSTER_UID}" ] || { echo "Unexpected Kubernetes cluster"; return 1; } + kubectl get --raw=/readyz --request-timeout=10s >/dev/null + [ "$(git -C "$REPO" rev-parse HEAD)" = "$DEPLOY_SHA" ] || return 1 + select_manifests + for entry in "${HELM_RELEASES[@]}"; do + IFS='|' read -r release chart namespace version values marker <<<"$entry" + [ -f "$REPO/$marker" ] || continue + [ -f "$REPO/$values" ] || [ -f "$CONFIG_REPO/$values" ] || { echo "Missing values: $values"; return 1; } + done + jq '{sha, selected, helm, removed}' "$DEPLOY_PLAN" + local cf + for cf in "${COMPOSE_STACKS[@]}"; do + compose "$cf" config --quiet + while IFS= read -r network; do + docker network inspect "$network" >/dev/null || return 1 + done < <(compose "$cf" config --format json | jq -r '.networks // {} | to_entries[] | select(.value.external == true) | .value.name') + python3 "$REPO/.gitea/workflows/compose-release.py" "$cf" + done + local image refs m k + refs="$( + for m in "${K8S_MANIFESTS[@]}"; do render_pinned <"$m" || return 1; done + for k in "${KUSTOMIZE_APPS[@]}"; do kubectl kustomize "$k" | render_pinned || return 1; done + )" || return 1 + refs="$(grep -oE 'gcr\.forust\.xyz/forust/[A-Za-z0-9._-]+@sha256:[0-9a-f]{64}' <<<"$refs" | sort -u || true)" + while IFS= read -r image; do + [ -n "$image" ] || continue + timeout 60s docker buildx imagetools inspect "$image" >/dev/null + done <<<"$refs" +} + +# Required pod Secrets, scoped to the resource namespace. TLS route Secrets are +# created by cert-manager and are not prerequisites for applying a Certificate. +check_referenced_secrets() { + local m k objects refs extracted ns name + local missing=() + refs="" + for m in ${K8S_MANIFESTS[@]+"${K8S_MANIFESTS[@]}"}; do + if skip_uninstalled_vmagent_crd "$m"; then + continue + fi + objects="$(kubectl create --dry-run=client --validate=false -f "$m" -o json)" || return 1 + extracted="$(printf '%s' "$objects" | jq -r -f "$REPO/.gitea/workflows/secret-references.jq")" || return 1 + refs+="$extracted"$'\n' + done + for k in ${KUSTOMIZE_APPS[@]+"${KUSTOMIZE_APPS[@]}"}; do + objects="$(kubectl kustomize "$k" | kubectl create --dry-run=client --validate=false -f - -o json)" || return 1 + extracted="$(printf '%s' "$objects" | jq -r -f "$REPO/.gitea/workflows/secret-references.jq")" || return 1 + refs+="$extracted"$'\n' + done + while read -r ns name; do + [ -n "${name:-}" ] || continue + if kubectl get secret "$name" -n "$ns" -o name >/dev/null 2>&1; then + echo " ok: $ns/$name" + else + echo " MISSING OR UNREADABLE: $ns/$name" + missing+=("$ns/$name") + fi + done < <(printf '%s' "$refs" | sort -u) + if [ "${#missing[@]}" -gt 0 ]; then + echo "ERROR: required pod Secrets are missing or unreadable:" + printf ' - %s\n' "${missing[@]}" + echo "Create them in the listed namespaces from the service's secret example." + return 1 fi - if [ -n "$DEPLOY_SHA" ]; then - log "Checking out the commit CI validated ($DEPLOY_SHA)" - git -C "$REPO" fetch origin --quiet "$DEPLOY_SHA" 2>/dev/null \ - || git -C "$REPO" fetch origin main - else - git -C "$REPO" fetch origin main +} + +# The VMAgent CRD is installed by the VictoriaMetrics Operator Helm release in +# stage_apply_k8s, after this preflight stage. Skip only its dry-run until then. +skip_uninstalled_vmagent_crd() { + local manifest="$1" + if [[ "$manifest" == "$REPO/prometheus-stack/k8s/vmagent.yaml" ]] \ + && ! kubectl get crd vmagents.operator.victoriametrics.com >/dev/null 2>&1; then + echo " skip: VMAgent CRD is installed by Helm during apply: ${manifest#"$REPO"/}" + return 0 fi - target="${DEPLOY_SHA:-origin/main}" - log "Workstation state" - echo " local: $(git -C "$REPO" rev-parse --short HEAD)" - echo " target: $(git -C "$REPO" rev-parse --short "$target")" - if [ -n "$(git -C "$REPO" status --porcelain --untracked-files=no)" ]; then - echo "ERROR: workstation has local tracked modifications, refusing reset:" - git -C "$REPO" status --porcelain --untracked-files=no - git -C "$REPO" diff --stat - echo "Fix it on the workstation (commit, or 'git restore .'), then re-run the deploy." - exit 1 + return 1 +} + +# Render one complete resource list so new namespaces can be identified across +# files and Kustomize apps. A missing undeclared namespace remains an error. +render_selected_resources() { + local m k + { + for m in "${K8S_MANIFESTS[@]}"; do + if skip_uninstalled_vmagent_crd "$m" >/dev/null; then continue; fi + kubectl create --dry-run=client --validate=false -f "$m" -o json || return 1 + done + for k in "${KUSTOMIZE_APPS[@]}"; do + kubectl kustomize "$k" | kubectl create --dry-run=client --validate=false -f - -o json || return 1 + done + } | jq -s '{apiVersion: "v1", kind: "List", items: [ .[] | if .kind == "List" then .items[] else . end ]}' +} + +validate_server_resources() { + local defer_new="$1" resources existing filtered + resources="$(render_selected_resources)" || return 1 + existing="$(kubectl get namespaces -o json)" || return 1 + filtered="$(jq --argjson existing "$existing" --argjson defer "$defer_new" ' + [.items[] | select(.kind == "Namespace") | .metadata.name] as $declared + | [$existing.items[].metadata.name] as $present + | .items |= map( + (.metadata.namespace // "default") as $ns + | if .kind == "Namespace" or ($present | index($ns)) != null then . + elif ($declared | index($ns)) == null then error("Undeclared missing namespace: " + $ns) + elif $defer then empty + else error("Namespace still missing after namespace apply: " + $ns) + end) + ' <<<"$resources")" || return 1 + if [ "$(jq '.items | length' <<<"$filtered")" -gt 0 ]; then + kubectl apply --dry-run=server -f - <<<"$filtered" >/dev/null fi - git -C "$REPO" reset --hard "$target" } stage_validate() { + check_prune_mode || return 1 cd "$REPO" select_manifests local m k cf - # Compose .env files and secret files are gitignored by design, so the - # workstation never has real values for the inactive stacks. This stage only - # runs the full check on active stacks; the general structure check for every - # committed Compose file (active or not) lives in the ci workflow, which has no - # .env at all. - # - # Active stacks are still validated with interpolation and env-file resolution - # off, so required-variable guards (:?) and missing local files do not fail the - # deploy. Normalization and consistency checks stay enabled. + # The deploy host has the local .env and secret files. Resolve them here so + # missing configuration fails before either apply job changes workloads. + # CI keeps the structure-only check for inactive stacks. # shellcheck source=compose-lint.sh source "$REPO/.gitea/workflows/compose-lint.sh" - local compose_validate_flags=() - mapfile -t compose_validate_flags < <(compose_safe_flags) log "Validate compose stacks" for cf in ${COMPOSE_STACKS[@]+"${COMPOSE_STACKS[@]}"}; do echo " config: $cf" - validate_compose_file "$cf" ${compose_validate_flags[@]+"${compose_validate_flags[@]}"} + compose "$cf" config --quiet done log "Validate k8s manifests (kubectl dry-run=client)" for m in ${K8S_MANIFESTS[@]+"${K8S_MANIFESTS[@]}"}; do + if skip_uninstalled_vmagent_crd "$m"; then + continue + fi kubectl apply --dry-run=client -f "$m" >/dev/null done for k in ${KUSTOMIZE_APPS[@]+"${KUSTOMIZE_APPS[@]}"}; do kubectl apply -k "$k" --dry-run=client >/dev/null done log "Validate k8s manifests (kubectl dry-run=server)" - for m in ${K8S_MANIFESTS[@]+"${K8S_MANIFESTS[@]}"}; do - kubectl apply --dry-run=server -f "$m" >/dev/null - done - for k in ${KUSTOMIZE_APPS[@]+"${KUSTOMIZE_APPS[@]}"}; do - kubectl apply -k "$k" --dry-run=server >/dev/null - done + validate_server_resources true log "Checking referenced Secrets exist" echo " (deploy never applies *secret*.yaml; create missing ones manually)" - local ref_secrets=() missing_secrets=() all_secrets s - if [ "${#K8S_MANIFESTS[@]}" -gt 0 ]; then - while IFS= read -r s; do - [ -n "$s" ] && ref_secrets+=("$s") - done < <( - { - grep -h -A1 -E 'secretRef:|secretKeyRef:' "${K8S_MANIFESTS[@]}" 2>/dev/null || true - grep -h -E 'secretName:' "${K8S_MANIFESTS[@]}" 2>/dev/null || true - } | grep -E 'name:' | sed -E 's/.*name:[[:space:]]*//' | tr -d '"'"'"' "'"'" | sed -E 's/[[:space:]]*#.*//' | awk 'NF' | sort -u || true - ) - fi - all_secrets="$(kubectl get secrets -A --no-headers -o custom-columns=:metadata.name 2>/dev/null || true)" - for s in ${ref_secrets[@]+"${ref_secrets[@]}"}; do - if printf '%s\n' "$all_secrets" | grep -qx "$s"; then - echo " ok: $s" - else - echo " MISSING: $s" - missing_secrets+=("$s") - fi + check_referenced_secrets +} + +selected_workload_refs() { + local m k + for m in "${K8S_MANIFESTS[@]}"; do + if skip_uninstalled_vmagent_crd "$m" >/dev/null; then continue; fi + kubectl create --dry-run=client --validate=false -f "$m" -o json | jq -r ' + (if .kind == "List" then .items[] else . end) | select(.kind | test("^(Deployment|StatefulSet|DaemonSet)$")) + | "\(.kind | ascii_downcase) \(.metadata.namespace // "default") \(.metadata.name)"' + done + for k in "${KUSTOMIZE_APPS[@]}"; do + kubectl kustomize "$k" | kubectl create --dry-run=client --validate=false -f - -o json | jq -r ' + (if .kind == "List" then .items[] else . end) | select(.kind | test("^(Deployment|StatefulSet|DaemonSet)$")) + | "\(.kind | ascii_downcase) \(.metadata.namespace // "default") \(.metadata.name)"' done - if [ "${#missing_secrets[@]}" -gt 0 ]; then - echo "ERROR: ${#missing_secrets[@]} referenced Secret(s) not found in the cluster:" - printf ' - %s\n' "${missing_secrets[@]}" - echo "Create them manually from the laptop, e.g.:" - echo " kubectl apply -f SERVICE/k8s/secrets.yaml # see SERVICE/k8s/secrets.yaml.example" - exit 1 - fi } stage_apply_k8s() { + check_prune_mode || return 1 cd "$REPO" select_manifests >/dev/null - local ns_files=() other_files=() m k prune_opts=() + local ns_files=() other_files=() m k for m in ${K8S_MANIFESTS[@]+"${K8S_MANIFESTS[@]}"}; do case "$m" in - */namespace.y?ml) ns_files+=("$m") ;; + */namespace.yaml|*/namespace.yml) ns_files+=("$m") ;; *) other_files+=("$m") ;; esac done - if [ "$APPLY_PRUNE" = "true" ]; then - prune_opts=(--prune -l app.kubernetes.io/managed-by=homelab-deploy) - fi # Record what is about to change, and publish it for the verify job, before # the first apply. Both are fatal on failure: see snapshot_dir. + selected_workload_refs >"$RUN_DIR/workload-refs" local snapshot snapshot="$(snapshot_dir)" || return 1 save_snapshot "$snapshot" || return 1 + touch "$snapshot/ready" if [ "${#ns_files[@]}" -gt 0 ]; then log "Applying namespaces (${#ns_files[@]} files)" for m in "${ns_files[@]}"; do - kubectl apply -f "$m" + record_apply kubectl "${m#"$REPO"/}" started + if ! kubectl apply -f "$m"; then + record_apply kubectl "${m#"$REPO"/}" failure + return 1 + fi + record_apply kubectl "${m#"$REPO"/}" success done fi - if [ -f "$REPO/prometheus-stack/k8s/active" ]; then - if [ ! -f "$REPO/prometheus-stack/k8s/grafana-values.yaml" ]; then + # Kustomize may declare namespaces inside its rendered resources too. + local namespace_resources + namespace_resources="$(render_selected_resources | jq '.items |= map(select(.kind == "Namespace"))')" || return 1 + if [ "$(jq '.items | length' <<<"$namespace_resources")" -gt 0 ]; then + kubectl apply -f - <<<"$namespace_resources" || return 1 + fi + # Complete the deferred server checks before Helm or application resources change. + validate_server_resources false || return 1 + if selected_service k8s prometheus-stack && [ -f "$REPO/prometheus-stack/k8s/active" ]; then + if [ ! -f "$CONFIG_REPO/prometheus-stack/k8s/grafana-values.yaml" ]; then echo "ERROR: prometheus-stack/k8s/grafana-values.yaml (gitignored) missing on workstation, restore it first." exit 1 fi @@ -790,20 +703,26 @@ stage_apply_k8s() { if [ "${#other_files[@]}" -gt 0 ]; then log "Applying resources (${#other_files[@]} files, our images pinned to digests)" for m in "${other_files[@]}"; do - if ! render_pinned <"$m" | kubectl apply "${prune_opts[@]}" -f -; then + log "Applying ${m#"$REPO"/}" + record_apply kubectl "${m#"$REPO"/}" started + if ! render_pinned <"$m" | kubectl apply -f -; then + record_apply kubectl "${m#"$REPO"/}" failure echo "ERROR: apply failed for ${m#"$REPO"/}" >&2 exit 1 fi + record_apply kubectl "${m#"$REPO"/}" success done fi for k in ${KUSTOMIZE_APPS[@]+"${KUSTOMIZE_APPS[@]}"}; do log "Applying kustomize app: ${k#"$REPO"/} (our images pinned to digests)" + record_apply kustomize "${k#"$REPO"/}" started if ! kubectl kustomize "$k" | render_pinned | kubectl apply -f -; then + record_apply kustomize "${k#"$REPO"/}" failure echo "ERROR: apply failed for kustomize app ${k#"$REPO"/}" >&2 exit 1 fi + record_apply kustomize "${k#"$REPO"/}" success done - restart_stale_images # No verification here on purpose. This stage may be killed at any point by # timeout-minutes, by the runner cancelling the job, or by a dropped SSH @@ -818,7 +737,7 @@ stage_apply_k8s() { # and rolls back the ones that never became healthy. stage_verify_k8s() { local pointer="$DEPLOY_SNAPSHOT_DIR/current" - local snapshot want have generations + local snapshot want have generations changed local -a touched=() if [ ! -s "$pointer" ]; then @@ -829,7 +748,7 @@ stage_verify_k8s() { return 1 fi snapshot="$(head -1 "$pointer")" - if [ ! -d "$snapshot" ]; then + if [ ! -d "$snapshot" ] || [ ! -f "$snapshot/ready" ]; then echo "ERROR: snapshot pointer refers to a missing directory: $snapshot" return 1 fi @@ -852,6 +771,13 @@ stage_verify_k8s() { fi echo " snapshot: $snapshot (commit ${have:0:12})" + local entry release chart namespace version values marker + for entry in "${HELM_RELEASES[@]}"; do + IFS='|' read -r release chart namespace version values marker <<<"$entry" + jq -e --arg name "$release" '.helm | index($name) != null' "$DEPLOY_PLAN" >/dev/null || continue + recover_pending_release "$release" "$namespace" || return 1 + done + generations="$snapshot/generations.before" if [ ! -s "$generations" ]; then # Without a baseline we cannot tell which workloads the apply touched, so @@ -860,9 +786,10 @@ stage_verify_k8s() { : >"$generations" fi + changed="$(changed_workloads "$generations")" || return 1 while read -r kind ns name; do [ -n "${kind:-}" ] && touched+=("$kind $ns $name") - done < <(changed_workloads "$generations") + done <<<"$changed" log "Verifying ${#touched[@]} changed workload(s) (timeout ${ROLLOUT_TIMEOUT}s each)" if [ "${#touched[@]}" -eq 0 ]; then @@ -900,25 +827,24 @@ stage_verify_k8s() { # actually be running. verify_compose_stack() { local cf="$1" - local expected running missing=() - expected="$(docker compose -f "$cf" config --services 2>/dev/null | sort || true)" - running="$(docker compose -f "$cf" ps --status running --services 2>/dev/null | sort || true)" + local expected running svc missing=() service_count=0 + expected="$(compose "$cf" config --format json | jq -r ' .services | to_entries[] | select(.value.restart != "no") | .key' | sort)" || return 1 + running="$(compose "$cf" ps --status running --services | sort)" || return 1 [ -n "$expected" ] || return 0 while IFS= read -r svc; do [ -n "$svc" ] || continue + service_count=$((service_count + 1)) # restart:"no" services are allowed to have exited. - if ! printf '%s\n' "$running" | grep -qx "$svc" \ - && ! docker compose -f "$cf" config 2>/dev/null \ - | grep -A5 "^ ${svc}:" | grep -qE 'restart:\s*"?no"?'; then + if ! printf '%s\n' "$running" | grep -qx "$svc"; then missing+=("$svc") fi done <<<"$expected" if [ "${#missing[@]}" -gt 0 ]; then echo " NOT RUNNING: ${missing[*]}" - docker compose -f "$cf" ps --all 2>/dev/null | sed 's/^/ /' || true + compose "$cf" ps --all 2>/dev/null | sed 's/^/ /' || true return 1 fi - echo " all ${#expected} service(s) running" + echo " all $service_count service(s) running" return 0 } @@ -999,10 +925,11 @@ traefik_routed_hosts() { # cases, so ask Traefik which routes it built and fail on the difference. stage_smoke() { cd "$REPO" + if [ -n "${DEPLOY_PLAN:-}" ] && jq -e '.full_smoke' "$DEPLOY_PLAN" >/dev/null; then + DEPLOY_SMOKE_ALL=true + fi select_manifests >/dev/null local -a hosts=() - # Not named failed: an array of that name already exists in restart_stale_images - # above, and a scalar shadowing an array is a trap rather than a shadow. local h code rc bad=0 while IFS= read -r h; do [ -n "$h" ] && hosts+=("$h") @@ -1010,8 +937,8 @@ stage_smoke() { if [ "${#hosts[@]}" -eq 0 ]; then # Nothing to probe means the extraction broke, not that the cluster is empty. - echo "ERROR: no public hostnames found in active manifests, refusing to report success" - return 1 + echo "No public routes in the selected components" + return 0 fi log "Probing ${#hosts[@]} public route(s)" @@ -1095,37 +1022,22 @@ stage_apply_compose() { cd "$REPO" select_manifests >/dev/null local cf - log "Redeploying docker compose stacks (${#COMPOSE_STACKS[@]} stacks)" - for cf in ${COMPOSE_STACKS[@]+"${COMPOSE_STACKS[@]}"}; do - echo " compose: $cf" - if grep -Eq '^\s+pull_policy:\s*build\b' "$cf"; then - docker compose -f "$cf" build - docker compose -f "$cf" push + for cf in "${COMPOSE_STACKS[@]}"; do + log "Applying Compose ${cf#"$REPO"/}" + record_apply compose "${cf#"$REPO"/}" started + if ! compose "$cf" up -d --wait --wait-timeout 180 --pull missing --remove-orphans; then + record_apply compose "${cf#"$REPO"/}" failure + return 1 fi - docker compose -f "$cf" up -d --pull always --remove-orphans + record_apply compose "${cf#"$REPO"/}" success + verify_compose_stack "$cf" done - - local -a broken=() - for cf in ${COMPOSE_STACKS[@]+"${COMPOSE_STACKS[@]}"}; do - echo " verifying: $cf" - if ! verify_compose_stack "$cf"; then - broken+=("$cf") - fi - done - if [ "${#broken[@]}" -gt 0 ]; then - echo - echo "ERROR: ${#broken[@]} compose stack(s) did not come up:" - printf ' - %s\n' "${broken[@]}" - echo "Compose stacks are not rolled back automatically: their images use mutable" - echo "':latest' tags, so there is no previous version to return to. Check the logs" - echo "above, then re-run the deploy once the cause is fixed." - return 1 - fi + echo "Compose recovery files: $RUN_DIR/compose-before (manual recovery only)" } run_stage() { case "${1:?stage required}" in - preflight) stage_preflight ;; + doctor) stage_doctor ;; validate) stage_validate ;; apply-k8s) stage_apply_k8s ;; verify-k8s) stage_verify_k8s ;; diff --git a/.gitea/workflows/deploy-plan.py b/.gitea/workflows/deploy-plan.py new file mode 100644 index 0000000..f3a8572 --- /dev/null +++ b/.gitea/workflows/deploy-plan.py @@ -0,0 +1,124 @@ +#!/usr/bin/env python3 +"""Calculate selected components against the last fully successful deploy.""" + +import hashlib +import json +import re +import subprocess +from pathlib import Path + + +def output(*args, **kwargs): + return subprocess.check_output(args, text=True, **kwargs).strip() # noqa: S603 + + +def tracked(repo): + return output('git', '-C', str(repo), 'ls-files').splitlines() + + +def helm_releases(repo): + text = (repo / '.gitea/workflows/deploy-lib.sh').read_text() + return [line.split('|') for line in re.findall(r'^ "([^"\n]+\|[^"\n]+)"$', text, re.MULTILINE)] + + +def inventory(repo): + files = tracked(repo) + k8s = sorted( + {f.split('/k8s/')[0] for f in files if '/k8s/' in f and (repo / f.split('/k8s/')[0] / 'k8s/active').is_file()} + ) + compose = sorted( + { + str(Path(f).parent) + for f in files + if Path(f).name in ('compose.yaml', 'compose.yml') and (repo / Path(f).parent / 'active').is_file() + } + ) + return {'k8s': k8s, 'compose': compose} + + +def file_hash(path): + return hashlib.sha256(path.read_bytes()).hexdigest() if path.is_file() else 'missing' + + +def make_plan(repo, config_repo, release, previous, mode, live_helm): + active = inventory(repo) + all_services = set(active['k8s'] + active['compose']) + helm_inputs = {} + helm_selected = [] + for name, chart, namespace, version, values, marker in helm_releases(repo): + if not (repo / marker).is_file(): + continue + value_path = repo / values if (repo / values).is_file() else config_repo / values + if not value_path.is_file(): + raise ValueError(f'Missing Helm values: {values}') + stamp = hashlib.sha256(f'{chart}|{version}|{file_hash(value_path)}'.encode()).hexdigest() + helm_inputs[name] = stamp + live = next((h for h in live_helm if h['name'] == name and h['namespace'] == namespace), None) + if ( + mode == 'full' + or previous is None + or previous.get('helm_inputs', {}).get(name) != stamp + or live is None + or live.get('status') != 'deployed' + or live.get('chart') != f'{chart.split("/")[-1]}-{version}' + ): + helm_selected.append(name) + local_inputs = {} + for service in all_services: + candidates = [config_repo / service / '.env'] + if service in active['compose']: + candidates.append(config_repo / '.env') + cfg = config_repo / service / 'config' + if cfg.is_dir(): + candidates.extend( + p for p in cfg.rglob('*') if p.is_file() and p.suffix in ('.yaml', '.yml', '.json', '.conf') + ) + local_inputs[service] = hashlib.sha256( + '\n'.join(f'{p.relative_to(config_repo)}:{file_hash(p)}' for p in sorted(candidates)).encode() + ).hexdigest() + if previous is None: + if mode == 'changed': + raise ValueError('No successful baseline; run deploy in full mode first') + changed = set(all_services) + removed = [] + else: + paths = output('git', '-C', str(repo), 'diff', '--name-only', previous['sha'], release['sha']).splitlines() + changed = {service for service in all_services for path in paths if path.startswith(service + '/')} + if any(path.startswith('.gitea/') for path in paths): + changed |= all_services + changed |= {s for s in all_services if previous.get('local_inputs', {}).get(s) != local_inputs[s]} + for file in tracked(repo): + owners = {service for service in all_services if file.startswith(service + '/')} + if not owners or not file.endswith(('.yaml', '.yml')): + continue + text = (repo / file).read_text() + if any( + image in text and previous.get('images', {}).get(image) != digest + for image, digest in release['images'].items() + ): + changed |= owners + removed = sorted( + set(previous.get('active', {}).get('k8s', []) + previous.get('active', {}).get('compose', [])) + - all_services + ) + removed += [path for path in paths if '/k8s/' in path and not (repo / path).exists()] + if mode == 'full': + changed = set(all_services) + dependencies = json.loads((repo / '.gitea/deploy-dependencies.json').read_text()) + while True: + expanded = changed | {dependent for service in changed for dependent in dependencies.get(service, [])} + if expanded == changed: + break + changed = expanded + return { + 'version': 1, + 'sha': release['sha'], + 'images': release['images'], + 'active': active, + 'selected': {kind: sorted(set(services) & changed) for kind, services in active.items()}, + 'helm': helm_selected, + 'helm_inputs': helm_inputs, + 'local_inputs': local_inputs, + 'removed': sorted(set(removed)), + 'full_smoke': mode == 'full' or 'traefik' in changed, + } diff --git a/.gitea/workflows/deploy-stage.sh b/.gitea/workflows/deploy-stage.sh new file mode 100755 index 0000000..6d630f3 --- /dev/null +++ b/.gitea/workflows/deploy-stage.sh @@ -0,0 +1,10 @@ +#!/usr/bin/env bash +set -euo pipefail +source "${REPO:?}/.gitea/workflows/deploy-lib.sh" +case "${1:?stage required}" in + workload-count) + select_manifests >/dev/null + selected_workload_refs | sort -u | wc -l + ;; + *) run_stage "$1" ;; +esac diff --git a/.gitea/workflows/deploy.yaml b/.gitea/workflows/deploy.yaml index 159ba25..0a55a99 100644 --- a/.gitea/workflows/deploy.yaml +++ b/.gitea/workflows/deploy.yaml @@ -1,197 +1,140 @@ name: deploy on: - # Deploy only what CI already validated. workflow_run is used instead of - # workflow_dispatch so a red lint/validate run can never reach the cluster. workflow_run: workflows: [ci] + branches: [main] types: [completed] workflow_dispatch: + inputs: + deploy_ref: + description: "Commit already checked by successful main CI (main or SHA)" + default: main + required: true + deploy_mode: + description: "First deploy requires full; plan changes no production resources" + type: choice + options: [changed, full, plan] + default: changed + refresh_images: + description: "Explicitly refresh mutable third-party Compose tags" + type: boolean + default: false -# The deploy jobs read the tree, then reach the cluster over SSH with the -# deploy key. The Actions token itself is not part of that path, so it gets -# read-only contents and no more. permissions: contents: read + actions: read concurrency: group: deploy-main - # Queue instead of cancelling. Cancelling a run kills the apply job mid-loop and - # takes the verify job down with it, so a superseded deploy would leave the - # cluster half-applied and unchecked — the exact failure the verify job exists - # to catch. kubectl apply and docker compose up are both idempotent, so letting - # the older run finish and then deploying the newer commit costs little. cancel-in-progress: false env: - DEPLOY_HOST: ${{ secrets.DEPLOY_HOST }} - DEPLOY_PORT: ${{ secrets.DEPLOY_PORT }} - DEPLOY_USER: ${{ secrets.DEPLOY_USER }} - DEPLOY_PATH: ${{ secrets.DEPLOY_PATH }} + DEPLOY_HOST: ${{ vars.DEPLOY_HOST || secrets.DEPLOY_HOST }} + DEPLOY_PORT: ${{ vars.DEPLOY_PORT || secrets.DEPLOY_PORT }} + DEPLOY_USER: ${{ vars.DEPLOY_USER || secrets.DEPLOY_USER }} DEPLOY_KEY: ${{ secrets.DEPLOY_SSH_KEY }} - APPLY_PRUNE: ${{ vars.APPLY_PRUNE }} - # workflow_run's own GITHUB_SHA points at the branch head, not at the commit the - # finished ci run checked. Pin the exact validated commit instead, so a push - # landing mid-deploy cannot make the workstation deploy something else. Also - # what the verify job checks the snapshot against. Empty for workflow_dispatch, - # which falls back to the current origin/main. - DEPLOY_SHA: ${{ github.event.workflow_run.head_sha }} + DEPLOY_KNOWN_HOSTS: ${{ vars.DEPLOY_KNOWN_HOSTS }} + DEPLOY_RUN_ID: ${{ github.run_id }}-${{ github.run_attempt || 1 }} + DEPLOY_MODE: ${{ inputs.deploy_mode || 'changed' }} + REFRESH_IMAGES: ${{ inputs.refresh_images && 'true' || 'false' }} jobs: - preflight: - # Autodeploy defaults to OFF: pushes deploy only when the AUTODEPLOY repo - # variable is set to 'true' (Settings -> Actions -> Variables). A manual - # Run workflow always bypasses the switch: dispatching it is the explicit - # intent to deploy. + gate: if: >- + github.ref == 'refs/heads/main' && (vars.AUTODEPLOY == 'true' || github.event_name == 'workflow_dispatch') && (github.event_name != 'workflow_run' || - (github.event.workflow_run.conclusion == 'success' && - github.event.workflow_run.head_branch == 'main')) - runs-on: [self-hosted, linux, arch, homelab, prod] + (github.event.workflow_run.conclusion == 'success' && github.event.workflow_run.head_branch == 'main')) + runs-on: homelab timeout-minutes: 10 + outputs: + sha: ${{ steps.release.outputs.sha }} steps: - name: Checkout repository uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4 - - - name: Fetch and reset workstation - shell: bash + with: + fetch-depth: 0 + - name: Check successful CI and download the exact commit release + id: release + env: + GITEA_TOKEN: ${{ github.token }} + DEPLOY_REF: ${{ inputs.deploy_ref || 'main' }} + EVENT_SHA: ${{ github.event.workflow_run.head_sha }} + run: python3 .gitea/workflows/release.py gate --ref "$DEPLOY_REF" --event-sha "$EVENT_SHA" + - name: Submit durable deploy to workstation + run: bash .gitea/workflows/ssh-run.sh start + - name: Write the request result + if: always() + env: + REQUEST_RESULT: ${{ job.status }} + CHECKED_SHA: ${{ steps.release.outputs.sha }} run: | - set -euo pipefail - ./.gitea/workflows/ssh-run.sh preflight + if [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then + printf '## Deploy request\n\n- Result: **%s**\n- Checked commit: %s\n- Mode: %s\n' "$REQUEST_RESULT" "${CHECKED_SHA:-not checked}" "$DEPLOY_MODE" >>"$GITHUB_STEP_SUMMARY" + if [ "$REQUEST_RESULT" != success ]; then + echo 'Open the failed step log. If SSH submission failed, check the remote controller state.' >>"$GITHUB_STEP_SUMMARY" + fi + fi - validate: - needs: [preflight] - runs-on: [self-hosted, linux, arch, homelab, prod] - timeout-minutes: 20 + apply: + needs: [gate] + runs-on: homelab + timeout-minutes: 120 steps: - - name: Checkout repository + - name: Checkout checked commit uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4 - - - name: Dry-run manifests and check Secrets - shell: bash + with: + ref: ${{ needs.gate.outputs.sha }} + - name: Follow validation and sequential Kubernetes / Compose apply + run: bash .gitea/workflows/ssh-run.sh apply + - name: Write the deploy result + if: always() run: | - set -euo pipefail - ./.gitea/workflows/ssh-run.sh validate + if [ -f .gitea/workflows/ssh-run.sh ]; then + bash .gitea/workflows/ssh-run.sh summary + elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then + echo 'Source checkout failed. The remote deploy state is unknown. Check the job log.' >>"$GITHUB_STEP_SUMMARY" + fi - apply-k8s: - needs: [validate] - runs-on: [self-hosted, linux, arch, homelab, prod] - # Apply only, no verification, so this is just the work itself: snapshot, - # then sequential `helm upgrade --install --wait --rollback-on-failure --timeout 10m`, then the apply loop. - # Verification has its own job and its own budget. - # - # 45 is roughly four times the measured cost of the stage, which is - # deliberately not raised on a theory: - # - # helm, healthy 3 no-op upgrades ~3-5 min - # helm, one release bad rollback-on-failure spends its 10m, ~10-15 min - # then rolls that one back - # apply loop ~40 manifests, 4 of which ~1 min - # resolve an image digest - # restart_stale_images 7.6s to find 8 workloads, ~0.5 min - # 9.8s to resolve their digests - # - # The helm figure is one release, not three: `set -e` aborts - # upgrade_helm_releases on the first failure, so a broken release costs - # 10m and the other two are never attempted. Multiplying 10m by three - # overstates the worst case by 20 minutes. - # - # The 45 minutes this was last raised to 45 were still not enough, and the - # job logs for those runs no longer exist, so what actually consumed the - # budget is not known - the two measurable candidates above account for - # ~15 of it. The unbounded `docker manifest inspect` against the registry's - # known hang mode is now bounded inside registry_digest (25s timeout, 3 - # attempts): a dead registry fails each owned image after ~85s instead of - # hanging the stage, and a blinking one is retried instead of failing the - # whole apply file. Still open: make the stage announce which manifest it - # is working on, so a killed run leaves a diagnosable last line. - timeout-minutes: 45 + verify: + needs: [gate, apply] + if: always() && needs.gate.result == 'success' + runs-on: homelab + timeout-minutes: 130 steps: - - name: Checkout repository + - name: Checkout checked commit uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4 - - - name: Apply Kubernetes manifests - shell: bash + with: + ref: ${{ needs.gate.outputs.sha }} + - name: Follow workload verification and recovery + run: bash .gitea/workflows/ssh-run.sh verify + - name: Write the deploy result + if: always() run: | - set -euo pipefail - ./.gitea/workflows/ssh-run.sh apply-k8s + if [ -f .gitea/workflows/ssh-run.sh ]; then + bash .gitea/workflows/ssh-run.sh summary + elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then + echo 'Source checkout failed. The remote deploy state is unknown. Check the job log.' >>"$GITHUB_STEP_SUMMARY" + fi - apply-compose: - needs: [validate] - runs-on: [self-hosted, linux, arch, homelab, prod] - timeout-minutes: 30 - steps: - - name: Checkout repository - uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4 - - - name: Redeploy docker compose stacks - shell: bash - run: | - set -euo pipefail - ./.gitea/workflows/ssh-run.sh apply-compose - - # Watches the workloads this deploy changed and rolls back the ones that never - # became healthy. Runs even when the apply jobs failed, timed out or were - # cancelled — that is the whole point of splitting it out. `always()` is what - # lets it start after a failed dependency; the needs on apply-compose are a - # barrier, so verification begins only once both applies are done. - verify-k8s: - needs: [apply-k8s, apply-compose] - if: >- - always() && - needs.apply-k8s.result != 'skipped' && - needs.apply-compose.result != 'skipped' - runs-on: [self-hosted, linux, arch, homelab, prod] - # Not raised, because the arithmetic does not close. - # - # 32 workloads are under management and the wave width is 8, so the verify - # itself is 4 waves of ROLLOUT_TIMEOUT (300s) = 20 minutes worst case, when - # every rollout times out rather than converging. That is already 20 of 30. - # - # The other 10 would have to absorb rollback, and rollback_workloads is a - # serial `while read` loop at 300s per failed workload. 10 minutes buys two. - # Any larger number is buying a bigger multiple of an unbounded term rather - # than covering a known cost: 60 minutes buys eight, and 60 minutes is - # therefore not a bound, it is a guess with two digits. - # - # The number becomes derivable the moment rollback uses the same wave width - # as the verify: 32 failures then cost 4 waves = 20 minutes instead of 160, - # and 45 covers verify plus rollback at full width. That change is to the - # recovery path and is not folded into a timeout edit. - timeout-minutes: 30 - steps: - - name: Checkout repository - uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4 - - - name: Verify workloads and roll back on failure - shell: bash - run: | - set -euo pipefail - ./.gitea/workflows/ssh-run.sh verify-k8s - - # Asks the public route of every active service whether it is actually - # serving, which the rollout check above structurally cannot: a pod can - # converge and still be crash-looping, or be listening on a port no Service - # points at, or answer 500. - # - # `always()` for the same reason verify-k8s has it, and it runs after that job - # specifically because a rollback is when a route most needs re-checking. The - # needs is a barrier, not a filter: whether verify-k8s passed, failed or was - # cancelled, the probes are what say whether the cluster is serving, and - # suppressing them on a rollback would hide the one run where the answer - # matters most. smoke: - needs: [verify-k8s] - if: always() && needs.verify-k8s.result != 'skipped' - runs-on: [self-hosted, linux, arch, homelab, prod] - timeout-minutes: 10 + needs: [gate, verify] + if: always() && needs.gate.result == 'success' + runs-on: homelab + timeout-minutes: 15 steps: - - name: Checkout repository + - name: Checkout checked commit uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4 - - - name: Probe the public route of every active service - shell: bash + with: + ref: ${{ needs.gate.outputs.sha }} + - name: Follow public route checks + run: bash .gitea/workflows/ssh-run.sh smoke + - name: Write the deploy result + if: always() run: | - set -euo pipefail - ./.gitea/workflows/ssh-run.sh smoke + if [ -f .gitea/workflows/ssh-run.sh ]; then + bash .gitea/workflows/ssh-run.sh summary + elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then + echo 'Source checkout failed. The remote deploy state is unknown. Check the job log.' >>"$GITHUB_STEP_SUMMARY" + fi diff --git a/.gitea/workflows/install-ci-tools.sh b/.gitea/workflows/install-ci-tools.sh index 988fe64..60c2d87 100755 --- a/.gitea/workflows/install-ci-tools.sh +++ b/.gitea/workflows/install-ci-tools.sh @@ -13,9 +13,14 @@ here="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" # shellcheck source=tool-versions.env . "$here/tool-versions.env" -TOOLS_DIR="${TOOLS_DIR:-${RUNNER_TEMP:-/tmp}/homelab-tools}" +TOOLS_DIR="${TOOLS_DIR:-${XDG_CACHE_HOME:-$HOME/.cache}/homelab-ci}" BIN_DIR="$TOOLS_DIR/bin" mkdir -p "$BIN_DIR" +# A runner may accept overlapping workflows even though each workflow is sequential. +exec 9>"$TOOLS_DIR/install.lock" +flock -w 300 9 +export UV_TOOL_DIR="$TOOLS_DIR/uv-tools" +export UV_CACHE_DIR="$TOOLS_DIR/uv-cache" # The just-installed tools must resolve inside this script too: callers only # prepend BIN_DIR to PATH after the script exits, so a bare `uv` below would # miss the binary install_uv just placed (exit 127 on a clean runner). @@ -48,7 +53,7 @@ esac fetch() { # fetch if command -v curl >/dev/null 2>&1; then - curl -sSLf --retry 3 -o "$2" "$1" + curl -sSLf --connect-timeout 15 --max-time 120 --retry 3 -o "$2" "$1" elif command -v wget >/dev/null 2>&1; then wget -q -O "$2" "$1" else @@ -88,10 +93,13 @@ installed_version() { # at_version at_version() { - case "$(installed_version "$1")" in - *"$2"*) return 0 ;; - *) return 1 ;; - esac + local version expected="${2#v}" + version="$(installed_version "$1")" + if [[ "$version" =~ (^|[^0-9.])v?([0-9]+(\.[0-9]+)+) ]]; then + [ "${BASH_REMATCH[2]}" = "$expected" ] + else + return 1 + fi } install_kubeconform() { @@ -120,6 +128,15 @@ install_shellcheck() { rm -rf "$tmp" } +install_jq() { + if at_version jq "${JQ_VERSION}"; then + return 0 + fi + fetch "https://github.com/jqlang/jq/releases/download/jq-${JQ_VERSION}/jq-linux-${goarch}" \ + "$BIN_DIR/jq" + chmod 0755 "$BIN_DIR/jq" +} + install_uv() { if at_version uv "${UV_VERSION}"; then return 0 @@ -167,6 +184,7 @@ install_pip_audit() { } install_prettier() { + install_node if at_version prettier "${PRETTIER_VERSION}"; then return 0 fi @@ -227,28 +245,41 @@ install_actionlint() { rm -rf "$tmp" } -wanted=("$@") -if [ "${#wanted[@]}" -eq 0 ]; then - wanted=(kubeconform shellcheck actionlint prettier ruff yamllint hadolint) -fi +main() { + wanted=("$@") + if [ "${#wanted[@]}" -eq 0 ]; then + wanted=(node jq kubeconform shellcheck actionlint prettier ruff yamllint hadolint) + fi -for tool in "${wanted[@]}"; do - case "$tool" in - kubeconform) install_kubeconform ;; - shellcheck) install_shellcheck ;; - actionlint) install_actionlint ;; - prettier) install_prettier ;; - ruff) install_ruff ;; - yamllint) install_yamllint ;; - pip-audit) install_pip_audit ;; - hadolint) install_hadolint ;; - node) install_node ;; - uv) install_uv ;; - *) - echo "install-ci-tools: unknown tool: $tool" >&2 - exit 1 - ;; - esac -done + for tool in "${wanted[@]}"; do + case "$tool" in + kubeconform) install_kubeconform ;; + shellcheck) install_shellcheck ;; + jq) install_jq ;; + actionlint) install_actionlint ;; + prettier) install_prettier ;; + ruff) install_ruff ;; + yamllint) install_yamllint ;; + pip-audit) install_pip_audit ;; + hadolint) install_hadolint ;; + node) install_node ;; + uv) install_uv ;; + *) + echo "install-ci-tools: unknown tool: $tool" >&2 + exit 1 + ;; + esac + done -printf '%s\n' "$BIN_DIR" + for old in "$BIN_DIR"/node-* "$BIN_DIR"/prettier-*; do + [ -d "$old" ] || continue + case "$(basename "$old")" in + "node-$NODE_VERSION"|"prettier-$PRETTIER_VERSION") ;; + *) rm -rf "$old" ;; + esac + done + if [ -x "$BIN_DIR/uv" ]; then "$BIN_DIR/uv" cache prune >/dev/null; fi + printf '%s\n' "$BIN_DIR" +} + +if [ "${BASH_SOURCE[0]}" = "$0" ]; then main "$@"; fi diff --git a/.gitea/workflows/release.py b/.gitea/workflows/release.py new file mode 100644 index 0000000..c6b1068 --- /dev/null +++ b/.gitea/workflows/release.py @@ -0,0 +1,519 @@ +#!/usr/bin/env python3 +"""CI release artifacts and the SHA-specific Gitea deployment gate (stdlib only).""" + +import argparse +import hashlib +import io +import itertools +import json +import os +import re +import shutil +import subprocess +import sys +import tempfile +import urllib.error +import urllib.parse +import urllib.request +import zipfile +from pathlib import Path + +SHA = re.compile(r'[0-9a-f]{40}') +DIGEST = re.compile(r'sha256:[0-9a-f]{64}') +IMAGES = { + 'error-pages': ('errorpages', 'errorpages/Dockerfile'), + 'forust-homepage': ('homepages', 'homepages/Dockerfile.forust'), + 'xdfnx-homepage': ('homepages', 'homepages/Dockerfile.xdfnx'), +} + + +def command(*args, **kwargs): + """Arguments are passed directly to the executable, never to a shell.""" + return subprocess.check_output(args, text=True, **kwargs).strip() # noqa: S603, S607 + + +def validate_release(data, sha=None): + if data.get('version') != 1 or not SHA.fullmatch(data.get('sha', '')): + raise ValueError('Invalid release version or SHA') + if sha is not None and data['sha'] != sha: + raise ValueError('Release SHA does not match the checked CI commit') + expected = {f'gcr.forust.xyz/forust/{name}' for name in IMAGES} + if set(data.get('images', {})) != expected: + raise ValueError('Release must contain all owned images') + if not all(DIGEST.fullmatch(value) for value in data['images'].values()): + raise ValueError('Release has an invalid image digest') + if set(data.get('inputs', {})) != expected or not all( + re.fullmatch(r'[0-9a-f]{64}', value) for value in data['inputs'].values() + ): + raise ValueError('Release has invalid build input fingerprints') + return data + + +class NoRedirect(urllib.request.HTTPRedirectHandler): + def redirect_request(self, _req, _fp, _code, _msg, _headers, _newurl): + return None + + +class Gitea: + def __init__(self): + self.origin = os.environ['GITHUB_SERVER_URL'].rstrip('/') + if urllib.parse.urlsplit(self.origin).scheme != 'https': + raise ValueError('Gitea API must use HTTPS') + self.repository = os.environ['GITHUB_REPOSITORY'] + if not re.fullmatch(r'[\w.-]+/[\w.-]+', self.repository): + raise ValueError('Invalid Gitea repository') + self.token = os.environ['GITEA_TOKEN'] + self.base = f'{self.origin}/api/v1/repos/{self.repository}' + + def request(self, url, *, archive=False): + if not url.startswith(self.base + '/'): + raise ValueError('Refusing to send the Actions token to another origin') + req = urllib.request.Request(url, headers={'Authorization': f'token {self.token}'}) # noqa: S310 -- HTTPS origin validated above + opener = urllib.request.build_opener(NoRedirect()) + try: + response = opener.open(req, timeout=30) # noqa: S310 + except urllib.error.HTTPError as error: + if not archive or error.code not in (301, 302, 303, 307, 308): + raise RuntimeError(f'Gitea API returned HTTP {error.code}') from None + target = urllib.parse.urljoin(url, error.headers['Location']) + if urllib.parse.urlsplit(target).scheme != 'https': + raise ValueError('Artifact redirect must use HTTPS') from None + # Signed storage redirects must never receive the Gitea token. + response = urllib.request.urlopen(target, timeout=30) # noqa: S310 + with response: + payload = response.read(8 * 1024 * 1024 + 1) + if len(payload) > 8 * 1024 * 1024: + raise ValueError('Gitea response exceeds 8 MiB') + return payload if archive else json.loads(payload) + + def pages(self, path, key, **params): + for page in range(1, 101): + query = urllib.parse.urlencode({**params, 'page': page, 'limit': 50}) + data = self.request(f'{self.base}/{path}?{query}') + entries = data[key] + yield from entries + if len(entries) < 50: + return + raise RuntimeError('Gitea pagination limit exceeded') + + def successful_runs(self, sha=None): + params = {'branch': 'main', 'status': 'success', 'exclude_pull_requests': 'true'} + if sha: + params['head_sha'] = sha + for run in self.pages('actions/workflows/ci.yaml/runs', 'workflow_runs', **params): + if ( + run.get('status') == 'completed' + and run.get('conclusion') == 'success' + and run.get('head_branch') == 'main' + and run.get('event') in ('push', 'workflow_dispatch') + and (run.get('repository') or {}).get('full_name') == self.repository + and (run.get('head_repository') or run.get('repository') or {}).get('full_name') == self.repository + and (sha is None or run.get('head_sha') == sha) + ): + yield run + + def release(self, run): + sha = run['head_sha'] + jobs = list(self.pages(f'actions/runs/{run["id"]}/jobs', 'jobs')) + # A green workflow with a skipped build must not authorize a deploy. + if not any(job.get('name') == 'build' and job.get('conclusion') == 'success' for job in jobs): + raise ValueError('CI build job did not succeed') + artifacts = self.request(f'{self.base}/actions/runs/{run["id"]}/artifacts')['artifacts'] + matching = [a for a in artifacts if a['name'] == f'release-{sha}' and not a.get('expired')] + if len(matching) != 1: + raise ValueError('CI release artifact is missing, expired or ambiguous; rerun CI') + blob = self.request(f'{self.base}/actions/artifacts/{matching[0]["id"]}/zip', archive=True) + with zipfile.ZipFile(io.BytesIO(blob)) as archive: + files = [entry for entry in archive.infolist() if not entry.is_dir()] + if len(files) != 1 or files[0].filename != 'release.json' or files[0].file_size > 256 * 1024: + raise ValueError('Unexpected release archive contents') + return validate_release(json.loads(archive.read(files[0])), sha) + + +def fingerprint(context, dockerfile): + tree = command('git', 'ls-tree', '-r', 'HEAD', '--', context, dockerfile, '.gitea/workflows/release.py') + return hashlib.sha256(tree.encode()).hexdigest() + + +def gate(output, requested_ref, event_sha): + command('git', 'fetch', '--quiet', 'origin', 'main') + if event_sha: + if not SHA.fullmatch(event_sha): + raise ValueError('Invalid workflow_run SHA') + sha = event_sha + else: + if requested_ref == 'main': + requested_ref = 'origin/main' + sha = command('git', 'rev-parse', '--verify', '--end-of-options', f'{requested_ref}^{{commit}}') + if not SHA.fullmatch(sha): + raise ValueError('Invalid deploy SHA') + command('git', 'merge-base', '--is-ancestor', sha, 'origin/main') + api = Gitea() + runs = list(api.successful_runs(sha)) + if not runs: + raise ValueError(f'No successful main CI for {sha}; run CI before deploying') + release = api.release(max(runs, key=lambda run: run['id'])) + output.write_text(json.dumps(release, indent=2) + '\n') + if os.environ.get('GITHUB_OUTPUT'): + with Path(os.environ['GITHUB_OUTPUT']).open('a') as stream: + stream.write(f'sha={sha}\n') + print(f'CI gate accepted {sha}') + + +def prepare_images(output): + sha = command('git', 'rev-parse', 'HEAD') + if sha != os.environ['GITHUB_SHA'] or not SHA.fullmatch(sha): + raise ValueError('Build checkout does not match GITHUB_SHA') + api = Gitea() + previous = None + for run in sorted(itertools.islice(api.successful_runs(), 50), key=lambda item: item['id'], reverse=True): + if str(run['id']) == os.environ.get('GITHUB_RUN_ID'): + continue + try: + previous = api.release(run) + break + except ValueError: + # Expired artifacts only cost a rebuild; mutable tags are never a fallback. + continue + targets = [] + for name, (context, dockerfile) in IMAGES.items(): + image = f'gcr.forust.xyz/forust/{name}' + inputs = fingerprint(context, dockerfile) + old_digest = (previous or {}).get('images', {}).get(image) + targets.append( + { + 'name': name, + 'image': image, + 'context': context, + 'dockerfile': dockerfile, + 'inputs': inputs, + 'reuse_digest': old_digest if (previous or {}).get('inputs', {}).get(image) == inputs else None, + } + ) + output.write_text(json.dumps({'sha': sha, 'targets': targets}, indent=2) + '\n') + if os.environ.get('GITHUB_OUTPUT'): + with Path(os.environ['GITHUB_OUTPUT']).open('a') as stream: + stream.write('matrix=' + json.dumps({'include': targets}, separators=(',', ':')) + '\n') + print(f'Prepared {len(targets)} image jobs; {sum(t["reuse_digest"] is None for t in targets)} require builds') + + +def checked_plan(path): + data = json.loads(path.read_text()) + sha = command('git', 'rev-parse', 'HEAD') + if data.get('sha') != sha or sha != os.environ['GITHUB_SHA'] or not SHA.fullmatch(sha): + raise ValueError('Image plan does not match the checked source commit') + targets = data.get('targets', []) + if sorted(t['name'] for t in targets) != sorted(IMAGES): + raise ValueError('Image plan must contain each owned image once') + for target in targets: + name = target['name'] + context, dockerfile = IMAGES[name] + if (target['context'], target['dockerfile'], target['image']) != ( + context, + dockerfile, + f'gcr.forust.xyz/forust/{name}', + ) or target['inputs'] != fingerprint(context, dockerfile): + raise ValueError('Image plan has invalid build inputs') + if target['reuse_digest'] is not None and not DIGEST.fullmatch(target['reuse_digest']): + raise ValueError('Image plan has an invalid reuse digest') + return data + + +def build_images(output, report, name, plan): + data = checked_plan(plan) + sha = data['sha'] + target = next(t for t in data['targets'] if t['name'] == name) + context, dockerfile = IMAGES[name] + docker_config = tempfile.mkdtemp(prefix='homelab-registry-') + builder_config = Path.home() / '.cache/homelab-ci/buildx' + builder_config.mkdir(parents=True, exist_ok=True) + env = {**os.environ, 'DOCKER_CONFIG': docker_config, 'BUILDX_CONFIG': str(builder_config)} + try: + report['phase'] = 'Registry login' + subprocess.run( # noqa: S603, S607 + [ + shutil.which('docker') or '/usr/bin/docker', + 'login', + 'gcr.forust.xyz', + '-u', + os.environ['REGISTRY_USERNAME'], + '--password-stdin', + ], + input=os.environ['REGISTRY_PASSWORD'], + text=True, + check=True, + env=env, + ) + report['phase'] = 'Prepare the builder' + builder = 'homelab-ci' + versions = dict( + re.findall(r'^([A-Z_]+)="([^"\n]+)"$', Path('.gitea/workflows/tool-versions.env').read_text(), re.MULTILINE) + ) + image = versions['BUILDKIT_IMAGE'] + signature = builder_config / 'homelab-ci-image' + exists = ( + subprocess.run( # noqa: S603 + [shutil.which('docker') or '/usr/bin/docker', 'buildx', 'inspect', builder], + capture_output=True, + env=env, + ).returncode + == 0 + ) + if exists and (not signature.exists() or signature.read_text().strip() != image): + command('docker', 'buildx', 'rm', '--keep-state', builder, env=env) + exists = False + if not exists: + command( + 'docker', + 'buildx', + 'create', + '--name', + builder, + '--driver', + 'docker-container', + '--driver-opt', + f'image={image}', + '--buildkitd-config', + '.gitea/runner/buildkitd.toml', + env=env, + ) + signature.write_text(image + '\n') + release = {'version': 1, 'sha': sha, 'images': {}, 'inputs': {}} + report['images'] = release['images'] + report['phase'] = f'Build or reuse {name}' + report['current'] = name + image = f'gcr.forust.xyz/forust/{name}' + inputs = target['inputs'] + old_digest = target['reuse_digest'] + exists = False + if old_digest: + exists = ( + subprocess.run( # noqa: S603, S607 + [ + shutil.which('docker') or '/usr/bin/docker', + 'buildx', + 'imagetools', + 'inspect', + f'{image}@{old_digest}', + ], + capture_output=True, + env=env, + timeout=60, + ).returncode + == 0 + ) + if exists: + print(f'Reuse {name}: inputs unchanged') + digest = old_digest + else: + print(f'Build {name}', flush=True) + metadata = Path(docker_config) / 'metadata.json' + command( + 'docker', + 'buildx', + 'build', + '--builder', + builder, + '--platform', + 'linux/amd64', + '--provenance=false', + '--cache-from', + f'type=registry,ref={image}:buildcache', + '--cache-to', + f'type=registry,ref={image}:buildcache,mode=max', + '--output', + f'type=image,name={image},push-by-digest=true,name-canonical=true,push=true', + '--metadata-file', + str(metadata), + '--file', + dockerfile, + context, + env=env, + ) + digest = json.loads(metadata.read_text())['containerimage.digest'] + if not isinstance(digest, str) or not DIGEST.fullmatch(digest): + raise ValueError('Image job returned an invalid digest') + release['images'][image] = digest + release['inputs'][image] = inputs + report['reused' if exists else 'built'].append(name) + output.write_text(json.dumps(release, indent=2) + '\n') + report['current'] = None + report['phase'] = 'Image result file saved' + finally: + # Cleanup errors must neither leak credentials nor mask the original build error. + try: + subprocess.run( # noqa: S603 + [ + shutil.which('docker') or '/usr/bin/docker', + 'buildx', + 'prune', + '--builder', + 'homelab-ci', + '--force', + '--max-used-space', + '1gb', + ], + env=env, + timeout=60, + ) + except (OSError, subprocess.TimeoutExpired): + print('CI builder cache cleanup deferred', flush=True) + finally: + shutil.rmtree(docker_config) + + +def write_summary(lines): + path = os.environ.get('GITHUB_STEP_SUMMARY') + if path: + try: + with Path(path).open('a') as stream: + stream.write('\n'.join(lines) + '\n\n') + except OSError: + print('WARNING: cannot write the job summary') + + +def check_summary(): + lines = [ + f'## {os.environ["SUMMARY_CHECK"]}', + '', + f'- Commit: `{os.environ.get("GITHUB_SHA", "unknown")}`', + f'- Result: **{os.environ["SUMMARY_RESULT"]}**', + ] + if os.environ.get('SUMMARY_FAILED_STEP'): + lines.append(f'- Failed step: {os.environ["SUMMARY_FAILED_STEP"]}') + if os.environ['SUMMARY_RESULT'] != 'success': + lines.append('- Open the failed step log for the error details.') + write_summary(lines) + + +def build(output, name, plan): + report = {'phase': 'Check the source commit', 'current': None, 'built': [], 'reused': [], 'images': {}} + result = 'failure' + try: + build_images(output, report, name, plan) + result = 'success' + finally: + lines = [ + f'## Image build result `{name}`', + '', + f'- Commit: `{os.environ.get("GITHUB_SHA", "unknown")}`', + '', + f'- Result: **{result}**', + f'- Last stage: {report["phase"]}', + ] + if result == 'failure': + lines.append('- This image job failed. The complete release cannot be published. Open the failed step log.') + if result == 'success': + lines.append('- This is one image result. The final build job must publish the complete release.') + if report['current']: + lines.append(f'- Image at the failure: `{report["current"]}`') + for title, key in (('Built', 'built'), ('Reused from successful CI', 'reused')): + lines.extend(['', f'### {title}']) + lines.extend(f'- `{name}`' for name in report[key]) + if not report[key]: + lines.append('- None') + lines.extend(['', '### Completed image digests']) + lines.extend(f'- `{image}@{digest}`' for image, digest in report['images'].items()) + if not report['images']: + lines.append('- None') + write_summary(lines) + + +def render(stream, destination): + release = validate_release(json.loads(Path(os.environ['RELEASE_FILE']).read_text()), os.environ['DEPLOY_SHA']) + image_line = re.compile( + r"^(\s*(?:-\s*)?image:\s*)(['\"]?)(gcr\.forust\.xyz/forust/[\w.-]+)(?::[\w.-]+|@sha256:[0-9a-f]{64})\2(\s*(?:#.*)?)$" + ) + rendered = [] + for line in stream: + match = image_line.fullmatch(line.rstrip('\n')) + if match: + prefix, quote, image, tail = match.groups() + if image not in release['images']: + raise ValueError(f'Owned image missing from checked release: {image}') + line = f'{prefix}{quote}{image}@{release["images"][image]}{quote}{tail}\n' + elif re.match(r'\s*(?:-\s*)?image:', line) and 'gcr.forust.xyz/forust/' in line: + raise ValueError('Unsupported owned image syntax; refusing to apply a mutable tag') + rendered.append(line) + destination.writelines(rendered) + + +def finalize_images(output, fragments, plan): + data = checked_plan(plan) + sha = data['sha'] + release = {'version': 1, 'sha': sha, 'images': {}, 'inputs': {}} + for name in IMAGES: + fragment = json.loads((fragments / f'image-{name}' / 'image.json').read_text()) + image = f'gcr.forust.xyz/forust/{name}' + if fragment.get('sha') != sha or fragment.get('version') != 1 or set(fragment.get('images', {})) != {image}: + raise ValueError('Image job artifact is missing or belongs to another commit') + target = next(t for t in data['targets'] if t['name'] == name) + if fragment.get('inputs') != {image: target['inputs']}: + raise ValueError('Image artifact does not match the build plan') + release['images'].update(fragment['images']) + release['inputs'].update(fragment['inputs']) + validate_release(release, sha) + # Only a complete set of successful image jobs can publish the release tags. + docker_config = tempfile.mkdtemp(prefix='homelab-registry-') + env = {**os.environ, 'DOCKER_CONFIG': docker_config} + try: + subprocess.run( # noqa: S603, S607 + [ + shutil.which('docker') or '/usr/bin/docker', + 'login', + 'gcr.forust.xyz', + '-u', + os.environ['REGISTRY_USERNAME'], + '--password-stdin', + ], + input=os.environ['REGISTRY_PASSWORD'], + text=True, + check=True, + env=env, + ) + for image, digest in release['images'].items(): + command( + 'docker', + 'buildx', + 'imagetools', + 'create', + '--prefer-index=false', + '--tag', + f'{image}:sha-{sha}', + f'{image}@{digest}', + env=env, + timeout=90, + ) + output.write_text(json.dumps(release, indent=2) + '\n') + finally: + shutil.rmtree(docker_config) + + +def main(): + parser = argparse.ArgumentParser(description=__doc__) + parser.add_argument('action', choices=('prepare', 'image', 'finalize', 'gate', 'render', 'check-summary')) + parser.add_argument('--output', type=Path, default=Path('release.json')) + parser.add_argument('--ref', default='main') + parser.add_argument('--event-sha', default='') + parser.add_argument('--image', choices=IMAGES) + parser.add_argument('--plan', type=Path, default=Path('build-plan.json')) + parser.add_argument('--fragments', type=Path, default=Path('artifacts')) + args = parser.parse_args() + if args.action == 'check-summary': + check_summary() + elif args.action == 'render': + render(sys.stdin, sys.stdout) + elif args.action == 'gate': + gate(args.output, args.ref, args.event_sha) + elif args.action == 'prepare': + prepare_images(args.output) + elif args.action == 'image': + if not args.image: + parser.error('--image is required') + build(args.output, args.image, args.plan) + else: + finalize_images(args.output, args.fragments, args.plan) + + +if __name__ == '__main__': + main() diff --git a/.gitea/workflows/renovate-ci.yaml b/.gitea/workflows/renovate-ci.yaml index 00dfee9..ea3b92b 100644 --- a/.gitea/workflows/renovate-ci.yaml +++ b/.gitea/workflows/renovate-ci.yaml @@ -1,10 +1,26 @@ name: renovate-ci on: - pull_request: + # Read the workflow from the trusted base branch. PR code runs only on the + # unprivileged runner selected below. + pull_request_target: + paths: + - "renovate/**" + - ".gitea/workflows/renovate-ci.yaml" + - ".gitea/workflows/sync-renovate-configmap.sh" + - ".gitea/workflows/compose-lint.sh" + - ".gitea/workflows/install-ci-tools.sh" + - ".gitea/workflows/tool-versions.env" push: branches: - main + paths: + - "renovate/**" + - ".gitea/workflows/renovate-ci.yaml" + - ".gitea/workflows/sync-renovate-configmap.sh" + - ".gitea/workflows/compose-lint.sh" + - ".gitea/workflows/install-ci-tools.sh" + - ".gitea/workflows/tool-versions.env" workflow_dispatch: permissions: @@ -12,37 +28,47 @@ permissions: jobs: validate-renovate: - runs-on: [self-hosted, linux, arch, homelab] + runs-on: ${{ github.event_name == 'push' && github.ref == 'refs/heads/main' && 'homelab' || 'homelab-pr' }} timeout-minutes: 20 steps: - name: Checkout repository uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4 + with: + ref: ${{ github.event_name == 'pull_request_target' && github.event.pull_request.head.sha || github.sha }} - # renovate/k8s/cronjob.yaml is the single source of truth for the image tag, - # so the same version that runs in the cluster is the one validated here. - - name: Resolve the deployed Renovate image + # renovate/k8s/cronjob.yaml is the single source of truth for the version. + - name: Resolve the deployed Renovate version id: image shell: bash run: | set -euo pipefail image="$(sed -n 's|.*image:[[:space:]]*\(renovate/renovate:[^[:space:]]*\).*|\1|p' \ renovate/k8s/cronjob.yaml | head -1)" - if [ -z "$image" ]; then - echo "::error::no renovate/renovate image found in renovate/k8s/cronjob.yaml" + if [[ ! "$image" =~ ^renovate/renovate:([0-9]+\.[0-9]+\.[0-9]+)$ ]]; then + echo "::error::expected a pinned renovate/renovate semantic version in renovate/k8s/cronjob.yaml" exit 1 fi - echo "using $image" - echo "image=$image" >> "$GITHUB_OUTPUT" + version="${BASH_REMATCH[1]}" + echo "using Renovate $version" + printf 'version=%s\n' "$version" >> "$GITHUB_OUTPUT" - - name: Validate Renovate repository config + - name: Prepare pinned validation tools shell: bash run: | set -euo pipefail - docker run --rm \ - -v "$PWD/renovate:/opt/renovate:ro" \ - -e RENOVATE_CONFIG_FILE=/opt/renovate/renovate.json \ - "${{ steps.image.outputs.image }}" \ - renovate-config-validator /opt/renovate/renovate.json + tools_dir="$(bash .gitea/workflows/install-ci-tools.sh kubeconform node)" + echo "$tools_dir" >> "$GITHUB_PATH" + + - name: Validate Renovate repository config + shell: bash + env: + RENOVATE_VERSION: ${{ steps.image.outputs.version }} + run: | + set -euo pipefail + npm_cache="$(mktemp -d "${RUNNER_TEMP:-/tmp}/renovate-npm-cache.XXXXXXXX")" + trap 'rm -rf "$npm_cache"' EXIT + NPM_CONFIG_CACHE="$npm_cache" RENOVATE_CONFIG_FILE="$PWD/renovate/renovate.json" \ + npm exec --yes --package="renovate@${RENOVATE_VERSION}" -- renovate-config-validator # The CronJob cannot read the repository, so renovate/k8s/configmap.yaml # carries an inlined copy of the config. Fail if it no longer matches. @@ -56,8 +82,6 @@ jobs: shell: bash run: | set -euo pipefail - tools_dir="$(bash .gitea/workflows/install-ci-tools.sh kubeconform)" - export PATH="$tools_dir:$PATH" kubeconform \ -strict \ -ignore-missing-schemas \ diff --git a/.gitea/workflows/renovate-run.yaml b/.gitea/workflows/renovate-run.yaml index 658a603..7e32fc9 100644 --- a/.gitea/workflows/renovate-run.yaml +++ b/.gitea/workflows/renovate-run.yaml @@ -32,11 +32,14 @@ concurrency: jobs: run-renovate: - runs-on: [self-hosted, linux, arch, homelab] + if: github.ref == 'refs/heads/main' + runs-on: homelab timeout-minutes: 60 steps: - name: Checkout repository uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4 + with: + ref: refs/heads/main # renovate/k8s/cronjob.yaml is the single source of truth for the image tag. # Reading it here means this workflow validates and runs the exact version @@ -48,21 +51,23 @@ jobs: set -euo pipefail image="$(sed -n 's|.*image:[[:space:]]*\(renovate/renovate:[^[:space:]]*\).*|\1|p' \ renovate/k8s/cronjob.yaml | head -1)" - if [ -z "$image" ]; then - echo "::error::no renovate/renovate image found in renovate/k8s/cronjob.yaml" + if [[ ! "$image" =~ ^renovate/renovate:[0-9]+\.[0-9]+\.[0-9]+$ ]]; then + echo "::error::expected a pinned renovate/renovate semantic version in renovate/k8s/cronjob.yaml" exit 1 fi echo "using $image" - echo "image=$image" >> "$GITHUB_OUTPUT" + printf 'image=%s\n' "$image" >> "$GITHUB_OUTPUT" - name: Validate Renovate config shell: bash + env: + RENOVATE_IMAGE: ${{ steps.image.outputs.image }} run: | set -euo pipefail docker run --rm \ -v "$PWD/renovate/renovate.json:/opt/renovate/renovate.json:ro" \ -e RENOVATE_CONFIG_FILE=/opt/renovate/renovate.json \ - "${{ steps.image.outputs.image }}" \ + "$RENOVATE_IMAGE" \ renovate-config-validator - name: Run Renovate @@ -73,6 +78,7 @@ jobs: RENOVATE_REPOSITORIES: ${{ inputs.repositories }} RENOVATE_DRY_RUN: ${{ inputs.dry_run && 'full' || '' }} LOG_LEVEL: ${{ inputs.log_level }} + RENOVATE_IMAGE: ${{ steps.image.outputs.image }} run: | set -euo pipefail @@ -89,4 +95,4 @@ jobs: -e RENOVATE_CONFIG_FILE=/opt/renovate/renovate.json \ -e RENOVATE_BASE_DIR=/tmp/renovate \ -e LOG_LEVEL="${LOG_LEVEL:-info}" \ - "${{ steps.image.outputs.image }}" + "$RENOVATE_IMAGE" diff --git a/.gitea/workflows/secret-references.jq b/.gitea/workflows/secret-references.jq new file mode 100644 index 0000000..a205f61 --- /dev/null +++ b/.gitea/workflows/secret-references.jq @@ -0,0 +1,13 @@ +# kubectl emits a List for files containing multiple resources. +(if .kind == "List" then .items[] else . end) +| (.metadata.namespace // "default") as $ns +| [ + (.. | objects + | (.secretRef? // empty), (.secretKeyRef? // empty), (.secret? // empty) + | select(.optional != true) + | .name // .secretName // empty), + (.. | objects | .imagePullSecrets[]?.name) + ] +| unique[] +| select(. != null and . != "") +| "\($ns) \(.)" diff --git a/.gitea/workflows/ssh-run.sh b/.gitea/workflows/ssh-run.sh index b2d8441..7f29be4 100755 --- a/.gitea/workflows/ssh-run.sh +++ b/.gitea/workflows/ssh-run.sh @@ -1,71 +1,70 @@ #!/usr/bin/env bash -# usage: ssh-run.sh -# Runs one deploy-lib.sh stage on the workstation over SSH. +# The SSH client submits once and follows durable stages on workstation. set -euo pipefail - : "${DEPLOY_HOST:?missing DEPLOY_HOST}" : "${DEPLOY_USER:?missing DEPLOY_USER}" : "${DEPLOY_KEY:?missing DEPLOY_SSH_KEY}" - -deploy_port="${DEPLOY_PORT:-22}" -deploy_path="${DEPLOY_PATH:-/srv/homelab}" -deploy_path="$(printf '%s' "$deploy_path" | tr -d '\"' | tr -d '\r' | xargs)" - -# The private key is written to a per-run directory that is removed on exit, so a -# failed or cancelled job cannot leave deploy credentials in the runner's temp -# directory. Do not use a fixed path: apply-k8s and apply-compose run in parallel. +: "${DEPLOY_KNOWN_HOSTS:?configure pinned DEPLOY_KNOWN_HOSTS}" +: "${DEPLOY_RUN_ID:?missing DEPLOY_RUN_ID}" +[[ "$DEPLOY_USER" =~ ^[A-Za-z_][A-Za-z0-9_.-]*$ ]] || exit 1 +[[ "$DEPLOY_HOST" =~ ^[A-Za-z0-9_.:-]+$ ]] || exit 1 +[[ "$DEPLOY_RUN_ID" =~ ^[0-9]+-[0-9]+$ ]] || exit 1 +[[ "${DEPLOY_PORT:-22}" =~ ^[0-9]+$ ]] || exit 1 key_dir="$(mktemp -d "${RUNNER_TEMP:-/tmp}/homelab-deploy-key.XXXXXXXX")" -trap 'rm -rf "$key_dir"' EXIT INT TERM - -ssh_key="$key_dir/deploy_key" -printf '%s\n' "$DEPLOY_KEY" > "$ssh_key" -chmod 600 "$ssh_key" - -# A connection that died silently used to hang until the job timeout, and the -# stage was never re-run: one flaky TCP session cost a whole 45-minute apply. -# ServerAlive* bounds how long a dead peer goes unnoticed, ConnectTimeout bounds -# setup. Only exit 255 - ssh's own transport failures - is retried. A stage that -# fails on its own merits exits with the remote's status, so a real failure -# still surfaces its own log instead of burning three attempts. The stages are -# declarative applies, so re-running one that had already committed is harmless. -ssh_opts=( - -i "$ssh_key" -p "$deploy_port" - -o BatchMode=yes -o StrictHostKeyChecking=accept-new - -o ConnectTimeout=15 - -o ServerAliveInterval=15 -o ServerAliveCountMax=4 -) - -rc=0 -# apply-k8s and apply-compose are separate workflow jobs so the graph stays -# intact for the verify job, but on a single node they must not run at once: -# host docker churn on top of cluster churn is what melts the node (load 40+, -# netbird/ssh die, helm is left pending-*). Serialize them on the workstation -# with a shared lock; whoever arrives second waits. -remote_cmd=(bash -se) -case "$1" in - apply-k8s | apply-compose) - remote_cmd=(flock -w 5400 /tmp/homelab-apply.lock bash -se) +trap 'rm -rf "$key_dir"' EXIT +chmod 700 "$key_dir" +printf '%s\n' "$DEPLOY_KEY" >"$key_dir/key" +printf '%s\n' "$DEPLOY_KNOWN_HOSTS" >"$key_dir/known_hosts" +chmod 600 "$key_dir/key" "$key_dir/known_hosts" +ssh_opts=(-i "$key_dir/key" -p "${DEPLOY_PORT:-22}" -o BatchMode=yes -o StrictHostKeyChecking=yes + -o "UserKnownHostsFile=$key_dir/known_hosts" -o ConnectTimeout=15 + -o ServerAliveInterval=15 -o ServerAliveCountMax=4) +controller=.local/lib/homelab-deploy/controller.py +case "${1:?start, apply, verify, smoke or summary required}" in + start) + python3 - <<'PY' >"$key_dir/request.json" +import json +import os +from pathlib import Path +release = json.loads(Path('release.json').read_text()) +print(json.dumps({'release': release, 'mode': os.environ.get('DEPLOY_MODE', 'changed'), + 'refresh_images': os.environ.get('REFRESH_IMAGES', 'false') == 'true'})) +PY + for attempt in 1 2 3; do + rc=0 + # shellcheck disable=SC2029 # The run ID and operation are validated local arguments, not remote variables. + ssh "${ssh_opts[@]}" "$DEPLOY_USER@$DEPLOY_HOST" python3 "$controller" start "$DEPLOY_RUN_ID" <"$key_dir/request.json" || rc=$? + [ "$rc" -eq 0 ] && exit 0 + [ "$rc" -eq 255 ] || exit "$rc" + sleep 5 + done + exit "$rc" ;; + apply|verify|smoke) + result=0 + for attempt in 1 2 3; do + rc=0 + # shellcheck disable=SC2029 # The run ID and operation are validated local arguments, not remote variables. + ssh "${ssh_opts[@]}" "$DEPLOY_USER@$DEPLOY_HOST" python3 "$controller" follow "$DEPLOY_RUN_ID" "$1" || rc=$? + [ "$rc" -eq 0 ] && break + [ "$rc" -eq 255 ] || { result="$rc"; break; } + echo "SSH disconnected; reconnecting to the existing deploy ($attempt/3)" + if [ "$attempt" -eq 3 ]; then result=255; break; fi + sleep 5 + done + exit "$result" + ;; + summary) + if [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then + rc=0 + # shellcheck disable=SC2029 # The run ID is validated above. + ssh "${ssh_opts[@]}" "$DEPLOY_USER@$DEPLOY_HOST" python3 "$controller" summary "$DEPLOY_RUN_ID" >"$key_dir/deploy-summary.md" || rc=$? + if [ "$rc" -eq 0 ]; then + cat "$key_dir/deploy-summary.md" >>"$GITHUB_STEP_SUMMARY" || echo "WARNING: cannot write the deploy summary" + else + echo 'Deploy summary is unavailable. The SSH connection failed or the controller did not respond. Check the job log.' >>"$GITHUB_STEP_SUMMARY" || true + fi + fi + ;; + *) echo "Unknown SSH operation: $1" >&2; exit 1 ;; esac -for attempt in 1 2 3; do - if [ "$attempt" -gt 1 ]; then - echo ":: warning::ssh transport failed, retrying (${attempt}/3)" - sleep $((attempt * 5)) - fi - rc=0 - # shellcheck disable=SC2029 # remote_cmd/ssh_opts expand on the client on purpose: they select the local ssh invocation, only the heredoc runs remotely. - ssh "${ssh_opts[@]}" "${DEPLOY_USER}@${DEPLOY_HOST}" \ - env "REPO=$deploy_path" "APPLY_PRUNE=${APPLY_PRUNE:-false}" \ - "DEPLOY_SHA=${DEPLOY_SHA:-}" "DEPLOY_SNAPSHOT_DIR=${DEPLOY_SNAPSHOT_DIR:-}" \ - "STAGE=$1" "${remote_cmd[@]}" <<'EOF' || rc=$? -source "$REPO/.gitea/workflows/deploy-lib.sh" -run_stage "$STAGE" -EOF - [ "$rc" -eq 0 ] && break - [ "$rc" -ne 255 ] && break -done - -if [ "$rc" -ne 0 ]; then - echo ":: error::stage $1 failed over ssh (exit $rc)" -fi -exit "$rc" diff --git a/.gitea/workflows/tool-versions.env b/.gitea/workflows/tool-versions.env index d010495..9de3c39 100644 --- a/.gitea/workflows/tool-versions.env +++ b/.gitea/workflows/tool-versions.env @@ -14,7 +14,7 @@ ACTIONLINT_VERSION="1.7.7" SHELLCHECK_VERSION="0.11.0" KUBECONFORM_VERSION="0.8.0" PRETTIER_VERSION="3.8.1" -RUFF_VERSION="0.16.8" +RUFF_VERSION="0.16.10" YAMLLINT_VERSION="1.38.0" HADOLINT_VERSION="2.14.0" # pip-audit reads the advisory database over the network, so a floating version @@ -31,3 +31,9 @@ UV_VERSION="0.12.17" # so the tree that gets tested is the tree that gets built. Renovate keeps this # in step with the Dockerfile's node: tag via the "node runtime" group. NODE_VERSION="22.23.3" + +# Secret-reference regression tests parse rendered Kubernetes objects. +JQ_VERSION="1.8.1" + +# BuildKit is the only auxiliary CI container; jobs themselves stay on the host. +BUILDKIT_IMAGE="moby/buildkit:v0.33.1" diff --git a/.gitignore b/.gitignore index d49fe11..9df6ca4 100644 --- a/.gitignore +++ b/.gitignore @@ -94,7 +94,6 @@ replacements.txt .idea # Temp files -edu_master/temp/ temp/* # Local-only tooling scratch space (pinned CI tools, verification scripts) tmp/ diff --git a/README.md b/README.md index 4829eef..f0f5ea5 100644 --- a/README.md +++ b/README.md @@ -2,8 +2,8 @@ Configuration for my homelab: Kubernetes manifests, Docker Compose stacks, and the Gitea Actions that build and deploy them. Most applications have both deployment -formats. Headscale, Nextcloud AIO, and the media stack run on Docker; Kubernetes -provides their ingress through Services and EndpointSlices. +formats. Headscale and Nextcloud AIO have Compose deployments with Kubernetes +ingress; the media stack has Compose and Kubernetes routing configuration. These files contain this lab's domains, IP addresses, storage paths, and private registry names. Running them on another machine takes some editing. @@ -12,7 +12,8 @@ registry names. Running them on another machine takes some editing. - [Service list](#services) — what each directory contains. - [Deployment workflow](.gitea/README.md) — selection, validation, and recovery. -- [Repository review](docs/repository-review.md) — confirmed problems and fix branches. +- [Repository review](docs/repository-review.md) — findings from the 6 October baseline and their status. +- [EDU ownership handoff](.gitea/EDU_HANDOFF.md) — the EDU workloads now live in their own repository. - [Shared PostgreSQL](postgres/README.md), [Traefik](traefik/README.md), and [cert-manager](cert-manager/README.md) — common dependencies. @@ -40,48 +41,47 @@ service is currently healthy or running. ## Services -| Service | Configuration | Selected by markers | -| ----------------------------------------------------------- | ---------------------------- | ------------------- | -| [AdGuard Home](adguardhome/README.md) | Kubernetes + Compose | Kubernetes | -| [Authentik](authentik/README.md) | Kubernetes + Compose | Kubernetes | -| [cert-manager](cert-manager/README.md) | Kubernetes / Helm | Manual | -| [Cloudflare DDNS](cfddns/README.md) | Kubernetes + Compose | Kubernetes | -| [Checkmk](checkmk/README.md) | Kubernetes + Compose | Manual | -| [Cloudflare Tunnel](cloudflared/README.md) | Kubernetes / Helm | Manual | -| [File converters](converters/README.md) | Kubernetes + Compose | Kubernetes | -| [CrowdSec](crowdsec/README.md) | Kubernetes / Helm | Manual | -| [Dockmon](dockmon/README.md) | Kubernetes + Compose | Manual | -| [Downtify](downtify/README.md) | Kubernetes + Compose | Manual | -| [EDU session keeper and Telegram bot](edu_master/README.md) | Kubernetes + Compose | Kubernetes | -| [Error pages](errorpages/README.md) | Kubernetes + Compose | Kubernetes | -| [Gitea](gitea/README.md) | Kubernetes + Compose | Kubernetes | -| [Glance](glance/README.md) | Kubernetes + Compose | Manual | -| [Headscale](headscale/README.md) | Compose + Kubernetes routing | Compose, Kubernetes | -| [Homarr](homarr/README.md) | Kubernetes + Compose | Manual | -| [Homepages](homepages/README.md) | Kubernetes + Compose | Kubernetes | -| [Immich](immich/README.md) | Kubernetes + Compose | Kubernetes | -| [Kener](kener/README.md) | Kubernetes + Compose | Manual | -| [Loki and Alloy](loki/README.md) | Kubernetes / Helm | Kubernetes | -| [MeTube](metube/README.md) | Kubernetes + Compose | Kubernetes | -| [n8n](n8n/README.md) | Kubernetes + Compose | Manual | -| [NetBird](netbird/README.md) | Kubernetes + Compose | Kubernetes | -| [NetBox](netbox/README.md) | Kubernetes + Compose | Kubernetes | -| [Netronome](netronome/README.md) | Kubernetes + Compose | Kubernetes | -| [Nextcloud AIO](nextcloud/README.md) | Compose + Kubernetes routing | Compose, Kubernetes | -| [Penpot](penpot/README.md) | Compose | Manual | -| [Portainer](portainer/README.md) | Kubernetes + Compose | Manual | -| [Shared PostgreSQL](postgres/README.md) | Kubernetes + Compose | Kubernetes | -| [Monitoring stack](prometheus-stack/README.md) | Kubernetes + Compose | Kubernetes | -| [RackPeek](rackpeek/README.md) | Kubernetes + Compose | Kubernetes | -| [Reloader](reloader/README.md) | Kubernetes / Helm | Kubernetes | -| [Renovate](renovate/README.md) | Kubernetes + Compose | Kubernetes | -| [SearXNG](searxng/README.md) | Kubernetes + Compose | Manual | -| [Media stack](streaming/README.md) | Compose + Kubernetes routing | Compose, Kubernetes | -| [Termix](termix/README.md) | Kubernetes + Compose | Manual | -| [Traefik](traefik/README.md) | Kubernetes + Compose | Kubernetes | -| [Uptime Kuma](uptime-kuma/README.md) | Kubernetes + Compose | Kubernetes | -| [Vaultwarden](vaultwarden/README.md) | Kubernetes + Compose | Kubernetes | -| [3x-ui](vpn/xui/README.md) | Kubernetes | Kubernetes | +| Service | Configuration | Selected by markers | +| ---------------------------------------------- | ---------------------------- | ------------------- | +| [AdGuard Home](adguardhome/README.md) | Kubernetes + Compose | Kubernetes | +| [Authentik](authentik/README.md) | Kubernetes + Compose | Kubernetes | +| [cert-manager](cert-manager/README.md) | Kubernetes / Helm | Manual | +| [Cloudflare DDNS](cfddns/README.md) | Kubernetes + Compose | Kubernetes | +| [Checkmk](checkmk/README.md) | Kubernetes + Compose | Manual | +| [Cloudflare Tunnel](cloudflared/README.md) | Kubernetes / Helm | Manual | +| [File converters](converters/README.md) | Kubernetes + Compose | Kubernetes | +| [CrowdSec](crowdsec/README.md) | Kubernetes / Helm | Manual | +| [Dockmon](dockmon/README.md) | Kubernetes + Compose | Manual | +| [Downtify](downtify/README.md) | Kubernetes + Compose | Manual | +| [Error pages](errorpages/README.md) | Kubernetes + Compose | Kubernetes | +| [Gitea](gitea/README.md) | Kubernetes + Compose | Kubernetes | +| [Glance](glance/README.md) | Kubernetes + Compose | Manual | +| [Headscale](headscale/README.md) | Compose + Kubernetes routing | Compose, Kubernetes | +| [Homarr](homarr/README.md) | Kubernetes + Compose | Manual | +| [Homepages](homepages/README.md) | Kubernetes + Compose | Kubernetes | +| [Immich](immich/README.md) | Kubernetes + Compose | Kubernetes | +| [Kener](kener/README.md) | Kubernetes + Compose | Manual | +| [Loki and Alloy](loki/README.md) | Kubernetes / Helm | Kubernetes | +| [MeTube](metube/README.md) | Kubernetes + Compose | Kubernetes | +| [n8n](n8n/README.md) | Kubernetes + Compose | Manual | +| [NetBird](netbird/README.md) | Kubernetes + Compose | Kubernetes | +| [NetBox](netbox/README.md) | Kubernetes + Compose | Kubernetes | +| [Netronome](netronome/README.md) | Kubernetes + Compose | Kubernetes | +| [Nextcloud AIO](nextcloud/README.md) | Compose + Kubernetes routing | Compose, Kubernetes | +| [Penpot](penpot/README.md) | Compose | Manual | +| [Portainer](portainer/README.md) | Kubernetes + Compose | Manual | +| [Shared PostgreSQL](postgres/README.md) | Kubernetes + Compose | Kubernetes | +| [Monitoring stack](prometheus-stack/README.md) | Kubernetes + Compose | Kubernetes | +| [RackPeek](rackpeek/README.md) | Kubernetes + Compose | Kubernetes | +| [Reloader](reloader/README.md) | Kubernetes / Helm | Kubernetes | +| [Renovate](renovate/README.md) | Kubernetes + Compose | Kubernetes | +| [SearXNG](searxng/README.md) | Kubernetes + Compose | Manual | +| [Media stack](streaming/README.md) | Compose + Kubernetes routing | Manual | +| [Termix](termix/README.md) | Kubernetes + Compose | Manual | +| [Traefik](traefik/README.md) | Kubernetes + Compose | Kubernetes | +| [Uptime Kuma](uptime-kuma/README.md) | Kubernetes + Compose | Kubernetes | +| [Vaultwarden](vaultwarden/README.md) | Kubernetes + Compose | Kubernetes | +| [3x-ui](vpn/xui/README.md) | Kubernetes | Kubernetes | ## Running a Compose stack @@ -136,16 +136,16 @@ role's password; see the database README. CI pins its tools in `.gitea/workflows/tool-versions.env`. Use the same versions: -```sh -tools_dir="$(bash .gitea/workflows/install-ci-tools.sh)" -export PATH="$tools_dir:$PATH" +```fish +set tools_dir (bash .gitea/workflows/install-ci-tools.sh) +set -gx PATH $tools_dir $PATH ruff check . ruff format --check . actionlint -config-file .gitea/actionlint.yaml .gitea/workflows/*.yaml .gitea/workflows/sync-renovate-configmap.sh --check ``` -The [workflow README](.gitea/README.md#checks) lists the rest of the checks. +The [workflow README](.gitea/README.md#ci) lists the rest of the checks. Structure checks do not establish that local Secrets, mounted files, storage, or external services are ready. diff --git a/adguardhome/compose.yaml b/adguardhome/compose.yaml index 4ce64d1..69b4e5d 100644 --- a/adguardhome/compose.yaml +++ b/adguardhome/compose.yaml @@ -31,7 +31,7 @@ services: - "traefik.http.routers.adguard-dev.entrypoints=websecure" - "traefik.http.routers.adguard-dev.tls=true" # DoH Router - - "traefik.http.routers.dns-over-https.rule=(Host(`dns.forust.xyz` || Host(`adguard.forust.xyz`)) && PathPrefix(`/dns-query`))" + - "traefik.http.routers.dns-over-https.rule=(Host(`dns.forust.xyz`) || Host(`adguard.forust.xyz`)) && PathPrefix(`/dns-query`)" - "traefik.http.routers.dns-over-https.entrypoints=websecure" - "traefik.http.routers.dns-over-https.tls.certresolver=letsencrypt" diff --git a/adguardhome/k8s/adguard.yaml b/adguardhome/k8s/adguard.yaml index 290dfc4..40db3ea 100644 --- a/adguardhome/k8s/adguard.yaml +++ b/adguardhome/k8s/adguard.yaml @@ -51,6 +51,8 @@ spec: apiVersion: apps/v1 kind: Deployment metadata: + annotations: + reloader.stakater.com/auto: "true" name: adguard-deployment namespace: adguard spec: @@ -64,8 +66,6 @@ spec: metadata: labels: app: adguard - annotations: - reloader.stakater.com/auto: "true" spec: containers: - name: adguard diff --git a/authentik/k8s/authentik.yaml b/authentik/k8s/authentik.yaml index 5cfcbdf..2afecb6 100644 --- a/authentik/k8s/authentik.yaml +++ b/authentik/k8s/authentik.yaml @@ -27,6 +27,8 @@ spec: apiVersion: apps/v1 kind: Deployment metadata: + annotations: + reloader.stakater.com/auto: "true" name: authentik-server-deployment namespace: authentik spec: @@ -63,6 +65,8 @@ spec: apiVersion: apps/v1 kind: Deployment metadata: + annotations: + reloader.stakater.com/auto: "true" name: authentik-worker-deployment namespace: authentik spec: diff --git a/cfddns/k8s/deployment.yaml b/cfddns/k8s/deployment.yaml index a5598ec..b10f2b0 100644 --- a/cfddns/k8s/deployment.yaml +++ b/cfddns/k8s/deployment.yaml @@ -1,6 +1,8 @@ apiVersion: apps/v1 kind: Deployment metadata: + annotations: + reloader.stakater.com/auto: "true" name: cfddns labels: app: cfddns diff --git a/checkmk/k8s/checkmk.yaml b/checkmk/k8s/checkmk.yaml index 3fb5fc2..ddb7095 100644 --- a/checkmk/k8s/checkmk.yaml +++ b/checkmk/k8s/checkmk.yaml @@ -17,6 +17,8 @@ spec: apiVersion: apps/v1 kind: Deployment metadata: + annotations: + reloader.stakater.com/auto: "true" name: checkmk-deployment namespace: checkmk spec: diff --git a/cloudflared/k8s/deployment.yaml b/cloudflared/k8s/deployment.yaml index c164c46..03bc695 100644 --- a/cloudflared/k8s/deployment.yaml +++ b/cloudflared/k8s/deployment.yaml @@ -1,6 +1,8 @@ apiVersion: apps/v1 kind: Deployment metadata: + annotations: + reloader.stakater.com/auto: "true" name: cloudflared labels: app: cloudflared @@ -18,7 +20,7 @@ spec: spec: containers: - name: cloudflared - image: cloudflare/cloudflared:2026.9.3 + image: cloudflare/cloudflared:2026.10.0 imagePullPolicy: IfNotPresent args: - tunnel diff --git a/converters/k8s/convertx.yaml b/converters/k8s/convertx.yaml index bc81cc1..d5924b8 100644 --- a/converters/k8s/convertx.yaml +++ b/converters/k8s/convertx.yaml @@ -13,6 +13,8 @@ spec: apiVersion: apps/v1 kind: Deployment metadata: + annotations: + reloader.stakater.com/auto: "true" name: convertx-deployment namespace: converters spec: diff --git a/docs/repository-review.md b/docs/repository-review.md index 937c573..0310fbf 100644 --- a/docs/repository-review.md +++ b/docs/repository-review.md @@ -1,39 +1,40 @@ -# Repository review +# Repository review (6 October 2026 baseline) -Reviewed the tracked tree at `cc9c3de` and read the live workstation state on -6 October 2026. Changes are split into documentation and individual fix branches, -all based on that main commit. The original local checkout and its uncommitted -monitoring changes were preserved. No deployment was performed. +This records the tracked tree at `cc9c3de` and the workstation state observed on +6 October 2026. It is a historical review, not a current runtime inventory. The +listed code fixes have since merged into `main`; EDU ownership has moved to the +separate repository described in [the handoff record](../.gitea/EDU_HANDOFF.md). +See the [CI and deployment guide](../.gitea/README.md) and +[runner and recovery guide](../.gitea/runner/README.md) for the current workflow. +No deployment was performed during the original review. -## Confirmed problems with prepared fixes +## Findings at the baseline and current status -| Priority | Problem and consequence | Fix branch | -| -------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------- | -| High | `APPLY_PRUNE=true` is passed to each individual manifest apply. Each invocation sees only that file's desired objects and can delete other resources selected by the shared label. | `fix/deploy-prune-guard` | -| High | Deploy validates Compose with interpolation and env/path resolution disabled. Required settings can pass validation and then fail during apply after other workloads have changed. | `fix/deploy-validation` | -| Medium | Secret validation is text-based and compares names across all namespaces. A Secret elsewhere can hide a missing local Secret; mounted Secrets are also missed. | `fix/deploy-validation` | -| Medium | Compose CI misses `postgres/shared-compose.yaml`, `netbird/client.compose.yaml`, and `renovate/renovate-compose.yaml`. | `fix/deploy-validation` | -| Medium | NetBird Compose mounts `entrypoint.sh`, but it is absent. Its README also calls a missing `setup.sh`; a fresh checkout cannot start this stack as documented. | `fix/netbird-compose-runtime` | -| Medium | Glance's CSS mount uses `glance-config`, whose keys do not include `user.css`. That key is in `glance-assets`; the pod's subPath mount cannot be prepared correctly. | `fix/glance-assets` | -| Medium | The shared PostgreSQL initializer requires `NETBOX_DB_PASSWORD`, but the Compose env example omits it. Following the example leaves first initialization incomplete. | `fix/postgres-env-example` | -| Medium | EDU's Compose env example uses old credential names and full URL variables, while the code reads `KEEPER_*` and paths under `EDU_URL_BASE`. | `fix/session-keeper-reliability` | -| Medium | Session keeper HTTP calls have no timeouts. Its Redis cookie never expires, probes only check existence, and its logs include cookies. A hung or failed refresh can leave a stale session appearing ready. | `fix/session-keeper-reliability` | -| Medium | AdGuard's DoH and SearXNG's Compose rules put Boolean expressions inside `Host(...)`. They are invalid router expressions despite valid YAML. | `fix/compose-router-rules` | +| Priority | Finding at the baseline | Current status | +| -------- | ---------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------- | +| High | Per-file `APPLY_PRUNE=true` could delete resources selected by a shared label. | The deploy workflow rejects unsafe pruning before applying resources. | +| High | Compose validation did not resolve the local configuration required at deploy time. | Preflight resolves the selected Compose configuration before apply. | +| Medium | Secret validation could miss namespace-specific and mounted Secret references. | Preflight checks rendered references in their namespaces, including mounted and projected Secrets. | +| Medium | Compose CI missed manual entry points such as `shared-compose.yaml` and `client.compose.yaml`. | CI checks all tracked Compose files. | +| Medium | NetBird Compose referenced missing setup and renderer files. | The setup and renderer files are now present; Compose remains a manual alternative to the active Kubernetes deployment. | +| Medium | Glance mounted its CSS from the wrong ConfigMap. | The mount now uses the ConfigMap that contains `user.css`. | +| Medium | The PostgreSQL env example omitted the required NetBox password. | The example now includes the required variable. | +| Medium | The former EDU code had stale Compose variable names and session reliability problems. | EDU workloads and their fixes moved out of this repository; see the handoff record. | +| Medium | AdGuard DoH and SearXNG Compose router expressions used invalid `Host(...)` syntax. | The router expressions now follow Traefik's rule syntax. | Traefik matchers should be combined as `Host(a) || Host(b)`; the rule syntax is described in the [Traefik rules documentation](https://doc.traefik.io/traefik/reference/routing-configuration/http/routing/rules-and-priority/). The fix retains the DoH path constraint for both hostnames. -The prune fix deliberately rejects the unsafe option. It does not introduce -automatic deletion under a different implementation. Prune defaults to false, -and no tracked resource currently carries the selector label, so this is a -latent defect rather than evidence of a live deletion incident. +The current deploy workflow deliberately rejects the unsafe prune option. It +does not introduce automatic deletion under a different implementation. The +baseline finding was a configuration risk, not evidence of a live deletion +incident. -The session fix bounds HTTP and Redis calls, validates required credentials, -sets a cookie lifetime of two refresh intervals, and marks success only after -publishing the verified cookie. With the default ten-minute interval, an outage -longer than twenty minutes will make the existing Redis-key readiness checks fail. -That is an intentional change from indefinite apparent readiness. +The former session fix bounded HTTP and Redis calls, validated credentials, set +a cookie lifetime of two refresh intervals, and marked success only after +publishing the verified cookie. The service is now owned by the EDU repository; +see that repository for its current implementation. The deployment fix extracts required pod Secret references from rendered JSON, checks their namespaces, includes init containers, image-pull credentials, and @@ -61,51 +62,34 @@ reviewed local commit. It has untracked host configuration and a separate | Default `local-path` has reclaim policy Delete, while many existing PVs have been changed to Retain. | Current retention is partly live state. Recreating a claim can get a different policy from the old PV. | | NetBird, NetBox media/reports/scripts, EDU Redis, Homarr, and VictoriaMetrics have Delete-policy PVs. | Deleting their claims can delete important state. Plan backup and retention changes before namespace cleanup. | -The monitoring files already modified in the user's local tree correspond to the -live migration. They are excluded from these branches. Reconcile that work before -using this review's baseline to deploy monitoring. +The VictoriaMetrics monitoring trial later merged into `main` in PR #95. The +first row above records the state before that change. Read +[`prometheus-stack/README.md`](../prometheus-stack/README.md) for the current +tracked monitoring configuration; the live observations in this section remain +a snapshot from 6 October. -## Remaining work +## Current recovery limits -These need recovery design or infrastructure decisions rather than a small -configuration correction: +The deployment controller and its recovery process changed after this review. +The current operator workflow is documented in the +[runner and recovery guide](../.gitea/runner/README.md). The remaining boundaries +are: -- **SSH apply retries can replace the rollback baseline.** `ssh-run.sh` retries - exit 255, including `apply-k8s`; every new invocation publishes a fresh snapshot. - If the first attempt already changed workloads, the retry snapshots that partial - state. Preserve a run-specific original baseline and verify it across retries. -- **Rollback can exceed the job budget.** Verification is parallel, but - `rollback_workloads` is serial with a five-minute limit per workload. The - thirty-minute job budget can expire before recovery finishes. Bound recovery - concurrency and account for both phases before choosing a new timeout. -- **Snapshot collection is allowed to fail.** Generation and workload snapshot - errors are warnings; verify can fall back to all workloads. A snapshot failure - must not permit unrelated workloads to be selected for automatic undo. -- **Rollback uses the previous revision, not the captured revision.** `rollout undo` - without an explicit revision cannot guarantee restoration to the snapshot after - retries or intervening rollouts. First deployments also have no previous revision. -- **Manual deploy dispatch bypasses the CI-success trigger.** Either validate the - target commit's successful CI run or document manual dispatch as an operator - override with its own required checks. -- **Direct Traefik API exposure is unauthenticated.** The latest local commit - explicitly added it for Homarr. Preserve that integration while choosing a - cluster-internal authenticated path or a verified network restriction; do not - simply disable an integration that is already in use. -- **Storage retention and backup are not reproducible as a whole.** Defaults and - several important PV policies are Delete. There is no repository-wide backup - schedule. Existing PVC StorageClass changes require migration rather than an - in-place YAML edit. -- **MeTube downloads are temporary on Kubernetes.** `/downloads` is a 20 GiB - emptyDir. Decide whether pod replacement should discard files or whether it - should use persistent storage. Compose uses a host directory instead. -- **First-time activation needs a bootstrap path.** Deploy validation dry-runs - namespaced resources before the apply stage creates namespaces and installs - selected charts. On a fresh cluster, missing namespaces and CRDs need separate - preparation; activation is not a complete installer. +- Kubernetes recovery can restore captured workload revisions. It does not + restore ConfigMaps, Secrets, database schemas, or persistent data. +- Compose recovery is manual. It uses saved resolved configuration, but it does + not restore volume data or reverse database migrations. +- Removed resources require manual review and removal; the deploy workflow does + not prune them automatically. +- Plan mode does not create namespaces. During apply, server validation for new + namespaces runs after namespace creation and chart installation; a failed + check can leave an empty namespace. +- Storage policy and backup coverage remain service-specific. Check the live PV, + PVC, and backup state before changing stateful workloads. ## Validation -Baseline lint checks passed for Python, shell, workflows, YAML, standard Compose +At the review baseline, lint checks passed for Python, shell, workflows, YAML, standard Compose files, and Kubernetes resources with available schemas. Kubeconform found 347 resources in 174 files: 201 valid, 146 skipped CRDs, zero invalid resources. That skip count matters: passing schema validation does not validate Traefik rule @@ -124,8 +108,8 @@ Fix validation covers: documented Traefik grammar. They were not exercised on the live proxy. - Prune rejection before any cluster invocation. -All seven fix branches and the documentation branch merged together without -conflicts in a disposable validation worktree. The combined tree passed the +At the time of review, all seven fix branches and the documentation branch +merged together in a disposable validation worktree. That combined tree passed the CI-equivalent local checks, Markdown formatting/lint and link checks, all 35 Compose structure checks, and 11 Python regression tests plus the shell validation regressions. CRD server-side validation and live rollout tests were @@ -135,17 +119,16 @@ Runtime tests use fixtures and mocks, not production credentials. Live checks re workload metadata, storage policies, chart versions, and container state only. They did not read Secret contents or change services. -## Reloader follow-up +## Reloader follow-up (baseline) -`fix/reloader-integration` adds the active marker and opt-in annotations to 28 +`fix/reloader-integration` added the active marker and opt-in annotations to application Deployments/StatefulSets that consume runtime ConfigMaps or Secrets. -It corrects AdGuard's misplaced pod-template annotation. The Helm settings use +It corrected AdGuard's misplaced pod-template annotation. The Helm settings use annotation-based reloads, keep global auto-reload disabled, and ignore Jobs and CronJobs. PostgreSQL workloads are excluded because their credential variables and init scripts are only effective on an empty data directory. -The controller was already running on workstation when inspected. Its live -configuration is unchanged by the branch: merge and deploy the integration to -apply the new policy and application annotations. Configuration reload behavior -was checked against the pinned chart, with Helm rendering and manifest validation; -no production configuration was changed to provoke a test restart. +The controller was running on the workstation when inspected. The original +review checked configuration against the pinned chart with Helm rendering and +manifest validation; it did not change production configuration to provoke a +test restart or confirm every application's live reload behavior. diff --git a/edu_master/.env.example b/edu_master/.env.example deleted file mode 100644 index ec9c514..0000000 --- a/edu_master/.env.example +++ /dev/null @@ -1,14 +0,0 @@ -EDU_LOGIN=your_edu_login_here -EDU_PASSWORD=your_edu_password_here -EDU_URL_LOGIN=https://edu.edu.vn.ua/user/login -EDU_URL_VERIFY=https://edu.edu.vn.ua/course/userlist -PHPSESSID_INTERVAL=10 -USER_AGENT="Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/142.0.0.0 Safari/537.36" -WEBINAR_URL=https://edu.edu.vn.ua/webinar/useractive -WEBINAR_CHECK_INTERVAL=60 -REDIS_HOST=redis -REDIS_PORT=6379 -PLAYWRIGHT_WS=ws://playwright-service:3000/ws -TZ=Europe/Kyiv -WEBINAR_TELEGRAM_TOKEN=your_telegram_bot_token_here -WEBINAR_ADMIN_ID=123456789 diff --git a/edu_master/PLAYWRIGHT_VERSION b/edu_master/PLAYWRIGHT_VERSION deleted file mode 100644 index 3ebf789..0000000 --- a/edu_master/PLAYWRIGHT_VERSION +++ /dev/null @@ -1 +0,0 @@ -1.56.0 diff --git a/edu_master/README.md b/edu_master/README.md deleted file mode 100644 index 0429ee3..0000000 --- a/edu_master/README.md +++ /dev/null @@ -1,52 +0,0 @@ -# EDU session keeper and Telegram bot - -Keeps an EDU login session in Redis and sends Telegram notifications for new webinars. The bot also serves diary and schedule commands. - -`phpsessid-bot/` logs into EDU and publishes `EDU_PHPSESSID` in Redis. -`webinar-checker/` uses that cookie through a remote Playwright browser and stores -subscribers, language preferences, and webinar history in Redis. - -Kubernetes runs in `edu-master`, with Redis data in `redis-data-pvc`. -`service.yaml`, `servicemonitor.yaml`, and `alerts.yaml` expose and monitor the -checker's metrics on port 8000. Its `/health` endpoint reflects recent checks. - -## Configuration - -Use the keys in `k8s/secrets.yaml.example` as the reference. The committed Compose -`.env.example` has stale names until `fix/session-keeper-reliability` is merged. -The code reads: - -| Variable | Purpose | -| ----------------------------------------------------- | ------------------------------------------------ | -| `KEEPER_LOGIN`, `KEEPER_PASSWORD` | EDU login credentials. | -| `KEEPER_INTERVAL` | Session refresh interval in minutes; default 10. | -| `EDU_URL_BASE` | EDU site origin. | -| `EDU_URL_LOGIN`, `EDU_URL_COURSES`, `EDU_URL_WEBINAR` | Paths under that origin. | -| `WEBINAR_TELEGRAM_TOKEN`, `WEBINAR_ADMIN_ID` | Telegram bot and administrator. | -| `WEBINAR_CHECK_INTERVAL` | Checker interval in seconds; default 60. | -| `REDIS_HOST`, `REDIS_PORT` | Redis connection. | -| `PLAYWRIGHT_WS` | Remote browser WebSocket endpoint. | - -Set the keeper keys explicitly in the Compose `.env`. Keep the Playwright Python -package, browser image, server command, and `PLAYWRIGHT_VERSION` file on matching -versions. The two Python images are built and published by CI. - -## Bot use - -Start a private chat with `/start` to subscribe. `/stop`, `/language`, `/diary`, -`/schedule`, and `/setclass` manage subscriptions and school views. The -administrator can manage the whitelist with `/adduser` and `/removeuser`. - -Back up Redis if subscriber settings and notification history matter. Session -cookies and Telegram tokens are credentials; keep them out of shared logs. - -## Inspect - -From the repository root: - -```sh -kubectl get pods,svc,pvc -n edu-master -kubectl get events -n edu-master --sort-by=.metadata.creationTimestamp -``` - -See the [repository README](../README.md) for deployment selection. diff --git a/edu_master/compose.yaml b/edu_master/compose.yaml deleted file mode 100644 index a87d9ff..0000000 --- a/edu_master/compose.yaml +++ /dev/null @@ -1,49 +0,0 @@ -services: - redis: - image: redis:8.10.2-alpine - restart: unless-stopped - volumes: - - redis-data:/data - healthcheck: - test: ["CMD", "redis-cli", "ping"] - interval: 5s - timeout: 3s - retries: 5 - - playwright-service: - image: mcr.microsoft.com/playwright:v1.56.0-jammy - restart: unless-stopped - command: npx -y playwright@1.56.0 run-server --port 3000 --path /ws - - session-keeper: - build: ./phpsessid-bot - image: gcr.forust.xyz/forust/session-keeper:prod - pull_policy: build - env_file: .env - restart: unless-stopped - depends_on: - redis: - condition: service_healthy - healthcheck: - test: ["CMD-SHELL", "redis-cli -h redis EXISTS EDU_PHPSESSID | grep -q 1"] - interval: 30s - timeout: 5s - retries: 10 - start_period: 60s - - webinar-checker: - build: ./webinar-checker - image: gcr.forust.xyz/forust/webinar-checker:prod - pull_policy: build - env_file: .env - restart: unless-stopped - depends_on: - redis: - condition: service_healthy - session-keeper: - condition: service_healthy - playwright-service: - condition: service_started - -volumes: - redis-data: diff --git a/edu_master/k8s/alerts.yaml b/edu_master/k8s/alerts.yaml deleted file mode 100644 index dce9760..0000000 --- a/edu_master/k8s/alerts.yaml +++ /dev/null @@ -1,96 +0,0 @@ -apiVersion: monitoring.coreos.com/v1 -kind: PrometheusRule -metadata: - name: edu-master-webinar - namespace: edu-master - labels: - release: prometheus-stack -spec: - groups: - - name: edu_master.webinar - rules: - # No successful webinar check for 5m (~2-3 missed 2-min checks). - # Catches: playwright hangs/timeouts, version skew, site changes, hung job. - # The last_success > 0 guard is mandatory: checker.py initialises - # last_success to 0, so without it `time() - 0` equals the current epoch - # and humanizeDuration renders ~20722d on every pod restart. Keep the - # duration expression on the left so $value stays the real gap. - - alert: WebinarCheckerNoSuccessfulCheck - expr: | - ((time() - webinar_check_last_success_timestamp_seconds) > 300) - and (webinar_check_last_success_timestamp_seconds > 0) - and (webinar_check_last_run_timestamp_seconds > 0) - for: 2m - labels: - severity: critical - annotations: - summary: "Webinar checker has no successful check for 5m" - description: "edu-master/webinar-checker: last successful webinar check was {{ $value | humanizeDuration }} ago. Checks are failing or hanging (see consecutive failures alert). Notifications about new webinars are NOT being sent." - - # Checks are running but none has ever succeeded since pod start. - # Split out from the rule above so a zeroed gauge never feeds - # humanizeDuration. - - alert: WebinarCheckerNeverSucceeded - expr: | - (webinar_check_last_success_timestamp_seconds == 0) - and (webinar_check_last_run_timestamp_seconds > 0) - for: 10m - labels: - severity: critical - annotations: - summary: "Webinar checker has never completed a successful check" - description: 'edu-master/webinar-checker: checks have been running for 10m but not one has ever succeeded since the pod started, so every check is failing. Check pod logs (Loki: {namespace="edu-master", container="webinar-checker"}).' - - # Fast path: 3 consecutive failures (~6+ min at 2-min interval). - - alert: WebinarCheckerConsecutiveFailures - expr: | - webinar_check_consecutive_failures >= 3 - for: 5m - labels: - severity: critical - annotations: - summary: "Webinar checker failing consecutively" - description: 'edu-master/webinar-checker: {{ $value }} consecutive webinar check failures (timeout / playwright error / page error). Check pod logs (Loki: {namespace="edu-master", container="webinar-checker"}).' - - # Metrics endpoint not scraped for 10m: pod down, metrics server dead, or ServiceMonitor broken. - - alert: WebinarCheckerScrapeDown - expr: | - absent(webinar_check_last_run_timestamp_seconds) == 1 - for: 10m - labels: - severity: critical - annotations: - summary: "Webinar checker metrics missing" - description: "edu-master/webinar-checker: no metrics series for 10m. Pod may be down, metrics server dead, or ServiceMonitor/Service broken. Webinar checks are unobserved." - - # EDU session lost: session-keeper down or credentials expired. Without PHPSESSID every check is skipped. - - alert: EduPhpsessidMissing - expr: | - edu_phpsessid_present == 0 - for: 10m - labels: - severity: critical - annotations: - summary: "EDU_PHPSESSID missing" - description: "edu-master: EDU_PHPSESSID absent from redis for 10m. Webinar/diari/schedule checks are all skipped. Check session-keeper logs and EDU credentials." - - # Hard deps: checker and playwright deployments unavailable. - - alert: WebinarCheckerDeploymentDown - expr: | - kube_deployment_status_replicas_unavailable{deployment="webinar-checker", namespace="edu-master"} > 0 - for: 10m - labels: - severity: critical - annotations: - summary: "Webinar checker deployment unavailable" - description: "edu-master/webinar-checker deployment has {{ $value }} unavailable replica(s) for 10m." - - - alert: PlaywrightServiceDown - expr: | - kube_deployment_status_replicas_unavailable{deployment="playwright-service", namespace="edu-master"} > 0 - for: 10m - labels: - severity: critical - annotations: - summary: "Playwright service unavailable" - description: "edu-master/playwright-service deployment has {{ $value }} unavailable replica(s) for 10m. All webinar/diari/schedule checks fail without it." diff --git a/edu_master/k8s/namespace.yaml b/edu_master/k8s/namespace.yaml deleted file mode 100644 index e241a25..0000000 --- a/edu_master/k8s/namespace.yaml +++ /dev/null @@ -1,4 +0,0 @@ -apiVersion: v1 -kind: Namespace -metadata: - name: edu-master diff --git a/edu_master/k8s/playwright.yaml b/edu_master/k8s/playwright.yaml deleted file mode 100644 index c071a05..0000000 --- a/edu_master/k8s/playwright.yaml +++ /dev/null @@ -1,69 +0,0 @@ -apiVersion: apps/v1 -kind: Deployment -metadata: - name: playwright-service - namespace: edu-master - labels: - app: edu-master-playwright -spec: - replicas: 1 - selector: - matchLabels: - app: edu-master-playwright - strategy: - type: Recreate - template: - metadata: - labels: - app: edu-master-playwright - spec: - containers: - - name: playwright - # renovate: datasource=docker depName=mcr.microsoft.com/playwright versioning=docker - image: mcr.microsoft.com/playwright:v1.56.0-jammy - imagePullPolicy: IfNotPresent - # p95 412M, max 478M over 7 days, no limit before. Request is set at p95 - # so the pod is not an eviction candidate; the limit stays above 2x the - # request because browser page lifetimes are unpredictable. - resources: - requests: - cpu: "200m" - memory: "416Mi" - limits: - memory: "1Gi" - command: - - npx - - -y - - playwright@1.56.0 - - run-server - - --port - - "3000" - - --path - - /ws - ports: - - containerPort: 3000 - readinessProbe: - tcpSocket: - port: 3000 - initialDelaySeconds: 5 - periodSeconds: 10 - timeoutSeconds: 3 - livenessProbe: - tcpSocket: - port: 3000 - initialDelaySeconds: 15 - periodSeconds: 20 - timeoutSeconds: 3 ---- -apiVersion: v1 -kind: Service -metadata: - name: playwright-service - namespace: edu-master -spec: - selector: - app: edu-master-playwright - ports: - - name: ws - port: 3000 - targetPort: 3000 diff --git a/edu_master/k8s/redis.yaml b/edu_master/k8s/redis.yaml deleted file mode 100644 index e0fca76..0000000 --- a/edu_master/k8s/redis.yaml +++ /dev/null @@ -1,75 +0,0 @@ -apiVersion: apps/v1 -kind: StatefulSet -metadata: - name: redis - namespace: edu-master - labels: - app: edu-master-redis -spec: - serviceName: redis - replicas: 1 - selector: - matchLabels: - app: edu-master-redis - template: - metadata: - labels: - app: edu-master-redis - spec: - containers: - - name: redis - image: redis:8.10.2-alpine - imagePullPolicy: IfNotPresent - ports: - - containerPort: 6379 - volumeMounts: - - name: redis-data - mountPath: /data - resources: - requests: - cpu: 25m - memory: 32Mi - limits: - cpu: 250m - memory: 128Mi - readinessProbe: - exec: - command: ["redis-cli", "ping"] - initialDelaySeconds: 5 - periodSeconds: 5 - timeoutSeconds: 3 - livenessProbe: - exec: - command: ["redis-cli", "ping"] - initialDelaySeconds: 10 - periodSeconds: 10 - timeoutSeconds: 3 - volumes: - - name: redis-data - persistentVolumeClaim: - claimName: redis-data-pvc ---- -apiVersion: v1 -kind: PersistentVolumeClaim -metadata: - name: redis-data-pvc - namespace: edu-master -spec: - accessModes: - - ReadWriteOnce - resources: - requests: - storage: 1Gi ---- -apiVersion: v1 -kind: Service -metadata: - name: redis - namespace: edu-master -spec: - selector: - app: edu-master-redis - ports: - - name: redis - port: 6379 - targetPort: 6379 diff --git a/edu_master/k8s/restore-seed-job.yaml.example b/edu_master/k8s/restore-seed-job.yaml.example deleted file mode 100644 index c85b078..0000000 --- a/edu_master/k8s/restore-seed-job.yaml.example +++ /dev/null @@ -1,50 +0,0 @@ -# One-time Job to migrate redis state from docker compose to k8s (maintenance window). -# The .example file is not applied by the deploy pipeline (mask *.example.yaml). -# -# Runbook: -# 1. docker compose -f /edu_master/compose.yaml stop # SIGTERM -> redis will flush dump.rdb -# 2. docker run --rm -v edu_master_redis-data:/data \ -# -v /tmp/edu-master-backup:/backup \ -# redis:alpine sh -c "cp /data/dump.rdb /backup/ && ls -la /backup" -# 3. kubectl apply -f edu_master/k8s/namespace.yaml -# 4. kubectl apply -f # seed must come BEFORE redis pod starts -# 5. kubectl apply -f edu_master/k8s/restore-seed-job.yaml.example -# kubectl wait --for=condition=complete job/redis-restore-seed -n edu-master --timeout=120s -# 6. kubectl delete job redis-restore-seed -n edu-master -# 7. kubectl apply -f edu_master/k8s/ -R # apply remaining manifests -apiVersion: batch/v1 -kind: Job -metadata: - name: redis-restore-seed - namespace: edu-master -spec: - backoffLimit: 2 - ttlSecondsAfterFinished: 3600 - template: - spec: - restartPolicy: Never - containers: - - name: seed - image: redis:alpine - command: - - /bin/sh - - -ec - - | - ls -la /backup - cp /backup/dump.rdb /data/dump.rdb - chmod 644 /data/dump.rdb - ls -la /data - volumeMounts: - - name: redis-data - mountPath: /data - - name: backup - mountPath: /backup - readOnly: true - volumes: - - name: redis-data - persistentVolumeClaim: - claimName: redis-data-pvc - - name: backup - hostPath: - path: /tmp/edu-master-backup - type: DirectoryOrCreate diff --git a/edu_master/k8s/secrets.yaml.example b/edu_master/k8s/secrets.yaml.example deleted file mode 100644 index 7f1cf9b..0000000 --- a/edu_master/k8s/secrets.yaml.example +++ /dev/null @@ -1,29 +0,0 @@ -apiVersion: v1 -kind: Secret -metadata: - name: edu-master-secrets - namespace: edu-master -type: Opaque -stringData: - # Session keeper credentials - KEEPER_LOGIN: "" - KEEPER_PASSWORD: "" - KEEPER_INTERVAL: "10" - # EDU links - EDU_URL_BASE: "https://edu.edu.vn.ua" - EDU_URL_LOGIN: "/user/login" - EDU_URL_COURSES: "/course/userlist" - EDU_URL_WEBINAR: "/webinar/useractive" - # Playwright - USER_AGENT: "" - PLAYWRIGHT_WS: "ws://playwright-service:3000/ws" - # Webinar-checker - WEBINAR_TELEGRAM_TOKEN: "" - WEBINAR_ADMIN_ID: "" - WEBINAR_CHECK_INTERVAL: "60" - # Prometheus metrics endpoint (scraped via ServiceMonitor, alerts in k8s/alerts.yaml) - METRICS_PORT: "8000" - # Database - REDIS_HOST: "redis" - REDIS_PORT: "6379" - TZ: "Europe/Kyiv" diff --git a/edu_master/k8s/service.yaml b/edu_master/k8s/service.yaml deleted file mode 100644 index a20026e..0000000 --- a/edu_master/k8s/service.yaml +++ /dev/null @@ -1,15 +0,0 @@ -apiVersion: v1 -kind: Service -metadata: - name: webinar-checker - namespace: edu-master - labels: - app: edu-master-webinar-checker -spec: - selector: - app: edu-master-webinar-checker - ports: - - name: metrics - port: 8000 - targetPort: metrics - protocol: TCP diff --git a/edu_master/k8s/session-keeper.yaml b/edu_master/k8s/session-keeper.yaml deleted file mode 100644 index d9f5cda..0000000 --- a/edu_master/k8s/session-keeper.yaml +++ /dev/null @@ -1,53 +0,0 @@ -apiVersion: apps/v1 -kind: Deployment -metadata: - name: session-keeper - namespace: edu-master - labels: - app: edu-master-session-keeper -spec: - replicas: 1 - selector: - matchLabels: - app: edu-master-session-keeper - strategy: - type: Recreate - template: - metadata: - labels: - app: edu-master-session-keeper - spec: - initContainers: - - name: wait-redis - image: redis:8.10.2-alpine - command: - - /bin/sh - - -ec - - | - i=0 - until redis-cli -h redis ping | grep -q PONG; do - i=$((i+1)) - [ "$i" -ge 300 ] && echo "TIMEOUT: redis not ready" && exit 1 - sleep 2 - done - echo "redis is ready" - containers: - - name: session-keeper - image: gcr.forust.xyz/forust/session-keeper:prod - envFrom: - - secretRef: - name: edu-master-secrets - resources: - requests: - cpu: 25m - memory: 32Mi - limits: - cpu: 250m - memory: 128Mi - readinessProbe: - exec: - command: ["/bin/sh", "-ec", "redis-cli -h redis EXISTS EDU_PHPSESSID | grep -q 1"] - initialDelaySeconds: 15 - periodSeconds: 30 - timeoutSeconds: 5 - failureThreshold: 10 diff --git a/edu_master/k8s/webinar-checker.yaml b/edu_master/k8s/webinar-checker.yaml deleted file mode 100644 index a8f275e..0000000 --- a/edu_master/k8s/webinar-checker.yaml +++ /dev/null @@ -1,75 +0,0 @@ -apiVersion: apps/v1 -kind: Deployment -metadata: - name: webinar-checker - namespace: edu-master - labels: - app: edu-master-webinar-checker -spec: - replicas: 1 - selector: - matchLabels: - app: edu-master-webinar-checker - strategy: - type: Recreate - template: - metadata: - labels: - app: edu-master-webinar-checker - spec: - # Enforces dependency order like compose depends_on: - # redis healthy -> session-keeper healthy (EXISTS EDU_PHPSESSID) -> playwright started - initContainers: - - name: wait-deps - image: redis:8.10.2-alpine - command: - - /bin/sh - - -ec - - | - i=0 - until redis-cli -h redis ping | grep -q PONG; do - i=$((i+1)) - [ "$i" -ge 300 ] && echo "TIMEOUT: redis not ready" && exit 1 - sleep 2 - done - echo "redis ok" - until [ "$(redis-cli -h redis EXISTS EDU_PHPSESSID)" = "1" ]; do - i=$((i+1)) - [ "$i" -ge 300 ] && echo "TIMEOUT: no PHPSESSID (session-keeper down?)" && exit 1 - sleep 2 - done - echo "PHPSESSID ok" - until nc -z playwright-service 3000; do - i=$((i+1)) - [ "$i" -ge 300 ] && echo "TIMEOUT: playwright-service not reachable" && exit 1 - sleep 2 - done - echo "playwright ok" - containers: - - name: webinar-checker - image: gcr.forust.xyz/forust/webinar-checker:prod - ports: - - name: metrics - containerPort: 8000 - protocol: TCP - readinessProbe: - httpGet: - path: /health - port: metrics - periodSeconds: 10 - timeoutSeconds: 3 - failureThreshold: 12 - initialDelaySeconds: 10 - envFrom: - - secretRef: - name: edu-master-secrets - env: - - name: TZ - value: "Europe/Kyiv" - resources: - requests: - cpu: "50m" - memory: "192Mi" - limits: - cpu: "600m" - memory: "384Mi" diff --git a/edu_master/phpsessid-bot/Dockerfile b/edu_master/phpsessid-bot/Dockerfile deleted file mode 100644 index 18a7a51..0000000 --- a/edu_master/phpsessid-bot/Dockerfile +++ /dev/null @@ -1,15 +0,0 @@ -FROM python:3.11-slim - -WORKDIR /app - -# Install system dependencies -RUN apt-get update && apt-get install -y --no-install-recommends redis-tools && rm -rf /var/lib/apt/lists/* - -# Install dependencies -RUN pip install --no-cache-dir requests==2.32.3 redis==5.2.1 - -# Copy application code -COPY . . - -# Run the bot -CMD ["python", "bot.py"] diff --git a/edu_master/phpsessid-bot/bot.py b/edu_master/phpsessid-bot/bot.py deleted file mode 100644 index 963a143..0000000 --- a/edu_master/phpsessid-bot/bot.py +++ /dev/null @@ -1,132 +0,0 @@ -import logging -import os -import time -from datetime import datetime - -import redis -import requests - -# Configure logging -logging.basicConfig(level=logging.INFO, format='%(asctime)s - %(levelname)s - %(message)s') -logger = logging.getLogger(__name__) - - -# Load configuration (adapted to .env keys) -def _env(key, default=None): - v = os.getenv(key, default) - if isinstance(v, str) and len(v) >= 2 and ((v[0] == '"' and v[-1] == '"') or (v[0] == "'" and v[-1] == "'")): - return v[1:-1] - return v - - -LOGIN = _env('KEEPER_LOGIN') -PASSWORD = _env('KEEPER_PASSWORD') - -EDU_BASE = _env('EDU_URL_BASE', 'https://edu.edu.vn.ua') -EDU_LOGIN_PATH = _env('EDU_URL_LOGIN', '/user/login') -EDU_COURSES_PATH = _env('EDU_URL_COURSES', '/course/userlist') -URL_LOGIN = f'{EDU_BASE.rstrip("/")}/{EDU_LOGIN_PATH.lstrip("/")}' -URL_VERIFY = f'{EDU_BASE.rstrip("/")}/{EDU_COURSES_PATH.lstrip("/")}' - -INTERVAL = int(_env('KEEPER_INTERVAL', 10)) -USER_AGENT = _env( - 'USER_AGENT', - 'Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/142.0.0.0 Safari/537.36', -) -REDIS_HOST = _env('REDIS_HOST', 'redis') -REDIS_PORT = int(_env('REDIS_PORT', 6379)) - -SUCCESS_FILE = '/tmp/last_success' # noqa: S108 - - -def touch_success_file(): - """Updates the timestamp of the success file for healthchecks.""" - try: - with open(SUCCESS_FILE, 'w') as f: - f.write(str(datetime.now().timestamp())) - except Exception as e: - logger.error(f'Failed to touch success file: {e}') - - -def main(): - logger.info('Starting Session Keeper Bot') - - # Connect to Redis - try: - redis_client = redis.Redis(host=REDIS_HOST, port=REDIS_PORT, decode_responses=True) - redis_client.ping() - logger.info(f'Connected to Redis at {REDIS_HOST}:{REDIS_PORT}') - except Exception as e: - logger.error(f'Failed to connect to Redis: {e}') - return - - session = requests.Session() - - # Set headers - headers = { - 'User-Agent': USER_AGENT, - 'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,image/apng,*/*;q=0.8,application/signed-exchange;v=b3;q=0.7', - 'Accept-Language': 'en-US,en;q=0.9', - 'Cache-Control': 'max-age=0', - 'Upgrade-Insecure-Requests': '1', - 'Sec-Fetch-Site': 'same-origin', - 'Sec-Fetch-Mode': 'navigate', - 'Sec-Fetch-User': '?1', - 'Sec-Fetch-Dest': 'document', - 'Sec-Ch-Ua': '"Not_A Brand";v="99", "Chromium";v="142"', - 'Sec-Ch-Ua-Mobile': '?0', - 'Sec-Ch-Ua-Platform': '"Linux"', - 'Accept-Encoding': 'gzip, deflate, br', - 'Priority': 'u=0, i', - } - session.headers.update(headers) - - while True: - try: - logger.info('Attempting login...') - - # Login payload - payload = {'login': LOGIN, 'password': PASSWORD} - - # Perform Login - # Note: The user request shows a POST to /user/login with form data - # We need to make sure we handle the PHPSESSID correctly. - # If we already have a PHPSESSID, requests will send it. - - login_response = session.post(URL_LOGIN, data=payload, allow_redirects=True) - - logger.info(f'Login Response Status: {login_response.status_code}') - logger.info(f'Cookies after login: {session.cookies.get_dict()}') - - # Verify Session - logger.info('Verifying session...') - verify_response = session.get(URL_VERIFY, allow_redirects=False) - - logger.info(f'Verify Response Status: {verify_response.status_code}') - - if verify_response.status_code == 200: - logger.info('Session verification SUCCESS (200 OK).') - touch_success_file() - - # Save PHPSESSID to Redis - phpsessid = session.cookies.get('PHPSESSID') - if phpsessid: - try: - redis_client.set('EDU_PHPSESSID', phpsessid) - logger.info(f'Saved PHPSESSID to Redis: {phpsessid}') - except Exception as e: - logger.error(f'Failed to save PHPSESSID to Redis: {e}') - elif verify_response.status_code == 302: - logger.warning('Session verification FAILED (302 Redirect). Session might be invalid.') - else: - logger.warning(f'Session verification returned unexpected status: {verify_response.status_code}') - - except Exception as e: - logger.error(f'An error occurred: {e}') - - logger.info(f'Sleeping for {INTERVAL} minutes...') - time.sleep(INTERVAL * 60) - - -if __name__ == '__main__': - main() diff --git a/edu_master/webinar-checker/Dockerfile b/edu_master/webinar-checker/Dockerfile deleted file mode 100644 index 3f8d5fa..0000000 --- a/edu_master/webinar-checker/Dockerfile +++ /dev/null @@ -1,13 +0,0 @@ -FROM python:3.11-slim - -WORKDIR /app - -# renovate: datasource=pypi depName=playwright versioning=pep440 -ARG PLAYWRIGHT_VERSION=1.56.0 - -# Install dependencies - PLAYWRIGHT_VERSION is single-source, renovate updates ARG above and all other places via regexManagers -RUN pip install --no-cache-dir pip==25.0.1 && pip install --no-cache-dir playwright==${PLAYWRIGHT_VERSION} redis==5.2.1 requests==2.32.3 "python-telegram-bot[job-queue]==21.10" - -COPY checker.py . - -CMD ["python", "checker.py"] diff --git a/edu_master/webinar-checker/checker.py b/edu_master/webinar-checker/checker.py deleted file mode 100644 index ae3de7a..0000000 --- a/edu_master/webinar-checker/checker.py +++ /dev/null @@ -1,1821 +0,0 @@ -import asyncio -import contextlib -import json -import logging -import os -import re -import tempfile -import threading -import time -from datetime import datetime, timedelta -from html import escape -from http.server import BaseHTTPRequestHandler, HTTPServer - -import redis -from playwright.async_api import async_playwright -from telegram import ChatMember, InlineKeyboardButton, InlineKeyboardMarkup, Update -from telegram.constants import ChatType -from telegram.ext import Application, CallbackQueryHandler, CommandHandler, ContextTypes - -# Logger -logging.basicConfig(level=logging.INFO, format='%(asctime)s - %(levelname)s - %(message)s') -logger = logging.getLogger(__name__) - -# Suppress HTTP request logs -logging.getLogger('urllib3').setLevel(logging.WARNING) -logging.getLogger('httpx').setLevel(logging.WARNING) -logging.getLogger('telegram.ext._application').setLevel(logging.WARNING) - - -# Load environment variables -def _env(key, default=None): - v = os.getenv(key, default) - if isinstance(v, str) and len(v) >= 2 and ((v[0] == '"' and v[-1] == '"') or (v[0] == "'" and v[-1] == "'")): - return v[1:-1] - return v - - -EDU_BASE = _env('EDU_URL_BASE', 'https://edu.edu.vn.ua') -EDU_WEBINAR_PATH = _env('EDU_URL_WEBINAR', '/webinar/useractive') -WEBINAR_URL = f'{EDU_BASE.rstrip("/")}/{EDU_WEBINAR_PATH.lstrip("/")}' -DIARY_URL = f'{EDU_BASE.rstrip("/")}/user/diary' -SCHEDULE_URL = f'{EDU_BASE.rstrip("/")}/lessons/table' - -WEBINAR_CHECK_INTERVAL = int(_env('WEBINAR_CHECK_INTERVAL', 60)) -REDIS_HOST = _env('REDIS_HOST', 'redis') -REDIS_PORT = int(_env('REDIS_PORT', 6379)) -PLAYWRIGHT_WS = _env('PLAYWRIGHT_WS', 'ws://playwright-service:3000/ws') -USER_AGENT = _env( - 'USER_AGENT', - 'Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/142.0.0.0 Safari/537.36', -) -WEBINAR_TELEGRAM_TOKEN = _env('WEBINAR_TELEGRAM_TOKEN') -ADMIN_ID = int(_env('WEBINAR_ADMIN_ID', '0')) -METRICS_PORT = int(_env('METRICS_PORT', '8000')) - -# --- Prometheus metrics (stdlib only, no extra deps) --- -# Scraped by prometheus-stack via ServiceMonitor (edu_master/k8s/servicemonitor.yaml). -# Critical alerts in edu_master/k8s/alerts.yaml fire to Telegram via Alertmanager. -_METRICS_LOCK = threading.Lock() -_METRICS = { - 'last_run': 0.0, # Unix ts of last check start - 'last_success': 0.0, # Unix ts of last successful check - 'last_duration': 0.0, # Duration of last check in seconds - 'success_total': 0, - 'failure_total': 0, - 'consecutive_failures': 0, - 'phpsessid_present': 1, # 1 if EDU_PHPSESSID found in redis, else 0 -} - - -def _metric_check_start(): - with _METRICS_LOCK: - _METRICS['last_run'] = time.time() - - -def _metric_check_ok(duration: float): - now = time.time() - with _METRICS_LOCK: - _METRICS['last_success'] = now - _METRICS['last_duration'] = duration - _METRICS['success_total'] += 1 - _METRICS['consecutive_failures'] = 0 - _METRICS['phpsessid_present'] = 1 - - -def _metric_check_fail(duration: float, phpsessid_missing: bool = False): - with _METRICS_LOCK: - _METRICS['last_duration'] = duration - _METRICS['failure_total'] += 1 - _METRICS['consecutive_failures'] += 1 - _METRICS['phpsessid_present'] = 0 if phpsessid_missing else 1 - - -def _metrics_render() -> bytes: - with _METRICS_LOCK: - m = dict(_METRICS) - lines = [ - '# HELP webinar_check_last_run_timestamp_seconds Unix timestamp of last webinar check start.', - '# TYPE webinar_check_last_run_timestamp_seconds gauge', - f'webinar_check_last_run_timestamp_seconds {m["last_run"]}', - '# HELP webinar_check_last_success_timestamp_seconds Unix timestamp of last successful webinar check.', - '# TYPE webinar_check_last_success_timestamp_seconds gauge', - f'webinar_check_last_success_timestamp_seconds {m["last_success"]}', - '# HELP webinar_check_last_duration_seconds Duration of last webinar check in seconds.', - '# TYPE webinar_check_last_duration_seconds gauge', - f'webinar_check_last_duration_seconds {m["last_duration"]}', - '# HELP webinar_check_success_total Total successful webinar checks.', - '# TYPE webinar_check_success_total counter', - f'webinar_check_success_total {m["success_total"]}', - '# HELP webinar_check_failure_total Total failed webinar checks (timeout, playwright error, page error).', - '# TYPE webinar_check_failure_total counter', - f'webinar_check_failure_total {m["failure_total"]}', - '# HELP webinar_check_consecutive_failures Consecutive failed webinar checks (reset on success).', - '# TYPE webinar_check_consecutive_failures gauge', - f'webinar_check_consecutive_failures {m["consecutive_failures"]}', - '# HELP edu_phpsessid_present 1 if EDU_PHPSESSID exists in redis, 0 otherwise.', - '# TYPE edu_phpsessid_present gauge', - f'edu_phpsessid_present {m["phpsessid_present"]}', - ] - return ('\n'.join(lines) + '\n').encode() - - -class _MetricsHandler(BaseHTTPRequestHandler): - def do_GET(self): - if self.path == '/metrics': - body = _metrics_render() - self.send_response(200) - self.send_header('Content-Type', 'text/plain; version=0.0.4') - self.send_header('Content-Length', str(len(body))) - self.end_headers() - self.wfile.write(body) - elif self.path in ('/healthz', '/health'): - body = b'ok\n' - self.send_response(200) - self.send_header('Content-Type', 'text/plain') - self.send_header('Content-Length', str(len(body))) - self.end_headers() - self.wfile.write(body) - else: - self.send_response(404) - self.end_headers() - - def log_message(self, *args): - pass # keep bot logs clean - - -def start_metrics_server(port: int = METRICS_PORT): - server = HTTPServer(('0.0.0.0', port), _MetricsHandler) # noqa: S104 - k8s ServiceMonitor scrapes pod IP - thread = threading.Thread(target=server.serve_forever, name='metrics-server', daemon=True) - thread.start() - logger.info(f'Metrics server listening on :{port}/metrics') - return server - - -# Redis Keys -KEY_WHITELIST = 'bot:whitelist' -KEY_WHITELIST_ENABLED = 'bot:whitelist_enabled' -KEY_SUBSCRIBERS = 'bot:subscribers' -KEY_PHPSESSID = 'EDU_PHPSESSID' -KEY_WEBINAR_HISTORY = 'bot:webinar_history' # Stores last 3 webinars - -# Marker shown by the site when there are no active online lessons -NO_WEBINAR_MARKER = 'Жодного онлайн уроку зараз' - -# Initialize Redis -try: - redis_client = redis.Redis(host=REDIS_HOST, port=REDIS_PORT, decode_responses=True) - redis_client.ping() - logger.info(f'Connected to Redis at {REDIS_HOST}:{REDIS_PORT}') -except Exception as e: - logger.error(f'Failed to connect to Redis: {e}') - exit(1) - -# --- Translations --- - -TRANSLATIONS = { - 'ru': { - 'welcome': '👋 Привет, {name}!\n\nЯ бот-уведомитель о вебинарах. Я буду сообщать вам, когда появится новый вебинар.\nВы подписаны на уведомления.', - 'welcome_admin': '\n\n👑 Режим администратора активен', - 'access_denied': '⛔ Доступ запрещен. Вас нет в белом списке.', - 'help_title': '🤖 Помощь по боту\n\n', - 'help_commands': '/start - Подписаться на уведомления\n/stop - Отписаться от уведомлений\n/help - Показать это сообщение\n/language - Сменить язык', - 'help_admin': '\nКоманды администратора:\n/adduser [user_id] - Добавить пользователя в белый список\n/removeuser [user_id] - Удалить пользователя из белого списка\nИли используйте панель ниже для управления настройками.', - 'admin_only': '⛔ Только для администратора!', - 'user_added': '✅ Пользователь {user_id} добавлен в белый список', - 'user_removed': '✅ Пользователь {user_id} удален из белого списка', - 'user_not_in_whitelist': '⚠️ Пользователь {user_id} не был в белом списке', - 'cannot_remove_admin': '❌ Невозможно удалить администратора из белого списка', - 'invalid_user_id': '❌ Неверный ID пользователя. Должно быть число.', - 'usage_adduser': 'Использование: /adduser [user_id]', - 'usage_removeuser': 'Использование: /removeuser [user_id]', - 'whitelist_enabled': '✅ Белый список включен', - 'whitelist_disabled': '🚫 Белый список отключен', - 'whitelist_title': '📋 Белый список:\n', - 'subscribers_title': '👥 Подписчики:\n', - 'empty': 'Пусто', - 'force_check_running': '🔄 Запускаю проверку...', - 'check_failed': '❌ Проверка не удалась. Смотрите логи.', - 'check_completed_none': '✅ Проверка завершена. Вебинаров не найдено.', - 'check_completed': '✅ Проверка завершена. Найдено {count} вебинар(ов)!', - 'toggle_whitelist_disable': '🔒 Отключить белый список', - 'toggle_whitelist_enable': '🔓 Включить белый список', - 'view_whitelist': '📋 Посмотреть белый список', - 'view_subscribers': '👥 Посмотреть подписчиков', - 'force_check': '🔄 Принудительная проверка', - 'webinar_found': '🎓 Новый вебинар!\n\n', - 'webinar_item': '📌 {name}\n🔗 https://edu.edu.vn.ua{url}', - 'select_language': '🌐 Выберите язык / Оберіть мову / Select language:', - 'language_changed': '✅ Язык изменен на {lang}', - 'flag_ru': '🇷🇺 Русский', - 'flag_uk': '🇺🇦 Українська', - 'flag_en': '🇬🇧 English', - 'history_cleared': '✅ История вебинаров очищена', - 'history_clear_failed': '❌ Ошибка при очистке истории', - 'today': '📌 Сегодня', - 'tomorrow': '📌 Завтра', - 'week': '📅 Эта неделя', - 'month': '📅 Весь месяц', - 'no_events': 'Нет событий', - 'no_lessons': 'Нет уроков', - 'time_unknown': 'неизвестно', - 'free_period': 'Свободно', - 'unsubscribed': '🔕 Вы отписались от уведомлений.', - 'diary_title': '📅 Дневник — выберите период:', - 'diary_week_title': '📅 Неделя {start} – {end}', - 'schedule_title': '📅 {weekday} — {class_num} класс', - 'loading_diary': '🔄 Загружаю дневник...', - 'diary_load_failed': '❌ Не удалось загрузить дневник.', - 'diary_no_session': '❌ Не удалось загрузить дневник. Нет сессии или ошибка.', - 'diary_next_month': '❌ Данные за следующий месяц недоступны. Перейдите на сайт.', - 'schedule_load_failed': '❌ Не удалось загрузить расписание.', - 'class_not_set': '❌ Класс не настроен. Используйте /setclass.', - 'class_not_found': '❌ Классы не найдены в расписании.', - 'select_class': '🎒 Выберите класс — сохранится и больше не спросится:', - 'admin_class_not_set': '⚠️ Администратор еще не настроил класс для этой группы. Используйте /setclass 11', - 'loading_schedule': '🔄 Загружаю расписание...', - 'invalid_class_num': '❌ Неверный номер класса.', - 'class_saved': '✅ Класс {class_num} сохранен.', - 'admin_only_msg': '⚠️ Только администратор может настроить класс для группы.', - 'truncation': '\n\n✂️ ...(обрезано)', - 'day_mon': 'Пн', - 'day_tue': 'Вт', - 'day_wed': 'Ср', - 'day_thu': 'Чт', - 'day_fri': 'Пт', - 'day_sat': 'Сб', - 'day_sun': 'Вс', - }, - 'uk': { - 'welcome': "👋 Привіт, {name}!\n\nЯ бот-сповіщувач про вебінари. Я повідомлятиму вас, коли з'явиться новий вебінар.\nВи підписані на сповіщення.", - 'welcome_admin': '\n\n👑 Режим адміністратора активний', - 'access_denied': '⛔ Доступ заборонено. Вас немає в білому списку.', - 'help_title': '🤖 Довідка по боту\n\n', - 'help_commands': '/start - Підписатися на сповіщення\n/stop - Відписатися від сповіщень\n/help - Показати це повідомлення\n/language - Змінити мову', - 'help_admin': '\nКоманди адміністратора:\n/adduser [user_id] - Додати користувача до білого списку\n/removeuser [user_id] - Видалити користувача з білого списку\nАбо використовуйте панель нижче для керування налаштуваннями.', - 'admin_only': '⛔ Тільки для адміністратора!', - 'user_added': '✅ Користувач {user_id} доданий до білого списку', - 'user_removed': '✅ Користувач {user_id} видалений з білого списку', - 'user_not_in_whitelist': '⚠️ Користувач {user_id} не був у білому списку', - 'cannot_remove_admin': '❌ Неможливо видалити адміністратора з білого списку', - 'invalid_user_id': '❌ Невірний ID користувача. Має бути число.', - 'usage_adduser': 'Використання: /adduser [user_id]', - 'usage_removeuser': 'Використання: /removeuser [user_id]', - 'whitelist_enabled': '✅ Білий список увімкнено', - 'whitelist_disabled': '🚫 Білий список вимкнено', - 'whitelist_title': '📋 Білий список:\n', - 'subscribers_title': '👥 Підписники:\n', - 'empty': 'Порожньо', - 'force_check_running': '🔄 Запускаю перевірку...', - 'check_failed': '❌ Перевірка не вдалася. Дивіться логи.', - 'check_completed_none': '✅ Перевірка завершена. Вебінарів не знайдено.', - 'check_completed': '✅ Перевірка завершена. Знайдено {count} вебінар(ів)!', - 'toggle_whitelist_disable': '🔒 Вимкнути білий список', - 'toggle_whitelist_enable': '🔓 Увімкнути білий список', - 'view_whitelist': '📋 Переглянути білий список', - 'view_subscribers': '👥 Переглянути підписників', - 'force_check': '🔄 Примусова перевірка', - 'webinar_found': '🎓 Новий вебінар!\n\n', - 'webinar_item': '📌 {name}\n🔗 https://edu.edu.vn.ua{url}', - 'select_language': '🌐 Виберіть мову / Выберите язык / Select language:', - 'language_changed': '✅ Мову змінено на {lang}', - 'flag_ru': '🇷🇺 Русский', - 'flag_uk': '🇺🇦 Українська', - 'flag_en': '🇬🇧 English', - 'history_cleared': '✅ Історія вебінарів очищена', - 'history_clear_failed': '❌ Помилка при очищенні історії', - 'today': '📌 Сьогодні', - 'tomorrow': '📌 Завтра', - 'week': '📅 Цей тиждень', - 'month': '📅 Весь місяць', - 'no_events': 'Немає подій', - 'no_lessons': 'Немає уроків', - 'time_unknown': 'невідомо', - 'free_period': 'Вільно', - 'unsubscribed': '🔕 Ви відписалися від сповіщень.', - 'diary_title': '📅 Щоденник — виберіть період:', - 'diary_week_title': '📅 Тиждень {start} – {end}', - 'schedule_title': '📅 {weekday} — {class_num} клас', - 'loading_diary': '🔄 Завантажую щоденник...', - 'diary_load_failed': '❌ Не вдалося завантажити щоденник.', - 'diary_no_session': '❌ Не вдалося завантажити щоденник. Немає сесії або помилка.', - 'diary_next_month': '❌ Дані за наступний місяць недоступні. Перейдіть на сайт.', - 'schedule_load_failed': '❌ Не вдалося завантажити розклад.', - 'class_not_set': '❌ Клас не налаштовано. Використайте /setclass.', - 'class_not_found': '❌ Класи не знайдені в розкладі.', - 'select_class': '🎒 Оберіть клас — збережеться і більше не питатиметься:', - 'admin_class_not_set': '⚠️ Адміністратор ще не налаштував клас для цієї групи. Використайте /setclass 11', - 'loading_schedule': '🔄 Завантажую розклад...', - 'invalid_class_num': '❌ Невірний номер класу.', - 'class_saved': '✅ Клас {class_num} збережено.', - 'admin_only_msg': '⚠️ Тільки адміністратор може налаштувати клас для групи.', - 'truncation': '\n\n✂️ ...(обрізано)', - 'day_mon': 'Пн', - 'day_tue': 'Вт', - 'day_wed': 'Ср', - 'day_thu': 'Чт', - 'day_fri': 'Пт', - 'day_sat': 'Сб', - 'day_sun': 'Нд', - }, - 'en': { - 'welcome': '👋 Hello, {name}!\n\nI am the Webinar Checker Bot. I will notify you when a new webinar appears.\nYou have been subscribed to notifications.', - 'welcome_admin': '\n\n👑 Admin Mode Active', - 'access_denied': '⛔ Access denied. You are not on the whitelist.', - 'help_title': '🤖 Bot Help\n\n', - 'help_commands': '/start - Subscribe to notifications\n/stop - Unsubscribe from notifications\n/help - Show this message\n/language - Change language', - 'help_admin': '\nAdmin Commands:\n/adduser [user_id] - Add user to whitelist\n/removeuser [user_id] - Remove user from whitelist\nOr use the panel below to manage settings.', - 'admin_only': '⛔ Admin only!', - 'user_added': '✅ User {user_id} added to whitelist', - 'user_removed': '✅ User {user_id} removed from whitelist', - 'user_not_in_whitelist': '⚠️ User {user_id} was not in whitelist', - 'cannot_remove_admin': '❌ Cannot remove admin from whitelist', - 'invalid_user_id': '❌ Invalid user ID. Must be a number.', - 'usage_adduser': 'Usage: /adduser [user_id]', - 'usage_removeuser': 'Usage: /removeuser [user_id]', - 'whitelist_enabled': '✅ Whitelist Enabled', - 'whitelist_disabled': '🚫 Whitelist Disabled', - 'whitelist_title': '📋 Whitelist:\n', - 'subscribers_title': '👥 Subscribers:\n', - 'empty': 'Empty', - 'force_check_running': '🔄 Running immediate check...', - 'check_failed': '❌ Check failed. See logs for details.', - 'check_completed_none': '✅ Check completed. No webinars found.', - 'check_completed': '✅ Check completed. Found {count} webinar(s)!', - 'toggle_whitelist_disable': '🔒 Disable Whitelist', - 'toggle_whitelist_enable': '🔓 Enable Whitelist', - 'view_whitelist': '📋 View Whitelist', - 'view_subscribers': '👥 View Subscribers', - 'force_check': '🔄 Force Check', - 'webinar_found': '🎓 New webinar found!\n\n', - 'webinar_item': '📌 {name}\n🔗 https://edu.edu.vn.ua{url}', - 'select_language': '🌐 Select language / Виберіть мову / Выберите язык:', - 'language_changed': '✅ Language changed to {lang}', - 'flag_ru': '🇷🇺 Русский', - 'flag_uk': '🇺🇦 Українська', - 'flag_en': '🇬🇧 English', - 'history_cleared': '✅ Webinar history cleared', - 'history_clear_failed': '❌ Error clearing history', - 'today': '📌 Today', - 'tomorrow': '📌 Tomorrow', - 'week': '📅 This week', - 'month': '📅 Whole month', - 'no_events': 'No events', - 'no_lessons': 'No lessons', - 'time_unknown': 'unknown', - 'free_period': 'Free', - 'unsubscribed': '🔕 You have unsubscribed from notifications.', - 'diary_title': '📅 Diary — choose a period:', - 'diary_week_title': '📅 Week {start} – {end}', - 'schedule_title': '📅 {weekday} — {class_num} class', - 'loading_diary': '🔄 Loading diary...', - 'diary_load_failed': '❌ Failed to load diary.', - 'diary_no_session': '❌ Failed to load diary. No session or error.', - 'diary_next_month': '❌ Next month data is not available. Please visit the website.', - 'schedule_load_failed': '❌ Failed to load schedule.', - 'class_not_set': '❌ Class is not set. Use /setclass.', - 'class_not_found': '❌ No classes found in the schedule.', - 'select_class': '🎒 Choose a class — it will be saved and not asked again:', - 'admin_class_not_set': '⚠️ Admin has not set a class for this group yet. Use /setclass 11', - 'loading_schedule': '🔄 Loading schedule...', - 'invalid_class_num': '❌ Invalid class number.', - 'class_saved': '✅ Class {class_num} saved.', - 'admin_only_msg': '⚠️ Only an admin can set the class for this group.', - 'truncation': '\n\n✂️ ...(truncated)', - 'day_mon': 'Mon', - 'day_tue': 'Tue', - 'day_wed': 'Wed', - 'day_thu': 'Thu', - 'day_fri': 'Fri', - 'day_sat': 'Sat', - 'day_sun': 'Sun', - }, -} - -# --- Language Helper Functions --- - - -LANG_NAME_MAP = {'ru': 'Русский', 'uk': 'Українська', 'en': 'English'} - - -def get_user_language(user_id: int) -> str: - """Get user's preferred language from Redis. Default: English.""" - lang = redis_client.get(f'user:{user_id}:language') - return lang if lang in ['ru', 'uk', 'en'] else 'en' - - -def set_user_language(user_id: int, lang: str): - """Save user's language preference to Redis.""" - if lang in ['ru', 'uk', 'en']: - redis_client.set(f'user:{user_id}:language', lang) - logger.info(f'User {user_id} language set to {lang}') - - -def t(user_id: int, key: str, **kwargs) -> str: - """Translate message for user with optional formatting. - - Fallback chain: user lang → en → uk → ru, otherwise return the key. - """ - lang = get_user_language(user_id) - message = None - for fallback in (lang, 'en', 'uk', 'ru'): - message = TRANSLATIONS.get(fallback, {}).get(key) - if message is not None: - break - if message is None: - return key - if kwargs: - return message.format(**kwargs) - return message - - -def _tr(lang: str, key: str, **kwargs) -> str: - """Translate by explicit lang code (for format_* helpers). - - Fallback chain: lang → en → uk → ru, otherwise return the key. - """ - message = None - for fallback in (lang, 'en', 'uk', 'ru'): - message = TRANSLATIONS.get(fallback, {}).get(key) - if message is not None: - break - if message is None: - return key - if kwargs: - return message.format(**kwargs) - return message - - -def get_chat_language(chat_id: int) -> str: - """Get group chat language from Redis. Default: Ukrainian.""" - lang = redis_client.get(f'chat:{chat_id}:language') - return lang if lang in ['ru', 'uk', 'en'] else 'uk' - - -def set_chat_language(chat_id: int, lang: str): - """Save group chat language preference to Redis.""" - if lang in ['ru', 'uk', 'en']: - redis_client.set(f'chat:{chat_id}:language', lang) - logger.info(f'Chat {chat_id} language set to {lang}') - - -def resolve_lang(chat, user_id=None) -> str: - """Resolve effective language: private → user lang, groups → chat lang (default 'uk').""" - chat_id = getattr(chat, 'id', chat) - chat_type = getattr(chat, 'type', None) - if chat_type is None: - try: - is_private = int(chat_id) > 0 - except (TypeError, ValueError): - is_private = True - if is_private: - uid = user_id if user_id is not None else chat_id - return get_user_language(int(uid)) - return get_chat_language(int(chat_id)) - if chat_type in (ChatType.PRIVATE, 'private'): - uid = user_id if user_id is not None else chat_id - return get_user_language(int(uid)) - return get_chat_language(chat_id) - - -def t_chat(chat, user_id, key: str, **kwargs) -> str: - """Translate using chat-resolved language (private → user, groups → chat 'uk' default).""" - return _tr(resolve_lang(chat, user_id), key, **kwargs) - - -def get_language_keyboard(user_id: int | None = None): - """Generate language selection keyboard.""" - lang = get_user_language(user_id) if user_id is not None else 'en' - keyboard = [ - [ - InlineKeyboardButton(_tr(lang, 'flag_ru'), callback_data='lang_ru'), - InlineKeyboardButton(_tr(lang, 'flag_uk'), callback_data='lang_uk'), - ], - [ - InlineKeyboardButton(_tr(lang, 'flag_en'), callback_data='lang_en'), - ], - ] - return InlineKeyboardMarkup(keyboard) - - -# --- Helper Functions --- - - -def is_whitelisted(user_id: int) -> bool: - """Check if user is allowed to use the bot.""" - if user_id == ADMIN_ID: - return True - - enabled = redis_client.get(KEY_WHITELIST_ENABLED) - if enabled == '0': # Whitelist disabled - return True - - return redis_client.sismember(KEY_WHITELIST, str(user_id)) - - -async def is_group_admin(update: Update, context: ContextTypes.DEFAULT_TYPE) -> bool: - """Check if the user is an administrator in the group.""" - user = update.effective_user - chat = update.effective_chat - - if chat.type in [ChatType.PRIVATE, 'private']: - return True - - try: - member = await context.bot.get_chat_member(chat.id, user.id) - return member.status in [ChatMember.OWNER, ChatMember.ADMINISTRATOR] - except Exception as e: - logger.error(f'Failed to check admin status: {e}') - return False - - -def get_admin_keyboard(user_id: int): - """Generate admin panel keyboard.""" - whitelist_enabled = redis_client.get(KEY_WHITELIST_ENABLED) != '0' - toggle_text = t(user_id, 'toggle_whitelist_disable') if whitelist_enabled else t(user_id, 'toggle_whitelist_enable') - - keyboard = [ - [InlineKeyboardButton(toggle_text, callback_data='toggle_whitelist')], - [InlineKeyboardButton(t(user_id, 'view_whitelist'), callback_data='view_whitelist')], - [InlineKeyboardButton(t(user_id, 'view_subscribers'), callback_data='view_subscribers')], - [InlineKeyboardButton(t(user_id, 'force_check'), callback_data='force_check')], - ] - return InlineKeyboardMarkup(keyboard) - - -# --- Diary Functions --- - -SCHEDULE_WEEKDAYS_FULL = { - 'uk': ['Понеділок', 'Вівторок', 'Середа', 'Четвер', "П'ятниця", 'Субота', 'Неділя'], - 'ru': ['Понедельник', 'Вторник', 'Среда', 'Четверг', 'Пятница', 'Суббота', 'Воскресенье'], - 'en': ['Monday', 'Tuesday', 'Wednesday', 'Thursday', 'Friday', 'Saturday', 'Sunday'], -} -# Canonical weekday names used in callback_data and schedule lookup (site language). -SCHEDULE_WEEKDAYS_CANONICAL = SCHEDULE_WEEKDAYS_FULL['uk'] -SCHEDULE_CACHE_TTL = 18000 # 5 hours - - -def _norm_day(s): - """Normalize weekday name for tolerant matching (case, spaces, apostrophes).""" - if not isinstance(s, str): - return '' - s = s.strip().casefold() - for _q in ('\u2019', '\u2018', '\u02bc', '`'): - s = s.replace(_q, "'") - s = re.sub(r'\s+', ' ', s) - return s.strip() - - -_SCHEDULE_WEEKDAYS_CANONICAL_NORM = [_norm_day(d) for d in SCHEDULE_WEEKDAYS_CANONICAL] - -DIARY_WEEKDAYS_SHORT = { - 'uk': ['Пн', 'Вт', 'Ср', 'Чт', 'Пт', 'Сб', 'Нд'], - 'ru': ['Пн', 'Вт', 'Ср', 'Чт', 'Пт', 'Сб', 'Вс'], - 'en': ['Mon', 'Tue', 'Wed', 'Thu', 'Fri', 'Sat', 'Sun'], -} - - -def get_diary_keyboard(user_id: int | None = None): - today = datetime.now() - lang = get_user_language(user_id) if user_id is not None else 'en' - today_label = _tr(lang, 'today') - tomorrow_label = _tr(lang, 'tomorrow') - week_label = _tr(lang, 'week') - month_label = _tr(lang, 'month') - keyboard = [ - [ - InlineKeyboardButton(f'{today_label} ({today.day}.{today.month:02d})', callback_data='diary_today'), - InlineKeyboardButton(tomorrow_label, callback_data='diary_tomorrow'), - ], - [ - InlineKeyboardButton(week_label, callback_data='diary_week'), - InlineKeyboardButton(month_label, callback_data='diary_month'), - ], - ] - return InlineKeyboardMarkup(keyboard) - - -def _parse_calendar_html(table_html: str) -> tuple: - """Parse calendar HTML table into (month_text, {day_num: {weekday, events}}). - - Each event is a dict: {'title': str, 'id': site event id | None, 'time': 'HH:MM' | None}. - """ - days = {} - weekdays = [] - rows = re.findall(r']*>(.*?)', table_html, re.DOTALL) - - month_text = '' - for r_idx, row in enumerate(rows): - cells = re.findall(r']*>(.*?)', row, re.DOTALL) - - if r_idx == 0: - # Month navigation row: extract "Травень 2026" from nav text - raw = re.sub(r'<[^>]+>', ' ', row).strip() - raw = re.sub(r'\s+', ' ', raw) - m = re.search(r'([А-Яа-яіїєґ\']+\s*:?\s*\d{4})', raw) - month_text = m.group(1).replace(' : ', ' ').strip() if m else raw - elif r_idx == 1: - # Day names row - for cell in cells: - name = re.sub(r'<[^>]+>', '', cell).strip() - if name: - weekdays.append(name) - else: - # Data rows: each cell = a day - for col_idx, cell in enumerate(cells): - # Extract day number — first number in the cell text - text = re.sub(r'<[^>]+>', ' ', cell).strip() - text = re.sub(r'\s+', ' ', text) - dm = re.match(r'(\d+)', text) - if not dm: - continue - day_num = dm.group(1) - - # Extract events: title attribute (full name) of ALL tags inside the cell. - # Each event keeps its site data-event-id and a 'time' slot (filled later - # from the AJAX popup by fetch_diary_data). 'time' is 'HH:MM' or None. - events = [] - for a_match in re.finditer(r']*>(.*?)', cell, re.DOTALL): - a_tag = a_match.group(0) - # Prefer the title attribute (contains full name, not truncated) - title_m = re.search(r'title\s*=\s*"([^"]*)"', a_tag) - et = title_m.group(1).strip() if title_m else re.sub(r'<[^>]+>', '', a_match.group(1)).strip() - if not et: - continue - id_m = re.search(r'data-event-id\s*=\s*"?(\d+)"?', a_tag) - events.append( - { - 'title': et, - 'id': id_m.group(1) if id_m else None, - 'time': None, - } - ) - - weekday = weekdays[col_idx] if col_idx < len(weekdays) else '' - days[day_num] = {'weekday': weekday, 'weekday_idx': col_idx, 'events': events} - - return month_text, days - - -async def _collect_event_times(page) -> dict: - """Read event times straight from the rendered calendar DOM. - - The diary page embeds `div.event-full-info[data-event-full-info-id]` - containing `span.data` (e.g. "2026-09-02 16:30:00") for every event, so no - AJAX popup clicks are needed. Returns {event_id: 'HH:MM'} for events that - have a date; events without one are simply skipped. - """ - times_by_id: dict[str, str] = {} - try: - raw_times = await page.evaluate( - """() => { - const out = {}; - for (const div of document.querySelectorAll( - 'div.event-full-info[data-event-full-info-id]' - )) { - const id = div.getAttribute('data-event-full-info-id'); - const date_span = div.querySelector('p.date span.data'); - if (id && date_span) { - out[id] = date_span.textContent.trim(); - } - } - return out; - }""" - ) - except Exception as e: - logger.warning(f'Failed to read diary event times from DOM: {e}') - return times_by_id - - for event_id, time_text in (raw_times or {}).items(): - if not time_text: - continue - m = re.search(r'(\d{1,2}:\d{2})', time_text) - if m: - times_by_id[event_id] = m.group(1) - - logger.info(f'Diary event times collected for {len(times_by_id)} events') - return times_by_id - - -async def fetch_diary_data(phpsessid: str) -> dict | None: - logger.info('Fetching diary data via Playwright...') - try: - async with asyncio.timeout(60): - async with async_playwright() as p: - browser = await asyncio.wait_for(p.chromium.connect(PLAYWRIGHT_WS), timeout=15) - try: - context_browser = await browser.new_context(user_agent=USER_AGENT) - await context_browser.add_cookies( - [{'name': 'PHPSESSID', 'value': phpsessid, 'domain': 'edu.edu.vn.ua', 'path': '/'}] - ) - page = await context_browser.new_page() - - try: - await asyncio.wait_for(page.goto(DIARY_URL, wait_until='domcontentloaded'), timeout=30) - await page.wait_for_selector('table.calendar', timeout=10000) - await page.wait_for_timeout(1500) - - table_html = await page.evaluate(""" - () => { - const t = document.querySelector('table.calendar'); - return t ? t.outerHTML : null; - } - """) - if not table_html: - logger.error('table.calendar not found in DOM') - return None - - # Debug: save HTML for troubleshooting - with contextlib.suppress(Exception), open('/tmp/diary_debug.html', 'w', encoding='utf-8') as f: # noqa: S108 - f.write(table_html) - - month_text, days = _parse_calendar_html(table_html) - - # Read event times by opening each event's AJAX popup. - times_by_id = await _collect_event_times(page) - if times_by_id: - for day_data in days.values(): - for ev in day_data.get('events', []): - eid = ev.get('id') - if eid and eid in times_by_id: - ev['time'] = times_by_id[eid] - - logger.info( - f'Diary parsed: month={month_text!r}, days_with_events={sum(1 for d in days.values() if d["events"])}/{len(days)}' - ) - - return {'monthFullText': month_text, 'days': days} - - except Exception as e: - logger.error(f'Error parsing diary: {e}') - return None - finally: - with contextlib.suppress(Exception): - await asyncio.wait_for(page.close(), timeout=5) - with contextlib.suppress(Exception): - await asyncio.wait_for(context_browser.close(), timeout=5) - finally: - with contextlib.suppress(Exception): - await asyncio.wait_for(browser.close(), timeout=5) - except TimeoutError: - logger.error('Diary fetch timed out (60s)') - return None - except Exception as e: - logger.error(f'Playwright error in diary fetch: {e}') - return None - - -def _parse_diary_month(text: str) -> str: - match = re.search(r'([А-Яа-яіїєґ\']+\s*:\s*\d{4})', text) - if match: - return match.group(1).replace(' : ', ' ').strip() - return text.strip() - - -def _format_event_line(event, lang: str) -> str: - """Render one diary event line: 📌 title, with time in parens when known. - - '08:00' is the site's placeholder for "no time specified" — show a localized - 'time_unknown' mark instead. Unknown/absent time renders as before, no parens. - """ - if isinstance(event, dict): - title = escape(str(event.get('title') or '')) - tme = event.get('time') - else: - title = escape(str(event)) - tme = None - if tme == '08:00': - return f'📌 {title} ({_tr(lang, "time_unknown")})' - if tme: - return f'📌 {title} ({escape(str(tme))})' - return f'📌 {title}' - - -def format_diary_day(data: dict, day_num: int, lang: str = 'en') -> str: - days = data.get('days', {}) - month_str = _parse_diary_month(data.get('monthFullText', '')) - day_data = days.get(str(day_num)) - lines = [f'📅 {day_num} {month_str}', '─' * 18] - if not day_data or not day_data.get('events'): - lines.append(_tr(lang, 'no_events')) - else: - for e in day_data['events']: - lines.append(_format_event_line(e, lang)) - lines.append(f'\n🔗 {DIARY_URL}') - return '\n'.join(lines) - - -def format_diary_week(data: dict, today: datetime, lang: str = 'en') -> str: - days = data.get('days', {}) - _parse_diary_month(data.get('monthFullText', '')) - monday = today - timedelta(days=today.weekday()) - friday = monday + timedelta(days=4) - start = f'{monday.day}.{monday.month}' - end = f'{friday.day}.{friday.month}' - short = DIARY_WEEKDAYS_SHORT.get(lang, DIARY_WEEKDAYS_SHORT['en']) - lines = [_tr(lang, 'diary_week_title', start=start, end=end) + '\n'] - for i in range(5): - d = monday + timedelta(days=i) - day_data = days.get(str(d.day)) - lines.append(f'─ {short[i]} {d.day}.{d.month} ─') - if not day_data or not day_data.get('events'): - lines.append(_tr(lang, 'no_events') + '\n') - else: - for e in day_data['events']: - lines.append(_format_event_line(e, lang)) - lines.append('') - lines.append(f'🔗 {DIARY_URL}') - return '\n'.join(lines) - - -def format_diary_month(data: dict, lang: str = 'en') -> str: - days = data.get('days', {}) - month_str = _parse_diary_month(data.get('monthFullText', '')) - lines = [f'📅 {month_str}\n'] - for day_num in sorted(days.keys(), key=int): - day_data = days[day_num] - if day_data.get('weekday_idx', 0) >= 5: - continue - if day_data.get('weekday', '').strip().lower() in ( - 'субота', - 'суббота', - 'saturday', - 'неділя', - 'воскресенье', - 'sunday', - ): - continue - events = day_data.get('events', []) - weekday = day_data.get('weekday', '') - lines.append(f'─ {weekday} {day_num} ─') - if not events: - lines.append(_tr(lang, 'no_events') + '\n') - else: - for e in events: - lines.append(_format_event_line(e, lang)) - lines.append('') - lines.append(f'🔗 {DIARY_URL}') - return '\n'.join(lines) - - -async def _get_diary_data(context: ContextTypes.DEFAULT_TYPE) -> dict | None: - cached = context.user_data.get('diary_cache') - now_ts = time.time() - if cached and (now_ts - cached.get('timestamp', 0)) < 300: - return cached['data'] - phpsessid = redis_client.get(KEY_PHPSESSID) - if not phpsessid: - return None - data = await fetch_diary_data(phpsessid) - if data: - context.user_data['diary_cache'] = {'data': data, 'timestamp': now_ts} - return data - - -# --- Schedule Functions --- - - -def get_schedule_day_keyboard(user_id: int | None = None): - lang = get_user_language(user_id) if user_id is not None else 'en' - full = SCHEDULE_WEEKDAYS_FULL.get(lang, SCHEDULE_WEEKDAYS_FULL['en']) - today = datetime.now() - tomorrow = today + timedelta(days=1) - today_label = _tr(lang, 'today') - tomorrow_label = _tr(lang, 'tomorrow') - short_keys = ['day_mon', 'day_tue', 'day_wed', 'day_thu', 'day_fri'] - keyboard = [ - [ - InlineKeyboardButton( - f'{today_label} ({full[today.weekday()]})', - callback_data=f'schedule_day_{SCHEDULE_WEEKDAYS_CANONICAL[today.weekday()]}', - ), - InlineKeyboardButton( - f'{tomorrow_label} ({full[tomorrow.weekday()]})', - callback_data=f'schedule_day_{SCHEDULE_WEEKDAYS_CANONICAL[tomorrow.weekday()]}', - ), - ], - [ - InlineKeyboardButton( - _tr(lang, short_keys[0]), callback_data=f'schedule_day_{SCHEDULE_WEEKDAYS_CANONICAL[0]}' - ), - InlineKeyboardButton( - _tr(lang, short_keys[1]), callback_data=f'schedule_day_{SCHEDULE_WEEKDAYS_CANONICAL[1]}' - ), - InlineKeyboardButton( - _tr(lang, short_keys[2]), callback_data=f'schedule_day_{SCHEDULE_WEEKDAYS_CANONICAL[2]}' - ), - ], - [ - InlineKeyboardButton( - _tr(lang, short_keys[3]), callback_data=f'schedule_day_{SCHEDULE_WEEKDAYS_CANONICAL[3]}' - ), - InlineKeyboardButton( - _tr(lang, short_keys[4]), callback_data=f'schedule_day_{SCHEDULE_WEEKDAYS_CANONICAL[4]}' - ), - ], - ] - return InlineKeyboardMarkup(keyboard) - - -def get_schedule_class_keyboard(classes: list): - rows = [] - for i in range(0, len(classes), 4): - row = [InlineKeyboardButton(str(c), callback_data=f'schedule_class_{c}') for c in classes[i : i + 4]] - rows.append(row) - return InlineKeyboardMarkup(rows) - - -def _get_chat_class(chat, user_id) -> str | None: - """Get stored schedule class for a private user or a group chat.""" - if chat.type == 'private': - return redis_client.get(f'user:{user_id}:schedule_class') - return redis_client.get(f'chat:{chat.id}:schedule_class') - - -def _set_chat_class(chat, user_id, class_num) -> None: - """Save schedule class for a private user or a group chat.""" - if chat.type == 'private': - redis_client.set(f'user:{user_id}:schedule_class', str(class_num)) - else: - redis_client.set(f'chat:{chat.id}:schedule_class', str(class_num)) - - -def _parse_schedule_cell(cell_html: str) -> list: - """Parse a single schedule cell, returning list of {subject, note, teacher}.""" - lessons = [] - parts = re.split(r'', cell_html, flags=re.IGNORECASE) - for part in parts: - part = re.sub(r'', '', part, flags=re.DOTALL) - part = re.sub(r'