Compare commits
70
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
033a69bc1e | ||
|
|
8a5d23b560 | ||
|
|
fb46fb48b5 | ||
|
|
4c7c53e0f2 | ||
|
|
ba934265ac | ||
|
|
c7155808d9 | ||
|
|
fc4d64bdb2 | ||
|
|
a6f7fc6030 | ||
|
|
11de1d1468 | ||
|
|
64962d1a63 | ||
|
|
b08a0a927d | ||
|
|
8203ba1b0b | ||
|
|
48033b5495 | ||
|
|
73d2af73e5 | ||
|
|
64253e005e | ||
|
|
69accd1752 | ||
|
|
0cf4b08a95 | ||
|
|
86df5d9048 | ||
|
|
d7441bbbc2 | ||
|
|
2be089e048 | ||
|
|
83b2e68371 | ||
|
|
597f64cbb0 | ||
|
|
f767f3ce1a | ||
|
|
95e4d8f875 | ||
|
|
1d93588e06 | ||
|
|
fb400eea6e | ||
|
|
0c76426c17 | ||
|
|
d74822cd27 | ||
|
|
b677d553b4 | ||
|
|
898463b759 | ||
|
|
9a76529be8 | ||
|
|
5f9354b9a8 | ||
|
|
fcd16f6128 | ||
|
|
2ac2a94bb4 | ||
|
|
3c7e358dd5 | ||
|
|
9be8fb6c69 | ||
|
|
7fdeacffb8 | ||
|
|
cef499de73 | ||
|
|
1b70a55300 | ||
|
|
3f2b4e9acf | ||
|
|
6057734a4f | ||
|
|
8eedf8b74b | ||
|
|
8c0e36a5c0 | ||
|
|
78cd15f12c | ||
|
|
18c633c242 | ||
|
|
4510394531 | ||
|
|
2545312db1 | ||
|
|
8ce0b809c0 | ||
|
|
506c04e15c | ||
|
|
bdcbb3d5af | ||
|
|
faac6febb8 | ||
|
|
695da1308c | ||
|
|
552b22cb66 | ||
|
|
3756e60c95 | ||
|
|
d0d217ddf4 | ||
|
|
24f84bab2f | ||
|
|
0b8a3c16a7 | ||
|
|
e9a68aae77 | ||
|
|
5359df5ed6 | ||
|
|
8aea0f0d13 | ||
|
|
65d4f2d482 | ||
|
|
2fe9b7632f | ||
|
|
e436d89eef | ||
|
|
8729cb5062 | ||
|
|
f1e4a1088d | ||
|
|
b357ef95d8 | ||
|
|
67422663b7 | ||
|
|
88705fec88 | ||
|
|
c098807aa4 | ||
|
|
e1eee2d3c7 |
No files matched your search
@@ -0,0 +1,50 @@
|
|||||||
|
# EDU ownership handoff
|
||||||
|
|
||||||
|
## Status
|
||||||
|
|
||||||
|
The EDU ownership handoff is complete. The homelab repository no longer owns
|
||||||
|
EDU workloads, images, routes, alerts, or deployment selection. The EDU
|
||||||
|
repository is the only deployment owner: [forust/edu-master](https://git.forust.xyz/forust/edu-master).
|
||||||
|
|
||||||
|
Homelab PRs #99 and #105 are merged. PR #105 removed the EDU subtree and its
|
||||||
|
build, deploy, rollback, verification, route-probe, and registry references.
|
||||||
|
It also added the serial image build matrix for the homelab services. This
|
||||||
|
handoff record is the only remaining EDU-specific file in homelab Git.
|
||||||
|
|
||||||
|
The dedicated workstation checkout is `/srv/edu-master`, at release
|
||||||
|
`4f2b2a0e37dc11ac2c75441a15076c178e219d37`. It contains `k8s/active`; root
|
||||||
|
`active` is absent. The old untracked `/srv/homelab/edu_master` checkout was
|
||||||
|
moved outside the homelab repository to
|
||||||
|
`/srv/edu-master-legacy-archive-20261007/edu_master`. Its private files remain
|
||||||
|
mode `0600` inside an archive directory with mode `0700`. The homelab deploy
|
||||||
|
checkout has no EDU marker or tracked EDU application/deployment files.
|
||||||
|
`AUTODEPLOY=false` remains in place for homelab deployment.
|
||||||
|
|
||||||
|
## Release evidence
|
||||||
|
|
||||||
|
EDU PR #4 merged after its review and CI checks. Main-push CI run 1652 passed
|
||||||
|
all validation and both image builds. Deploy run 1653 passed for the exact main
|
||||||
|
SHA above.
|
||||||
|
|
||||||
|
The workstation rollout completed for both Deployments. The deployment
|
||||||
|
verified `/health` and `/live` with HTTP 200, Redis AUTH, session TTL of 1058
|
||||||
|
seconds, a delivery backlog of zero, and all nine EDU vmalert rules with
|
||||||
|
matching expressions and healthy evaluation.
|
||||||
|
|
||||||
|
The images now run by digest:
|
||||||
|
|
||||||
|
- Session keeper: `sha256:998dea51aa3015fd9cabefb0f53b030157a650c3bef72e02fe84f17d5762613d`
|
||||||
|
- Webinar checker: `sha256:92f3c1fa2bb7f9b4680a9fc76a5b33dfbea8ef3dd9c6490ebc45876fd4c54461`
|
||||||
|
|
||||||
|
Redis StatefulSet was unchanged. PVC `redis-data-pvc` remains bound to PV
|
||||||
|
`pvc-a4f2a79a-363a-4c12-ae91-92cdfc2a0d2e` with capacity 1 GiB. The existing
|
||||||
|
runtime Secret and Fernet key were preserved during the handoff. Notification
|
||||||
|
delivery was verified before closeout, as confirmed by the operator. The
|
||||||
|
deployment did not record downtime.
|
||||||
|
|
||||||
|
The release rollback snapshot is
|
||||||
|
`/home/forust/.local/state/edu-master-deploy/20261007T180541Z-4f2b2a0e37dc11ac2c75441a15076c178e219d37`.
|
||||||
|
The handoff data snapshot remains at
|
||||||
|
`/home/forust/.local/state/edu-master-deploy/handoff-20261007T080838Z`.
|
||||||
|
Both snapshots are outside Git. Do not restore old Redis data unless recovery
|
||||||
|
requires it. Never delete or recreate the Redis PVC.
|
||||||
@@ -7,4 +7,5 @@ self-hosted-runner:
|
|||||||
labels:
|
labels:
|
||||||
- arch
|
- arch
|
||||||
- homelab
|
- homelab
|
||||||
|
- homelab-pr
|
||||||
- prod
|
- prod
|
||||||
@@ -0,0 +1,3 @@
|
|||||||
|
{
|
||||||
|
"postgres": ["authentik", "gitea", "immich", "n8n", "netbox", "netronome"]
|
||||||
|
}
|
||||||
@@ -0,0 +1,159 @@
|
|||||||
|
# Homelab CI/CD
|
||||||
|
|
||||||
|
The native Gitea runners run on **vps**; production runs on **workstation**.
|
||||||
|
Main-branch checks and image builds use `homelab:host`. Pull request and
|
||||||
|
non-main checks use `homelab-pr:host` under a separate account without Docker
|
||||||
|
access. The `homelab-pr` runner is registered at User scope for `forust`, so
|
||||||
|
any repository under that account can schedule jobs that request this label.
|
||||||
|
Each runner accepts one job at a time; the build waits for every check to pass.
|
||||||
|
CI and deploy runs also show a summary with
|
||||||
|
the release SHA, image build or reuse results, deploy mode, selected services,
|
||||||
|
and image digests. Failed runs keep a summary of completed image builds, stage
|
||||||
|
results, apply results, and recorded Kubernetes recovery. The final deploy
|
||||||
|
summary is in the smoke job; earlier jobs show the state observed at that time.
|
||||||
|
Apply success is separate from health and recovery. Update the installed
|
||||||
|
workstation controller with `setup-workstation.sh` when no deploy is running.
|
||||||
|
No job images or Kubernetes credentials are needed on the VPS. Builds use one
|
||||||
|
pinned BuildKit helper container. CI and deploy are separate workflows.
|
||||||
|
|
||||||
|
## Runner installation
|
||||||
|
|
||||||
|
Install Docker Engine with Compose and Buildx, Git, Python 3.11+, Bash, curl,
|
||||||
|
GNU tar/xz, flock and systemd using the host's package manager. Keep the existing
|
||||||
|
Gitea runner 3.0.2 binary at `/usr/local/bin/gitea-runner`.
|
||||||
|
|
||||||
|
From this checkout on the VPS:
|
||||||
|
|
||||||
|
```sh
|
||||||
|
sudo bash .gitea/runner/setup-runner.sh
|
||||||
|
```
|
||||||
|
|
||||||
|
The installer reuses `/var/lib/gitea-runner/.runner` and the existing service.
|
||||||
|
For a new host, install the same runner binary and register as `gitea-runner`
|
||||||
|
using the registration token interactively, label `homelab:host`, and working
|
||||||
|
directory `/var/lib/gitea-runner`; then rerun the installer. Tokens never belong
|
||||||
|
in this repository or command-line examples.
|
||||||
|
|
||||||
|
Pinned tools live in the runner user's `~/.cache/homelab-ci`; CI repairs version
|
||||||
|
drift there. Installations are locked. Buildx uses only the `homelab-ci` builder,
|
||||||
|
pushes directly to the registry, and caps retained local cache at 1 GiB with a
|
||||||
|
2 GiB free-space target. This is not a hard limit on peak build disk usage.
|
||||||
|
Nothing runs `docker system prune`, removes unrelated images, or deletes volumes.
|
||||||
|
|
||||||
|
### Pull request runner
|
||||||
|
|
||||||
|
Install the unprivileged host runner on the VPS:
|
||||||
|
|
||||||
|
```sh
|
||||||
|
sudo bash .gitea/runner/setup-pr-runner.sh
|
||||||
|
```
|
||||||
|
|
||||||
|
Get a registration token from the user Actions runner settings. Run the
|
||||||
|
installer in a terminal. It asks for the token without echoing it, registers the
|
||||||
|
runner as `homelab-pr` with label `homelab-pr:host`, then enables the service.
|
||||||
|
The work directory is `/var/lib/gitea-pr-runner`. Confirm that Gitea lists the
|
||||||
|
runner as User scope before merging the workflow change. An unmatched label can
|
||||||
|
fall back to the default job image.
|
||||||
|
|
||||||
|
Renovate PR validation uses `pull_request_target`, which reads the workflow from
|
||||||
|
the base branch. It checks out the PR head only after runner selection and runs
|
||||||
|
that code on `homelab-pr`. Keep this workflow read-only and do not add secrets.
|
||||||
|
|
||||||
|
The PR runner has a separate home and tool cache. Do not add it to the `docker`
|
||||||
|
group or give it access to `/var/run/docker.sock`. It runs repository code from
|
||||||
|
pull requests, so keep its registration and permissions separate from the
|
||||||
|
trusted `homelab` runner. This separates users and host permissions, but both
|
||||||
|
runners still share the VPS kernel and network. Use a disposable VM if PRs from
|
||||||
|
untrusted external authors must be fully isolated.
|
||||||
|
|
||||||
|
## Workstation setup
|
||||||
|
|
||||||
|
As the existing SSH deploy user on workstation:
|
||||||
|
|
||||||
|
```sh
|
||||||
|
sudo loginctl enable-linger forust
|
||||||
|
bash .gitea/runner/setup-workstation.sh
|
||||||
|
```
|
||||||
|
|
||||||
|
The controller uses `/srv/homelab` as the persistent configuration tree and makes
|
||||||
|
a detached source worktree for each SHA. It never resets `/srv/homelab`, moves
|
||||||
|
local configuration, renames Compose projects, or changes volume names.
|
||||||
|
The installer records the current Kubernetes context and cluster UID in
|
||||||
|
`~/.config/homelab-deploy/environment`. Check these before installing.
|
||||||
|
|
||||||
|
Configure Gitea Actions Variables:
|
||||||
|
|
||||||
|
- `DEPLOY_HOST`, `DEPLOY_USER`, `DEPLOY_PORT`: the existing VPS-to-workstation SSH endpoint.
|
||||||
|
- `DEPLOY_KNOWN_HOSTS`: workstation's verified SSH host key entry for that endpoint.
|
||||||
|
- `AUTODEPLOY`: `false` initially; `true` enables deployment after successful main CI.
|
||||||
|
|
||||||
|
Keep `DEPLOY_SSH_KEY`, `REGISTRY_USERNAME` and `REGISTRY_PASSWORD` in Actions
|
||||||
|
Secrets. Legacy endpoint secrets remain accepted during migration. The Actions
|
||||||
|
token must have repository read and Actions read access for release downloads.
|
||||||
|
The deploy user's existing Docker registry authentication remains necessary.
|
||||||
|
|
||||||
|
## Releases and deployment
|
||||||
|
|
||||||
|
CI publishes `release-<full SHA>` as a Gitea artifact with all three owned image
|
||||||
|
digests and build input fingerprints. Unchanged images are reused only from a
|
||||||
|
successful main CI artifact, never from `:prod`. Expired artifacts cause CI to
|
||||||
|
rebuild images; they block deployment until CI is rerun.
|
||||||
|
|
||||||
|
Run deploy from main with `deploy_ref=main` or a checked SHA:
|
||||||
|
|
||||||
|
- `full`: required for the first baseline; reconcile all active components.
|
||||||
|
- `changed`: compare with the last fully successful production deploy.
|
||||||
|
- `plan`: validate configuration and show selection without changing production resources.
|
||||||
|
- `refresh_images=true`: explicitly refresh mutable third-party Compose tags.
|
||||||
|
|
||||||
|
The manual and automatic paths both require successful CI, a successful build
|
||||||
|
job and the exact SHA's release artifact. PRs cannot publish images or deploy.
|
||||||
|
Removed resources are reported and require explicit removal; no automatic prune.
|
||||||
|
Service dependencies are listed in `.gitea/deploy-dependencies.json`.
|
||||||
|
|
||||||
|
A workstation user systemd service holds the deploy lock across validation,
|
||||||
|
sequential apply, verification and smoke checks. SSH clients only submit/follow:
|
||||||
|
disconnecting or cancelling the Actions client does not kill production apply.
|
||||||
|
Retrying the same run ID does not start another apply. `ExecStopPost` recovers
|
||||||
|
interrupted runs before the unit finishes. Kubernetes rolls back to captured
|
||||||
|
revisions; configuration and persistent data are not reverted.
|
||||||
|
|
||||||
|
## Status and recovery
|
||||||
|
|
||||||
|
`--retry` repeats failed verification and smoke checks, never apply. Recovery
|
||||||
|
keeps a failed deploy out of the successful baseline, even after rollback.
|
||||||
|
|
||||||
|
On workstation (replace the numeric ID with Actions run ID and attempt):
|
||||||
|
|
||||||
|
```sh
|
||||||
|
python3 ~/.local/lib/homelab-deploy/controller.py status 123-1
|
||||||
|
python3 ~/.local/lib/homelab-deploy/controller.py recover 123-1 --retry
|
||||||
|
journalctl --user -u homelab-deploy@123-1
|
||||||
|
```
|
||||||
|
|
||||||
|
Runs live in `~/.local/state/homelab-deploy/runs`. Compose stores resolved configs
|
||||||
|
with restricted permissions; these may contain credentials and must never be
|
||||||
|
uploaded as CI artifacts. Stage logs print the exact manual recovery command
|
||||||
|
using `compose-before/<stack>.json`, the original project directory and project
|
||||||
|
name. Compose does not automatically roll back, and Nextcloud AIO's child
|
||||||
|
containers remain managed by AIO. Preserve its own backups for data recovery.
|
||||||
|
|
||||||
|
The controller retains twenty successful/planned runs and preserves failures.
|
||||||
|
Update the workstation dispatcher only when no deploy is running.
|
||||||
|
|
||||||
|
## Validation and migration rollback
|
||||||
|
|
||||||
|
```sh
|
||||||
|
python3 -m unittest discover -s tests -v
|
||||||
|
bash .gitea/tests/deploy-validation.sh
|
||||||
|
```
|
||||||
|
|
||||||
|
Test on a separate namespace before the initial production `full` run. Check a
|
||||||
|
failed rollout, interrupted SSH and repeated run ID, and verify that an isolated
|
||||||
|
service change does not upgrade unrelated Helm releases or Compose stacks.
|
||||||
|
|
||||||
|
To roll back the migration, disable autodeploy and finish or recover the remote
|
||||||
|
run first. Restore the runner config/unit from `.before-<timestamp>` backups,
|
||||||
|
reload systemd and restart the runner. Restore the prior workflows from Git.
|
||||||
|
Production data and persistent volumes stay where they were. Do not remove run
|
||||||
|
state or Compose recovery files until recovery is confirmed.
|
||||||
@@ -0,0 +1,11 @@
|
|||||||
|
[worker.oci]
|
||||||
|
gc = true
|
||||||
|
reservedSpace = "256MB"
|
||||||
|
maxUsedSpace = "1GB"
|
||||||
|
minFreeSpace = "2GB"
|
||||||
|
|
||||||
|
[[worker.oci.gcpolicy]]
|
||||||
|
reservedSpace = "256MB"
|
||||||
|
maxUsedSpace = "1GB"
|
||||||
|
minFreeSpace = "2GB"
|
||||||
|
all = true
|
||||||
@@ -0,0 +1,10 @@
|
|||||||
|
runner:
|
||||||
|
file: /var/lib/gitea-runner/.runner
|
||||||
|
capacity: 1
|
||||||
|
timeout: 5h
|
||||||
|
labels:
|
||||||
|
- homelab:host
|
||||||
|
cache:
|
||||||
|
enabled: false
|
||||||
|
container:
|
||||||
|
docker_host: unix:///var/run/docker.sock
|
||||||
@@ -0,0 +1,18 @@
|
|||||||
|
[Unit]
|
||||||
|
Description=Gitea Actions runner
|
||||||
|
After=network-online.target docker.service
|
||||||
|
Wants=network-online.target
|
||||||
|
|
||||||
|
[Service]
|
||||||
|
User=gitea-runner
|
||||||
|
Group=gitea-runner
|
||||||
|
SupplementaryGroups=docker
|
||||||
|
WorkingDirectory=/var/lib/gitea-runner
|
||||||
|
Environment=PATH=/var/lib/gitea-runner/.cache/homelab-ci/bin:/usr/local/bin:/usr/bin:/bin
|
||||||
|
ExecStart=/usr/local/bin/gitea-runner daemon --config /etc/gitea-runner/config.yaml
|
||||||
|
Restart=on-failure
|
||||||
|
RestartSec=5
|
||||||
|
UMask=0077
|
||||||
|
|
||||||
|
[Install]
|
||||||
|
WantedBy=multi-user.target
|
||||||
@@ -0,0 +1,12 @@
|
|||||||
|
[Unit]
|
||||||
|
Description=Homelab deploy %i
|
||||||
|
|
||||||
|
[Service]
|
||||||
|
Type=exec
|
||||||
|
EnvironmentFile=%h/.config/homelab-deploy/environment
|
||||||
|
ExecStart=/usr/bin/python3 %h/.local/lib/homelab-deploy/controller.py execute %i
|
||||||
|
ExecStopPost=/usr/bin/python3 %h/.local/lib/homelab-deploy/controller.py recover %i
|
||||||
|
RuntimeMaxSec=5h
|
||||||
|
TimeoutStopSec=135min
|
||||||
|
KillMode=control-group
|
||||||
|
UMask=0077
|
||||||
@@ -0,0 +1,8 @@
|
|||||||
|
runner:
|
||||||
|
file: /var/lib/gitea-pr-runner/.runner
|
||||||
|
capacity: 1
|
||||||
|
timeout: 5h
|
||||||
|
labels:
|
||||||
|
- homelab-pr:host
|
||||||
|
cache:
|
||||||
|
enabled: false
|
||||||
@@ -0,0 +1,27 @@
|
|||||||
|
[Unit]
|
||||||
|
Description=Gitea Actions untrusted pull request runner
|
||||||
|
After=network-online.target
|
||||||
|
Wants=network-online.target
|
||||||
|
|
||||||
|
[Service]
|
||||||
|
User=gitea-pr-runner
|
||||||
|
Group=gitea-pr-runner
|
||||||
|
WorkingDirectory=/var/lib/gitea-pr-runner
|
||||||
|
Environment=HOME=/var/lib/gitea-pr-runner
|
||||||
|
Environment=PATH=/var/lib/gitea-pr-runner/.cache/homelab-ci/bin:/usr/local/bin:/usr/bin:/bin
|
||||||
|
ExecStart=/usr/local/bin/gitea-runner daemon --config /etc/gitea-pr-runner/config.yaml
|
||||||
|
Restart=on-failure
|
||||||
|
RestartSec=5
|
||||||
|
NoNewPrivileges=yes
|
||||||
|
PrivateTmp=yes
|
||||||
|
ProtectSystem=full
|
||||||
|
ProtectHome=yes
|
||||||
|
ProtectKernelTunables=yes
|
||||||
|
ProtectKernelModules=yes
|
||||||
|
ProtectControlGroups=yes
|
||||||
|
RestrictSUIDSGID=yes
|
||||||
|
LockPersonality=yes
|
||||||
|
UMask=0077
|
||||||
|
|
||||||
|
[Install]
|
||||||
|
WantedBy=multi-user.target
|
||||||
Executable
+56
@@ -0,0 +1,56 @@
|
|||||||
|
#!/usr/bin/env bash
|
||||||
|
# Install a native runner for untrusted PR jobs without Docker access.
|
||||||
|
set -euo pipefail
|
||||||
|
here="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
||||||
|
[ "$(id -u)" -eq 0 ] || { echo 'Run with sudo on the runner host' >&2; exit 1; }
|
||||||
|
for tool in cp cut date getent id install runuser systemctl useradd; do
|
||||||
|
command -v "$tool" >/dev/null || { echo "Install missing prerequisite: $tool" >&2; exit 1; }
|
||||||
|
done
|
||||||
|
command -v /usr/local/bin/gitea-runner >/dev/null || {
|
||||||
|
echo 'Install gitea-runner 3.0.2 at /usr/local/bin/gitea-runner first' >&2
|
||||||
|
exit 1
|
||||||
|
}
|
||||||
|
|
||||||
|
id gitea-pr-runner >/dev/null 2>&1 || \
|
||||||
|
useradd --system --create-home --home-dir /var/lib/gitea-pr-runner --shell /usr/bin/bash gitea-pr-runner
|
||||||
|
runner_home="$(getent passwd gitea-pr-runner | cut -d: -f6)"
|
||||||
|
[ "$runner_home" = /var/lib/gitea-pr-runner ] || {
|
||||||
|
echo 'Unexpected PR runner home; inspect the existing service first' >&2
|
||||||
|
exit 1
|
||||||
|
}
|
||||||
|
case " $(id -nG gitea-pr-runner) " in
|
||||||
|
*' docker '*)
|
||||||
|
echo 'The PR runner account must not belong to the docker group' >&2
|
||||||
|
exit 1
|
||||||
|
;;
|
||||||
|
esac
|
||||||
|
|
||||||
|
install -d -m 0755 /etc/gitea-pr-runner
|
||||||
|
stamp="$(date -u +%Y%m%dT%H%M%SZ)"
|
||||||
|
for existing in /etc/gitea-pr-runner/config.yaml /etc/systemd/system/gitea-pr-runner.service; do
|
||||||
|
[ ! -f "$existing" ] || cp -p "$existing" "$existing.before-$stamp"
|
||||||
|
done
|
||||||
|
install -m 0644 "$here/pr-config.yaml" /etc/gitea-pr-runner/config.yaml
|
||||||
|
install -m 0644 "$here/pr-runner.service" /etc/systemd/system/gitea-pr-runner.service
|
||||||
|
|
||||||
|
if [ ! -f /var/lib/gitea-pr-runner/.runner ]; then
|
||||||
|
read -r -s -p 'Enter the Gitea repository runner registration token: ' runner_token
|
||||||
|
printf '\n'
|
||||||
|
[ -n "$runner_token" ] || { echo 'Runner token is required' >&2; exit 1; }
|
||||||
|
export GITEA_RUNNER_REGISTRATION_TOKEN="$runner_token"
|
||||||
|
unset runner_token
|
||||||
|
runuser --preserve-environment -u gitea-pr-runner -- \
|
||||||
|
/usr/local/bin/gitea-runner register \
|
||||||
|
--config /etc/gitea-pr-runner/config.yaml \
|
||||||
|
--instance https://gitea.forust.xyz \
|
||||||
|
--name homelab-pr \
|
||||||
|
--labels homelab-pr:host \
|
||||||
|
--no-interactive
|
||||||
|
unset GITEA_RUNNER_REGISTRATION_TOKEN
|
||||||
|
fi
|
||||||
|
chmod 0600 /var/lib/gitea-pr-runner/.runner
|
||||||
|
|
||||||
|
systemctl daemon-reload
|
||||||
|
systemctl enable --now gitea-pr-runner.service
|
||||||
|
systemctl restart gitea-pr-runner.service
|
||||||
|
echo "PR runner ready. Configuration backups: *.before-$stamp"
|
||||||
Executable
+48
@@ -0,0 +1,48 @@
|
|||||||
|
#!/usr/bin/env bash
|
||||||
|
# Native host runner, with pinned user-space tools and no extra CI images.
|
||||||
|
set -euo pipefail
|
||||||
|
here="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
||||||
|
[ "$(id -u)" -eq 0 ] || { echo 'Run with sudo on the runner host' >&2; exit 1; }
|
||||||
|
for tool in docker curl python3 git tar xz flock runuser systemctl; do
|
||||||
|
command -v "$tool" >/dev/null || { echo "Install missing prerequisite: $tool" >&2; exit 1; }
|
||||||
|
done
|
||||||
|
docker info >/dev/null
|
||||||
|
docker compose version >/dev/null
|
||||||
|
docker buildx version >/dev/null
|
||||||
|
id gitea-runner >/dev/null 2>&1 || useradd --system --create-home --home-dir /var/lib/gitea-runner --shell /usr/bin/bash gitea-runner
|
||||||
|
# Reuse the established service account and runner registration.
|
||||||
|
runner_home="$(getent passwd gitea-runner | cut -d: -f6)"
|
||||||
|
[ "$runner_home" = /var/lib/gitea-runner ] || { echo 'Unexpected runner home; inspect the existing service first' >&2; exit 1; }
|
||||||
|
runuser -u gitea-runner -- docker info >/dev/null || { echo "The runner user needs access to Docker before setup" >&2; exit 1; }
|
||||||
|
command -v gitea-runner >/dev/null || { echo 'Install gitea-runner 3.0.2 at /usr/local/bin/gitea-runner first' >&2; exit 1; }
|
||||||
|
mkdir -p /etc/gitea-runner
|
||||||
|
stamp="$(date -u +%Y%m%dT%H%M%SZ)"
|
||||||
|
for existing in /etc/gitea-runner/config.yaml /etc/systemd/system/gitea-runner.service; do
|
||||||
|
[ ! -f "$existing" ] || cp -p "$existing" "$existing.before-$stamp"
|
||||||
|
done
|
||||||
|
scratch="$(mktemp -d)"
|
||||||
|
trap 'rm -rf "$scratch"' EXIT
|
||||||
|
chmod 755 "$scratch"
|
||||||
|
install -m 0644 "$here/../workflows/install-ci-tools.sh" "$here/../workflows/tool-versions.env" "$scratch/"
|
||||||
|
runuser -u gitea-runner -- bash "$scratch/install-ci-tools.sh"
|
||||||
|
install -m 0644 "$here/config.yaml" /etc/gitea-runner/config.yaml
|
||||||
|
python3 - <<'PYLABELS'
|
||||||
|
import json
|
||||||
|
from pathlib import Path
|
||||||
|
registration = Path('/var/lib/gitea-runner/.runner')
|
||||||
|
if registration.exists():
|
||||||
|
labels = json.loads(registration.read_text()).get('labels', [])
|
||||||
|
labels = [label for label in labels if isinstance(label, str) and label.split(':')[0] != 'homelab']
|
||||||
|
labels.append('homelab:host')
|
||||||
|
config = Path('/etc/gitea-runner/config.yaml')
|
||||||
|
config.write_text(config.read_text().replace(' - homelab:host', '\n'.join(' - ' + json.dumps(label) for label in labels)))
|
||||||
|
PYLABELS
|
||||||
|
install -m 0644 "$here/gitea-runner.service" /etc/systemd/system/gitea-runner.service
|
||||||
|
if [ ! -f /var/lib/gitea-runner/.runner ]; then
|
||||||
|
echo 'Register once as gitea-runner with homelab:host before starting the service.'
|
||||||
|
exit 0
|
||||||
|
fi
|
||||||
|
systemctl daemon-reload
|
||||||
|
systemctl enable --now gitea-runner.service
|
||||||
|
systemctl restart gitea-runner.service
|
||||||
|
echo "Runner ready. Configuration backups: *.before-$stamp"
|
||||||
Executable
+32
@@ -0,0 +1,32 @@
|
|||||||
|
#!/usr/bin/env bash
|
||||||
|
# Run as the existing deploy user on workstation. Never resets the working tree.
|
||||||
|
set -euo pipefail
|
||||||
|
here="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
||||||
|
repo="${HOMELAB_REPO:-/srv/homelab}"
|
||||||
|
for tool in python3 git kubectl helm docker flock timeout; do
|
||||||
|
command -v "$tool" >/dev/null || { echo "Install missing dependency: $tool" >&2; exit 1; }
|
||||||
|
done
|
||||||
|
[ -d "$repo/.git" ] || { echo "Missing deploy checkout: $repo" >&2; exit 1; }
|
||||||
|
[[ "$repo" =~ ^/[A-Za-z0-9_./-]+$ ]] || { echo 'Deploy path must be absolute and contain no whitespace' >&2; exit 1; }
|
||||||
|
if [ "$(loginctl show-user "$USER" -p Linger --value)" != yes ]; then
|
||||||
|
echo "Run once: sudo loginctl enable-linger $USER" >&2
|
||||||
|
exit 1
|
||||||
|
fi
|
||||||
|
config="${XDG_CONFIG_HOME:-$HOME/.config}/homelab-deploy"
|
||||||
|
mkdir -p "$config" "$HOME/.local/lib/homelab-deploy" "$HOME/.config/systemd/user"
|
||||||
|
chmod 700 "$config"
|
||||||
|
if [ ! -f "$config/environment" ]; then
|
||||||
|
context="$(kubectl config current-context)"
|
||||||
|
cluster_uid="$(kubectl get namespace kube-system -o jsonpath='{.metadata.uid}')"
|
||||||
|
printf 'HOMELAB_REPO=%s\nKUBE_CONTEXT=%s\nEXPECTED_CLUSTER_UID=%s\n' "$repo" "$context" "$cluster_uid" >"$config/environment"
|
||||||
|
chmod 600 "$config/environment"
|
||||||
|
fi
|
||||||
|
# Do not replace a dispatcher while an existing deploy uses it.
|
||||||
|
if systemctl --user list-units 'homelab-deploy@*' --state=running --no-legend | grep -q .; then
|
||||||
|
echo 'An existing deploy is running; wait before updating the controller' >&2
|
||||||
|
exit 1
|
||||||
|
fi
|
||||||
|
install -m 0755 "$here/../workflows/deploy-controller.py" "$HOME/.local/lib/homelab-deploy/controller.py"
|
||||||
|
install -m 0644 "$here/homelab-deploy@.service" "$HOME/.config/systemd/user/homelab-deploy@.service"
|
||||||
|
systemctl --user daemon-reload
|
||||||
|
echo 'Controller ready. Run a checked main SHA in full mode for the initial baseline.'
|
||||||
@@ -73,6 +73,20 @@ fi
|
|||||||
grep -q 'MISSING OR UNREADABLE: app/credentials' "$scratch/secrets.log"
|
grep -q 'MISSING OR UNREADABLE: app/credentials' "$scratch/secrets.log"
|
||||||
# API/rendering errors must not produce an empty reference list and pass.
|
# API/rendering errors must not produce an empty reference list and pass.
|
||||||
kubectl() { return 1; }
|
kubectl() { return 1; }
|
||||||
|
if ! skip_uninstalled_vmagent_crd "$REPO/prometheus-stack/k8s/vmagent.yaml"; then
|
||||||
|
echo 'VMAgent preflight did not skip an uninstalled CRD' >&2
|
||||||
|
exit 1
|
||||||
|
fi
|
||||||
|
kubectl() { return 0; }
|
||||||
|
if skip_uninstalled_vmagent_crd "$REPO/prometheus-stack/k8s/vmagent.yaml"; then
|
||||||
|
echo 'VMAgent preflight skipped an installed CRD' >&2
|
||||||
|
exit 1
|
||||||
|
fi
|
||||||
|
if skip_uninstalled_vmagent_crd "$REPO/prometheus-stack/k8s/victoria.yaml"; then
|
||||||
|
echo 'VMAgent preflight skipped an unrelated manifest' >&2
|
||||||
|
exit 1
|
||||||
|
fi
|
||||||
|
kubectl() { return 1; }
|
||||||
if check_referenced_secrets >"$scratch/secrets.log"; then
|
if check_referenced_secrets >"$scratch/secrets.log"; then
|
||||||
echo 'Secret check accepted a failed manifest render' >&2
|
echo 'Secret check accepted a failed manifest render' >&2
|
||||||
exit 1
|
exit 1
|
||||||
|
|||||||
+356
-386
@@ -1,38 +1,25 @@
|
|||||||
name: ci
|
name: ci
|
||||||
|
"on":
|
||||||
on:
|
|
||||||
push:
|
push:
|
||||||
branches:
|
branches:
|
||||||
- "**"
|
- main
|
||||||
pull_request:
|
pull_request: null
|
||||||
workflow_dispatch:
|
workflow_dispatch: null
|
||||||
|
|
||||||
# Every job here is checkout plus local tools. The token needs to read the tree
|
|
||||||
# and nothing else, and saying so keeps a future step that reaches for the API
|
|
||||||
# from quietly holding a token that can write to the repository.
|
|
||||||
permissions:
|
permissions:
|
||||||
contents: read
|
contents: read
|
||||||
|
actions: read
|
||||||
concurrency:
|
concurrency:
|
||||||
group: ci-${{ github.ref }}
|
group: ci-${{ github.ref }}
|
||||||
cancel-in-progress: ${{ github.ref != 'refs/heads/main' }}
|
cancel-in-progress: ${{ github.ref != 'refs/heads/main' }}
|
||||||
|
|
||||||
env:
|
|
||||||
REGISTRY: gcr.forust.xyz
|
|
||||||
|
|
||||||
jobs:
|
jobs:
|
||||||
lint-compose:
|
compose:
|
||||||
runs-on: [self-hosted, linux, arch, homelab]
|
name: Compose
|
||||||
timeout-minutes: 10
|
runs-on: ${{ github.ref == 'refs/heads/main' && 'homelab' || 'homelab-pr' }}
|
||||||
|
timeout-minutes: 15
|
||||||
steps:
|
steps:
|
||||||
- name: Checkout repository
|
- name: Checkout repository
|
||||||
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
|
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262
|
||||||
|
id: source
|
||||||
# Structure check for every committed Compose file, active or not.
|
|
||||||
# Interpolation, env-file and bind-mount resolution are all switched off,
|
|
||||||
# because inactive stacks have no .env here and would only fail on their
|
|
||||||
# ${VAR:?} guards. Active stacks get the full check with interpolation in
|
|
||||||
# the deploy workflow, where the real .env files live.
|
|
||||||
- name: Validate Compose files
|
- name: Validate Compose files
|
||||||
shell: bash
|
shell: bash
|
||||||
run: |
|
run: |
|
||||||
@@ -61,35 +48,77 @@ jobs:
|
|||||||
exit 1
|
exit 1
|
||||||
fi
|
fi
|
||||||
echo "checked ${#files[@]} Compose file(s)"
|
echo "checked ${#files[@]} Compose file(s)"
|
||||||
|
id: check
|
||||||
lint-actionlint:
|
- name: Write the job result
|
||||||
runs-on: [self-hosted, linux, arch, homelab]
|
if: always()
|
||||||
timeout-minutes: 10
|
env:
|
||||||
|
SUMMARY_CHECK: Compose
|
||||||
|
SUMMARY_RESULT: ${{ job.status }}
|
||||||
|
SUMMARY_FAILED_STEP:
|
||||||
|
${{ steps.check.conclusion == 'failure' && 'Check or image build' || steps.source.conclusion == 'failure'
|
||||||
|
&& 'Source checkout' || '' }}
|
||||||
|
shell: bash
|
||||||
|
run: |
|
||||||
|
if [ -f .gitea/workflows/release.py ]; then
|
||||||
|
python3 .gitea/workflows/release.py check-summary
|
||||||
|
elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
|
||||||
|
printf '## %s\n\n- Result: **%s**\n- Failed step: %s\n' "$SUMMARY_CHECK" "$SUMMARY_RESULT" "$SUMMARY_FAILED_STEP" >>"$GITHUB_STEP_SUMMARY" || true
|
||||||
|
fi
|
||||||
|
workflows:
|
||||||
|
name: Workflows
|
||||||
|
runs-on: ${{ github.ref == 'refs/heads/main' && 'homelab' || 'homelab-pr' }}
|
||||||
|
timeout-minutes: 15
|
||||||
steps:
|
steps:
|
||||||
- name: Checkout repository
|
- name: Checkout repository
|
||||||
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
|
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262
|
||||||
|
id: source
|
||||||
|
- name: Prepare pinned tools
|
||||||
|
shell: bash
|
||||||
|
run: |
|
||||||
|
set -euo pipefail
|
||||||
|
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh actionlint shellcheck)"
|
||||||
|
echo "$tools_dir" >> "$GITHUB_PATH"
|
||||||
|
id: tools
|
||||||
- name: Lint Gitea Actions workflows with actionlint
|
- name: Lint Gitea Actions workflows with actionlint
|
||||||
shell: bash
|
shell: bash
|
||||||
run: |
|
run: |
|
||||||
set -euo pipefail
|
set -euo pipefail
|
||||||
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh actionlint)"
|
|
||||||
export PATH="$tools_dir:$PATH"
|
|
||||||
actionlint -config-file .gitea/actionlint.yaml -color .gitea/workflows/*.yaml
|
actionlint -config-file .gitea/actionlint.yaml -color .gitea/workflows/*.yaml
|
||||||
|
id: check
|
||||||
lint-shellcheck:
|
- name: Write the job result
|
||||||
runs-on: [self-hosted, linux, arch, homelab]
|
if: always()
|
||||||
timeout-minutes: 10
|
env:
|
||||||
|
SUMMARY_CHECK: Workflows
|
||||||
|
SUMMARY_RESULT: ${{ job.status }}
|
||||||
|
SUMMARY_FAILED_STEP:
|
||||||
|
${{ steps.check.conclusion == 'failure' && 'Check or image build' || steps.tools.conclusion == 'failure'
|
||||||
|
&& 'Tool setup' || steps.source.conclusion == 'failure' && 'Source checkout' || '' }}
|
||||||
|
shell: bash
|
||||||
|
run: |
|
||||||
|
if [ -f .gitea/workflows/release.py ]; then
|
||||||
|
python3 .gitea/workflows/release.py check-summary
|
||||||
|
elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
|
||||||
|
printf '## %s\n\n- Result: **%s**\n- Failed step: %s\n' "$SUMMARY_CHECK" "$SUMMARY_RESULT" "$SUMMARY_FAILED_STEP" >>"$GITHUB_STEP_SUMMARY" || true
|
||||||
|
fi
|
||||||
|
shell:
|
||||||
|
name: Shell
|
||||||
|
runs-on: ${{ github.ref == 'refs/heads/main' && 'homelab' || 'homelab-pr' }}
|
||||||
|
timeout-minutes: 15
|
||||||
steps:
|
steps:
|
||||||
- name: Checkout repository
|
- name: Checkout repository
|
||||||
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
|
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262
|
||||||
|
id: source
|
||||||
- name: Lint shell scripts with ShellCheck
|
- name: Prepare pinned tools
|
||||||
shell: bash
|
shell: bash
|
||||||
run: |
|
run: |
|
||||||
set -euo pipefail
|
set -euo pipefail
|
||||||
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh shellcheck jq)"
|
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh shellcheck jq)"
|
||||||
export PATH="$tools_dir:$PATH"
|
echo "$tools_dir" >> "$GITHUB_PATH"
|
||||||
|
id: tools
|
||||||
|
- name: Lint shell scripts with ShellCheck
|
||||||
|
shell: bash
|
||||||
|
run: |
|
||||||
|
set -euo pipefail
|
||||||
mapfile -t scripts < <(
|
mapfile -t scripts < <(
|
||||||
git ls-files '*.sh' ':(glob)**/*.bash'
|
git ls-files '*.sh' ':(glob)**/*.bash'
|
||||||
)
|
)
|
||||||
@@ -99,20 +128,41 @@ jobs:
|
|||||||
fi
|
fi
|
||||||
shellcheck --external-sources --source-path=SCRIPTDIR --severity=style "${scripts[@]}"
|
shellcheck --external-sources --source-path=SCRIPTDIR --severity=style "${scripts[@]}"
|
||||||
bash .gitea/tests/deploy-validation.sh
|
bash .gitea/tests/deploy-validation.sh
|
||||||
|
id: check
|
||||||
lint-prettier:
|
- name: Write the job result
|
||||||
runs-on: [self-hosted, linux, arch, homelab]
|
if: always()
|
||||||
timeout-minutes: 10
|
env:
|
||||||
|
SUMMARY_CHECK: Shell
|
||||||
|
SUMMARY_RESULT: ${{ job.status }}
|
||||||
|
SUMMARY_FAILED_STEP:
|
||||||
|
${{ steps.check.conclusion == 'failure' && 'Check or image build' || steps.tools.conclusion == 'failure'
|
||||||
|
&& 'Tool setup' || steps.source.conclusion == 'failure' && 'Source checkout' || '' }}
|
||||||
|
shell: bash
|
||||||
|
run: |
|
||||||
|
if [ -f .gitea/workflows/release.py ]; then
|
||||||
|
python3 .gitea/workflows/release.py check-summary
|
||||||
|
elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
|
||||||
|
printf '## %s\n\n- Result: **%s**\n- Failed step: %s\n' "$SUMMARY_CHECK" "$SUMMARY_RESULT" "$SUMMARY_FAILED_STEP" >>"$GITHUB_STEP_SUMMARY" || true
|
||||||
|
fi
|
||||||
|
formatting:
|
||||||
|
name: Formatting
|
||||||
|
runs-on: ${{ github.ref == 'refs/heads/main' && 'homelab' || 'homelab-pr' }}
|
||||||
|
timeout-minutes: 15
|
||||||
steps:
|
steps:
|
||||||
- name: Checkout repository
|
- name: Checkout repository
|
||||||
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
|
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262
|
||||||
|
id: source
|
||||||
- name: Check formatting with Prettier
|
- name: Prepare pinned tools
|
||||||
shell: bash
|
shell: bash
|
||||||
run: |
|
run: |
|
||||||
set -euo pipefail
|
set -euo pipefail
|
||||||
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh prettier)"
|
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh prettier)"
|
||||||
export PATH="$tools_dir:$PATH"
|
echo "$tools_dir" >> "$GITHUB_PATH"
|
||||||
|
id: tools
|
||||||
|
- name: Check formatting with Prettier
|
||||||
|
shell: bash
|
||||||
|
run: |
|
||||||
|
set -euo pipefail
|
||||||
|
|
||||||
mapfile -t prettier_files < <(
|
mapfile -t prettier_files < <(
|
||||||
git ls-files \
|
git ls-files \
|
||||||
@@ -126,36 +176,79 @@ jobs:
|
|||||||
fi
|
fi
|
||||||
|
|
||||||
prettier --check --ignore-unknown "${prettier_files[@]}"
|
prettier --check --ignore-unknown "${prettier_files[@]}"
|
||||||
|
id: check
|
||||||
lint-ruff:
|
- name: Write the job result
|
||||||
runs-on: [self-hosted, linux, arch, homelab]
|
if: always()
|
||||||
timeout-minutes: 10
|
env:
|
||||||
|
SUMMARY_CHECK: Formatting
|
||||||
|
SUMMARY_RESULT: ${{ job.status }}
|
||||||
|
SUMMARY_FAILED_STEP:
|
||||||
|
${{ steps.check.conclusion == 'failure' && 'Check or image build' || steps.tools.conclusion == 'failure'
|
||||||
|
&& 'Tool setup' || steps.source.conclusion == 'failure' && 'Source checkout' || '' }}
|
||||||
|
shell: bash
|
||||||
|
run: |
|
||||||
|
if [ -f .gitea/workflows/release.py ]; then
|
||||||
|
python3 .gitea/workflows/release.py check-summary
|
||||||
|
elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
|
||||||
|
printf '## %s\n\n- Result: **%s**\n- Failed step: %s\n' "$SUMMARY_CHECK" "$SUMMARY_RESULT" "$SUMMARY_FAILED_STEP" >>"$GITHUB_STEP_SUMMARY" || true
|
||||||
|
fi
|
||||||
|
python:
|
||||||
|
name: Python and tests
|
||||||
|
runs-on: ${{ github.ref == 'refs/heads/main' && 'homelab' || 'homelab-pr' }}
|
||||||
|
timeout-minutes: 15
|
||||||
steps:
|
steps:
|
||||||
- name: Checkout repository
|
- name: Checkout repository
|
||||||
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
|
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262
|
||||||
|
id: source
|
||||||
|
- name: Prepare pinned tools
|
||||||
|
shell: bash
|
||||||
|
run: |
|
||||||
|
set -euo pipefail
|
||||||
|
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh ruff jq)"
|
||||||
|
echo "$tools_dir" >> "$GITHUB_PATH"
|
||||||
|
id: tools
|
||||||
- name: Lint and format-check Python with Ruff
|
- name: Lint and format-check Python with Ruff
|
||||||
shell: bash
|
shell: bash
|
||||||
run: |
|
run: |
|
||||||
set -euo pipefail
|
set -euo pipefail
|
||||||
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh ruff)"
|
ruff check . .gitea/workflows
|
||||||
export PATH="$tools_dir:$PATH"
|
ruff format --check . .gitea/workflows
|
||||||
ruff check .
|
python3 -m unittest discover -s tests -v
|
||||||
ruff format --check .
|
id: check
|
||||||
|
- name: Write the job result
|
||||||
lint-yaml:
|
if: always()
|
||||||
runs-on: [self-hosted, linux, arch, homelab]
|
env:
|
||||||
timeout-minutes: 10
|
SUMMARY_CHECK: Python and tests
|
||||||
|
SUMMARY_RESULT: ${{ job.status }}
|
||||||
|
SUMMARY_FAILED_STEP:
|
||||||
|
${{ steps.check.conclusion == 'failure' && 'Check or image build' || steps.tools.conclusion == 'failure'
|
||||||
|
&& 'Tool setup' || steps.source.conclusion == 'failure' && 'Source checkout' || '' }}
|
||||||
|
shell: bash
|
||||||
|
run: |
|
||||||
|
if [ -f .gitea/workflows/release.py ]; then
|
||||||
|
python3 .gitea/workflows/release.py check-summary
|
||||||
|
elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
|
||||||
|
printf '## %s\n\n- Result: **%s**\n- Failed step: %s\n' "$SUMMARY_CHECK" "$SUMMARY_RESULT" "$SUMMARY_FAILED_STEP" >>"$GITHUB_STEP_SUMMARY" || true
|
||||||
|
fi
|
||||||
|
yaml:
|
||||||
|
name: YAML
|
||||||
|
runs-on: ${{ github.ref == 'refs/heads/main' && 'homelab' || 'homelab-pr' }}
|
||||||
|
timeout-minutes: 15
|
||||||
steps:
|
steps:
|
||||||
- name: Checkout repository
|
- name: Checkout repository
|
||||||
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
|
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262
|
||||||
|
id: source
|
||||||
- name: Lint YAML syntax
|
- name: Prepare pinned tools
|
||||||
shell: bash
|
shell: bash
|
||||||
run: |
|
run: |
|
||||||
set -euo pipefail
|
set -euo pipefail
|
||||||
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh yamllint)"
|
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh yamllint)"
|
||||||
export PATH="$tools_dir:$PATH"
|
echo "$tools_dir" >> "$GITHUB_PATH"
|
||||||
|
id: tools
|
||||||
|
- name: Lint YAML syntax
|
||||||
|
shell: bash
|
||||||
|
run: |
|
||||||
|
set -euo pipefail
|
||||||
|
|
||||||
mapfile -t yaml_files < <(
|
mapfile -t yaml_files < <(
|
||||||
git ls-files '*.yaml' '*.yml' \
|
git ls-files '*.yaml' '*.yml' \
|
||||||
@@ -169,20 +262,41 @@ jobs:
|
|||||||
fi
|
fi
|
||||||
|
|
||||||
yamllint -c .yamllint "${yaml_files[@]}"
|
yamllint -c .yamllint "${yaml_files[@]}"
|
||||||
|
id: check
|
||||||
lint-dockerfiles:
|
- name: Write the job result
|
||||||
runs-on: [self-hosted, linux, arch, homelab]
|
if: always()
|
||||||
timeout-minutes: 10
|
env:
|
||||||
|
SUMMARY_CHECK: YAML
|
||||||
|
SUMMARY_RESULT: ${{ job.status }}
|
||||||
|
SUMMARY_FAILED_STEP:
|
||||||
|
${{ steps.check.conclusion == 'failure' && 'Check or image build' || steps.tools.conclusion == 'failure'
|
||||||
|
&& 'Tool setup' || steps.source.conclusion == 'failure' && 'Source checkout' || '' }}
|
||||||
|
shell: bash
|
||||||
|
run: |
|
||||||
|
if [ -f .gitea/workflows/release.py ]; then
|
||||||
|
python3 .gitea/workflows/release.py check-summary
|
||||||
|
elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
|
||||||
|
printf '## %s\n\n- Result: **%s**\n- Failed step: %s\n' "$SUMMARY_CHECK" "$SUMMARY_RESULT" "$SUMMARY_FAILED_STEP" >>"$GITHUB_STEP_SUMMARY" || true
|
||||||
|
fi
|
||||||
|
dockerfiles:
|
||||||
|
name: Dockerfiles
|
||||||
|
runs-on: ${{ github.ref == 'refs/heads/main' && 'homelab' || 'homelab-pr' }}
|
||||||
|
timeout-minutes: 15
|
||||||
steps:
|
steps:
|
||||||
- name: Checkout repository
|
- name: Checkout repository
|
||||||
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
|
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262
|
||||||
|
id: source
|
||||||
- name: Lint Dockerfiles
|
- name: Prepare pinned tools
|
||||||
shell: bash
|
shell: bash
|
||||||
run: |
|
run: |
|
||||||
set -euo pipefail
|
set -euo pipefail
|
||||||
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh hadolint)"
|
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh hadolint)"
|
||||||
export PATH="$tools_dir:$PATH"
|
echo "$tools_dir" >> "$GITHUB_PATH"
|
||||||
|
id: tools
|
||||||
|
- name: Lint Dockerfiles
|
||||||
|
shell: bash
|
||||||
|
run: |
|
||||||
|
set -euo pipefail
|
||||||
|
|
||||||
mapfile -t dockerfiles < <(
|
mapfile -t dockerfiles < <(
|
||||||
git ls-files ':(glob)**/Dockerfile' ':(glob)**/Dockerfile.*'
|
git ls-files ':(glob)**/Dockerfile' ':(glob)**/Dockerfile.*'
|
||||||
@@ -194,20 +308,41 @@ jobs:
|
|||||||
fi
|
fi
|
||||||
|
|
||||||
hadolint -c .hadolint.yaml "${dockerfiles[@]}"
|
hadolint -c .hadolint.yaml "${dockerfiles[@]}"
|
||||||
|
id: check
|
||||||
validate:
|
- name: Write the job result
|
||||||
runs-on: [self-hosted, linux, arch, homelab]
|
if: always()
|
||||||
timeout-minutes: 20
|
env:
|
||||||
|
SUMMARY_CHECK: Dockerfiles
|
||||||
|
SUMMARY_RESULT: ${{ job.status }}
|
||||||
|
SUMMARY_FAILED_STEP:
|
||||||
|
${{ steps.check.conclusion == 'failure' && 'Check or image build' || steps.tools.conclusion == 'failure'
|
||||||
|
&& 'Tool setup' || steps.source.conclusion == 'failure' && 'Source checkout' || '' }}
|
||||||
|
shell: bash
|
||||||
|
run: |
|
||||||
|
if [ -f .gitea/workflows/release.py ]; then
|
||||||
|
python3 .gitea/workflows/release.py check-summary
|
||||||
|
elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
|
||||||
|
printf '## %s\n\n- Result: **%s**\n- Failed step: %s\n' "$SUMMARY_CHECK" "$SUMMARY_RESULT" "$SUMMARY_FAILED_STEP" >>"$GITHUB_STEP_SUMMARY" || true
|
||||||
|
fi
|
||||||
|
kubernetes:
|
||||||
|
name: Kubernetes
|
||||||
|
runs-on: ${{ github.ref == 'refs/heads/main' && 'homelab' || 'homelab-pr' }}
|
||||||
|
timeout-minutes: 15
|
||||||
steps:
|
steps:
|
||||||
- name: Checkout repository
|
- name: Checkout repository
|
||||||
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
|
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262
|
||||||
|
id: source
|
||||||
- name: Validate Kubernetes manifests against JSON schemas
|
- name: Prepare pinned tools
|
||||||
shell: bash
|
shell: bash
|
||||||
run: |
|
run: |
|
||||||
set -euo pipefail
|
set -euo pipefail
|
||||||
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh kubeconform)"
|
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh kubeconform)"
|
||||||
export PATH="$tools_dir:$PATH"
|
echo "$tools_dir" >> "$GITHUB_PATH"
|
||||||
|
id: tools
|
||||||
|
- name: Validate Kubernetes manifests against JSON schemas
|
||||||
|
shell: bash
|
||||||
|
run: |
|
||||||
|
set -euo pipefail
|
||||||
|
|
||||||
mapfile -t manifests < <(
|
mapfile -t manifests < <(
|
||||||
git ls-files ':(glob)**/k8s/**/*.yaml' ':(glob)**/k8s/**/*.yml' \
|
git ls-files ':(glob)**/k8s/**/*.yaml' ':(glob)**/k8s/**/*.yml' \
|
||||||
@@ -224,327 +359,162 @@ jobs:
|
|||||||
-ignore-missing-schemas \
|
-ignore-missing-schemas \
|
||||||
-summary \
|
-summary \
|
||||||
"${manifests[@]}"
|
"${manifests[@]}"
|
||||||
|
id: check
|
||||||
# kubeconform has no schemas for CRDs, so every IngressRoute, Certificate,
|
- name: Write the job result
|
||||||
# PrometheusRule, Middleware, ServersTransport and ServiceMonitor is silently
|
if: always()
|
||||||
# skipped above. The live API server knows the real CRD schemas (and runs the
|
env:
|
||||||
# cert-manager / Traefik admission webhooks), so validate there too.
|
SUMMARY_CHECK: Kubernetes
|
||||||
#
|
SUMMARY_RESULT: ${{ job.status }}
|
||||||
# Only services marked with a k8s/active marker are checked: server-side
|
SUMMARY_FAILED_STEP:
|
||||||
# dry-run needs the target namespace to exist, and inactive services are not
|
${{ steps.check.conclusion == 'failure' && 'Check or image build' || steps.tools.conclusion == 'failure'
|
||||||
# deployed. Services being enabled for the first time are still covered by
|
&& 'Tool setup' || steps.source.conclusion == 'failure' && 'Source checkout' || '' }}
|
||||||
# the JSON-schema pass above.
|
|
||||||
#
|
|
||||||
# Main pushes only. `--dry-run=server` persists nothing, but it does execute
|
|
||||||
# the admission webhooks of the production API server, so anyone able to open
|
|
||||||
# a pull request would be able to run arbitrary manifest content through
|
|
||||||
# cert-manager and Traefik. A pull request has nothing to gain from it either:
|
|
||||||
# only main is ever deployed, and this job runs to completion before the
|
|
||||||
# deploy workflow is allowed to start, so a bad CRD is still caught before
|
|
||||||
# anything reaches the cluster -- just on the push rather than on the PR.
|
|
||||||
- name: Note the server-side check is not running here
|
|
||||||
if: github.event_name == 'pull_request' || github.ref != 'refs/heads/main'
|
|
||||||
shell: bash
|
shell: bash
|
||||||
run: |
|
run: |
|
||||||
echo "::notice::Skipping the server-side dry-run. It executes the cert-manager and" \
|
if [ -f .gitea/workflows/release.py ]; then
|
||||||
"Traefik admission webhooks against the production API server, so it is limited" \
|
python3 .gitea/workflows/release.py check-summary
|
||||||
"to pushes to main. CRDs are still schema-checked by kubeconform above, and the" \
|
elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
|
||||||
"server-side pass still runs on main before the deploy."
|
printf '## %s\n\n- Result: **%s**\n- Failed step: %s\n' "$SUMMARY_CHECK" "$SUMMARY_RESULT" "$SUMMARY_FAILED_STEP" >>"$GITHUB_STEP_SUMMARY" || true
|
||||||
|
|
||||||
- name: Validate active manifests against the live API server
|
|
||||||
if: github.event_name != 'pull_request' && github.ref == 'refs/heads/main'
|
|
||||||
shell: bash
|
|
||||||
run: |
|
|
||||||
set -euo pipefail
|
|
||||||
|
|
||||||
if ! kubectl get --raw='/readyz' --request-timeout=10s >/dev/null 2>&1; then
|
|
||||||
echo "::warning::Cluster unreachable — skipped server-side validation of CRDs (IngressRoute, Certificate, PrometheusRule). Review manifest changes manually."
|
|
||||||
exit 0
|
|
||||||
fi
|
fi
|
||||||
|
image-plan:
|
||||||
mapfile -t k8s_dirs < <(
|
needs: [compose, workflows, shell, formatting, python, yaml, dockerfiles, kubernetes]
|
||||||
git ls-files '*.yaml' '*.yml' \
|
if: github.event_name != 'pull_request' && github.ref == 'refs/heads/main'
|
||||||
| grep -E '(^|/)k8s/' \
|
runs-on: homelab
|
||||||
| sed -E 's#((^|.*/)k8s)/.*#\1#' \
|
timeout-minutes: 10
|
||||||
| sort -u
|
|
||||||
)
|
|
||||||
|
|
||||||
manifests=()
|
|
||||||
kustomize_apps=()
|
|
||||||
for dir in "${k8s_dirs[@]}"; do
|
|
||||||
if [ ! -f "${dir}/active" ]; then
|
|
||||||
echo "skip (no k8s/active): ${dir}"
|
|
||||||
continue
|
|
||||||
fi
|
|
||||||
if [ -f "${dir}/overlays/prod/kustomization.yaml" ]; then
|
|
||||||
kustomize_apps+=("${dir}/overlays/prod")
|
|
||||||
elif [ -f "${dir}/base/kustomization.yaml" ]; then
|
|
||||||
kustomize_apps+=("${dir}/base")
|
|
||||||
else
|
|
||||||
while IFS= read -r f; do
|
|
||||||
[ -n "$f" ] && manifests+=("$f")
|
|
||||||
done < <(
|
|
||||||
git ls-files "${dir}/*.yaml" "${dir}/*.yml" \
|
|
||||||
| grep -Ev '(^|/)(kustomization\.ya?ml|.*\.example\.ya?ml|.*values\.ya?ml|patch-.*\.ya?ml)$'
|
|
||||||
)
|
|
||||||
fi
|
|
||||||
done
|
|
||||||
|
|
||||||
echo "server-side dry-run: ${#manifests[@]} manifests, ${#kustomize_apps[@]} kustomize apps"
|
|
||||||
failed=0
|
|
||||||
for m in ${manifests[@]+"${manifests[@]}"}; do
|
|
||||||
if ! out="$(kubectl apply --dry-run=server -f "$m" 2>&1)"; then
|
|
||||||
failed=1
|
|
||||||
echo "::error file=${m}::$(printf '%s' "$out" | head -1)"
|
|
||||||
fi
|
|
||||||
done
|
|
||||||
for k in ${kustomize_apps[@]+"${kustomize_apps[@]}"}; do
|
|
||||||
if ! out="$(kubectl apply -k "$k" --dry-run=server 2>&1)"; then
|
|
||||||
failed=1
|
|
||||||
echo "::error file=${k}::$(printf '%s' "$out" | head -1)"
|
|
||||||
fi
|
|
||||||
done
|
|
||||||
|
|
||||||
if [ "$failed" -ne 0 ]; then
|
|
||||||
echo "Server-side validation failed. The API server (or an admission webhook) rejected these manifests."
|
|
||||||
exit 1
|
|
||||||
fi
|
|
||||||
echo "server-side dry-run: all active manifests accepted by the API server"
|
|
||||||
|
|
||||||
build:
|
|
||||||
needs:
|
|
||||||
# The panel's scan-deps/test-backend/test-frontend jobs gated here until
|
|
||||||
# userbot moved to its own repo; upstream's code is upstream's gate now.
|
|
||||||
# The rule is unchanged: publishing and passing the checks are the same
|
|
||||||
# gate, so a commit that fails any of these still cannot move :prod.
|
|
||||||
[lint-actionlint, lint-shellcheck, lint-compose, lint-prettier, lint-ruff, lint-yaml, lint-dockerfiles, validate]
|
|
||||||
if: github.event_name != 'pull_request' && (github.ref_name == 'main' || github.ref_name == 'dev') && !startsWith(github.ref_name, 'renovate/')
|
|
||||||
runs-on: [self-hosted, linux, arch, homelab]
|
|
||||||
timeout-minutes: 60
|
|
||||||
outputs:
|
outputs:
|
||||||
services: ${{ steps.services.outputs.services }}
|
matrix: ${{ steps.plan.outputs.matrix }}
|
||||||
steps:
|
steps:
|
||||||
- name: Checkout repository
|
- name: Checkout repository
|
||||||
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
|
id: source
|
||||||
|
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262
|
||||||
with:
|
with:
|
||||||
fetch-depth: 0
|
fetch-depth: 0
|
||||||
|
- name: Detect build inputs against successful CI
|
||||||
|
id: plan
|
||||||
|
env:
|
||||||
|
GITEA_TOKEN: ${{ github.token }}
|
||||||
|
run: python3 .gitea/workflows/release.py prepare --output build-plan.json
|
||||||
|
- name: Store the image plan
|
||||||
|
id: artifact
|
||||||
|
uses: actions/upload-artifact@ea165f8d65b6e75b540449e92b4886f43607fa02 # v4.6.2
|
||||||
|
with:
|
||||||
|
name: build-plan
|
||||||
|
path: build-plan.json
|
||||||
|
if-no-files-found: error
|
||||||
|
retention-days: 30
|
||||||
|
|
||||||
- name: Detect changed docker-built services
|
- name: Write the plan result
|
||||||
id: services
|
if: always()
|
||||||
|
env:
|
||||||
|
SUMMARY_CHECK: Image plan
|
||||||
|
SUMMARY_RESULT: ${{ job.status }}
|
||||||
|
SUMMARY_FAILED_STEP: >-
|
||||||
|
${{ steps.plan.conclusion == 'failure' && 'Build input detection' ||
|
||||||
|
steps.artifact.conclusion == 'failure' && 'Plan upload' ||
|
||||||
|
steps.source.conclusion == 'failure' && 'Source checkout' || '' }}
|
||||||
shell: bash
|
shell: bash
|
||||||
run: |
|
run: |
|
||||||
set -euo pipefail
|
if [ -f .gitea/workflows/release.py ]; then
|
||||||
base="${{ github.event.before }}"
|
python3 .gitea/workflows/release.py check-summary
|
||||||
if [ -z "$base" ] || [ "$base" = "0000000000000000000000000000000000000000" ]; then
|
elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
|
||||||
base="$(git rev-list --max-parents=0 HEAD)"
|
printf '## Image plan\n\nResult: %s\n' "$SUMMARY_RESULT" >>"$GITHUB_STEP_SUMMARY" || true
|
||||||
fi
|
fi
|
||||||
|
|
||||||
# A failed diff used to leave changed_files empty, which reads exactly
|
images:
|
||||||
# like "nothing to build": the job went green having built nothing and
|
name: Image (${{ matrix.name }})
|
||||||
# the tag never moved. The status is checked, not assumed.
|
needs: [image-plan]
|
||||||
if ! changed="$(git diff --name-only "$base" "${GITHUB_SHA}")"; then
|
if: needs.image-plan.result == 'success'
|
||||||
echo "::error::cannot diff ${base}..${GITHUB_SHA}"
|
runs-on: homelab
|
||||||
exit 1
|
timeout-minutes: 60
|
||||||
fi
|
strategy:
|
||||||
mapfile -t changed_files <<<"$changed"
|
max-parallel: 1
|
||||||
|
fail-fast: false
|
||||||
services=()
|
matrix: ${{ fromJSON(needs.image-plan.outputs.matrix || '{"include":[{"name":"inactive"}]}') }}
|
||||||
|
steps:
|
||||||
add_service() {
|
- name: Checkout repository
|
||||||
local name="$1"
|
id: source
|
||||||
local seen=0
|
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262
|
||||||
for existing in "${services[@]}"; do
|
- name: Download the checked image plan
|
||||||
if [ "$existing" = "$name" ]; then
|
id: inputs
|
||||||
seen=1
|
uses: actions/download-artifact@d3f86a106a0bac45b974a628896c90dbdf5c8093 # v4.3.0
|
||||||
break
|
with:
|
||||||
fi
|
name: build-plan
|
||||||
done
|
- name: Build or reuse this image
|
||||||
if [ "$seen" -eq 0 ]; then
|
id: check
|
||||||
services+=("$name")
|
env:
|
||||||
fi
|
IMAGE_NAME: ${{ matrix.name }}
|
||||||
}
|
REGISTRY_USERNAME: ${{ secrets.REGISTRY_USERNAME }}
|
||||||
|
REGISTRY_PASSWORD: ${{ secrets.REGISTRY_PASSWORD }}
|
||||||
for file in "${changed_files[@]}"; do
|
run: python3 .gitea/workflows/release.py image --image "$IMAGE_NAME" --output image.json
|
||||||
case "$file" in
|
- name: Store the image result
|
||||||
errorpages/*)
|
id: artifact
|
||||||
add_service errorpages
|
uses: actions/upload-artifact@ea165f8d65b6e75b540449e92b4886f43607fa02 # v4.6.2
|
||||||
;;
|
with:
|
||||||
homepages/*)
|
name: image-${{ matrix.name }}
|
||||||
add_service homepages
|
path: image.json
|
||||||
;;
|
if-no-files-found: error
|
||||||
edu_master/phpsessid-bot/*|edu_master/webinar-checker/*|edu_master/compose.yaml)
|
retention-days: 30
|
||||||
add_service edu_master
|
- name: Write the job result
|
||||||
;;
|
if: always()
|
||||||
esac
|
env:
|
||||||
done
|
SUMMARY_CHECK: Image (${{ matrix.name }})
|
||||||
|
SUMMARY_RESULT: ${{ job.status }}
|
||||||
if [ "${#services[@]}" -eq 0 ]; then
|
SUMMARY_FAILED_STEP: >-
|
||||||
echo "No docker-built services changed."
|
${{ steps.check.conclusion == 'failure' && 'Build or tag images' ||
|
||||||
echo "services=" >> "$GITHUB_OUTPUT"
|
steps.artifact.conclusion == 'failure' && 'Artifact upload' ||
|
||||||
exit 0
|
steps.inputs.conclusion == 'failure' && 'Artifact download' ||
|
||||||
fi
|
steps.source.conclusion == 'failure' && 'Source checkout' || '' }}
|
||||||
|
|
||||||
printf '%s\n' "${services[@]}" | tee /tmp/services.txt
|
|
||||||
echo "services=$(paste -sd, /tmp/services.txt)" >> "$GITHUB_OUTPUT"
|
|
||||||
|
|
||||||
- name: Log in to registry
|
|
||||||
# The pin step below also writes (manifest PUTs), and it runs on every
|
|
||||||
# main push — including manifest-only ones where services is empty. A
|
|
||||||
# stale persistent login on the old runner used to mask this; a clean
|
|
||||||
# runner pushes anonymously and gets 401.
|
|
||||||
if: steps.services.outputs.services != '' || github.ref_name == 'main'
|
|
||||||
shell: bash
|
shell: bash
|
||||||
# Through env, not by substitution into the script. A secret written
|
run: |
|
||||||
# into a run: block is pasted into the shell source before bash parses
|
if [ -f .gitea/workflows/release.py ]; then
|
||||||
# it, so a password containing a quote, a backtick or $(...) becomes
|
python3 .gitea/workflows/release.py check-summary
|
||||||
# code that runs. Masking the value in the log does not prevent that.
|
elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
|
||||||
|
printf '## %s\n\n- Result: **%s**\n- Failed step: %s\n' "$SUMMARY_CHECK" "$SUMMARY_RESULT" "$SUMMARY_FAILED_STEP" >>"$GITHUB_STEP_SUMMARY" || true
|
||||||
|
fi
|
||||||
|
|
||||||
|
# Retain the build job name required by the immutable release deployment gate.
|
||||||
|
build:
|
||||||
|
needs: [image-plan, images]
|
||||||
|
runs-on: homelab
|
||||||
|
timeout-minutes: 15
|
||||||
|
steps:
|
||||||
|
- name: Checkout repository
|
||||||
|
id: source
|
||||||
|
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262
|
||||||
|
- name: Download all image results
|
||||||
|
id: inputs
|
||||||
|
uses: actions/download-artifact@d3f86a106a0bac45b974a628896c90dbdf5c8093 # v4.3.0
|
||||||
|
with:
|
||||||
|
path: artifacts
|
||||||
|
- name: Pin SHA tags and write the complete release
|
||||||
|
id: check
|
||||||
env:
|
env:
|
||||||
REGISTRY_USERNAME: ${{ secrets.REGISTRY_USERNAME }}
|
REGISTRY_USERNAME: ${{ secrets.REGISTRY_USERNAME }}
|
||||||
REGISTRY_PASSWORD: ${{ secrets.REGISTRY_PASSWORD }}
|
REGISTRY_PASSWORD: ${{ secrets.REGISTRY_PASSWORD }}
|
||||||
run: |
|
run: >-
|
||||||
set -euo pipefail
|
python3 .gitea/workflows/release.py finalize
|
||||||
printf '%s' "$REGISTRY_PASSWORD" | docker login "${REGISTRY}" \
|
--plan artifacts/build-plan/build-plan.json
|
||||||
-u "$REGISTRY_USERNAME" \
|
- name: Store commit release
|
||||||
--password-stdin
|
id: artifact
|
||||||
|
uses: actions/upload-artifact@ea165f8d65b6e75b540449e92b4886f43607fa02 # v4.6.2
|
||||||
- name: Build and push changed images
|
with:
|
||||||
if: steps.services.outputs.services != ''
|
name: release-${{ github.sha }}
|
||||||
|
path: release.json
|
||||||
|
if-no-files-found: error
|
||||||
|
retention-days: 30
|
||||||
|
- name: Write the job result
|
||||||
|
if: always()
|
||||||
|
env:
|
||||||
|
SUMMARY_CHECK: Image release and SHA tags
|
||||||
|
SUMMARY_RESULT: ${{ job.status }}
|
||||||
|
SUMMARY_FAILED_STEP: >-
|
||||||
|
${{ steps.check.conclusion == 'failure' && 'Build or tag images' ||
|
||||||
|
steps.artifact.conclusion == 'failure' && 'Artifact upload' ||
|
||||||
|
steps.inputs.conclusion == 'failure' && 'Artifact download' ||
|
||||||
|
steps.source.conclusion == 'failure' && 'Source checkout' || '' }}
|
||||||
shell: bash
|
shell: bash
|
||||||
run: |
|
run: |
|
||||||
# This step was the one run: block in the workflow without it, and it
|
if [ -f .gitea/workflows/release.py ]; then
|
||||||
# is the one that cannot afford it: a docker push that failed partway
|
python3 .gitea/workflows/release.py check-summary
|
||||||
# through the loop used to be followed by more pushes, the loop's exit
|
elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
|
||||||
# status came from the last one, and the job went green with half the
|
printf '## %s\n\n- Result: **%s**\n- Failed step: %s\n' "$SUMMARY_CHECK" "$SUMMARY_RESULT" "$SUMMARY_FAILED_STEP" >>"$GITHUB_STEP_SUMMARY" || true
|
||||||
# images missing from the registry.
|
|
||||||
set -euo pipefail
|
|
||||||
IFS=, read -r -a services <<< "${{ steps.services.outputs.services }}"
|
|
||||||
|
|
||||||
# Tags for this push. The commit-pinned name is the point of this
|
|
||||||
# step: the deploy resolves it in preference to :prod, so a deploy
|
|
||||||
# that sat in the queue behind a later push still gets the build of
|
|
||||||
# the commit CI validated, instead of whatever :prod points at by the
|
|
||||||
# time it runs. See render_pinned in deploy-lib.sh.
|
|
||||||
commit_tag=""
|
|
||||||
if [ "${GITHUB_REF_NAME}" = "main" ]; then
|
|
||||||
commit_tag="sha-${GITHUB_SHA:0:12}"
|
|
||||||
fi
|
fi
|
||||||
|
|
||||||
set_tags() {
|
|
||||||
tags=()
|
|
||||||
case "${GITHUB_REF_NAME}" in
|
|
||||||
main) tags+=("main" "prod") ;;
|
|
||||||
dev) tags+=("dev") ;;
|
|
||||||
esac
|
|
||||||
if [ -n "$commit_tag" ]; then
|
|
||||||
tags+=("$commit_tag")
|
|
||||||
fi
|
|
||||||
}
|
|
||||||
|
|
||||||
for service in "${services[@]}"; do
|
|
||||||
case "$service" in
|
|
||||||
errorpages)
|
|
||||||
image="${REGISTRY}/forust/error-pages"
|
|
||||||
set_tags
|
|
||||||
build_args=()
|
|
||||||
for tag in "${tags[@]}"; do
|
|
||||||
build_args+=(-t "${image}:${tag}")
|
|
||||||
done
|
|
||||||
docker build \
|
|
||||||
--cache-from "type=registry,ref=${image}:buildcache" \
|
|
||||||
--cache-to "type=registry,ref=${image}:buildcache,mode=max" \
|
|
||||||
"${build_args[@]}" errorpages
|
|
||||||
for tag in "${tags[@]}"; do
|
|
||||||
docker push "${image}:${tag}"
|
|
||||||
done
|
|
||||||
;;
|
|
||||||
homepages)
|
|
||||||
for variant in forust xdfnx; do
|
|
||||||
case "$variant" in
|
|
||||||
forust)
|
|
||||||
image="${REGISTRY}/forust/forust-homepage"
|
|
||||||
;;
|
|
||||||
xdfnx)
|
|
||||||
image="${REGISTRY}/forust/xdfnx-homepage"
|
|
||||||
;;
|
|
||||||
esac
|
|
||||||
set_tags
|
|
||||||
build_args=()
|
|
||||||
for tag in "${tags[@]}"; do
|
|
||||||
build_args+=(-t "${image}:${tag}")
|
|
||||||
done
|
|
||||||
docker build \
|
|
||||||
--cache-from "type=registry,ref=${image}:buildcache" \
|
|
||||||
--cache-to "type=registry,ref=${image}:buildcache,mode=max" \
|
|
||||||
"${build_args[@]}" -f "homepages/Dockerfile.${variant}" homepages
|
|
||||||
for tag in "${tags[@]}"; do
|
|
||||||
docker push "${image}:${tag}"
|
|
||||||
done
|
|
||||||
done
|
|
||||||
;;
|
|
||||||
edu_master)
|
|
||||||
for variant in session-keeper webinar-checker; do
|
|
||||||
case "$variant" in
|
|
||||||
session-keeper)
|
|
||||||
context="edu_master/phpsessid-bot"
|
|
||||||
image="${REGISTRY}/forust/session-keeper"
|
|
||||||
;;
|
|
||||||
webinar-checker)
|
|
||||||
context="edu_master/webinar-checker"
|
|
||||||
image="${REGISTRY}/forust/webinar-checker"
|
|
||||||
;;
|
|
||||||
esac
|
|
||||||
set_tags
|
|
||||||
build_args=()
|
|
||||||
for tag in "${tags[@]}"; do
|
|
||||||
build_args+=(-t "${image}:${tag}")
|
|
||||||
done
|
|
||||||
docker build \
|
|
||||||
--cache-from "type=registry,ref=${image}:buildcache" \
|
|
||||||
--cache-to "type=registry,ref=${image}:buildcache,mode=max" \
|
|
||||||
"${build_args[@]}" "$context"
|
|
||||||
for tag in "${tags[@]}"; do
|
|
||||||
docker push "${image}:${tag}"
|
|
||||||
done
|
|
||||||
done
|
|
||||||
;;
|
|
||||||
esac
|
|
||||||
done
|
|
||||||
|
|
||||||
# Every image the tree names has to carry the commit-pinned name, not only
|
|
||||||
# the ones this push rebuilt. A push that touches nothing but manifests
|
|
||||||
# builds nothing, and its deploy would then find no commit-pinned tag to
|
|
||||||
# resolve and quietly fall back to the moving :prod - which is the whole
|
|
||||||
# failure the commit-pinned name exists to remove.
|
|
||||||
#
|
|
||||||
# Re-tagging copies the manifest list and transfers no layers, so pinning
|
|
||||||
# six images that already exist costs six registry writes.
|
|
||||||
#
|
|
||||||
# The list is derived from the tree rather than written out here, so an
|
|
||||||
# image added to a manifest is covered without a second place to update.
|
|
||||||
- name: Pin the commit name on the images this push did not rebuild
|
|
||||||
if: github.ref_name == 'main'
|
|
||||||
shell: bash
|
|
||||||
run: |
|
|
||||||
set -euo pipefail
|
|
||||||
commit_tag="sha-${GITHUB_SHA:0:12}"
|
|
||||||
mapfile -t repos < <(
|
|
||||||
git grep -hoE 'gcr\.forust\.xyz/forust/[A-Za-z0-9._-]+' -- '*.yaml' '*.yml' \
|
|
||||||
| sort -u
|
|
||||||
)
|
|
||||||
if [ "${#repos[@]}" -eq 0 ]; then
|
|
||||||
echo "No own images referenced by the tree."
|
|
||||||
exit 0
|
|
||||||
fi
|
|
||||||
echo "pinning ${#repos[@]} image(s) to $commit_tag"
|
|
||||||
for repo in "${repos[@]}"; do
|
|
||||||
if docker buildx imagetools inspect "$repo:$commit_tag" >/dev/null 2>&1; then
|
|
||||||
echo " already built by this push: ${repo##*/}"
|
|
||||||
continue
|
|
||||||
fi
|
|
||||||
if ! docker buildx imagetools inspect "$repo:prod" >/dev/null 2>&1; then
|
|
||||||
echo " WARNING: ${repo##*/} has no :prod to pin and no build produced it"
|
|
||||||
continue
|
|
||||||
fi
|
|
||||||
docker buildx imagetools create --tag "$repo:$commit_tag" "$repo:prod"
|
|
||||||
echo " pinned ${repo##*/}"
|
|
||||||
done
|
|
||||||
@@ -0,0 +1,103 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
"""Resolve Compose images without changing project names or local bind paths."""
|
||||||
|
|
||||||
|
import json
|
||||||
|
import os
|
||||||
|
import re
|
||||||
|
import subprocess
|
||||||
|
import sys
|
||||||
|
from pathlib import Path
|
||||||
|
|
||||||
|
|
||||||
|
def output(*args, **kwargs):
|
||||||
|
return subprocess.check_output(args, text=True, **kwargs).strip() # noqa: S603
|
||||||
|
|
||||||
|
|
||||||
|
def resolve(reference):
|
||||||
|
if '@sha256:' in reference:
|
||||||
|
return reference
|
||||||
|
descriptor = json.loads(
|
||||||
|
output('docker', 'buildx', 'imagetools', 'inspect', reference, '--format', '{{json .Manifest}}')
|
||||||
|
)
|
||||||
|
digest = descriptor['digest']
|
||||||
|
if not re.fullmatch(r'sha256:[0-9a-f]{64}', digest):
|
||||||
|
raise ValueError(f'Invalid registry digest for {reference}')
|
||||||
|
# Strip tag only from the final path segment (registry ports are preserved).
|
||||||
|
repository = reference.rsplit('/', 1)
|
||||||
|
repository[-1] = repository[-1].split(':')[0]
|
||||||
|
return '/'.join(repository) + '@' + digest
|
||||||
|
|
||||||
|
|
||||||
|
def prepare(source_file):
|
||||||
|
config_repo = Path(os.environ['CONFIG_REPO'])
|
||||||
|
source_repo = Path(os.environ['REPO'])
|
||||||
|
directory = Path(os.environ['RUN_DIR'])
|
||||||
|
relative = source_file.relative_to(source_repo)
|
||||||
|
project_dir = config_repo / relative.parent
|
||||||
|
base = ['docker', 'compose', '--project-directory', str(project_dir), '-f', str(source_file)]
|
||||||
|
config = json.loads(output(*base, 'config', '--format', 'json', cwd=config_repo))
|
||||||
|
project = config['name']
|
||||||
|
previous_file = directory / 'previous.json'
|
||||||
|
previous = json.loads(previous_file.read_text()) if previous_file.exists() else {}
|
||||||
|
images_file = directory / 'compose-images.json'
|
||||||
|
locks = json.loads(images_file.read_text()) if images_file.exists() else previous.get('compose-images', {})
|
||||||
|
release = json.loads((directory / 'release.json').read_text())
|
||||||
|
before = json.loads(json.dumps(config))
|
||||||
|
for service, settings in config['services'].items():
|
||||||
|
reference = settings.get('image')
|
||||||
|
nextcloud_aio_master = project == 'nextcloud' and service == 'nextcloud-aio-mastercontainer'
|
||||||
|
if not reference or settings.get('build'):
|
||||||
|
raise ValueError(f'{project}/{service}: Compose deploy requires a published image')
|
||||||
|
image_repo = reference.split('@')[0].rsplit('/', 1)
|
||||||
|
image_repo[-1] = image_repo[-1].split(':')[0]
|
||||||
|
image_repo = '/'.join(image_repo)
|
||||||
|
# Nextcloud AIO validates the mastercontainer image reference and rejects
|
||||||
|
# a digest. Keep its configured tag so AIO can start and manage its stack.
|
||||||
|
if nextcloud_aio_master:
|
||||||
|
pinned = reference
|
||||||
|
elif image_repo in release['images']:
|
||||||
|
pinned = image_repo + '@' + release['images'][image_repo]
|
||||||
|
elif os.environ.get('REFRESH_IMAGES') != 'true' and reference in locks:
|
||||||
|
pinned = locks[reference]
|
||||||
|
else:
|
||||||
|
pinned = resolve(reference)
|
||||||
|
settings['image'] = pinned
|
||||||
|
locks[reference] = pinned
|
||||||
|
# Capture what is running, not the current value of its mutable tag.
|
||||||
|
ids = output(
|
||||||
|
'docker',
|
||||||
|
'ps',
|
||||||
|
'-aq',
|
||||||
|
'--filter',
|
||||||
|
f'label=com.docker.compose.project={project}',
|
||||||
|
'--filter',
|
||||||
|
f'label=com.docker.compose.service={service}',
|
||||||
|
).splitlines()
|
||||||
|
actual = set()
|
||||||
|
for container in ids:
|
||||||
|
image_id = output('docker', 'inspect', container, '--format', '{{.Image}}')
|
||||||
|
digests = json.loads(output('docker', 'image', 'inspect', image_id, '--format', '{{json .RepoDigests}}'))
|
||||||
|
actual.add(next((d for d in digests or [] if d.split('@')[0] == image_repo), image_id))
|
||||||
|
if len(actual) > 1:
|
||||||
|
raise ValueError(f'{project}/{service}: mixed running images, cannot capture one recovery config')
|
||||||
|
# AIO also rejects a digest in its recovery config. Preserve its tag in
|
||||||
|
# both deploy and recovery files.
|
||||||
|
if nextcloud_aio_master:
|
||||||
|
before['services'][service]['image'] = reference
|
||||||
|
else:
|
||||||
|
before['services'][service]['image'] = next(iter(actual)) if actual else reference
|
||||||
|
for name, data in (('compose', config), ('compose-before', before)):
|
||||||
|
folder = directory / name
|
||||||
|
folder.mkdir(mode=0o700, exist_ok=True)
|
||||||
|
destination = folder / f'{relative.parent.name}.json'
|
||||||
|
destination.write_text(json.dumps(data, indent=2) + '\n')
|
||||||
|
destination.chmod(0o600)
|
||||||
|
images_file.write_text(json.dumps(locks, indent=2) + '\n')
|
||||||
|
print(f'Compose {project}: images pinned; local paths preserved')
|
||||||
|
print(
|
||||||
|
f'Recovery: docker compose --project-directory {project_dir} -p {project} -f {directory}/compose-before/{relative.parent.name}.json up -d --pull never'
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == '__main__':
|
||||||
|
prepare(Path(sys.argv[1]))
|
||||||
@@ -0,0 +1,423 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
"""Durable workstation deployment controller. Install with setup-workstation.sh."""
|
||||||
|
|
||||||
|
import argparse
|
||||||
|
import contextlib
|
||||||
|
import fcntl
|
||||||
|
import importlib.util
|
||||||
|
import json
|
||||||
|
import math
|
||||||
|
import os
|
||||||
|
import re
|
||||||
|
import shutil
|
||||||
|
import subprocess
|
||||||
|
import sys
|
||||||
|
import time
|
||||||
|
from pathlib import Path
|
||||||
|
|
||||||
|
STATE = Path(os.environ.get('HOMELAB_STATE', Path.home() / '.local/state/homelab-deploy'))
|
||||||
|
CONFIG_REPO = Path(os.environ.get('HOMELAB_REPO', '/srv/homelab'))
|
||||||
|
RUN_ID = re.compile(r'[0-9]+-[0-9]+')
|
||||||
|
|
||||||
|
|
||||||
|
def command(*args, **kwargs):
|
||||||
|
return subprocess.check_output(args, text=True, **kwargs).strip() # noqa: S603, S607
|
||||||
|
|
||||||
|
|
||||||
|
def atomic_json(path, data):
|
||||||
|
temporary = path.with_suffix('.tmp')
|
||||||
|
temporary.write_text(json.dumps(data, indent=2) + '\n')
|
||||||
|
temporary.chmod(0o600)
|
||||||
|
temporary.replace(path)
|
||||||
|
|
||||||
|
|
||||||
|
@contextlib.contextmanager
|
||||||
|
def lock(name):
|
||||||
|
STATE.mkdir(mode=0o700, parents=True, exist_ok=True)
|
||||||
|
with (STATE / name).open('a') as stream:
|
||||||
|
fcntl.flock(stream, fcntl.LOCK_EX)
|
||||||
|
yield
|
||||||
|
|
||||||
|
|
||||||
|
def load_module(name, path):
|
||||||
|
spec = importlib.util.spec_from_file_location(name, path)
|
||||||
|
module = importlib.util.module_from_spec(spec)
|
||||||
|
spec.loader.exec_module(module)
|
||||||
|
return module
|
||||||
|
|
||||||
|
|
||||||
|
def run_directory(run_id):
|
||||||
|
if not RUN_ID.fullmatch(run_id):
|
||||||
|
raise ValueError('Run ID must be numeric workflow-id and attempt')
|
||||||
|
return STATE / 'runs' / run_id
|
||||||
|
|
||||||
|
|
||||||
|
def start(run_id):
|
||||||
|
payload = sys.stdin.buffer.read(256 * 1024 + 1)
|
||||||
|
if len(payload) > 256 * 1024:
|
||||||
|
raise ValueError('Deploy request exceeds 256 KiB')
|
||||||
|
request = json.loads(payload)
|
||||||
|
sha = request['release']['sha']
|
||||||
|
if not re.fullmatch(r'[0-9a-f]{40}', sha) or request['mode'] not in ('changed', 'full', 'plan'):
|
||||||
|
raise ValueError('Invalid deploy SHA or mode')
|
||||||
|
if not isinstance(request['refresh_images'], bool):
|
||||||
|
raise ValueError('refresh_images must be boolean')
|
||||||
|
directory = run_directory(run_id)
|
||||||
|
with lock('prepare.lock'):
|
||||||
|
if (directory / 'request.json').exists():
|
||||||
|
if json.loads((directory / 'request.json').read_text()) != request:
|
||||||
|
raise ValueError('Run ID already belongs to a different request')
|
||||||
|
else:
|
||||||
|
directory.mkdir(mode=0o700, parents=True, exist_ok=True)
|
||||||
|
command('git', '-C', str(CONFIG_REPO), 'fetch', '--quiet', 'origin', 'main')
|
||||||
|
command('git', '-C', str(CONFIG_REPO), 'merge-base', '--is-ancestor', sha, 'origin/main')
|
||||||
|
if not (directory / 'source').exists():
|
||||||
|
command('git', '-C', str(CONFIG_REPO), 'worktree', 'add', '--detach', str(directory / 'source'), sha)
|
||||||
|
if command('git', '-C', str(directory / 'source'), 'rev-parse', 'HEAD') != sha:
|
||||||
|
raise ValueError('Prepared source does not match deploy SHA')
|
||||||
|
release_module = load_module('release', directory / 'source/.gitea/workflows/release.py')
|
||||||
|
release_module.validate_release(request['release'], sha)
|
||||||
|
atomic_json(directory / 'release.json', request['release'])
|
||||||
|
atomic_json(directory / 'request.json', request)
|
||||||
|
if not (directory / 'status.json').exists():
|
||||||
|
atomic_json(directory / 'status.json', {'state': 'queued', 'stages': {}})
|
||||||
|
# Starting an existing active or finished ID is idempotent; never re-apply it.
|
||||||
|
if json.loads((directory / 'status.json').read_text())['state'] == 'queued':
|
||||||
|
command('systemctl', '--user', 'start', '--no-block', f'homelab-deploy@{run_id}.service')
|
||||||
|
print(f'Accepted deploy {run_id} ({sha})')
|
||||||
|
|
||||||
|
|
||||||
|
def environment(directory):
|
||||||
|
request = json.loads((directory / 'request.json').read_text())
|
||||||
|
return {
|
||||||
|
**os.environ,
|
||||||
|
'REPO': str(directory / 'source'),
|
||||||
|
'CONFIG_REPO': str(CONFIG_REPO),
|
||||||
|
'RUN_DIR': str(directory),
|
||||||
|
'DEPLOY_SHA': request['release']['sha'],
|
||||||
|
'RELEASE_FILE': str(directory / 'release.json'),
|
||||||
|
'DEPLOY_PLAN': str(directory / 'plan.json'),
|
||||||
|
'DEPLOY_SNAPSHOT_DIR': str(directory / 'snapshot'),
|
||||||
|
'REFRESH_IMAGES': str(request['refresh_images']).lower(),
|
||||||
|
'ROLLOUT_PARALLELISM': '4',
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def stage(directory, name, budget):
|
||||||
|
status = json.loads((directory / 'status.json').read_text())
|
||||||
|
if name in status['stages'] and status['stages'][name].get('result') in ('success', 'failure'):
|
||||||
|
return status['stages'][name]['result'] == 'success'
|
||||||
|
started = time.time()
|
||||||
|
status['stages'][name] = {'result': 'running', 'started': started}
|
||||||
|
atomic_json(directory / 'status.json', status)
|
||||||
|
script = directory / 'source/.gitea/workflows/deploy-stage.sh'
|
||||||
|
with (directory / f'{name}.log').open('a') as log:
|
||||||
|
# timeout kills the whole stage process group, including children, before recovery.
|
||||||
|
result = subprocess.run( # noqa: S603, S607
|
||||||
|
[
|
||||||
|
shutil.which('timeout') or '/usr/bin/timeout',
|
||||||
|
'--signal=TERM',
|
||||||
|
'--kill-after=30s',
|
||||||
|
str(budget),
|
||||||
|
'bash',
|
||||||
|
str(script),
|
||||||
|
name,
|
||||||
|
],
|
||||||
|
env=environment(directory),
|
||||||
|
stdout=log,
|
||||||
|
stderr=subprocess.STDOUT,
|
||||||
|
check=False,
|
||||||
|
).returncode
|
||||||
|
status = json.loads((directory / 'status.json').read_text())
|
||||||
|
status['stages'][name].update(
|
||||||
|
result='success' if result == 0 else 'failure', exit_code=result, seconds=round(time.time() - started)
|
||||||
|
)
|
||||||
|
atomic_json(directory / 'status.json', status)
|
||||||
|
return result == 0
|
||||||
|
|
||||||
|
|
||||||
|
def make_plan(directory):
|
||||||
|
source = directory / 'source'
|
||||||
|
planner = load_module('deploy_plan', source / '.gitea/workflows/deploy-plan.py')
|
||||||
|
request = json.loads((directory / 'request.json').read_text())
|
||||||
|
previous = json.loads((STATE / 'last-success.json').read_text()) if (STATE / 'last-success.json').exists() else None
|
||||||
|
# Helm 4 lists every release status by default and removed the --all flag.
|
||||||
|
helm = json.loads(command('helm', 'list', '-A', '-o', 'json'))
|
||||||
|
plan = planner.make_plan(source, CONFIG_REPO, request['release'], previous, request['mode'], helm)
|
||||||
|
if request['refresh_images']:
|
||||||
|
plan['selected']['compose'] = plan['active']['compose']
|
||||||
|
atomic_json(directory / 'plan.json', plan)
|
||||||
|
if previous:
|
||||||
|
atomic_json(directory / 'previous.json', previous)
|
||||||
|
# Local config is deliberately separate from the immutable Git source.
|
||||||
|
return plan
|
||||||
|
|
||||||
|
|
||||||
|
def finish_success(directory, plan):
|
||||||
|
# Repeating finalization after a crash is safe while holding deploy.lock.
|
||||||
|
plan['run_id'] = directory.name
|
||||||
|
path = directory / 'compose-images.json'
|
||||||
|
previous = directory / 'previous.json'
|
||||||
|
plan['compose-images'] = (
|
||||||
|
json.loads(path.read_text())
|
||||||
|
if path.exists()
|
||||||
|
else json.loads(previous.read_text()).get('compose-images', {})
|
||||||
|
if previous.exists()
|
||||||
|
else {}
|
||||||
|
)
|
||||||
|
atomic_json(STATE / 'last-success.json', plan)
|
||||||
|
status = json.loads((directory / 'status.json').read_text())
|
||||||
|
status['state'] = 'success'
|
||||||
|
atomic_json(directory / 'status.json', status)
|
||||||
|
try:
|
||||||
|
retain_completed(directory)
|
||||||
|
except (OSError, subprocess.CalledProcessError) as error:
|
||||||
|
print(f'Retention deferred: {error}', flush=True)
|
||||||
|
|
||||||
|
|
||||||
|
def recover(directory, retry=False):
|
||||||
|
status = json.loads((directory / 'status.json').read_text())
|
||||||
|
if status['state'] in ('success', 'planned'):
|
||||||
|
return
|
||||||
|
completed = ('doctor', 'validate', 'apply-k8s', 'apply-compose', 'verify-k8s', 'smoke')
|
||||||
|
if all(status['stages'].get(name, {}).get('result') == 'success' for name in completed):
|
||||||
|
finish_success(directory, json.loads((directory / 'plan.json').read_text()))
|
||||||
|
return
|
||||||
|
if retry:
|
||||||
|
for name in ('verify-k8s', 'smoke'):
|
||||||
|
if status['stages'].get(name, {}).get('result') == 'failure':
|
||||||
|
del status['stages'][name]
|
||||||
|
atomic_json(directory / 'status.json', status)
|
||||||
|
snapshot = directory / 'snapshot/current'
|
||||||
|
if snapshot.exists():
|
||||||
|
stage(directory, 'verify-k8s', 7200)
|
||||||
|
stage(directory, 'smoke', 600)
|
||||||
|
status = json.loads((directory / 'status.json').read_text())
|
||||||
|
status['state'] = 'failure'
|
||||||
|
atomic_json(directory / 'status.json', status)
|
||||||
|
|
||||||
|
|
||||||
|
def execute(run_id):
|
||||||
|
directory = run_directory(run_id)
|
||||||
|
with lock('deploy.lock'):
|
||||||
|
status = json.loads((directory / 'status.json').read_text())
|
||||||
|
if status['state'] != 'queued':
|
||||||
|
return
|
||||||
|
# A crashed predecessor must be recovered before another apply begins.
|
||||||
|
for other in (STATE / 'runs').iterdir():
|
||||||
|
if (
|
||||||
|
other != directory
|
||||||
|
and (other / 'status.json').exists()
|
||||||
|
and json.loads((other / 'status.json').read_text())['state'] == 'running'
|
||||||
|
):
|
||||||
|
raise ValueError(f'Interrupted deploy {other.name}; run recover first')
|
||||||
|
status['state'] = 'running'
|
||||||
|
atomic_json(directory / 'status.json', status)
|
||||||
|
phase = 'plan'
|
||||||
|
try:
|
||||||
|
plan = make_plan(directory)
|
||||||
|
print(
|
||||||
|
json.dumps({'selected': plan['selected'], 'helm': plan['helm'], 'manual_removals': plan['removed']}),
|
||||||
|
flush=True,
|
||||||
|
)
|
||||||
|
phase = 'doctor'
|
||||||
|
if not stage(directory, 'doctor', 600):
|
||||||
|
raise RuntimeError('Preflight failed')
|
||||||
|
phase = 'validate'
|
||||||
|
if not stage(directory, 'validate', 1200):
|
||||||
|
raise RuntimeError('Validation failed')
|
||||||
|
if json.loads((directory / 'request.json').read_text())['mode'] == 'plan':
|
||||||
|
status = json.loads((directory / 'status.json').read_text())
|
||||||
|
status['state'] = 'planned'
|
||||||
|
atomic_json(directory / 'status.json', status)
|
||||||
|
return
|
||||||
|
# Budget includes both rollout checks and rollback waves, plus API overhead.
|
||||||
|
phase = 'Recovery budget'
|
||||||
|
count = int(
|
||||||
|
command(
|
||||||
|
'bash',
|
||||||
|
str(directory / 'source/.gitea/workflows/deploy-stage.sh'),
|
||||||
|
'workload-count',
|
||||||
|
env=environment(directory),
|
||||||
|
)
|
||||||
|
)
|
||||||
|
verify_budget = max(600, 2 * math.ceil(count / 4) * 300 + 120)
|
||||||
|
if verify_budget > 7200:
|
||||||
|
raise ValueError('More than two hours of recovery required; split this deploy')
|
||||||
|
phase = 'apply-k8s'
|
||||||
|
k8s_ok = stage(directory, 'apply-k8s', 2700)
|
||||||
|
phase = 'apply-compose'
|
||||||
|
compose_ok = stage(directory, 'apply-compose', 1800) if k8s_ok else False
|
||||||
|
phase = 'verify-k8s'
|
||||||
|
verify_ok = stage(directory, 'verify-k8s', verify_budget)
|
||||||
|
phase = 'smoke'
|
||||||
|
smoke_ok = stage(directory, 'smoke', 600)
|
||||||
|
if not all((k8s_ok, compose_ok, verify_ok, smoke_ok)):
|
||||||
|
raise RuntimeError('Deploy failed; inspect stage logs and recovery report')
|
||||||
|
phase = 'Save the successful baseline'
|
||||||
|
finish_success(directory, plan)
|
||||||
|
except Exception as error:
|
||||||
|
status = json.loads((directory / 'status.json').read_text())
|
||||||
|
status['failure_stage'] = next(
|
||||||
|
(name for name, result in status['stages'].items() if result.get('result') == 'failure'), phase
|
||||||
|
)
|
||||||
|
atomic_json(directory / 'status.json', status)
|
||||||
|
with (directory / 'controller.log').open('a') as stream:
|
||||||
|
stream.write(f'{error}\n')
|
||||||
|
recover(directory)
|
||||||
|
raise
|
||||||
|
|
||||||
|
|
||||||
|
def retain_completed(current):
|
||||||
|
finished = []
|
||||||
|
for directory in (STATE / 'runs').iterdir():
|
||||||
|
status_file = directory / 'status.json'
|
||||||
|
if status_file.exists() and json.loads(status_file.read_text())['state'] in ('success', 'planned'):
|
||||||
|
finished.append(directory)
|
||||||
|
for directory in sorted(finished, key=lambda p: p.stat().st_mtime, reverse=True)[20:]:
|
||||||
|
if directory == current:
|
||||||
|
continue
|
||||||
|
command('git', '-C', str(CONFIG_REPO), 'worktree', 'remove', '--force', str(directory / 'source'))
|
||||||
|
shutil.rmtree(directory)
|
||||||
|
|
||||||
|
|
||||||
|
def follow(run_id, phase):
|
||||||
|
directory = run_directory(run_id)
|
||||||
|
groups = {
|
||||||
|
'apply': ('doctor', 'validate', 'apply-k8s', 'apply-compose'),
|
||||||
|
'verify': ('verify-k8s',),
|
||||||
|
'smoke': ('smoke',),
|
||||||
|
}
|
||||||
|
names = groups[phase]
|
||||||
|
offsets = {}
|
||||||
|
while True:
|
||||||
|
status = json.loads((directory / 'status.json').read_text())
|
||||||
|
for name in (*names, 'controller'):
|
||||||
|
path = directory / f'{name}.log'
|
||||||
|
if path.exists():
|
||||||
|
with path.open() as stream:
|
||||||
|
stream.seek(offsets.get(name, 0))
|
||||||
|
content = stream.read()
|
||||||
|
if content:
|
||||||
|
print(content, end='', flush=True)
|
||||||
|
offsets[name] = stream.tell()
|
||||||
|
stages = status['stages']
|
||||||
|
if all(stages.get(name, {}).get('result') in ('success', 'failure') for name in names):
|
||||||
|
return all(stages[name]['result'] == 'success' for name in names)
|
||||||
|
if status['state'] in ('success', 'failure', 'planned'):
|
||||||
|
return status['state'] in ('success', 'planned')
|
||||||
|
time.sleep(3)
|
||||||
|
|
||||||
|
|
||||||
|
def summary(run_id):
|
||||||
|
directory = run_directory(run_id)
|
||||||
|
request = json.loads((directory / 'request.json').read_text())
|
||||||
|
release = request['release']
|
||||||
|
plan_file = directory / 'plan.json'
|
||||||
|
lines = [
|
||||||
|
f'## Deploy `{release["sha"]}`',
|
||||||
|
'',
|
||||||
|
f'- Mode: `{request["mode"]}`',
|
||||||
|
f'- Refresh third-party images: `{request["refresh_images"]}`',
|
||||||
|
]
|
||||||
|
status = json.loads((directory / 'status.json').read_text())
|
||||||
|
if status.get('failure_stage'):
|
||||||
|
lines.append(f'- Failed stage: **{status["failure_stage"]}**')
|
||||||
|
lines.extend(
|
||||||
|
[
|
||||||
|
'',
|
||||||
|
f'- Observed run state: **{status["state"]}**',
|
||||||
|
'',
|
||||||
|
'### Stage results',
|
||||||
|
'| Stage | Result | Exit code |',
|
||||||
|
'| --- | --- | --- |',
|
||||||
|
]
|
||||||
|
)
|
||||||
|
for name in ('doctor', 'validate', 'apply-k8s', 'apply-compose', 'verify-k8s', 'smoke'):
|
||||||
|
stage_result = status['stages'].get(name, {})
|
||||||
|
lines.append(f'| {name} | {stage_result.get("result", "not started")} | {stage_result.get("exit_code", "—")} |')
|
||||||
|
lines.extend(['', '### Apply and Helm recovery results'])
|
||||||
|
events_file = directory / 'apply-events.jsonl'
|
||||||
|
events = []
|
||||||
|
if events_file.exists():
|
||||||
|
for line in events_file.read_text().splitlines():
|
||||||
|
try:
|
||||||
|
events.append(json.loads(line))
|
||||||
|
except json.JSONDecodeError:
|
||||||
|
lines.append('- An operation record is incomplete. Check the stage log.')
|
||||||
|
latest = {(event['action'], event['target']): event['result'] for event in events}
|
||||||
|
lines.extend(f'- `{action}` `{target}`: **{result}**' for (action, target), result in latest.items())
|
||||||
|
if not latest:
|
||||||
|
lines.append('- No apply results were recorded.')
|
||||||
|
lines.append('- A completed apply does not confirm health. See verification and smoke results.')
|
||||||
|
lines.extend(['', '### Kubernetes recovery'])
|
||||||
|
pointer = directory / 'snapshot/current'
|
||||||
|
failed = Path(pointer.read_text().strip()) / 'failed-workloads' if pointer.exists() else None
|
||||||
|
if failed and failed.exists():
|
||||||
|
contents = failed.read_text()
|
||||||
|
counts = dict(re.findall(r'^(ROLLED_BACK|UNRECOVERED)=([0-9]+)$', contents, re.MULTILINE))
|
||||||
|
if not contents.strip():
|
||||||
|
lines.append('- No failed workloads were recorded. See the verification result above.')
|
||||||
|
elif counts:
|
||||||
|
lines.append(f'- Workloads restored: **{counts.get("ROLLED_BACK", "unknown")}**')
|
||||||
|
lines.append(f'- Workloads that need manual recovery: **{counts.get("UNRECOVERED", "unknown")}**')
|
||||||
|
else:
|
||||||
|
lines.append('- Rollback has no recorded result yet. Check the verification log.')
|
||||||
|
else:
|
||||||
|
lines.append('- No workload rollback was recorded. This does not confirm health.')
|
||||||
|
lines.append('- Compose requires manual recovery. Use the saved command in the apply log.')
|
||||||
|
if not plan_file.exists():
|
||||||
|
lines.extend(['', 'Plan was not created. Check the controller log.'])
|
||||||
|
print('\n'.join(lines))
|
||||||
|
return
|
||||||
|
plan = json.loads(plan_file.read_text())
|
||||||
|
lines.extend(['', '### Selected services'])
|
||||||
|
count = 0
|
||||||
|
for kind, services in plan['selected'].items():
|
||||||
|
for service in services:
|
||||||
|
lines.append(f'- `{kind}`: `{service}`')
|
||||||
|
count += 1
|
||||||
|
if not count:
|
||||||
|
lines.append('- None')
|
||||||
|
lines.extend(['', '### Selected Helm releases'])
|
||||||
|
lines.extend(f'- `{release}`' for release in plan.get('helm', []))
|
||||||
|
if not plan.get('helm'):
|
||||||
|
lines.append('- None')
|
||||||
|
lines.extend(['', '### Images pinned in the checked release'])
|
||||||
|
lines.extend(f'- `{image}@{digest}`' for image, digest in sorted(release['images'].items()))
|
||||||
|
lines.extend(['', '### Removed resources requiring manual review'])
|
||||||
|
lines.extend(f'- `{item}`' for item in plan.get('removed', []))
|
||||||
|
if not plan.get('removed'):
|
||||||
|
lines.append('- None')
|
||||||
|
print('\n'.join(lines))
|
||||||
|
|
||||||
|
|
||||||
|
def main():
|
||||||
|
os.umask(0o077)
|
||||||
|
parser = argparse.ArgumentParser(description=__doc__)
|
||||||
|
parser.add_argument('action', choices=('start', 'execute', 'recover', 'status', 'follow', 'summary'))
|
||||||
|
parser.add_argument('run_id')
|
||||||
|
parser.add_argument('phase', nargs='?', choices=('apply', 'verify', 'smoke'))
|
||||||
|
parser.add_argument('--retry', action='store_true', help='Retry failed recovery checks; never repeat apply')
|
||||||
|
args = parser.parse_args()
|
||||||
|
directory = run_directory(args.run_id)
|
||||||
|
if args.action == 'start':
|
||||||
|
start(args.run_id)
|
||||||
|
elif args.action == 'execute':
|
||||||
|
execute(args.run_id)
|
||||||
|
elif args.action == 'recover':
|
||||||
|
with lock('deploy.lock'):
|
||||||
|
recover(directory, retry=args.retry)
|
||||||
|
elif args.action == 'status':
|
||||||
|
print((directory / 'status.json').read_text())
|
||||||
|
if (directory / 'plan.json').exists():
|
||||||
|
plan = json.loads((directory / 'plan.json').read_text())
|
||||||
|
print(json.dumps({k: plan[k] for k in ('sha', 'selected', 'helm', 'removed')}, indent=2))
|
||||||
|
elif args.action == 'summary':
|
||||||
|
summary(args.run_id)
|
||||||
|
elif not follow(args.run_id, args.phase):
|
||||||
|
sys.exit(1)
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == '__main__':
|
||||||
|
main()
|
||||||
+286
-409
@@ -1,5 +1,5 @@
|
|||||||
#!/usr/bin/env bash
|
#!/usr/bin/env bash
|
||||||
# Shared stages for the deploy workflow. Runs on the workstation, invoked as:
|
# Workstation deploy stages; invoked by the durable controller against pinned source.
|
||||||
# REPO=/srv/homelab APPLY_PRUNE=false bash -se <<'EOF'
|
# REPO=/srv/homelab APPLY_PRUNE=false bash -se <<'EOF'
|
||||||
# source "$REPO/.gitea/workflows/deploy-lib.sh"
|
# source "$REPO/.gitea/workflows/deploy-lib.sh"
|
||||||
# run_stage "$STAGE"
|
# run_stage "$STAGE"
|
||||||
@@ -8,8 +8,8 @@ set -euo pipefail
|
|||||||
|
|
||||||
: "${REPO:?REPO must be set}"
|
: "${REPO:?REPO must be set}"
|
||||||
APPLY_PRUNE="${APPLY_PRUNE:-false}"
|
APPLY_PRUNE="${APPLY_PRUNE:-false}"
|
||||||
# Commit CI validated. Empty for a manual workflow_dispatch, which falls back to
|
CONFIG_REPO="${CONFIG_REPO:-$REPO}"
|
||||||
# the current origin/main.
|
# Exact SHA accepted by the CI gate for both manual and automatic deploys.
|
||||||
DEPLOY_SHA="${DEPLOY_SHA:-}"
|
DEPLOY_SHA="${DEPLOY_SHA:-}"
|
||||||
# Handoff point between the apply stage (writes) and the verify stage (reads).
|
# Handoff point between the apply stage (writes) and the verify stage (reads).
|
||||||
# Under the deploy user's own XDG state directory rather than /var/backups: the
|
# Under the deploy user's own XDG state directory rather than /var/backups: the
|
||||||
@@ -20,17 +20,36 @@ DEPLOY_SNAPSHOT_DIR="${DEPLOY_SNAPSHOT_DIR:-${XDG_STATE_HOME:-$HOME/.local/state
|
|||||||
# Per-workload rollout budget and how many workloads to watch at once. The whole
|
# Per-workload rollout budget and how many workloads to watch at once. The whole
|
||||||
# apply job has its own timeout-minutes as a backstop.
|
# apply job has its own timeout-minutes as a backstop.
|
||||||
ROLLOUT_TIMEOUT="${ROLLOUT_TIMEOUT:-300}"
|
ROLLOUT_TIMEOUT="${ROLLOUT_TIMEOUT:-300}"
|
||||||
ROLLOUT_PARALLELISM="${ROLLOUT_PARALLELISM:-8}"
|
ROLLOUT_PARALLELISM="${ROLLOUT_PARALLELISM:-4}"
|
||||||
WORKLOAD_KINDS="deployments.apps,statefulsets.apps,daemonsets.apps"
|
WORKLOAD_KINDS="deployments.apps,statefulsets.apps,daemonsets.apps"
|
||||||
|
|
||||||
log() {
|
log() {
|
||||||
echo "== $* =="
|
echo "== $* =="
|
||||||
}
|
}
|
||||||
|
|
||||||
|
# Store operation results without command output or local configuration values.
|
||||||
|
record_apply() {
|
||||||
|
[ -n "${RUN_DIR:-}" ] || return 0
|
||||||
|
jq -cn --arg action "$1" --arg target "$2" --arg result "$3" \
|
||||||
|
'{action: $action, target: $target, result: $result}' >>"$RUN_DIR/apply-events.jsonl" \
|
||||||
|
|| echo 'WARNING: cannot record an apply result' >&2
|
||||||
|
return 0
|
||||||
|
}
|
||||||
|
|
||||||
warn() {
|
warn() {
|
||||||
echo "WARNING: $*" >&2
|
echo "WARNING: $*" >&2
|
||||||
}
|
}
|
||||||
|
|
||||||
|
# Prune needs the complete desired set in one invocation. Per-file pruning
|
||||||
|
# treats resources from the other files as absent and can delete them.
|
||||||
|
check_prune_mode() {
|
||||||
|
if [ "$APPLY_PRUNE" = "true" ]; then
|
||||||
|
echo "ERROR: APPLY_PRUNE=true is unsupported by the per-file deploy loop." >&2
|
||||||
|
echo "Disable it; remove obsolete resources explicitly after review." >&2
|
||||||
|
return 1
|
||||||
|
fi
|
||||||
|
}
|
||||||
|
|
||||||
collect_k8s() {
|
collect_k8s() {
|
||||||
git -C "$REPO" ls-files -- "$1" \
|
git -C "$REPO" ls-files -- "$1" \
|
||||||
| grep -E '\.ya?ml$' \
|
| grep -E '\.ya?ml$' \
|
||||||
@@ -50,6 +69,25 @@ kustomize_overlay() {
|
|||||||
fi
|
fi
|
||||||
}
|
}
|
||||||
|
|
||||||
|
selected_service() {
|
||||||
|
local kind="$1" service="$2" section=selected
|
||||||
|
[ -n "${DEPLOY_PLAN:-}" ] || return 0
|
||||||
|
if [ "${DEPLOY_SMOKE_ALL:-false}" = true ]; then section=active; fi
|
||||||
|
jq -e --arg kind "$kind" --arg service "$service" --arg section "$section" \
|
||||||
|
'.[$section][$kind] | index($service) != null' "$DEPLOY_PLAN" >/dev/null
|
||||||
|
}
|
||||||
|
|
||||||
|
# Resolve .env and relative binds on the persistent workstation tree. Locked
|
||||||
|
# JSON configs keep the same Compose project name and volume names.
|
||||||
|
compose() {
|
||||||
|
local cf="$1" locked project_dir
|
||||||
|
project_dir="$CONFIG_REPO/$(basename "$(dirname "$cf")")"
|
||||||
|
shift
|
||||||
|
locked="${RUN_DIR:-/nonexistent}/compose/$(basename "$(dirname "$cf")").json"
|
||||||
|
if [ -f "$locked" ]; then cf="$locked"; fi
|
||||||
|
(cd "$CONFIG_REPO" && docker compose --project-directory "$project_dir" -f "$cf" "$@")
|
||||||
|
}
|
||||||
|
|
||||||
select_manifests() {
|
select_manifests() {
|
||||||
K8S_MANIFESTS=()
|
K8S_MANIFESTS=()
|
||||||
KUSTOMIZE_APPS=()
|
KUSTOMIZE_APPS=()
|
||||||
@@ -57,6 +95,7 @@ select_manifests() {
|
|||||||
local kd_rel kd overlay cf_rel cf f
|
local kd_rel kd overlay cf_rel cf f
|
||||||
while IFS= read -r kd_rel; do
|
while IFS= read -r kd_rel; do
|
||||||
kd="$REPO/$kd_rel"
|
kd="$REPO/$kd_rel"
|
||||||
|
selected_service k8s "${kd_rel%/k8s}" || continue
|
||||||
if [ ! -f "$kd/active" ]; then
|
if [ ! -f "$kd/active" ]; then
|
||||||
echo "skip (no k8s/active): $kd_rel"
|
echo "skip (no k8s/active): $kd_rel"
|
||||||
continue
|
continue
|
||||||
@@ -78,6 +117,7 @@ select_manifests() {
|
|||||||
)
|
)
|
||||||
while IFS= read -r cf_rel; do
|
while IFS= read -r cf_rel; do
|
||||||
cf="$REPO/$cf_rel"
|
cf="$REPO/$cf_rel"
|
||||||
|
selected_service compose "$(dirname "$cf_rel")" || continue
|
||||||
if [ -f "$(dirname "$cf")/active" ]; then
|
if [ -f "$(dirname "$cf")/active" ]; then
|
||||||
echo "compose: $cf_rel"
|
echo "compose: $cf_rel"
|
||||||
COMPOSE_STACKS+=("$cf")
|
COMPOSE_STACKS+=("$cf")
|
||||||
@@ -94,16 +134,10 @@ select_manifests() {
|
|||||||
# on failure roll them back to the revision that was running before, so a bad
|
# on failure roll them back to the revision that was running before, so a bad
|
||||||
# push to main cannot leave a service crash-looping.
|
# push to main cannot leave a service crash-looping.
|
||||||
#
|
#
|
||||||
# Verification lives in its own workflow job, not at the end of the apply stage.
|
# The workstation controller runs apply and verification as separate durable
|
||||||
# Inside a single process it is worthless exactly when it is needed most: a job
|
# stages. Runner jobs only follow their logs. ExecStopPost recovers interrupted
|
||||||
# killed by timeout-minutes or cancelled mid-apply never reaches the rollback
|
# runs using the per-run snapshot, even when the SSH connection has gone away.
|
||||||
# code, and leaves a half-applied cluster behind. Split out, the apply job can
|
|
||||||
# die in any way and the verify job still runs.
|
|
||||||
#
|
#
|
||||||
# That split needs a handoff point on the workstation, because the two stages are
|
|
||||||
# separate processes on separate runner jobs: DEPLOY_SNAPSHOT_DIR/current, written
|
|
||||||
# before anything is applied, read by the verify stage afterwards.
|
|
||||||
|
|
||||||
# Creates this run's snapshot directory and publishes it as the handoff point for
|
# Creates this run's snapshot directory and publishes it as the handoff point for
|
||||||
# the verify stage. Fails hard by design: a deploy that cannot record what it is
|
# the verify stage. Fails hard by design: a deploy that cannot record what it is
|
||||||
# about to change must not start, because then nothing can be rolled back for it
|
# about to change must not start, because then nothing can be rolled back for it
|
||||||
@@ -133,18 +167,36 @@ snapshot_dir() {
|
|||||||
}
|
}
|
||||||
|
|
||||||
save_snapshot() {
|
save_snapshot() {
|
||||||
local dir="$1"
|
local dir="$1" releases revision status
|
||||||
log "Saving pre-apply snapshot to $dir"
|
log "Saving pre-apply snapshot to $dir"
|
||||||
workload_generations >"$dir/generations.before" 2>/dev/null \
|
workload_generations >"$dir/generations.before" || return 1
|
||||||
|| warn "could not snapshot workload generations"
|
kubectl get "$WORKLOAD_KINDS" -A -o json >"$dir/workloads.json" || return 1
|
||||||
kubectl get "$WORKLOAD_KINDS" -A -o yaml >"$dir/workloads.yaml" 2>/dev/null \
|
kubectl get controllerrevisions.apps -A -o json >"$dir/controller-revisions.json" || return 1
|
||||||
|| warn "could not snapshot workloads"
|
jq --slurpfile revisions "$dir/controller-revisions.json" '
|
||||||
for release in prometheus-stack loki alloy; do
|
[.items[] | . as $w | {
|
||||||
if helm status "$release" -n prometheus >/dev/null 2>&1; then
|
kind: (.kind | ascii_downcase), namespace: .metadata.namespace, name: .metadata.name, uid: .metadata.uid,
|
||||||
{
|
revision: (if .kind == "Deployment" then (.metadata.annotations["deployment.kubernetes.io/revision"] // "0" | tonumber)
|
||||||
echo "revision: $(helm history "$release" -n prometheus -o json 2>/dev/null)"
|
else ([$revisions[0].items[] | select(.metadata.namespace == $w.metadata.namespace)
|
||||||
helm get values "$release" -n prometheus --all 2>/dev/null
|
| select(any(.metadata.ownerReferences[]?; .uid == $w.metadata.uid))
|
||||||
} >"$dir/helm-$release.txt"
|
| select($w.kind != "StatefulSet" or .metadata.name == $w.status.currentRevision) | .revision] | max // 0) end)
|
||||||
|
}]' "$dir/workloads.json" >"$dir/revisions.json" || return 1
|
||||||
|
# Helm 4 lists every release status by default and removed the --all flag.
|
||||||
|
releases="$(helm list -A -o json)" || return 1
|
||||||
|
for entry in "${HELM_RELEASES[@]}"; do
|
||||||
|
IFS='|' read -r release _ namespace _ _ _ <<<"$entry"
|
||||||
|
if ! jq -e --arg r "$release" --arg n "$namespace" \
|
||||||
|
'any(.[]; .name == $r and .namespace == $n)' <<<"$releases" >/dev/null; then
|
||||||
|
continue
|
||||||
|
fi
|
||||||
|
helm status "$release" -n "$namespace" -o json >"$dir/helm-$release.json" || return 1
|
||||||
|
status="$(jq -r '.info.status' "$dir/helm-$release.json")"
|
||||||
|
if [ "$status" != deployed ]; then
|
||||||
|
# Never capture a pending/failed revision as the recovery target.
|
||||||
|
helm history "$release" -n "$namespace" -o json >"$dir/helm-$release.history.json" || return 1
|
||||||
|
revision="$(jq '[.[] | select(.status == "deployed" or .status == "superseded") | .revision] | max // 0' \
|
||||||
|
"$dir/helm-$release.history.json")"
|
||||||
|
jq --argjson revision "$revision" '.version = $revision' "$dir/helm-$release.json" >"$dir/helm-$release.tmp"
|
||||||
|
mv "$dir/helm-$release.tmp" "$dir/helm-$release.json"
|
||||||
fi
|
fi
|
||||||
done
|
done
|
||||||
# The verify stage compares this against the commit it is deploying, to refuse
|
# The verify stage compares this against the commit it is deploying, to refuse
|
||||||
@@ -168,294 +220,24 @@ workload_generations() {
|
|||||||
# moved since the snapshot, i.e. the ones this apply actually touched.
|
# moved since the snapshot, i.e. the ones this apply actually touched.
|
||||||
changed_workloads() {
|
changed_workloads() {
|
||||||
local before="$1"
|
local before="$1"
|
||||||
local ns name kind gen old
|
local ns name kind gen old current
|
||||||
|
current="$(workload_generations)" || return 1
|
||||||
while read -r ns name kind gen; do
|
while read -r ns name kind gen; do
|
||||||
[ -n "${gen:-}" ] || continue
|
[ -n "${gen:-}" ] || continue
|
||||||
old="$(awk -v want_ns="$ns" -v want_name="$name" \
|
old="$(awk -v want_ns="$ns" -v want_name="$name" -v want_kind="$kind" \
|
||||||
'$1 == want_ns && $2 == want_name { print $4; exit }' "$before" 2>/dev/null || true)"
|
'$1 == want_ns && $2 == want_name && $3 == want_kind { print $4; exit }' "$before" 2>/dev/null || true)"
|
||||||
|
if [ -n "${RUN_DIR:-}" ] && ! grep -qxF "$kind $ns $name" "$RUN_DIR/workload-refs"; then
|
||||||
|
continue
|
||||||
|
fi
|
||||||
if [ "$old" != "$gen" ]; then
|
if [ "$old" != "$gen" ]; then
|
||||||
printf '%s %s %s\n' "$kind" "$ns" "$name"
|
printf '%s %s %s\n' "$kind" "$ns" "$name"
|
||||||
fi
|
fi
|
||||||
done < <(workload_generations)
|
done <<<"$current"
|
||||||
}
|
}
|
||||||
|
|
||||||
# Prints "<ns> <kind>/<name> <image>" for every workload this repository owns that
|
# Resolve owned image references exclusively from the checked CI artifact.
|
||||||
# runs an image from our own registry.
|
|
||||||
#
|
|
||||||
# The repository is the scope, deliberately. The cluster also holds workloads on
|
|
||||||
# our registry that no manifest here declares (they are applied out of band), and
|
|
||||||
# those are somebody else's to deploy. Walking the manifests rather than the
|
|
||||||
# cluster means those can never be restarted by this pipeline, now or later.
|
|
||||||
owned_registry_workloads() {
|
|
||||||
local kd_rel f
|
|
||||||
while IFS= read -r kd_rel; do
|
|
||||||
[ -f "$REPO/$kd_rel/active" ] || continue
|
|
||||||
while IFS= read -r f; do
|
|
||||||
[ -n "$f" ] || continue
|
|
||||||
# A file that does not mention the registry cannot declare a workload on it,
|
|
||||||
# and parsing costs ~2.5s per file against a millisecond for the grep. The
|
|
||||||
# filter keeps this at a handful of parses instead of one per manifest.
|
|
||||||
grep -q 'gcr\.forust\.xyz/forust/' "$REPO/$f" 2>/dev/null || continue
|
|
||||||
# kubectl prints a bare object for a single-document file and a List for a
|
|
||||||
# multi-document one, so normalise both shapes before filtering.
|
|
||||||
kubectl apply --dry-run=client -f "$REPO/$f" -o json 2>/dev/null \
|
|
||||||
| jq -r '
|
|
||||||
(if .items then .items[] else . end)
|
|
||||||
| select(.kind | test("^(Deployment|StatefulSet|DaemonSet)$"))
|
|
||||||
| select(any((.spec.template.spec.containers // [])[]?;
|
|
||||||
(.image // "") | test("^gcr\\.forust\\.xyz/forust/")))
|
|
||||||
| (.metadata.namespace // "default") as $ns
|
|
||||||
| ([.spec.template.spec.containers[].image
|
|
||||||
| select(test("^gcr\\.forust\\.xyz/forust/"))][0]) as $img
|
|
||||||
| "\($ns) \(.kind | ascii_downcase)/\(.metadata.name) \($img)"
|
|
||||||
' 2>/dev/null || true
|
|
||||||
done < <(collect_k8s "$kd_rel" || true)
|
|
||||||
done < <(
|
|
||||||
git -C "$REPO" ls-files '*.yaml' '*.yml' \
|
|
||||||
| grep -E '(^|/)k8s/' \
|
|
||||||
| sed -E 's#((^|.*/)k8s)/.*#\1#' \
|
|
||||||
| sort -u
|
|
||||||
)
|
|
||||||
}
|
|
||||||
|
|
||||||
# Prints the digest an image tag resolves to for this cluster's architecture, or
|
|
||||||
# nothing when it cannot be resolved.
|
|
||||||
#
|
|
||||||
# Only the manifest entry matching the node architecture counts. A multi-arch tag
|
|
||||||
# also carries `unknown/unknown` entries for the build attestation, and a pod's
|
|
||||||
# imageID is always the per-platform digest, so comparing the wrong entry would
|
|
||||||
# mark every workload stale forever and restart the whole cluster on every deploy.
|
|
||||||
registry_digest() {
|
|
||||||
local arch
|
|
||||||
arch="$(kubectl get nodes -o jsonpath='{.items[0].status.nodeInfo.architecture}' 2>/dev/null || true)"
|
|
||||||
[ -n "$arch" ] || arch=amd64
|
|
||||||
# The || true is load-bearing. Every caller runs under set -euo pipefail, and
|
|
||||||
# pipefail reports the rightmost non-zero stage, so a ref the registry does not
|
|
||||||
# have would abort the caller at the assignment instead of yielding an empty
|
|
||||||
# string. The callers check for empty themselves and report it by name.
|
|
||||||
#
|
|
||||||
# Retried with a hard timeout because the registry has a known hang mode (and
|
|
||||||
# a known blink mode: a single failed lookup aborts the whole apply file in
|
|
||||||
# render_pinned). A short sleep between attempts lets a restarting registry
|
|
||||||
# come back instead of failing the deploy on one bad second.
|
|
||||||
local attempt=0 digest=""
|
|
||||||
while [ "$attempt" -lt 3 ]; do
|
|
||||||
digest="$(timeout 25s docker manifest inspect "$1" 2>/dev/null \
|
|
||||||
| jq -r --arg arch "$arch" '
|
|
||||||
.manifests[]?
|
|
||||||
| select(.platform.os == "linux" and .platform.architecture == $arch)
|
|
||||||
| .digest
|
|
||||||
' 2>/dev/null \
|
|
||||||
| head -1 || true)"
|
|
||||||
[ -n "$digest" ] && break
|
|
||||||
attempt=$((attempt + 1))
|
|
||||||
if [ "$attempt" -lt 3 ]; then
|
|
||||||
echo "WARNING: registry lookup of $1 failed (attempt $attempt/3), retrying in 5s" >&2
|
|
||||||
sleep 5
|
|
||||||
fi
|
|
||||||
done
|
|
||||||
printf '%s' "$digest"
|
|
||||||
}
|
|
||||||
|
|
||||||
# The commit this deploy is for: what CI validated, or - on a manual dispatch,
|
|
||||||
# whatever stage_preflight just checked out.
|
|
||||||
deploy_commit() {
|
|
||||||
local c="${DEPLOY_SHA:-}"
|
|
||||||
[ -n "$c" ] || c="$(git -C "$REPO" rev-parse HEAD 2>/dev/null || true)"
|
|
||||||
printf '%.12s' "${c:-}"
|
|
||||||
}
|
|
||||||
|
|
||||||
# Resolves one of our image refs to the digest THIS commit's build produced.
|
|
||||||
#
|
|
||||||
# A manifest naming `:prod` names a pointer, not a version, and the deploy
|
|
||||||
# resolves it when the apply runs - which is not when CI ran it. Deploy runs are
|
|
||||||
# queued rather than cancelled (see deploy.yaml), so two pushes in a row leave
|
|
||||||
# the first deploy resolving the second push's build: the right manifests with
|
|
||||||
# the wrong code, and nothing anywhere reports it. ci therefore publishes every
|
|
||||||
# image it ships under `sha-<commit12>`, a name that cannot move, and that is
|
|
||||||
# the name resolved here.
|
|
||||||
#
|
|
||||||
# The fallback to the plain tag is for an image this pipeline never built. It
|
|
||||||
# reports itself, because a fallback nobody sees is the failure this removes.
|
|
||||||
pinned_digest() {
|
|
||||||
local ref="$1" commit pinned
|
|
||||||
commit="$(deploy_commit)"
|
|
||||||
if [ -n "$commit" ]; then
|
|
||||||
pinned="$(registry_digest "${ref%:*}:sha-$commit")"
|
|
||||||
if [ -n "$pinned" ]; then
|
|
||||||
printf '%s' "$pinned"
|
|
||||||
return 0
|
|
||||||
fi
|
|
||||||
fi
|
|
||||||
pinned="$(registry_digest "$ref")"
|
|
||||||
if [ -n "$pinned" ]; then
|
|
||||||
echo "WARNING: ${ref} carries no sha-${commit:-<unknown>} tag; resolved the moving tag instead" >&2
|
|
||||||
fi
|
|
||||||
printf '%s' "$pinned"
|
|
||||||
}
|
|
||||||
|
|
||||||
# Rewrites our own images to immutable digests on the way into the cluster.
|
|
||||||
# Reads a manifest stream on stdin, writes the pinned stream to stdout.
|
|
||||||
#
|
|
||||||
# A digest is not knowable when a manifest is written, so it is never committed:
|
|
||||||
# git keeps a readable `:prod` tag and the exact bytes are chosen here, at apply
|
|
||||||
# time, from the tag ci published for the commit being deployed. That is what
|
|
||||||
# makes rollback mean something. `kubectl rollout undo` restores the previous
|
|
||||||
# ReplicaSet's pod template verbatim, and a template naming a digest restores the
|
|
||||||
# exact bytes that were serving before. A template naming a moving tag does not —
|
|
||||||
# the tag has already moved by the time the rollback runs, so the "rollback"
|
|
||||||
# re-pulls the very image that just failed and the cluster stays broken.
|
|
||||||
#
|
|
||||||
# imagePullPolicy is deliberately left alone. The manifests no longer set it, and a
|
|
||||||
# reference that is not `:latest` defaults to IfNotPresent, which is what the
|
|
||||||
# Kubernetes docs ask for alongside a digest: the bytes under a digest cannot
|
|
||||||
# change, so pulling again buys nothing.
|
|
||||||
#
|
|
||||||
# An image that cannot be resolved is fatal. Carrying on would quietly apply a
|
|
||||||
# mutable tag again, which is the exact failure this function exists to remove.
|
|
||||||
render_pinned() {
|
render_pinned() {
|
||||||
local src refs map ref digest missing=0
|
python3 "$REPO/.gitea/workflows/release.py" render
|
||||||
src="$(mktemp)"
|
|
||||||
refs="$(mktemp)"
|
|
||||||
map="$(mktemp)"
|
|
||||||
|
|
||||||
cat >"$src"
|
|
||||||
grep -oE 'gcr\.forust\.xyz/forust/[A-Za-z0-9._-]+:[A-Za-z0-9._-]+' "$src" | sort -u >"$refs" || true
|
|
||||||
|
|
||||||
while read -r ref; do
|
|
||||||
[ -n "$ref" ] || continue
|
|
||||||
digest="$(pinned_digest "$ref")"
|
|
||||||
if [ -z "$digest" ]; then
|
|
||||||
echo "ERROR: cannot resolve ${ref} in the registry; applying nothing." >&2
|
|
||||||
echo " The build job has to push that tag before the deploy resolves it." >&2
|
|
||||||
missing=$((missing + 1))
|
|
||||||
continue
|
|
||||||
fi
|
|
||||||
printf '%s\t%s\n' "$ref" "$digest" >>"$map"
|
|
||||||
done <"$refs"
|
|
||||||
if [ "$missing" -gt 0 ]; then
|
|
||||||
rm -f "$src" "$refs" "$map"
|
|
||||||
return 1
|
|
||||||
fi
|
|
||||||
|
|
||||||
awk -v mapfile="$map" '
|
|
||||||
BEGIN {
|
|
||||||
while ((getline line < mapfile) > 0) {
|
|
||||||
i = index(line, "\t")
|
|
||||||
d[substr(line, 1, i - 1)] = substr(line, i + 1)
|
|
||||||
}
|
|
||||||
}
|
|
||||||
{
|
|
||||||
if (match($0, /^[[:space:]]*image:[[:space:]]*gcr\.forust\.xyz\/forust\/[A-Za-z0-9._-]+:[A-Za-z0-9._-]+[[:space:]]*$/)) {
|
|
||||||
name = $0
|
|
||||||
sub(/^[[:space:]]*image:[[:space:]]*/, "", name)
|
|
||||||
sub(/[[:space:]]*$/, "", name)
|
|
||||||
if (name in d) {
|
|
||||||
pad = $0
|
|
||||||
sub(/image:.*/, "", pad)
|
|
||||||
# Drop the tag: the canonical form used in the docs is repo@sha256:...,
|
|
||||||
# and leaving :prod next to the digest reads like it still matters.
|
|
||||||
repo = name
|
|
||||||
sub(/:[A-Za-z0-9._-]+$/, "", repo)
|
|
||||||
print pad "image: " repo "@" d[name]
|
|
||||||
next
|
|
||||||
}
|
|
||||||
}
|
|
||||||
print
|
|
||||||
}
|
|
||||||
' "$src"
|
|
||||||
rm -f "$src" "$refs" "$map"
|
|
||||||
}
|
|
||||||
|
|
||||||
# Restarts every owned workload whose running image is not the one its tag
|
|
||||||
# resolves to now.
|
|
||||||
#
|
|
||||||
# This used to be how a rebuild reached the cluster at all: the manifests pinned
|
|
||||||
# `:latest`, so a rebuild left the pod template byte-identical, `kubectl apply`
|
|
||||||
# decided there was nothing to do, and the cluster served the previous build
|
|
||||||
# indefinitely. The apply now pins digests via render_pinned, so a rebuild moves
|
|
||||||
# the pod template and rolls out on its own.
|
|
||||||
#
|
|
||||||
# What is left is the drift check: a hand-run `kubectl set image`, or anything
|
|
||||||
# else that edits a live workload behind the deploy's back, is the only way to end
|
|
||||||
# up serving a digest the tag has moved past. It stays idempotent, so a redeploy
|
|
||||||
# that changed no image still does not bounce healthy services.
|
|
||||||
#
|
|
||||||
# The container is matched on its repository rather than on the exact reference:
|
|
||||||
# once render_pinned has run, a pod's status reports `repo@sha256:...` while this
|
|
||||||
# still reads the repository's `:prod` tag out of the manifest.
|
|
||||||
restart_stale_images() {
|
|
||||||
local ns target image want selector running entry one
|
|
||||||
local unchecked=0
|
|
||||||
local -A digests=()
|
|
||||||
local -a stale=()
|
|
||||||
while read -r ns target image; do
|
|
||||||
[ -n "${target:-}" ] || continue
|
|
||||||
if [ -z "${digests[$image]:-}" ]; then
|
|
||||||
digests[$image]="$(pinned_digest "$image")"
|
|
||||||
fi
|
|
||||||
want="${digests[$image]}"
|
|
||||||
if [ -z "$want" ]; then
|
|
||||||
warn "cannot resolve ${image##*/} in the registry, leaving $target alone"
|
|
||||||
unchecked=$((unchecked + 1))
|
|
||||||
continue
|
|
||||||
fi
|
|
||||||
selector="$(kubectl get "$target" -n "$ns" -o jsonpath='{.spec.selector.matchLabels}' 2>/dev/null \
|
|
||||||
| jq -r 'to_entries | map("\(.key)=\(.value)") | join(",")' 2>/dev/null)"
|
|
||||||
if [ -z "$selector" ]; then
|
|
||||||
warn "cannot read the pod selector of $target, skipping"
|
|
||||||
unchecked=$((unchecked + 1))
|
|
||||||
continue
|
|
||||||
fi
|
|
||||||
running="$(kubectl get pods -n "$ns" -l "$selector" -o json 2>/dev/null \
|
|
||||||
| jq -r --arg repo "${image%%:*}" '
|
|
||||||
.items[] | .status.containerStatuses[]?
|
|
||||||
| select(.image == $repo
|
|
||||||
or (.image | startswith($repo + ":"))
|
|
||||||
or (.image | startswith($repo + "@")))
|
|
||||||
| .imageID
|
|
||||||
' 2>/dev/null)"
|
|
||||||
if [ -z "$running" ]; then
|
|
||||||
# Scaled to zero. Nothing is serving stale code, and imagePullPolicy
|
|
||||||
# resolves the tag when it is scaled back up.
|
|
||||||
continue
|
|
||||||
fi
|
|
||||||
entry=""
|
|
||||||
while IFS= read -r one; do
|
|
||||||
[ -n "$one" ] || continue
|
|
||||||
entry="${one##*@}"
|
|
||||||
if [ "$entry" != "$want" ]; then
|
|
||||||
stale+=("$ns $target")
|
|
||||||
break
|
|
||||||
fi
|
|
||||||
done <<<"$running"
|
|
||||||
done < <(owned_registry_workloads)
|
|
||||||
if [ "${#stale[@]}" -eq 0 ]; then
|
|
||||||
if [ "$unchecked" -gt 0 ]; then
|
|
||||||
# Say so plainly. Reporting "everything is current" after checking nothing
|
|
||||||
# would tell the operator the deploy is fine when it may not be.
|
|
||||||
warn "No workload needed a restart, but $unchecked could not be checked"
|
|
||||||
else
|
|
||||||
log "All owned workloads already run the image their tag points at"
|
|
||||||
fi
|
|
||||||
return 0
|
|
||||||
fi
|
|
||||||
log "Restarting ${#stale[@]} workload(s) running an image their tag has moved past"
|
|
||||||
for ref in "${stale[@]}"; do
|
|
||||||
log " $ref"
|
|
||||||
done
|
|
||||||
local failed=()
|
|
||||||
for ref in "${stale[@]}"; do
|
|
||||||
ns="${ref%% *}"
|
|
||||||
target="${ref#* }"
|
|
||||||
if ! kubectl rollout restart "$target" -n "$ns" >/dev/null 2>&1; then
|
|
||||||
failed+=("$ref")
|
|
||||||
fi
|
|
||||||
done
|
|
||||||
if [ "${#failed[@]}" -gt 0 ]; then
|
|
||||||
warn "could not restart: ${failed[*]}"
|
|
||||||
return 1
|
|
||||||
fi
|
|
||||||
}
|
}
|
||||||
|
|
||||||
# verify_workloads <failed-file> <kind> <ns> <name> ...
|
# verify_workloads <failed-file> <kind> <ns> <name> ...
|
||||||
@@ -499,31 +281,44 @@ verify_workloads() {
|
|||||||
# settle. Prints a report and returns non-zero if any workload is still unhealthy,
|
# settle. Prints a report and returns non-zero if any workload is still unhealthy,
|
||||||
# so the operator knows manual recovery is required.
|
# so the operator knows manual recovery is required.
|
||||||
rollback_workloads() {
|
rollback_workloads() {
|
||||||
local failed_file="$1"
|
local failed_file="$1" snapshot kind ns name index=0 running=0 pid revision uid
|
||||||
local kind ns name unrecovered=()
|
local -a pids=()
|
||||||
local -a recovered=()
|
snapshot="$(cat "$DEPLOY_SNAPSHOT_DIR/current")"
|
||||||
while read -r kind ns name; do
|
while read -r kind ns name; do
|
||||||
[ -n "${kind:-}" ] || continue
|
[[ "$kind" =~ ^(deployment|statefulset|daemonset)$ ]] || continue
|
||||||
# Helm-owned workloads are already rolled back by the release's --rollback-on-failure
|
index=$((index + 1))
|
||||||
# upgrade. `rollout undo` here would step back to the revision Helm just
|
(
|
||||||
# escaped (the failed one), so leave them for the operator instead.
|
if kubectl get "$kind/$name" -n "$ns" -o jsonpath='{.metadata.annotations}' | grep -q 'meta.helm.sh/release-name'; then
|
||||||
if kubectl get "${kind}/${name}" -n "$ns" -o jsonpath='{.metadata.annotations}' 2>/dev/null | grep -q 'meta.helm.sh/release-name'; then
|
echo " skip (Helm recovery owns this workload): $kind/$ns/$name"
|
||||||
echo " skip (helm-managed, needs manual check): ${kind}/${ns}/${name}"
|
exit 1
|
||||||
unrecovered+=("${kind}/${ns}/${name} (helm-managed)")
|
fi
|
||||||
continue
|
revision="$(jq -r --arg ns "$ns" --arg name "$name" --arg kind "$kind" \
|
||||||
fi
|
'.[] | select(.namespace == $ns and .name == $name and .kind == $kind) | .revision' "$snapshot/revisions.json")"
|
||||||
if kubectl rollout undo "${kind}/${name}" -n "$ns" >/dev/null 2>&1 \
|
uid="$(jq -r --arg ns "$ns" --arg name "$name" --arg kind "$kind" \
|
||||||
&& kubectl rollout status "${kind}/${name}" -n "$ns" --timeout="${ROLLOUT_TIMEOUT}s" >/dev/null 2>&1; then
|
'.[] | select(.namespace == $ns and .name == $name and .kind == $kind) | .uid' "$snapshot/revisions.json")"
|
||||||
echo " rolled back: ${kind}/${ns}/${name}"
|
if [[ ! "$revision" =~ ^[1-9][0-9]*$ ]] || [ "$uid" != "$(kubectl get "$kind/$name" -n "$ns" -o jsonpath='{.metadata.uid}')" ]; then
|
||||||
recovered+=("${kind}/${ns}/${name}")
|
echo " no safe previous revision: $kind/$ns/$name (new or replaced workload)"
|
||||||
else
|
exit 1
|
||||||
echo " NOT RECOVERED: ${kind}/${ns}/${name}"
|
fi
|
||||||
unrecovered+=("${kind}/${ns}/${name}")
|
kubectl rollout undo "$kind/$name" -n "$ns" --to-revision="$revision" \
|
||||||
|
&& kubectl rollout status "$kind/$name" -n "$ns" --timeout="${ROLLOUT_TIMEOUT}s"
|
||||||
|
) >"$snapshot/rollback-$index.log" 2>&1 &
|
||||||
|
pids+=($!)
|
||||||
|
running=$((running + 1))
|
||||||
|
if [ "$running" -ge "$ROLLOUT_PARALLELISM" ]; then
|
||||||
|
wait -n 2>/dev/null || true
|
||||||
|
running=$((running - 1))
|
||||||
fi
|
fi
|
||||||
done <"$failed_file"
|
done <"$failed_file"
|
||||||
echo "ROLLED_BACK=${#recovered[@]}" >>"$failed_file"
|
local recovered=0 unrecovered=0 i=0
|
||||||
echo "UNRECOVERED=${#unrecovered[@]}" >>"$failed_file"
|
for pid in "${pids[@]}"; do
|
||||||
[ "${#unrecovered[@]}" -eq 0 ]
|
i=$((i + 1))
|
||||||
|
if wait "$pid"; then recovered=$((recovered + 1)); else unrecovered=$((unrecovered + 1)); fi
|
||||||
|
cat "$snapshot/rollback-$i.log"
|
||||||
|
done
|
||||||
|
echo "ROLLED_BACK=$recovered" >>"$failed_file"
|
||||||
|
echo "UNRECOVERED=$unrecovered" >>"$failed_file"
|
||||||
|
[ "$unrecovered" -eq 0 ]
|
||||||
}
|
}
|
||||||
|
|
||||||
# Helm releases owned by this stage, one line each:
|
# Helm releases owned by this stage, one line each:
|
||||||
@@ -536,6 +331,7 @@ rollback_workloads() {
|
|||||||
# have to be declared as custom.regex managers in renovate/renovate.json.
|
# have to be declared as custom.regex managers in renovate/renovate.json.
|
||||||
HELM_RELEASES=(
|
HELM_RELEASES=(
|
||||||
"prometheus-stack|prometheus-community/kube-prometheus-stack|prometheus|86.2.3|prometheus-stack/k8s/grafana-values.yaml|prometheus-stack/k8s/active"
|
"prometheus-stack|prometheus-community/kube-prometheus-stack|prometheus|86.2.3|prometheus-stack/k8s/grafana-values.yaml|prometheus-stack/k8s/active"
|
||||||
|
"victoria-operator|victoriametrics/victoria-metrics-operator|prometheus|0.68.1|prometheus-stack/k8s/victoria-operator-values.yaml|prometheus-stack/k8s/active"
|
||||||
"loki|grafana/loki|prometheus|7.3.0|loki/k8s/loki-values.yaml|loki/k8s/active"
|
"loki|grafana/loki|prometheus|7.3.0|loki/k8s/loki-values.yaml|loki/k8s/active"
|
||||||
"alloy|grafana/alloy|prometheus|1.12.1|loki/k8s/alloy-values.yaml|loki/k8s/active"
|
"alloy|grafana/alloy|prometheus|1.12.1|loki/k8s/alloy-values.yaml|loki/k8s/active"
|
||||||
"reloader|stakater/reloader|reloader|2.2.17|reloader/k8s/reloader-values.yaml|reloader/k8s/active"
|
"reloader|stakater/reloader|reloader|2.2.17|reloader/k8s/reloader-values.yaml|reloader/k8s/active"
|
||||||
@@ -547,6 +343,7 @@ helm_repo_for() {
|
|||||||
prometheus-community/*) echo "prometheus-community https://prometheus-community.github.io/helm-charts" ;;
|
prometheus-community/*) echo "prometheus-community https://prometheus-community.github.io/helm-charts" ;;
|
||||||
grafana/*) echo "grafana https://grafana.github.io/helm-charts" ;;
|
grafana/*) echo "grafana https://grafana.github.io/helm-charts" ;;
|
||||||
stakater/*) echo "stakater https://stakater.github.io/stakater-charts" ;;
|
stakater/*) echo "stakater https://stakater.github.io/stakater-charts" ;;
|
||||||
|
victoriametrics/*) echo "victoriametrics https://victoriametrics.github.io/helm-charts" ;;
|
||||||
esac
|
esac
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -556,8 +353,12 @@ helm_repo_for() {
|
|||||||
helm_release_status() {
|
helm_release_status() {
|
||||||
local out
|
local out
|
||||||
if ! out="$(helm status "$1" -n "$2" 2>&1)"; then
|
if ! out="$(helm status "$1" -n "$2" 2>&1)"; then
|
||||||
echo "not-found"
|
if [[ "$out" == *"release: not found"* ]]; then
|
||||||
return 0
|
echo "not-found"
|
||||||
|
return 0
|
||||||
|
fi
|
||||||
|
printf 'ERROR: cannot read Helm status: %s\n' "$out" >&2
|
||||||
|
return 1
|
||||||
fi
|
fi
|
||||||
awk '/^STATUS:/{print $2}' <<<"$out" | tr '[:upper:]' '[:lower:]'
|
awk '/^STATUS:/{print $2}' <<<"$out" | tr '[:upper:]' '[:lower:]'
|
||||||
}
|
}
|
||||||
@@ -569,20 +370,35 @@ helm_release_status() {
|
|||||||
# pending (deployed, failed, not-found). Returns non-zero when the release is
|
# pending (deployed, failed, not-found). Returns non-zero when the release is
|
||||||
# still not recoverable, so the pipeline fails loud instead of wedging.
|
# still not recoverable, so the pipeline fails loud instead of wedging.
|
||||||
recover_pending_release() {
|
recover_pending_release() {
|
||||||
local release="$1" namespace="$2" status
|
local release="$1" namespace="$2" status revision snapshot
|
||||||
status="$(helm_release_status "$release" "$namespace")"
|
status="$(helm_release_status "$release" "$namespace")" || return 1
|
||||||
case "$status" in
|
case "$status" in
|
||||||
pending-upgrade|pending-rollback|pending-install)
|
pending-upgrade|pending-rollback|pending-install)
|
||||||
log "Release $release is $status, rolling back to the last deployed revision"
|
log "Release $release is $status, rolling back to the last deployed revision"
|
||||||
if ! helm rollback "$release" -n "$namespace" --wait --timeout 10m >/dev/null 2>&1; then
|
revision=""
|
||||||
|
if [ -s "$DEPLOY_SNAPSHOT_DIR/current" ]; then
|
||||||
|
snapshot="$(cat "$DEPLOY_SNAPSHOT_DIR/current")"
|
||||||
|
if [ -s "$snapshot/helm-$release.json" ]; then
|
||||||
|
revision="$(jq -r '.version' "$snapshot/helm-$release.json")"
|
||||||
|
fi
|
||||||
|
fi
|
||||||
|
if [[ ! "$revision" =~ ^[1-9][0-9]*$ ]]; then
|
||||||
|
echo "ERROR: no captured Helm revision for $release; manual recovery required"
|
||||||
|
return 1
|
||||||
|
fi
|
||||||
|
record_apply helm-rollback "$namespace/$release" started
|
||||||
|
if ! helm rollback "$release" "$revision" -n "$namespace" --wait --timeout 10m; then
|
||||||
|
record_apply helm-rollback "$namespace/$release" failure
|
||||||
echo "WARN: helm rollback of $release did not complete"
|
echo "WARN: helm rollback of $release did not complete"
|
||||||
return 1
|
return 1
|
||||||
fi
|
fi
|
||||||
status="$(helm_release_status "$release" "$namespace")"
|
status="$(helm_release_status "$release" "$namespace")" || return 1
|
||||||
if [ "$status" != "deployed" ]; then
|
if [ "$status" != "deployed" ]; then
|
||||||
|
record_apply helm-rollback "$namespace/$release" failure
|
||||||
echo "WARN: $release is $status after rollback"
|
echo "WARN: $release is $status after rollback"
|
||||||
return 1
|
return 1
|
||||||
fi
|
fi
|
||||||
|
record_apply helm-rollback "$namespace/$release" success
|
||||||
;;
|
;;
|
||||||
esac
|
esac
|
||||||
return 0
|
return 0
|
||||||
@@ -614,11 +430,16 @@ upgrade_helm_releases() {
|
|||||||
local entry release chart namespace version values marker repo
|
local entry release chart namespace version values marker repo
|
||||||
for entry in ${HELM_RELEASES[@]+"${HELM_RELEASES[@]}"}; do
|
for entry in ${HELM_RELEASES[@]+"${HELM_RELEASES[@]}"}; do
|
||||||
IFS='|' read -r release chart namespace version values marker <<<"$entry"
|
IFS='|' read -r release chart namespace version values marker <<<"$entry"
|
||||||
|
if [ -n "${DEPLOY_PLAN:-}" ] && ! jq -e --arg name "$release" '.helm | index($name) != null' "$DEPLOY_PLAN" >/dev/null; then
|
||||||
|
echo "skip (unchanged Helm release): $release"
|
||||||
|
continue
|
||||||
|
fi
|
||||||
|
if [ ! -f "$REPO/$values" ] && [ -f "$CONFIG_REPO/$values" ]; then values="$CONFIG_REPO/$values"; else values="$REPO/$values"; fi
|
||||||
if [ ! -f "$REPO/$marker" ]; then
|
if [ ! -f "$REPO/$marker" ]; then
|
||||||
echo "skip (no $marker): $release"
|
echo "skip (no $marker): $release"
|
||||||
continue
|
continue
|
||||||
fi
|
fi
|
||||||
if [ ! -f "$REPO/$values" ]; then
|
if [ ! -f "$values" ]; then
|
||||||
echo "ERROR: $values is gitignored but missing on the workstation, restore it first."
|
echo "ERROR: $values is gitignored but missing on the workstation, restore it first."
|
||||||
return 1
|
return 1
|
||||||
fi
|
fi
|
||||||
@@ -627,8 +448,8 @@ upgrade_helm_releases() {
|
|||||||
echo "ERROR: no Helm repository configured for chart $chart"
|
echo "ERROR: no Helm repository configured for chart $chart"
|
||||||
return 1
|
return 1
|
||||||
fi
|
fi
|
||||||
helm repo add "${repo%% *}" "${repo#* }" >/dev/null 2>&1 || true
|
helm repo add "${repo%% *}" "${repo#* }" >/dev/null
|
||||||
helm repo update "${repo%% *}" >/dev/null 2>&1 || true
|
helm repo update "${repo%% *}" >/dev/null
|
||||||
log "Upgrading $release ($chart $version)"
|
log "Upgrading $release ($chart $version)"
|
||||||
wait_for_calm "helm $release"
|
wait_for_calm "helm $release"
|
||||||
# A previous run with --rollback-on-failure whose own rollback never finished leaves the
|
# A previous run with --rollback-on-failure whose own rollback never finished leaves the
|
||||||
@@ -641,11 +462,13 @@ upgrade_helm_releases() {
|
|||||||
# --rollback-on-failure (+ --wait) rolls the release back when the upgrade
|
# --rollback-on-failure (+ --wait) rolls the release back when the upgrade
|
||||||
# times out or the workloads it touches never become ready, so a bad chart
|
# times out or the workloads it touches never become ready, so a bad chart
|
||||||
# bump is not left half applied. (--atomic was this combo; deprecated.)
|
# bump is not left half applied. (--atomic was this combo; deprecated.)
|
||||||
|
record_apply helm-upgrade "$namespace/$release" started
|
||||||
if ! helm upgrade --install "$release" "$chart" \
|
if ! helm upgrade --install "$release" "$chart" \
|
||||||
--namespace "$namespace" \
|
--namespace "$namespace" \
|
||||||
--version "$version" \
|
--version "$version" \
|
||||||
--values "$REPO/$values" \
|
--values "$values" \
|
||||||
--wait --rollback-on-failure --cleanup-on-fail --timeout 10m; then
|
--wait --rollback-on-failure --cleanup-on-fail --timeout 10m; then
|
||||||
|
record_apply helm-upgrade "$namespace/$release" failure
|
||||||
echo "WARN: upgrade of $release failed, checking release state"
|
echo "WARN: upgrade of $release failed, checking release state"
|
||||||
# --rollback-on-failure already attempted its own rollback; finish the job when that
|
# --rollback-on-failure already attempted its own rollback; finish the job when that
|
||||||
# rollback never completed, otherwise the release stays pending-* and
|
# rollback never completed, otherwise the release stays pending-* and
|
||||||
@@ -655,35 +478,49 @@ upgrade_helm_releases() {
|
|||||||
else
|
else
|
||||||
echo "ERROR: upgrade of $release failed (release is back on its previous revision)."
|
echo "ERROR: upgrade of $release failed (release is back on its previous revision)."
|
||||||
fi
|
fi
|
||||||
|
record_apply helm-recovery-state "$namespace/$release" "$(helm_release_status "$release" "$namespace" || echo unknown)"
|
||||||
return 1
|
return 1
|
||||||
fi
|
fi
|
||||||
|
record_apply helm-upgrade "$namespace/$release" success
|
||||||
done
|
done
|
||||||
}
|
}
|
||||||
|
|
||||||
stage_preflight() {
|
stage_doctor() {
|
||||||
if [ ! -d "$REPO/.git" ]; then
|
local tool entry release chart namespace version values marker
|
||||||
echo "Repository not found at $REPO"
|
for tool in git docker kubectl helm jq curl timeout flock python3; do
|
||||||
exit 1
|
command -v "$tool" >/dev/null || { echo "Missing workstation tool: $tool"; return 1; }
|
||||||
fi
|
done
|
||||||
if [ -n "$DEPLOY_SHA" ]; then
|
docker compose version >/dev/null
|
||||||
log "Checking out the commit CI validated ($DEPLOY_SHA)"
|
docker buildx version >/dev/null
|
||||||
git -C "$REPO" fetch origin --quiet "$DEPLOY_SHA" 2>/dev/null \
|
[ "$(kubectl config current-context)" = "${KUBE_CONTEXT:?configure KUBE_CONTEXT}" ] || { echo "Unexpected Kubernetes context"; return 1; }
|
||||||
|| git -C "$REPO" fetch origin main
|
[ "$(kubectl get namespace kube-system -o jsonpath='{.metadata.uid}')" = "${EXPECTED_CLUSTER_UID:?configure EXPECTED_CLUSTER_UID}" ] || { echo "Unexpected Kubernetes cluster"; return 1; }
|
||||||
else
|
kubectl get --raw=/readyz --request-timeout=10s >/dev/null
|
||||||
git -C "$REPO" fetch origin main
|
[ "$(git -C "$REPO" rev-parse HEAD)" = "$DEPLOY_SHA" ] || return 1
|
||||||
fi
|
select_manifests
|
||||||
target="${DEPLOY_SHA:-origin/main}"
|
for entry in "${HELM_RELEASES[@]}"; do
|
||||||
log "Workstation state"
|
IFS='|' read -r release chart namespace version values marker <<<"$entry"
|
||||||
echo " local: $(git -C "$REPO" rev-parse --short HEAD)"
|
[ -f "$REPO/$marker" ] || continue
|
||||||
echo " target: $(git -C "$REPO" rev-parse --short "$target")"
|
[ -f "$REPO/$values" ] || [ -f "$CONFIG_REPO/$values" ] || { echo "Missing values: $values"; return 1; }
|
||||||
if [ -n "$(git -C "$REPO" status --porcelain --untracked-files=no)" ]; then
|
done
|
||||||
echo "ERROR: workstation has local tracked modifications, refusing reset:"
|
jq '{sha, selected, helm, removed}' "$DEPLOY_PLAN"
|
||||||
git -C "$REPO" status --porcelain --untracked-files=no
|
local cf
|
||||||
git -C "$REPO" diff --stat
|
for cf in "${COMPOSE_STACKS[@]}"; do
|
||||||
echo "Fix it on the workstation (commit, or 'git restore .'), then re-run the deploy."
|
compose "$cf" config --quiet
|
||||||
exit 1
|
while IFS= read -r network; do
|
||||||
fi
|
docker network inspect "$network" >/dev/null || return 1
|
||||||
git -C "$REPO" reset --hard "$target"
|
done < <(compose "$cf" config --format json | jq -r '.networks // {} | to_entries[] | select(.value.external == true) | .value.name')
|
||||||
|
python3 "$REPO/.gitea/workflows/compose-release.py" "$cf"
|
||||||
|
done
|
||||||
|
local image refs m k
|
||||||
|
refs="$(
|
||||||
|
for m in "${K8S_MANIFESTS[@]}"; do render_pinned <"$m" || return 1; done
|
||||||
|
for k in "${KUSTOMIZE_APPS[@]}"; do kubectl kustomize "$k" | render_pinned || return 1; done
|
||||||
|
)" || return 1
|
||||||
|
refs="$(grep -oE 'gcr\.forust\.xyz/forust/[A-Za-z0-9._-]+@sha256:[0-9a-f]{64}' <<<"$refs" | sort -u || true)"
|
||||||
|
while IFS= read -r image; do
|
||||||
|
[ -n "$image" ] || continue
|
||||||
|
timeout 60s docker buildx imagetools inspect "$image" >/dev/null
|
||||||
|
done <<<"$refs"
|
||||||
}
|
}
|
||||||
|
|
||||||
# Required pod Secrets, scoped to the resource namespace. TLS route Secrets are
|
# Required pod Secrets, scoped to the resource namespace. TLS route Secrets are
|
||||||
@@ -693,6 +530,9 @@ check_referenced_secrets() {
|
|||||||
local missing=()
|
local missing=()
|
||||||
refs=""
|
refs=""
|
||||||
for m in ${K8S_MANIFESTS[@]+"${K8S_MANIFESTS[@]}"}; do
|
for m in ${K8S_MANIFESTS[@]+"${K8S_MANIFESTS[@]}"}; do
|
||||||
|
if skip_uninstalled_vmagent_crd "$m"; then
|
||||||
|
continue
|
||||||
|
fi
|
||||||
objects="$(kubectl create --dry-run=client --validate=false -f "$m" -o json)" || return 1
|
objects="$(kubectl create --dry-run=client --validate=false -f "$m" -o json)" || return 1
|
||||||
extracted="$(printf '%s' "$objects" | jq -r -f "$REPO/.gitea/workflows/secret-references.jq")" || return 1
|
extracted="$(printf '%s' "$objects" | jq -r -f "$REPO/.gitea/workflows/secret-references.jq")" || return 1
|
||||||
refs+="$extracted"$'\n'
|
refs+="$extracted"$'\n'
|
||||||
@@ -719,7 +559,20 @@ check_referenced_secrets() {
|
|||||||
fi
|
fi
|
||||||
}
|
}
|
||||||
|
|
||||||
|
# The VMAgent CRD is installed by the VictoriaMetrics Operator Helm release in
|
||||||
|
# stage_apply_k8s, after this preflight stage. Skip only its dry-run until then.
|
||||||
|
skip_uninstalled_vmagent_crd() {
|
||||||
|
local manifest="$1"
|
||||||
|
if [[ "$manifest" == "$REPO/prometheus-stack/k8s/vmagent.yaml" ]] \
|
||||||
|
&& ! kubectl get crd vmagents.operator.victoriametrics.com >/dev/null 2>&1; then
|
||||||
|
echo " skip: VMAgent CRD is installed by Helm during apply: ${manifest#"$REPO"/}"
|
||||||
|
return 0
|
||||||
|
fi
|
||||||
|
return 1
|
||||||
|
}
|
||||||
|
|
||||||
stage_validate() {
|
stage_validate() {
|
||||||
|
check_prune_mode || return 1
|
||||||
cd "$REPO"
|
cd "$REPO"
|
||||||
select_manifests
|
select_manifests
|
||||||
local m k cf
|
local m k cf
|
||||||
@@ -731,10 +584,13 @@ stage_validate() {
|
|||||||
log "Validate compose stacks"
|
log "Validate compose stacks"
|
||||||
for cf in ${COMPOSE_STACKS[@]+"${COMPOSE_STACKS[@]}"}; do
|
for cf in ${COMPOSE_STACKS[@]+"${COMPOSE_STACKS[@]}"}; do
|
||||||
echo " config: $cf"
|
echo " config: $cf"
|
||||||
validate_compose_file "$cf"
|
compose "$cf" config --quiet
|
||||||
done
|
done
|
||||||
log "Validate k8s manifests (kubectl dry-run=client)"
|
log "Validate k8s manifests (kubectl dry-run=client)"
|
||||||
for m in ${K8S_MANIFESTS[@]+"${K8S_MANIFESTS[@]}"}; do
|
for m in ${K8S_MANIFESTS[@]+"${K8S_MANIFESTS[@]}"}; do
|
||||||
|
if skip_uninstalled_vmagent_crd "$m"; then
|
||||||
|
continue
|
||||||
|
fi
|
||||||
kubectl apply --dry-run=client -f "$m" >/dev/null
|
kubectl apply --dry-run=client -f "$m" >/dev/null
|
||||||
done
|
done
|
||||||
for k in ${KUSTOMIZE_APPS[@]+"${KUSTOMIZE_APPS[@]}"}; do
|
for k in ${KUSTOMIZE_APPS[@]+"${KUSTOMIZE_APPS[@]}"}; do
|
||||||
@@ -742,6 +598,9 @@ stage_validate() {
|
|||||||
done
|
done
|
||||||
log "Validate k8s manifests (kubectl dry-run=server)"
|
log "Validate k8s manifests (kubectl dry-run=server)"
|
||||||
for m in ${K8S_MANIFESTS[@]+"${K8S_MANIFESTS[@]}"}; do
|
for m in ${K8S_MANIFESTS[@]+"${K8S_MANIFESTS[@]}"}; do
|
||||||
|
if skip_uninstalled_vmagent_crd "$m"; then
|
||||||
|
continue
|
||||||
|
fi
|
||||||
kubectl apply --dry-run=server -f "$m" >/dev/null
|
kubectl apply --dry-run=server -f "$m" >/dev/null
|
||||||
done
|
done
|
||||||
for k in ${KUSTOMIZE_APPS[@]+"${KUSTOMIZE_APPS[@]}"}; do
|
for k in ${KUSTOMIZE_APPS[@]+"${KUSTOMIZE_APPS[@]}"}; do
|
||||||
@@ -752,34 +611,54 @@ stage_validate() {
|
|||||||
check_referenced_secrets
|
check_referenced_secrets
|
||||||
}
|
}
|
||||||
|
|
||||||
|
selected_workload_refs() {
|
||||||
|
local m k
|
||||||
|
for m in "${K8S_MANIFESTS[@]}"; do
|
||||||
|
if skip_uninstalled_vmagent_crd "$m" >/dev/null; then continue; fi
|
||||||
|
kubectl create --dry-run=client --validate=false -f "$m" -o json | jq -r '
|
||||||
|
(if .kind == "List" then .items[] else . end) | select(.kind | test("^(Deployment|StatefulSet|DaemonSet)$"))
|
||||||
|
| "\(.kind | ascii_downcase) \(.metadata.namespace // "default") \(.metadata.name)"'
|
||||||
|
done
|
||||||
|
for k in "${KUSTOMIZE_APPS[@]}"; do
|
||||||
|
kubectl kustomize "$k" | kubectl create --dry-run=client --validate=false -f - -o json | jq -r '
|
||||||
|
(if .kind == "List" then .items[] else . end) | select(.kind | test("^(Deployment|StatefulSet|DaemonSet)$"))
|
||||||
|
| "\(.kind | ascii_downcase) \(.metadata.namespace // "default") \(.metadata.name)"'
|
||||||
|
done
|
||||||
|
}
|
||||||
|
|
||||||
stage_apply_k8s() {
|
stage_apply_k8s() {
|
||||||
|
check_prune_mode || return 1
|
||||||
cd "$REPO"
|
cd "$REPO"
|
||||||
select_manifests >/dev/null
|
select_manifests >/dev/null
|
||||||
local ns_files=() other_files=() m k prune_opts=()
|
local ns_files=() other_files=() m k
|
||||||
for m in ${K8S_MANIFESTS[@]+"${K8S_MANIFESTS[@]}"}; do
|
for m in ${K8S_MANIFESTS[@]+"${K8S_MANIFESTS[@]}"}; do
|
||||||
case "$m" in
|
case "$m" in
|
||||||
*/namespace.y?ml) ns_files+=("$m") ;;
|
*/namespace.yaml|*/namespace.yml) ns_files+=("$m") ;;
|
||||||
*) other_files+=("$m") ;;
|
*) other_files+=("$m") ;;
|
||||||
esac
|
esac
|
||||||
done
|
done
|
||||||
if [ "$APPLY_PRUNE" = "true" ]; then
|
|
||||||
prune_opts=(--prune -l app.kubernetes.io/managed-by=homelab-deploy)
|
|
||||||
fi
|
|
||||||
|
|
||||||
# Record what is about to change, and publish it for the verify job, before
|
# Record what is about to change, and publish it for the verify job, before
|
||||||
# the first apply. Both are fatal on failure: see snapshot_dir.
|
# the first apply. Both are fatal on failure: see snapshot_dir.
|
||||||
|
selected_workload_refs >"$RUN_DIR/workload-refs"
|
||||||
local snapshot
|
local snapshot
|
||||||
snapshot="$(snapshot_dir)" || return 1
|
snapshot="$(snapshot_dir)" || return 1
|
||||||
save_snapshot "$snapshot" || return 1
|
save_snapshot "$snapshot" || return 1
|
||||||
|
touch "$snapshot/ready"
|
||||||
|
|
||||||
if [ "${#ns_files[@]}" -gt 0 ]; then
|
if [ "${#ns_files[@]}" -gt 0 ]; then
|
||||||
log "Applying namespaces (${#ns_files[@]} files)"
|
log "Applying namespaces (${#ns_files[@]} files)"
|
||||||
for m in "${ns_files[@]}"; do
|
for m in "${ns_files[@]}"; do
|
||||||
kubectl apply -f "$m"
|
record_apply kubectl "${m#"$REPO"/}" started
|
||||||
|
if ! kubectl apply -f "$m"; then
|
||||||
|
record_apply kubectl "${m#"$REPO"/}" failure
|
||||||
|
return 1
|
||||||
|
fi
|
||||||
|
record_apply kubectl "${m#"$REPO"/}" success
|
||||||
done
|
done
|
||||||
fi
|
fi
|
||||||
if [ -f "$REPO/prometheus-stack/k8s/active" ]; then
|
if selected_service k8s prometheus-stack && [ -f "$REPO/prometheus-stack/k8s/active" ]; then
|
||||||
if [ ! -f "$REPO/prometheus-stack/k8s/grafana-values.yaml" ]; then
|
if [ ! -f "$CONFIG_REPO/prometheus-stack/k8s/grafana-values.yaml" ]; then
|
||||||
echo "ERROR: prometheus-stack/k8s/grafana-values.yaml (gitignored) missing on workstation, restore it first."
|
echo "ERROR: prometheus-stack/k8s/grafana-values.yaml (gitignored) missing on workstation, restore it first."
|
||||||
exit 1
|
exit 1
|
||||||
fi
|
fi
|
||||||
@@ -789,20 +668,26 @@ stage_apply_k8s() {
|
|||||||
if [ "${#other_files[@]}" -gt 0 ]; then
|
if [ "${#other_files[@]}" -gt 0 ]; then
|
||||||
log "Applying resources (${#other_files[@]} files, our images pinned to digests)"
|
log "Applying resources (${#other_files[@]} files, our images pinned to digests)"
|
||||||
for m in "${other_files[@]}"; do
|
for m in "${other_files[@]}"; do
|
||||||
if ! render_pinned <"$m" | kubectl apply "${prune_opts[@]}" -f -; then
|
log "Applying ${m#"$REPO"/}"
|
||||||
|
record_apply kubectl "${m#"$REPO"/}" started
|
||||||
|
if ! render_pinned <"$m" | kubectl apply -f -; then
|
||||||
|
record_apply kubectl "${m#"$REPO"/}" failure
|
||||||
echo "ERROR: apply failed for ${m#"$REPO"/}" >&2
|
echo "ERROR: apply failed for ${m#"$REPO"/}" >&2
|
||||||
exit 1
|
exit 1
|
||||||
fi
|
fi
|
||||||
|
record_apply kubectl "${m#"$REPO"/}" success
|
||||||
done
|
done
|
||||||
fi
|
fi
|
||||||
for k in ${KUSTOMIZE_APPS[@]+"${KUSTOMIZE_APPS[@]}"}; do
|
for k in ${KUSTOMIZE_APPS[@]+"${KUSTOMIZE_APPS[@]}"}; do
|
||||||
log "Applying kustomize app: ${k#"$REPO"/} (our images pinned to digests)"
|
log "Applying kustomize app: ${k#"$REPO"/} (our images pinned to digests)"
|
||||||
|
record_apply kustomize "${k#"$REPO"/}" started
|
||||||
if ! kubectl kustomize "$k" | render_pinned | kubectl apply -f -; then
|
if ! kubectl kustomize "$k" | render_pinned | kubectl apply -f -; then
|
||||||
|
record_apply kustomize "${k#"$REPO"/}" failure
|
||||||
echo "ERROR: apply failed for kustomize app ${k#"$REPO"/}" >&2
|
echo "ERROR: apply failed for kustomize app ${k#"$REPO"/}" >&2
|
||||||
exit 1
|
exit 1
|
||||||
fi
|
fi
|
||||||
|
record_apply kustomize "${k#"$REPO"/}" success
|
||||||
done
|
done
|
||||||
restart_stale_images
|
|
||||||
|
|
||||||
# No verification here on purpose. This stage may be killed at any point by
|
# No verification here on purpose. This stage may be killed at any point by
|
||||||
# timeout-minutes, by the runner cancelling the job, or by a dropped SSH
|
# timeout-minutes, by the runner cancelling the job, or by a dropped SSH
|
||||||
@@ -817,7 +702,7 @@ stage_apply_k8s() {
|
|||||||
# and rolls back the ones that never became healthy.
|
# and rolls back the ones that never became healthy.
|
||||||
stage_verify_k8s() {
|
stage_verify_k8s() {
|
||||||
local pointer="$DEPLOY_SNAPSHOT_DIR/current"
|
local pointer="$DEPLOY_SNAPSHOT_DIR/current"
|
||||||
local snapshot want have generations
|
local snapshot want have generations changed
|
||||||
local -a touched=()
|
local -a touched=()
|
||||||
|
|
||||||
if [ ! -s "$pointer" ]; then
|
if [ ! -s "$pointer" ]; then
|
||||||
@@ -828,7 +713,7 @@ stage_verify_k8s() {
|
|||||||
return 1
|
return 1
|
||||||
fi
|
fi
|
||||||
snapshot="$(head -1 "$pointer")"
|
snapshot="$(head -1 "$pointer")"
|
||||||
if [ ! -d "$snapshot" ]; then
|
if [ ! -d "$snapshot" ] || [ ! -f "$snapshot/ready" ]; then
|
||||||
echo "ERROR: snapshot pointer refers to a missing directory: $snapshot"
|
echo "ERROR: snapshot pointer refers to a missing directory: $snapshot"
|
||||||
return 1
|
return 1
|
||||||
fi
|
fi
|
||||||
@@ -851,6 +736,13 @@ stage_verify_k8s() {
|
|||||||
fi
|
fi
|
||||||
echo " snapshot: $snapshot (commit ${have:0:12})"
|
echo " snapshot: $snapshot (commit ${have:0:12})"
|
||||||
|
|
||||||
|
local entry release chart namespace version values marker
|
||||||
|
for entry in "${HELM_RELEASES[@]}"; do
|
||||||
|
IFS='|' read -r release chart namespace version values marker <<<"$entry"
|
||||||
|
jq -e --arg name "$release" '.helm | index($name) != null' "$DEPLOY_PLAN" >/dev/null || continue
|
||||||
|
recover_pending_release "$release" "$namespace" || return 1
|
||||||
|
done
|
||||||
|
|
||||||
generations="$snapshot/generations.before"
|
generations="$snapshot/generations.before"
|
||||||
if [ ! -s "$generations" ]; then
|
if [ ! -s "$generations" ]; then
|
||||||
# Without a baseline we cannot tell which workloads the apply touched, so
|
# Without a baseline we cannot tell which workloads the apply touched, so
|
||||||
@@ -859,9 +751,10 @@ stage_verify_k8s() {
|
|||||||
: >"$generations"
|
: >"$generations"
|
||||||
fi
|
fi
|
||||||
|
|
||||||
|
changed="$(changed_workloads "$generations")" || return 1
|
||||||
while read -r kind ns name; do
|
while read -r kind ns name; do
|
||||||
[ -n "${kind:-}" ] && touched+=("$kind $ns $name")
|
[ -n "${kind:-}" ] && touched+=("$kind $ns $name")
|
||||||
done < <(changed_workloads "$generations")
|
done <<<"$changed"
|
||||||
|
|
||||||
log "Verifying ${#touched[@]} changed workload(s) (timeout ${ROLLOUT_TIMEOUT}s each)"
|
log "Verifying ${#touched[@]} changed workload(s) (timeout ${ROLLOUT_TIMEOUT}s each)"
|
||||||
if [ "${#touched[@]}" -eq 0 ]; then
|
if [ "${#touched[@]}" -eq 0 ]; then
|
||||||
@@ -900,21 +793,19 @@ stage_verify_k8s() {
|
|||||||
verify_compose_stack() {
|
verify_compose_stack() {
|
||||||
local cf="$1"
|
local cf="$1"
|
||||||
local expected running missing=()
|
local expected running missing=()
|
||||||
expected="$(docker compose -f "$cf" config --services 2>/dev/null | sort || true)"
|
expected="$(compose "$cf" config --format json | jq -r ' .services | to_entries[] | select(.value.restart != "no") | .key' | sort)" || return 1
|
||||||
running="$(docker compose -f "$cf" ps --status running --services 2>/dev/null | sort || true)"
|
running="$(compose "$cf" ps --status running --services | sort)" || return 1
|
||||||
[ -n "$expected" ] || return 0
|
[ -n "$expected" ] || return 0
|
||||||
while IFS= read -r svc; do
|
while IFS= read -r svc; do
|
||||||
[ -n "$svc" ] || continue
|
[ -n "$svc" ] || continue
|
||||||
# restart:"no" services are allowed to have exited.
|
# restart:"no" services are allowed to have exited.
|
||||||
if ! printf '%s\n' "$running" | grep -qx "$svc" \
|
if ! printf '%s\n' "$running" | grep -qx "$svc"; then
|
||||||
&& ! docker compose -f "$cf" config 2>/dev/null \
|
|
||||||
| grep -A5 "^ ${svc}:" | grep -qE 'restart:\s*"?no"?'; then
|
|
||||||
missing+=("$svc")
|
missing+=("$svc")
|
||||||
fi
|
fi
|
||||||
done <<<"$expected"
|
done <<<"$expected"
|
||||||
if [ "${#missing[@]}" -gt 0 ]; then
|
if [ "${#missing[@]}" -gt 0 ]; then
|
||||||
echo " NOT RUNNING: ${missing[*]}"
|
echo " NOT RUNNING: ${missing[*]}"
|
||||||
docker compose -f "$cf" ps --all 2>/dev/null | sed 's/^/ /' || true
|
compose "$cf" ps --all 2>/dev/null | sed 's/^/ /' || true
|
||||||
return 1
|
return 1
|
||||||
fi
|
fi
|
||||||
echo " all ${#expected} service(s) running"
|
echo " all ${#expected} service(s) running"
|
||||||
@@ -998,10 +889,11 @@ traefik_routed_hosts() {
|
|||||||
# cases, so ask Traefik which routes it built and fail on the difference.
|
# cases, so ask Traefik which routes it built and fail on the difference.
|
||||||
stage_smoke() {
|
stage_smoke() {
|
||||||
cd "$REPO"
|
cd "$REPO"
|
||||||
|
if [ -n "${DEPLOY_PLAN:-}" ] && jq -e '.full_smoke' "$DEPLOY_PLAN" >/dev/null; then
|
||||||
|
DEPLOY_SMOKE_ALL=true
|
||||||
|
fi
|
||||||
select_manifests >/dev/null
|
select_manifests >/dev/null
|
||||||
local -a hosts=()
|
local -a hosts=()
|
||||||
# Not named failed: an array of that name already exists in restart_stale_images
|
|
||||||
# above, and a scalar shadowing an array is a trap rather than a shadow.
|
|
||||||
local h code rc bad=0
|
local h code rc bad=0
|
||||||
while IFS= read -r h; do
|
while IFS= read -r h; do
|
||||||
[ -n "$h" ] && hosts+=("$h")
|
[ -n "$h" ] && hosts+=("$h")
|
||||||
@@ -1009,8 +901,8 @@ stage_smoke() {
|
|||||||
|
|
||||||
if [ "${#hosts[@]}" -eq 0 ]; then
|
if [ "${#hosts[@]}" -eq 0 ]; then
|
||||||
# Nothing to probe means the extraction broke, not that the cluster is empty.
|
# Nothing to probe means the extraction broke, not that the cluster is empty.
|
||||||
echo "ERROR: no public hostnames found in active manifests, refusing to report success"
|
echo "No public routes in the selected components"
|
||||||
return 1
|
return 0
|
||||||
fi
|
fi
|
||||||
|
|
||||||
log "Probing ${#hosts[@]} public route(s)"
|
log "Probing ${#hosts[@]} public route(s)"
|
||||||
@@ -1094,37 +986,22 @@ stage_apply_compose() {
|
|||||||
cd "$REPO"
|
cd "$REPO"
|
||||||
select_manifests >/dev/null
|
select_manifests >/dev/null
|
||||||
local cf
|
local cf
|
||||||
log "Redeploying docker compose stacks (${#COMPOSE_STACKS[@]} stacks)"
|
for cf in "${COMPOSE_STACKS[@]}"; do
|
||||||
for cf in ${COMPOSE_STACKS[@]+"${COMPOSE_STACKS[@]}"}; do
|
log "Applying Compose ${cf#"$REPO"/}"
|
||||||
echo " compose: $cf"
|
record_apply compose "${cf#"$REPO"/}" started
|
||||||
if grep -Eq '^\s+pull_policy:\s*build\b' "$cf"; then
|
if ! compose "$cf" up -d --wait --wait-timeout 180 --pull missing --remove-orphans; then
|
||||||
docker compose -f "$cf" build
|
record_apply compose "${cf#"$REPO"/}" failure
|
||||||
docker compose -f "$cf" push
|
return 1
|
||||||
fi
|
fi
|
||||||
docker compose -f "$cf" up -d --pull always --remove-orphans
|
record_apply compose "${cf#"$REPO"/}" success
|
||||||
|
verify_compose_stack "$cf"
|
||||||
done
|
done
|
||||||
|
echo "Compose recovery files: $RUN_DIR/compose-before (manual recovery only)"
|
||||||
local -a broken=()
|
|
||||||
for cf in ${COMPOSE_STACKS[@]+"${COMPOSE_STACKS[@]}"}; do
|
|
||||||
echo " verifying: $cf"
|
|
||||||
if ! verify_compose_stack "$cf"; then
|
|
||||||
broken+=("$cf")
|
|
||||||
fi
|
|
||||||
done
|
|
||||||
if [ "${#broken[@]}" -gt 0 ]; then
|
|
||||||
echo
|
|
||||||
echo "ERROR: ${#broken[@]} compose stack(s) did not come up:"
|
|
||||||
printf ' - %s\n' "${broken[@]}"
|
|
||||||
echo "Compose stacks are not rolled back automatically: their images use mutable"
|
|
||||||
echo "':latest' tags, so there is no previous version to return to. Check the logs"
|
|
||||||
echo "above, then re-run the deploy once the cause is fixed."
|
|
||||||
return 1
|
|
||||||
fi
|
|
||||||
}
|
}
|
||||||
|
|
||||||
run_stage() {
|
run_stage() {
|
||||||
case "${1:?stage required}" in
|
case "${1:?stage required}" in
|
||||||
preflight) stage_preflight ;;
|
doctor) stage_doctor ;;
|
||||||
validate) stage_validate ;;
|
validate) stage_validate ;;
|
||||||
apply-k8s) stage_apply_k8s ;;
|
apply-k8s) stage_apply_k8s ;;
|
||||||
verify-k8s) stage_verify_k8s ;;
|
verify-k8s) stage_verify_k8s ;;
|
||||||
|
|||||||
@@ -0,0 +1,124 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
"""Calculate selected components against the last fully successful deploy."""
|
||||||
|
|
||||||
|
import hashlib
|
||||||
|
import json
|
||||||
|
import re
|
||||||
|
import subprocess
|
||||||
|
from pathlib import Path
|
||||||
|
|
||||||
|
|
||||||
|
def output(*args, **kwargs):
|
||||||
|
return subprocess.check_output(args, text=True, **kwargs).strip() # noqa: S603
|
||||||
|
|
||||||
|
|
||||||
|
def tracked(repo):
|
||||||
|
return output('git', '-C', str(repo), 'ls-files').splitlines()
|
||||||
|
|
||||||
|
|
||||||
|
def helm_releases(repo):
|
||||||
|
text = (repo / '.gitea/workflows/deploy-lib.sh').read_text()
|
||||||
|
return [line.split('|') for line in re.findall(r'^ "([^"\n]+\|[^"\n]+)"$', text, re.MULTILINE)]
|
||||||
|
|
||||||
|
|
||||||
|
def inventory(repo):
|
||||||
|
files = tracked(repo)
|
||||||
|
k8s = sorted(
|
||||||
|
{f.split('/k8s/')[0] for f in files if '/k8s/' in f and (repo / f.split('/k8s/')[0] / 'k8s/active').is_file()}
|
||||||
|
)
|
||||||
|
compose = sorted(
|
||||||
|
{
|
||||||
|
str(Path(f).parent)
|
||||||
|
for f in files
|
||||||
|
if Path(f).name in ('compose.yaml', 'compose.yml') and (repo / Path(f).parent / 'active').is_file()
|
||||||
|
}
|
||||||
|
)
|
||||||
|
return {'k8s': k8s, 'compose': compose}
|
||||||
|
|
||||||
|
|
||||||
|
def file_hash(path):
|
||||||
|
return hashlib.sha256(path.read_bytes()).hexdigest() if path.is_file() else 'missing'
|
||||||
|
|
||||||
|
|
||||||
|
def make_plan(repo, config_repo, release, previous, mode, live_helm):
|
||||||
|
active = inventory(repo)
|
||||||
|
all_services = set(active['k8s'] + active['compose'])
|
||||||
|
helm_inputs = {}
|
||||||
|
helm_selected = []
|
||||||
|
for name, chart, namespace, version, values, marker in helm_releases(repo):
|
||||||
|
if not (repo / marker).is_file():
|
||||||
|
continue
|
||||||
|
value_path = repo / values if (repo / values).is_file() else config_repo / values
|
||||||
|
if not value_path.is_file():
|
||||||
|
raise ValueError(f'Missing Helm values: {values}')
|
||||||
|
stamp = hashlib.sha256(f'{chart}|{version}|{file_hash(value_path)}'.encode()).hexdigest()
|
||||||
|
helm_inputs[name] = stamp
|
||||||
|
live = next((h for h in live_helm if h['name'] == name and h['namespace'] == namespace), None)
|
||||||
|
if (
|
||||||
|
mode == 'full'
|
||||||
|
or previous is None
|
||||||
|
or previous.get('helm_inputs', {}).get(name) != stamp
|
||||||
|
or live is None
|
||||||
|
or live.get('status') != 'deployed'
|
||||||
|
or live.get('chart') != f'{chart.split("/")[-1]}-{version}'
|
||||||
|
):
|
||||||
|
helm_selected.append(name)
|
||||||
|
local_inputs = {}
|
||||||
|
for service in all_services:
|
||||||
|
candidates = [config_repo / service / '.env']
|
||||||
|
if service in active['compose']:
|
||||||
|
candidates.append(config_repo / '.env')
|
||||||
|
cfg = config_repo / service / 'config'
|
||||||
|
if cfg.is_dir():
|
||||||
|
candidates.extend(
|
||||||
|
p for p in cfg.rglob('*') if p.is_file() and p.suffix in ('.yaml', '.yml', '.json', '.conf')
|
||||||
|
)
|
||||||
|
local_inputs[service] = hashlib.sha256(
|
||||||
|
'\n'.join(f'{p.relative_to(config_repo)}:{file_hash(p)}' for p in sorted(candidates)).encode()
|
||||||
|
).hexdigest()
|
||||||
|
if previous is None:
|
||||||
|
if mode == 'changed':
|
||||||
|
raise ValueError('No successful baseline; run deploy in full mode first')
|
||||||
|
changed = set(all_services)
|
||||||
|
removed = []
|
||||||
|
else:
|
||||||
|
paths = output('git', '-C', str(repo), 'diff', '--name-only', previous['sha'], release['sha']).splitlines()
|
||||||
|
changed = {path.split('/')[0] for path in paths}
|
||||||
|
if any(path.startswith('.gitea/') for path in paths):
|
||||||
|
changed |= all_services
|
||||||
|
changed |= {s for s in all_services if previous.get('local_inputs', {}).get(s) != local_inputs[s]}
|
||||||
|
for file in tracked(repo):
|
||||||
|
service = file.split('/')[0]
|
||||||
|
if service not in all_services or not file.endswith(('.yaml', '.yml')):
|
||||||
|
continue
|
||||||
|
text = (repo / file).read_text()
|
||||||
|
if any(
|
||||||
|
image in text and previous.get('images', {}).get(image) != digest
|
||||||
|
for image, digest in release['images'].items()
|
||||||
|
):
|
||||||
|
changed.add(service)
|
||||||
|
removed = sorted(
|
||||||
|
set(previous.get('active', {}).get('k8s', []) + previous.get('active', {}).get('compose', []))
|
||||||
|
- all_services
|
||||||
|
)
|
||||||
|
removed += [path for path in paths if '/k8s/' in path and not (repo / path).exists()]
|
||||||
|
if mode == 'full':
|
||||||
|
changed = set(all_services)
|
||||||
|
dependencies = json.loads((repo / '.gitea/deploy-dependencies.json').read_text())
|
||||||
|
while True:
|
||||||
|
expanded = changed | {dependent for service in changed for dependent in dependencies.get(service, [])}
|
||||||
|
if expanded == changed:
|
||||||
|
break
|
||||||
|
changed = expanded
|
||||||
|
return {
|
||||||
|
'version': 1,
|
||||||
|
'sha': release['sha'],
|
||||||
|
'images': release['images'],
|
||||||
|
'active': active,
|
||||||
|
'selected': {kind: sorted(set(services) & changed) for kind, services in active.items()},
|
||||||
|
'helm': helm_selected,
|
||||||
|
'helm_inputs': helm_inputs,
|
||||||
|
'local_inputs': local_inputs,
|
||||||
|
'removed': sorted(set(removed)),
|
||||||
|
'full_smoke': mode == 'full' or 'traefik' in changed,
|
||||||
|
}
|
||||||
Executable
+10
@@ -0,0 +1,10 @@
|
|||||||
|
#!/usr/bin/env bash
|
||||||
|
set -euo pipefail
|
||||||
|
source "${REPO:?}/.gitea/workflows/deploy-lib.sh"
|
||||||
|
case "${1:?stage required}" in
|
||||||
|
workload-count)
|
||||||
|
select_manifests >/dev/null
|
||||||
|
selected_workload_refs | sort -u | wc -l
|
||||||
|
;;
|
||||||
|
*) run_stage "$1" ;;
|
||||||
|
esac
|
||||||
+100
-157
@@ -1,197 +1,140 @@
|
|||||||
name: deploy
|
name: deploy
|
||||||
|
|
||||||
on:
|
on:
|
||||||
# Deploy only what CI already validated. workflow_run is used instead of
|
|
||||||
# workflow_dispatch so a red lint/validate run can never reach the cluster.
|
|
||||||
workflow_run:
|
workflow_run:
|
||||||
workflows: [ci]
|
workflows: [ci]
|
||||||
|
branches: [main]
|
||||||
types: [completed]
|
types: [completed]
|
||||||
workflow_dispatch:
|
workflow_dispatch:
|
||||||
|
inputs:
|
||||||
|
deploy_ref:
|
||||||
|
description: "Commit already checked by successful main CI (main or SHA)"
|
||||||
|
default: main
|
||||||
|
required: true
|
||||||
|
deploy_mode:
|
||||||
|
description: "First deploy requires full; plan changes no production resources"
|
||||||
|
type: choice
|
||||||
|
options: [changed, full, plan]
|
||||||
|
default: changed
|
||||||
|
refresh_images:
|
||||||
|
description: "Explicitly refresh mutable third-party Compose tags"
|
||||||
|
type: boolean
|
||||||
|
default: false
|
||||||
|
|
||||||
# The deploy jobs read the tree, then reach the cluster over SSH with the
|
|
||||||
# deploy key. The Actions token itself is not part of that path, so it gets
|
|
||||||
# read-only contents and no more.
|
|
||||||
permissions:
|
permissions:
|
||||||
contents: read
|
contents: read
|
||||||
|
actions: read
|
||||||
|
|
||||||
concurrency:
|
concurrency:
|
||||||
group: deploy-main
|
group: deploy-main
|
||||||
# Queue instead of cancelling. Cancelling a run kills the apply job mid-loop and
|
|
||||||
# takes the verify job down with it, so a superseded deploy would leave the
|
|
||||||
# cluster half-applied and unchecked — the exact failure the verify job exists
|
|
||||||
# to catch. kubectl apply and docker compose up are both idempotent, so letting
|
|
||||||
# the older run finish and then deploying the newer commit costs little.
|
|
||||||
cancel-in-progress: false
|
cancel-in-progress: false
|
||||||
|
|
||||||
env:
|
env:
|
||||||
DEPLOY_HOST: ${{ secrets.DEPLOY_HOST }}
|
DEPLOY_HOST: ${{ vars.DEPLOY_HOST || secrets.DEPLOY_HOST }}
|
||||||
DEPLOY_PORT: ${{ secrets.DEPLOY_PORT }}
|
DEPLOY_PORT: ${{ vars.DEPLOY_PORT || secrets.DEPLOY_PORT }}
|
||||||
DEPLOY_USER: ${{ secrets.DEPLOY_USER }}
|
DEPLOY_USER: ${{ vars.DEPLOY_USER || secrets.DEPLOY_USER }}
|
||||||
DEPLOY_PATH: ${{ secrets.DEPLOY_PATH }}
|
|
||||||
DEPLOY_KEY: ${{ secrets.DEPLOY_SSH_KEY }}
|
DEPLOY_KEY: ${{ secrets.DEPLOY_SSH_KEY }}
|
||||||
APPLY_PRUNE: ${{ vars.APPLY_PRUNE }}
|
DEPLOY_KNOWN_HOSTS: ${{ vars.DEPLOY_KNOWN_HOSTS }}
|
||||||
# workflow_run's own GITHUB_SHA points at the branch head, not at the commit the
|
DEPLOY_RUN_ID: ${{ github.run_id }}-${{ github.run_attempt || 1 }}
|
||||||
# finished ci run checked. Pin the exact validated commit instead, so a push
|
DEPLOY_MODE: ${{ inputs.deploy_mode || 'changed' }}
|
||||||
# landing mid-deploy cannot make the workstation deploy something else. Also
|
REFRESH_IMAGES: ${{ inputs.refresh_images && 'true' || 'false' }}
|
||||||
# what the verify job checks the snapshot against. Empty for workflow_dispatch,
|
|
||||||
# which falls back to the current origin/main.
|
|
||||||
DEPLOY_SHA: ${{ github.event.workflow_run.head_sha }}
|
|
||||||
|
|
||||||
jobs:
|
jobs:
|
||||||
preflight:
|
gate:
|
||||||
# Autodeploy defaults to OFF: pushes deploy only when the AUTODEPLOY repo
|
|
||||||
# variable is set to 'true' (Settings -> Actions -> Variables). A manual
|
|
||||||
# Run workflow always bypasses the switch: dispatching it is the explicit
|
|
||||||
# intent to deploy.
|
|
||||||
if: >-
|
if: >-
|
||||||
|
github.ref == 'refs/heads/main' &&
|
||||||
(vars.AUTODEPLOY == 'true' || github.event_name == 'workflow_dispatch') &&
|
(vars.AUTODEPLOY == 'true' || github.event_name == 'workflow_dispatch') &&
|
||||||
(github.event_name != 'workflow_run' ||
|
(github.event_name != 'workflow_run' ||
|
||||||
(github.event.workflow_run.conclusion == 'success' &&
|
(github.event.workflow_run.conclusion == 'success' && github.event.workflow_run.head_branch == 'main'))
|
||||||
github.event.workflow_run.head_branch == 'main'))
|
runs-on: homelab
|
||||||
runs-on: [self-hosted, linux, arch, homelab, prod]
|
|
||||||
timeout-minutes: 10
|
timeout-minutes: 10
|
||||||
|
outputs:
|
||||||
|
sha: ${{ steps.release.outputs.sha }}
|
||||||
steps:
|
steps:
|
||||||
- name: Checkout repository
|
- name: Checkout repository
|
||||||
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
|
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
|
||||||
|
with:
|
||||||
- name: Fetch and reset workstation
|
fetch-depth: 0
|
||||||
shell: bash
|
- name: Check successful CI and download the exact commit release
|
||||||
|
id: release
|
||||||
|
env:
|
||||||
|
GITEA_TOKEN: ${{ github.token }}
|
||||||
|
DEPLOY_REF: ${{ inputs.deploy_ref || 'main' }}
|
||||||
|
EVENT_SHA: ${{ github.event.workflow_run.head_sha }}
|
||||||
|
run: python3 .gitea/workflows/release.py gate --ref "$DEPLOY_REF" --event-sha "$EVENT_SHA"
|
||||||
|
- name: Submit durable deploy to workstation
|
||||||
|
run: bash .gitea/workflows/ssh-run.sh start
|
||||||
|
- name: Write the request result
|
||||||
|
if: always()
|
||||||
|
env:
|
||||||
|
REQUEST_RESULT: ${{ job.status }}
|
||||||
|
CHECKED_SHA: ${{ steps.release.outputs.sha }}
|
||||||
run: |
|
run: |
|
||||||
set -euo pipefail
|
if [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
|
||||||
./.gitea/workflows/ssh-run.sh preflight
|
printf '## Deploy request\n\n- Result: **%s**\n- Checked commit: %s\n- Mode: %s\n' "$REQUEST_RESULT" "${CHECKED_SHA:-not checked}" "$DEPLOY_MODE" >>"$GITHUB_STEP_SUMMARY"
|
||||||
|
if [ "$REQUEST_RESULT" != success ]; then
|
||||||
|
echo 'Open the failed step log. If SSH submission failed, check the remote controller state.' >>"$GITHUB_STEP_SUMMARY"
|
||||||
|
fi
|
||||||
|
fi
|
||||||
|
|
||||||
validate:
|
apply:
|
||||||
needs: [preflight]
|
needs: [gate]
|
||||||
runs-on: [self-hosted, linux, arch, homelab, prod]
|
runs-on: homelab
|
||||||
timeout-minutes: 20
|
timeout-minutes: 120
|
||||||
steps:
|
steps:
|
||||||
- name: Checkout repository
|
- name: Checkout checked commit
|
||||||
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
|
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
|
||||||
|
with:
|
||||||
- name: Dry-run manifests and check Secrets
|
ref: ${{ needs.gate.outputs.sha }}
|
||||||
shell: bash
|
- name: Follow validation and sequential Kubernetes / Compose apply
|
||||||
|
run: bash .gitea/workflows/ssh-run.sh apply
|
||||||
|
- name: Write the deploy result
|
||||||
|
if: always()
|
||||||
run: |
|
run: |
|
||||||
set -euo pipefail
|
if [ -f .gitea/workflows/ssh-run.sh ]; then
|
||||||
./.gitea/workflows/ssh-run.sh validate
|
bash .gitea/workflows/ssh-run.sh summary
|
||||||
|
elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
|
||||||
|
echo 'Source checkout failed. The remote deploy state is unknown. Check the job log.' >>"$GITHUB_STEP_SUMMARY"
|
||||||
|
fi
|
||||||
|
|
||||||
apply-k8s:
|
verify:
|
||||||
needs: [validate]
|
needs: [gate, apply]
|
||||||
runs-on: [self-hosted, linux, arch, homelab, prod]
|
if: always() && needs.gate.result == 'success'
|
||||||
# Apply only, no verification, so this is just the work itself: snapshot,
|
runs-on: homelab
|
||||||
# then sequential `helm upgrade --install --wait --rollback-on-failure --timeout 10m`, then the apply loop.
|
timeout-minutes: 130
|
||||||
# Verification has its own job and its own budget.
|
|
||||||
#
|
|
||||||
# 45 is roughly four times the measured cost of the stage, which is
|
|
||||||
# deliberately not raised on a theory:
|
|
||||||
#
|
|
||||||
# helm, healthy 3 no-op upgrades ~3-5 min
|
|
||||||
# helm, one release bad rollback-on-failure spends its 10m, ~10-15 min
|
|
||||||
# then rolls that one back
|
|
||||||
# apply loop ~40 manifests, 4 of which ~1 min
|
|
||||||
# resolve an image digest
|
|
||||||
# restart_stale_images 7.6s to find 8 workloads, ~0.5 min
|
|
||||||
# 9.8s to resolve their digests
|
|
||||||
#
|
|
||||||
# The helm figure is one release, not three: `set -e` aborts
|
|
||||||
# upgrade_helm_releases on the first failure, so a broken release costs
|
|
||||||
# 10m and the other two are never attempted. Multiplying 10m by three
|
|
||||||
# overstates the worst case by 20 minutes.
|
|
||||||
#
|
|
||||||
# The 45 minutes this was last raised to 45 were still not enough, and the
|
|
||||||
# job logs for those runs no longer exist, so what actually consumed the
|
|
||||||
# budget is not known - the two measurable candidates above account for
|
|
||||||
# ~15 of it. The unbounded `docker manifest inspect` against the registry's
|
|
||||||
# known hang mode is now bounded inside registry_digest (25s timeout, 3
|
|
||||||
# attempts): a dead registry fails each owned image after ~85s instead of
|
|
||||||
# hanging the stage, and a blinking one is retried instead of failing the
|
|
||||||
# whole apply file. Still open: make the stage announce which manifest it
|
|
||||||
# is working on, so a killed run leaves a diagnosable last line.
|
|
||||||
timeout-minutes: 45
|
|
||||||
steps:
|
steps:
|
||||||
- name: Checkout repository
|
- name: Checkout checked commit
|
||||||
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
|
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
|
||||||
|
with:
|
||||||
- name: Apply Kubernetes manifests
|
ref: ${{ needs.gate.outputs.sha }}
|
||||||
shell: bash
|
- name: Follow workload verification and recovery
|
||||||
|
run: bash .gitea/workflows/ssh-run.sh verify
|
||||||
|
- name: Write the deploy result
|
||||||
|
if: always()
|
||||||
run: |
|
run: |
|
||||||
set -euo pipefail
|
if [ -f .gitea/workflows/ssh-run.sh ]; then
|
||||||
./.gitea/workflows/ssh-run.sh apply-k8s
|
bash .gitea/workflows/ssh-run.sh summary
|
||||||
|
elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
|
||||||
|
echo 'Source checkout failed. The remote deploy state is unknown. Check the job log.' >>"$GITHUB_STEP_SUMMARY"
|
||||||
|
fi
|
||||||
|
|
||||||
apply-compose:
|
|
||||||
needs: [validate]
|
|
||||||
runs-on: [self-hosted, linux, arch, homelab, prod]
|
|
||||||
timeout-minutes: 30
|
|
||||||
steps:
|
|
||||||
- name: Checkout repository
|
|
||||||
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
|
|
||||||
|
|
||||||
- name: Redeploy docker compose stacks
|
|
||||||
shell: bash
|
|
||||||
run: |
|
|
||||||
set -euo pipefail
|
|
||||||
./.gitea/workflows/ssh-run.sh apply-compose
|
|
||||||
|
|
||||||
# Watches the workloads this deploy changed and rolls back the ones that never
|
|
||||||
# became healthy. Runs even when the apply jobs failed, timed out or were
|
|
||||||
# cancelled — that is the whole point of splitting it out. `always()` is what
|
|
||||||
# lets it start after a failed dependency; the needs on apply-compose are a
|
|
||||||
# barrier, so verification begins only once both applies are done.
|
|
||||||
verify-k8s:
|
|
||||||
needs: [apply-k8s, apply-compose]
|
|
||||||
if: >-
|
|
||||||
always() &&
|
|
||||||
needs.apply-k8s.result != 'skipped' &&
|
|
||||||
needs.apply-compose.result != 'skipped'
|
|
||||||
runs-on: [self-hosted, linux, arch, homelab, prod]
|
|
||||||
# Not raised, because the arithmetic does not close.
|
|
||||||
#
|
|
||||||
# 32 workloads are under management and the wave width is 8, so the verify
|
|
||||||
# itself is 4 waves of ROLLOUT_TIMEOUT (300s) = 20 minutes worst case, when
|
|
||||||
# every rollout times out rather than converging. That is already 20 of 30.
|
|
||||||
#
|
|
||||||
# The other 10 would have to absorb rollback, and rollback_workloads is a
|
|
||||||
# serial `while read` loop at 300s per failed workload. 10 minutes buys two.
|
|
||||||
# Any larger number is buying a bigger multiple of an unbounded term rather
|
|
||||||
# than covering a known cost: 60 minutes buys eight, and 60 minutes is
|
|
||||||
# therefore not a bound, it is a guess with two digits.
|
|
||||||
#
|
|
||||||
# The number becomes derivable the moment rollback uses the same wave width
|
|
||||||
# as the verify: 32 failures then cost 4 waves = 20 minutes instead of 160,
|
|
||||||
# and 45 covers verify plus rollback at full width. That change is to the
|
|
||||||
# recovery path and is not folded into a timeout edit.
|
|
||||||
timeout-minutes: 30
|
|
||||||
steps:
|
|
||||||
- name: Checkout repository
|
|
||||||
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
|
|
||||||
|
|
||||||
- name: Verify workloads and roll back on failure
|
|
||||||
shell: bash
|
|
||||||
run: |
|
|
||||||
set -euo pipefail
|
|
||||||
./.gitea/workflows/ssh-run.sh verify-k8s
|
|
||||||
|
|
||||||
# Asks the public route of every active service whether it is actually
|
|
||||||
# serving, which the rollout check above structurally cannot: a pod can
|
|
||||||
# converge and still be crash-looping, or be listening on a port no Service
|
|
||||||
# points at, or answer 500.
|
|
||||||
#
|
|
||||||
# `always()` for the same reason verify-k8s has it, and it runs after that job
|
|
||||||
# specifically because a rollback is when a route most needs re-checking. The
|
|
||||||
# needs is a barrier, not a filter: whether verify-k8s passed, failed or was
|
|
||||||
# cancelled, the probes are what say whether the cluster is serving, and
|
|
||||||
# suppressing them on a rollback would hide the one run where the answer
|
|
||||||
# matters most.
|
|
||||||
smoke:
|
smoke:
|
||||||
needs: [verify-k8s]
|
needs: [gate, verify]
|
||||||
if: always() && needs.verify-k8s.result != 'skipped'
|
if: always() && needs.gate.result == 'success'
|
||||||
runs-on: [self-hosted, linux, arch, homelab, prod]
|
runs-on: homelab
|
||||||
timeout-minutes: 10
|
timeout-minutes: 15
|
||||||
steps:
|
steps:
|
||||||
- name: Checkout repository
|
- name: Checkout checked commit
|
||||||
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
|
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
|
||||||
|
with:
|
||||||
- name: Probe the public route of every active service
|
ref: ${{ needs.gate.outputs.sha }}
|
||||||
shell: bash
|
- name: Follow public route checks
|
||||||
|
run: bash .gitea/workflows/ssh-run.sh smoke
|
||||||
|
- name: Write the deploy result
|
||||||
|
if: always()
|
||||||
run: |
|
run: |
|
||||||
set -euo pipefail
|
if [ -f .gitea/workflows/ssh-run.sh ]; then
|
||||||
./.gitea/workflows/ssh-run.sh smoke
|
bash .gitea/workflows/ssh-run.sh summary
|
||||||
|
elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
|
||||||
|
echo 'Source checkout failed. The remote deploy state is unknown. Check the job log.' >>"$GITHUB_STEP_SUMMARY"
|
||||||
|
fi
|
||||||
@@ -13,9 +13,14 @@ here="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
|||||||
# shellcheck source=tool-versions.env
|
# shellcheck source=tool-versions.env
|
||||||
. "$here/tool-versions.env"
|
. "$here/tool-versions.env"
|
||||||
|
|
||||||
TOOLS_DIR="${TOOLS_DIR:-${RUNNER_TEMP:-/tmp}/homelab-tools}"
|
TOOLS_DIR="${TOOLS_DIR:-${XDG_CACHE_HOME:-$HOME/.cache}/homelab-ci}"
|
||||||
BIN_DIR="$TOOLS_DIR/bin"
|
BIN_DIR="$TOOLS_DIR/bin"
|
||||||
mkdir -p "$BIN_DIR"
|
mkdir -p "$BIN_DIR"
|
||||||
|
# A runner may accept overlapping workflows even though each workflow is sequential.
|
||||||
|
exec 9>"$TOOLS_DIR/install.lock"
|
||||||
|
flock -w 300 9
|
||||||
|
export UV_TOOL_DIR="$TOOLS_DIR/uv-tools"
|
||||||
|
export UV_CACHE_DIR="$TOOLS_DIR/uv-cache"
|
||||||
# The just-installed tools must resolve inside this script too: callers only
|
# The just-installed tools must resolve inside this script too: callers only
|
||||||
# prepend BIN_DIR to PATH after the script exits, so a bare `uv` below would
|
# prepend BIN_DIR to PATH after the script exits, so a bare `uv` below would
|
||||||
# miss the binary install_uv just placed (exit 127 on a clean runner).
|
# miss the binary install_uv just placed (exit 127 on a clean runner).
|
||||||
@@ -48,7 +53,7 @@ esac
|
|||||||
fetch() {
|
fetch() {
|
||||||
# fetch <url> <dest>
|
# fetch <url> <dest>
|
||||||
if command -v curl >/dev/null 2>&1; then
|
if command -v curl >/dev/null 2>&1; then
|
||||||
curl -sSLf --retry 3 -o "$2" "$1"
|
curl -sSLf --connect-timeout 15 --max-time 120 --retry 3 -o "$2" "$1"
|
||||||
elif command -v wget >/dev/null 2>&1; then
|
elif command -v wget >/dev/null 2>&1; then
|
||||||
wget -q -O "$2" "$1"
|
wget -q -O "$2" "$1"
|
||||||
else
|
else
|
||||||
@@ -88,10 +93,13 @@ installed_version() {
|
|||||||
|
|
||||||
# at_version <command> <expected>
|
# at_version <command> <expected>
|
||||||
at_version() {
|
at_version() {
|
||||||
case "$(installed_version "$1")" in
|
local version expected="${2#v}"
|
||||||
*"$2"*) return 0 ;;
|
version="$(installed_version "$1")"
|
||||||
*) return 1 ;;
|
if [[ "$version" =~ (^|[^0-9.])v?([0-9]+(\.[0-9]+)+) ]]; then
|
||||||
esac
|
[ "${BASH_REMATCH[2]}" = "$expected" ]
|
||||||
|
else
|
||||||
|
return 1
|
||||||
|
fi
|
||||||
}
|
}
|
||||||
|
|
||||||
install_kubeconform() {
|
install_kubeconform() {
|
||||||
@@ -176,6 +184,7 @@ install_pip_audit() {
|
|||||||
}
|
}
|
||||||
|
|
||||||
install_prettier() {
|
install_prettier() {
|
||||||
|
install_node
|
||||||
if at_version prettier "${PRETTIER_VERSION}"; then
|
if at_version prettier "${PRETTIER_VERSION}"; then
|
||||||
return 0
|
return 0
|
||||||
fi
|
fi
|
||||||
@@ -236,29 +245,41 @@ install_actionlint() {
|
|||||||
rm -rf "$tmp"
|
rm -rf "$tmp"
|
||||||
}
|
}
|
||||||
|
|
||||||
wanted=("$@")
|
main() {
|
||||||
if [ "${#wanted[@]}" -eq 0 ]; then
|
wanted=("$@")
|
||||||
wanted=(kubeconform shellcheck actionlint prettier ruff yamllint hadolint)
|
if [ "${#wanted[@]}" -eq 0 ]; then
|
||||||
fi
|
wanted=(node jq kubeconform shellcheck actionlint prettier ruff yamllint hadolint)
|
||||||
|
fi
|
||||||
|
|
||||||
for tool in "${wanted[@]}"; do
|
for tool in "${wanted[@]}"; do
|
||||||
case "$tool" in
|
case "$tool" in
|
||||||
kubeconform) install_kubeconform ;;
|
kubeconform) install_kubeconform ;;
|
||||||
shellcheck) install_shellcheck ;;
|
shellcheck) install_shellcheck ;;
|
||||||
jq) install_jq ;;
|
jq) install_jq ;;
|
||||||
actionlint) install_actionlint ;;
|
actionlint) install_actionlint ;;
|
||||||
prettier) install_prettier ;;
|
prettier) install_prettier ;;
|
||||||
ruff) install_ruff ;;
|
ruff) install_ruff ;;
|
||||||
yamllint) install_yamllint ;;
|
yamllint) install_yamllint ;;
|
||||||
pip-audit) install_pip_audit ;;
|
pip-audit) install_pip_audit ;;
|
||||||
hadolint) install_hadolint ;;
|
hadolint) install_hadolint ;;
|
||||||
node) install_node ;;
|
node) install_node ;;
|
||||||
uv) install_uv ;;
|
uv) install_uv ;;
|
||||||
*)
|
*)
|
||||||
echo "install-ci-tools: unknown tool: $tool" >&2
|
echo "install-ci-tools: unknown tool: $tool" >&2
|
||||||
exit 1
|
exit 1
|
||||||
;;
|
;;
|
||||||
esac
|
esac
|
||||||
done
|
done
|
||||||
|
|
||||||
printf '%s\n' "$BIN_DIR"
|
for old in "$BIN_DIR"/node-* "$BIN_DIR"/prettier-*; do
|
||||||
|
[ -d "$old" ] || continue
|
||||||
|
case "$(basename "$old")" in
|
||||||
|
"node-$NODE_VERSION"|"prettier-$PRETTIER_VERSION") ;;
|
||||||
|
*) rm -rf "$old" ;;
|
||||||
|
esac
|
||||||
|
done
|
||||||
|
if [ -x "$BIN_DIR/uv" ]; then "$BIN_DIR/uv" cache prune >/dev/null; fi
|
||||||
|
printf '%s\n' "$BIN_DIR"
|
||||||
|
}
|
||||||
|
|
||||||
|
if [ "${BASH_SOURCE[0]}" = "$0" ]; then main "$@"; fi
|
||||||
@@ -0,0 +1,516 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
"""CI release artifacts and the SHA-specific Gitea deployment gate (stdlib only)."""
|
||||||
|
|
||||||
|
import argparse
|
||||||
|
import hashlib
|
||||||
|
import io
|
||||||
|
import itertools
|
||||||
|
import json
|
||||||
|
import os
|
||||||
|
import re
|
||||||
|
import shutil
|
||||||
|
import subprocess
|
||||||
|
import sys
|
||||||
|
import tempfile
|
||||||
|
import urllib.error
|
||||||
|
import urllib.parse
|
||||||
|
import urllib.request
|
||||||
|
import zipfile
|
||||||
|
from pathlib import Path
|
||||||
|
|
||||||
|
SHA = re.compile(r'[0-9a-f]{40}')
|
||||||
|
DIGEST = re.compile(r'sha256:[0-9a-f]{64}')
|
||||||
|
IMAGES = {
|
||||||
|
'error-pages': ('errorpages', 'errorpages/Dockerfile'),
|
||||||
|
'forust-homepage': ('homepages', 'homepages/Dockerfile.forust'),
|
||||||
|
'xdfnx-homepage': ('homepages', 'homepages/Dockerfile.xdfnx'),
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def command(*args, **kwargs):
|
||||||
|
"""Arguments are passed directly to the executable, never to a shell."""
|
||||||
|
return subprocess.check_output(args, text=True, **kwargs).strip() # noqa: S603, S607
|
||||||
|
|
||||||
|
|
||||||
|
def validate_release(data, sha=None):
|
||||||
|
if data.get('version') != 1 or not SHA.fullmatch(data.get('sha', '')):
|
||||||
|
raise ValueError('Invalid release version or SHA')
|
||||||
|
if sha is not None and data['sha'] != sha:
|
||||||
|
raise ValueError('Release SHA does not match the checked CI commit')
|
||||||
|
expected = {f'gcr.forust.xyz/forust/{name}' for name in IMAGES}
|
||||||
|
if set(data.get('images', {})) != expected:
|
||||||
|
raise ValueError('Release must contain all owned images')
|
||||||
|
if not all(DIGEST.fullmatch(value) for value in data['images'].values()):
|
||||||
|
raise ValueError('Release has an invalid image digest')
|
||||||
|
if set(data.get('inputs', {})) != expected or not all(
|
||||||
|
re.fullmatch(r'[0-9a-f]{64}', value) for value in data['inputs'].values()
|
||||||
|
):
|
||||||
|
raise ValueError('Release has invalid build input fingerprints')
|
||||||
|
return data
|
||||||
|
|
||||||
|
|
||||||
|
class NoRedirect(urllib.request.HTTPRedirectHandler):
|
||||||
|
def redirect_request(self, _req, _fp, _code, _msg, _headers, _newurl):
|
||||||
|
return None
|
||||||
|
|
||||||
|
|
||||||
|
class Gitea:
|
||||||
|
def __init__(self):
|
||||||
|
self.origin = os.environ['GITHUB_SERVER_URL'].rstrip('/')
|
||||||
|
if urllib.parse.urlsplit(self.origin).scheme != 'https':
|
||||||
|
raise ValueError('Gitea API must use HTTPS')
|
||||||
|
self.repository = os.environ['GITHUB_REPOSITORY']
|
||||||
|
if not re.fullmatch(r'[\w.-]+/[\w.-]+', self.repository):
|
||||||
|
raise ValueError('Invalid Gitea repository')
|
||||||
|
self.token = os.environ['GITEA_TOKEN']
|
||||||
|
self.base = f'{self.origin}/api/v1/repos/{self.repository}'
|
||||||
|
|
||||||
|
def request(self, url, *, archive=False):
|
||||||
|
if not url.startswith(self.base + '/'):
|
||||||
|
raise ValueError('Refusing to send the Actions token to another origin')
|
||||||
|
req = urllib.request.Request(url, headers={'Authorization': f'token {self.token}'}) # noqa: S310 -- HTTPS origin validated above
|
||||||
|
opener = urllib.request.build_opener(NoRedirect())
|
||||||
|
try:
|
||||||
|
response = opener.open(req, timeout=30) # noqa: S310
|
||||||
|
except urllib.error.HTTPError as error:
|
||||||
|
if not archive or error.code not in (301, 302, 303, 307, 308):
|
||||||
|
raise RuntimeError(f'Gitea API returned HTTP {error.code}') from None
|
||||||
|
target = urllib.parse.urljoin(url, error.headers['Location'])
|
||||||
|
if urllib.parse.urlsplit(target).scheme != 'https':
|
||||||
|
raise ValueError('Artifact redirect must use HTTPS') from None
|
||||||
|
# Signed storage redirects must never receive the Gitea token.
|
||||||
|
response = urllib.request.urlopen(target, timeout=30) # noqa: S310
|
||||||
|
with response:
|
||||||
|
payload = response.read(8 * 1024 * 1024 + 1)
|
||||||
|
if len(payload) > 8 * 1024 * 1024:
|
||||||
|
raise ValueError('Gitea response exceeds 8 MiB')
|
||||||
|
return payload if archive else json.loads(payload)
|
||||||
|
|
||||||
|
def pages(self, path, key, **params):
|
||||||
|
for page in range(1, 101):
|
||||||
|
query = urllib.parse.urlencode({**params, 'page': page, 'limit': 50})
|
||||||
|
data = self.request(f'{self.base}/{path}?{query}')
|
||||||
|
entries = data[key]
|
||||||
|
yield from entries
|
||||||
|
if len(entries) < 50:
|
||||||
|
return
|
||||||
|
raise RuntimeError('Gitea pagination limit exceeded')
|
||||||
|
|
||||||
|
def successful_runs(self, sha=None):
|
||||||
|
params = {'branch': 'main', 'status': 'success', 'exclude_pull_requests': 'true'}
|
||||||
|
if sha:
|
||||||
|
params['head_sha'] = sha
|
||||||
|
for run in self.pages('actions/workflows/ci.yaml/runs', 'workflow_runs', **params):
|
||||||
|
if (
|
||||||
|
run.get('status') == 'completed'
|
||||||
|
and run.get('conclusion') == 'success'
|
||||||
|
and run.get('head_branch') == 'main'
|
||||||
|
and run.get('event') in ('push', 'workflow_dispatch')
|
||||||
|
and (run.get('repository') or {}).get('full_name') == self.repository
|
||||||
|
and (run.get('head_repository') or run.get('repository') or {}).get('full_name') == self.repository
|
||||||
|
and (sha is None or run.get('head_sha') == sha)
|
||||||
|
):
|
||||||
|
yield run
|
||||||
|
|
||||||
|
def release(self, run):
|
||||||
|
sha = run['head_sha']
|
||||||
|
jobs = list(self.pages(f'actions/runs/{run["id"]}/jobs', 'jobs'))
|
||||||
|
# A green workflow with a skipped build must not authorize a deploy.
|
||||||
|
if not any(job.get('name') == 'build' and job.get('conclusion') == 'success' for job in jobs):
|
||||||
|
raise ValueError('CI build job did not succeed')
|
||||||
|
artifacts = self.request(f'{self.base}/actions/runs/{run["id"]}/artifacts')['artifacts']
|
||||||
|
matching = [a for a in artifacts if a['name'] == f'release-{sha}' and not a.get('expired')]
|
||||||
|
if len(matching) != 1:
|
||||||
|
raise ValueError('CI release artifact is missing, expired or ambiguous; rerun CI')
|
||||||
|
blob = self.request(f'{self.base}/actions/artifacts/{matching[0]["id"]}/zip', archive=True)
|
||||||
|
with zipfile.ZipFile(io.BytesIO(blob)) as archive:
|
||||||
|
files = [entry for entry in archive.infolist() if not entry.is_dir()]
|
||||||
|
if len(files) != 1 or files[0].filename != 'release.json' or files[0].file_size > 256 * 1024:
|
||||||
|
raise ValueError('Unexpected release archive contents')
|
||||||
|
return validate_release(json.loads(archive.read(files[0])), sha)
|
||||||
|
|
||||||
|
|
||||||
|
def fingerprint(context, dockerfile):
|
||||||
|
tree = command('git', 'ls-tree', '-r', 'HEAD', '--', context, dockerfile, '.gitea/workflows/release.py')
|
||||||
|
return hashlib.sha256(tree.encode()).hexdigest()
|
||||||
|
|
||||||
|
|
||||||
|
def gate(output, requested_ref, event_sha):
|
||||||
|
command('git', 'fetch', '--quiet', 'origin', 'main')
|
||||||
|
if event_sha:
|
||||||
|
if not SHA.fullmatch(event_sha):
|
||||||
|
raise ValueError('Invalid workflow_run SHA')
|
||||||
|
sha = event_sha
|
||||||
|
else:
|
||||||
|
if requested_ref == 'main':
|
||||||
|
requested_ref = 'origin/main'
|
||||||
|
sha = command('git', 'rev-parse', '--verify', '--end-of-options', f'{requested_ref}^{{commit}}')
|
||||||
|
if not SHA.fullmatch(sha):
|
||||||
|
raise ValueError('Invalid deploy SHA')
|
||||||
|
command('git', 'merge-base', '--is-ancestor', sha, 'origin/main')
|
||||||
|
api = Gitea()
|
||||||
|
runs = list(api.successful_runs(sha))
|
||||||
|
if not runs:
|
||||||
|
raise ValueError(f'No successful main CI for {sha}; run CI before deploying')
|
||||||
|
release = api.release(max(runs, key=lambda run: run['id']))
|
||||||
|
output.write_text(json.dumps(release, indent=2) + '\n')
|
||||||
|
if os.environ.get('GITHUB_OUTPUT'):
|
||||||
|
with Path(os.environ['GITHUB_OUTPUT']).open('a') as stream:
|
||||||
|
stream.write(f'sha={sha}\n')
|
||||||
|
print(f'CI gate accepted {sha}')
|
||||||
|
|
||||||
|
|
||||||
|
def prepare_images(output):
|
||||||
|
sha = command('git', 'rev-parse', 'HEAD')
|
||||||
|
if sha != os.environ['GITHUB_SHA'] or not SHA.fullmatch(sha):
|
||||||
|
raise ValueError('Build checkout does not match GITHUB_SHA')
|
||||||
|
api = Gitea()
|
||||||
|
previous = None
|
||||||
|
for run in sorted(itertools.islice(api.successful_runs(), 50), key=lambda item: item['id'], reverse=True):
|
||||||
|
if str(run['id']) == os.environ.get('GITHUB_RUN_ID'):
|
||||||
|
continue
|
||||||
|
try:
|
||||||
|
previous = api.release(run)
|
||||||
|
break
|
||||||
|
except ValueError:
|
||||||
|
# Expired artifacts only cost a rebuild; mutable tags are never a fallback.
|
||||||
|
continue
|
||||||
|
targets = []
|
||||||
|
for name, (context, dockerfile) in IMAGES.items():
|
||||||
|
image = f'gcr.forust.xyz/forust/{name}'
|
||||||
|
inputs = fingerprint(context, dockerfile)
|
||||||
|
old_digest = (previous or {}).get('images', {}).get(image)
|
||||||
|
targets.append(
|
||||||
|
{
|
||||||
|
'name': name,
|
||||||
|
'image': image,
|
||||||
|
'context': context,
|
||||||
|
'dockerfile': dockerfile,
|
||||||
|
'inputs': inputs,
|
||||||
|
'reuse_digest': old_digest if (previous or {}).get('inputs', {}).get(image) == inputs else None,
|
||||||
|
}
|
||||||
|
)
|
||||||
|
output.write_text(json.dumps({'sha': sha, 'targets': targets}, indent=2) + '\n')
|
||||||
|
if os.environ.get('GITHUB_OUTPUT'):
|
||||||
|
with Path(os.environ['GITHUB_OUTPUT']).open('a') as stream:
|
||||||
|
stream.write('matrix=' + json.dumps({'include': targets}, separators=(',', ':')) + '\n')
|
||||||
|
print(f'Prepared {len(targets)} image jobs; {sum(t["reuse_digest"] is None for t in targets)} require builds')
|
||||||
|
|
||||||
|
|
||||||
|
def checked_plan(path):
|
||||||
|
data = json.loads(path.read_text())
|
||||||
|
sha = command('git', 'rev-parse', 'HEAD')
|
||||||
|
if data.get('sha') != sha or sha != os.environ['GITHUB_SHA'] or not SHA.fullmatch(sha):
|
||||||
|
raise ValueError('Image plan does not match the checked source commit')
|
||||||
|
targets = data.get('targets', [])
|
||||||
|
if sorted(t['name'] for t in targets) != sorted(IMAGES):
|
||||||
|
raise ValueError('Image plan must contain each owned image once')
|
||||||
|
for target in targets:
|
||||||
|
name = target['name']
|
||||||
|
context, dockerfile = IMAGES[name]
|
||||||
|
if (target['context'], target['dockerfile'], target['image']) != (
|
||||||
|
context,
|
||||||
|
dockerfile,
|
||||||
|
f'gcr.forust.xyz/forust/{name}',
|
||||||
|
) or target['inputs'] != fingerprint(context, dockerfile):
|
||||||
|
raise ValueError('Image plan has invalid build inputs')
|
||||||
|
if target['reuse_digest'] is not None and not DIGEST.fullmatch(target['reuse_digest']):
|
||||||
|
raise ValueError('Image plan has an invalid reuse digest')
|
||||||
|
return data
|
||||||
|
|
||||||
|
|
||||||
|
def build_images(output, report, name, plan):
|
||||||
|
data = checked_plan(plan)
|
||||||
|
sha = data['sha']
|
||||||
|
target = next(t for t in data['targets'] if t['name'] == name)
|
||||||
|
context, dockerfile = IMAGES[name]
|
||||||
|
docker_config = tempfile.mkdtemp(prefix='homelab-registry-')
|
||||||
|
builder_config = Path.home() / '.cache/homelab-ci/buildx'
|
||||||
|
builder_config.mkdir(parents=True, exist_ok=True)
|
||||||
|
env = {**os.environ, 'DOCKER_CONFIG': docker_config, 'BUILDX_CONFIG': str(builder_config)}
|
||||||
|
try:
|
||||||
|
report['phase'] = 'Registry login'
|
||||||
|
subprocess.run( # noqa: S603, S607
|
||||||
|
[
|
||||||
|
shutil.which('docker') or '/usr/bin/docker',
|
||||||
|
'login',
|
||||||
|
'gcr.forust.xyz',
|
||||||
|
'-u',
|
||||||
|
os.environ['REGISTRY_USERNAME'],
|
||||||
|
'--password-stdin',
|
||||||
|
],
|
||||||
|
input=os.environ['REGISTRY_PASSWORD'],
|
||||||
|
text=True,
|
||||||
|
check=True,
|
||||||
|
env=env,
|
||||||
|
)
|
||||||
|
report['phase'] = 'Prepare the builder'
|
||||||
|
builder = 'homelab-ci'
|
||||||
|
versions = dict(
|
||||||
|
re.findall(r'^([A-Z_]+)="([^"\n]+)"$', Path('.gitea/workflows/tool-versions.env').read_text(), re.MULTILINE)
|
||||||
|
)
|
||||||
|
image = versions['BUILDKIT_IMAGE']
|
||||||
|
signature = builder_config / 'homelab-ci-image'
|
||||||
|
exists = (
|
||||||
|
subprocess.run( # noqa: S603
|
||||||
|
[shutil.which('docker') or '/usr/bin/docker', 'buildx', 'inspect', builder],
|
||||||
|
capture_output=True,
|
||||||
|
env=env,
|
||||||
|
).returncode
|
||||||
|
== 0
|
||||||
|
)
|
||||||
|
if exists and (not signature.exists() or signature.read_text().strip() != image):
|
||||||
|
command('docker', 'buildx', 'rm', '--keep-state', builder, env=env)
|
||||||
|
exists = False
|
||||||
|
if not exists:
|
||||||
|
command(
|
||||||
|
'docker',
|
||||||
|
'buildx',
|
||||||
|
'create',
|
||||||
|
'--name',
|
||||||
|
builder,
|
||||||
|
'--driver',
|
||||||
|
'docker-container',
|
||||||
|
'--driver-opt',
|
||||||
|
f'image={image}',
|
||||||
|
'--buildkitd-config',
|
||||||
|
'.gitea/runner/buildkitd.toml',
|
||||||
|
env=env,
|
||||||
|
)
|
||||||
|
signature.write_text(image + '\n')
|
||||||
|
release = {'version': 1, 'sha': sha, 'images': {}, 'inputs': {}}
|
||||||
|
report['images'] = release['images']
|
||||||
|
report['phase'] = f'Build or reuse {name}'
|
||||||
|
report['current'] = name
|
||||||
|
image = f'gcr.forust.xyz/forust/{name}'
|
||||||
|
inputs = target['inputs']
|
||||||
|
old_digest = target['reuse_digest']
|
||||||
|
exists = False
|
||||||
|
if old_digest:
|
||||||
|
exists = (
|
||||||
|
subprocess.run( # noqa: S603, S607
|
||||||
|
[
|
||||||
|
shutil.which('docker') or '/usr/bin/docker',
|
||||||
|
'buildx',
|
||||||
|
'imagetools',
|
||||||
|
'inspect',
|
||||||
|
f'{image}@{old_digest}',
|
||||||
|
],
|
||||||
|
capture_output=True,
|
||||||
|
env=env,
|
||||||
|
timeout=60,
|
||||||
|
).returncode
|
||||||
|
== 0
|
||||||
|
)
|
||||||
|
if exists:
|
||||||
|
print(f'Reuse {name}: inputs unchanged')
|
||||||
|
digest = old_digest
|
||||||
|
report['reused'].append(name)
|
||||||
|
else:
|
||||||
|
print(f'Build {name}', flush=True)
|
||||||
|
metadata = Path(docker_config) / 'metadata.json'
|
||||||
|
command(
|
||||||
|
'docker',
|
||||||
|
'buildx',
|
||||||
|
'build',
|
||||||
|
'--builder',
|
||||||
|
builder,
|
||||||
|
'--platform',
|
||||||
|
'linux/amd64',
|
||||||
|
'--provenance=false',
|
||||||
|
'--cache-from',
|
||||||
|
f'type=registry,ref={image}:buildcache',
|
||||||
|
'--cache-to',
|
||||||
|
f'type=registry,ref={image}:buildcache,mode=max',
|
||||||
|
'--output',
|
||||||
|
f'type=image,name={image},push-by-digest=true,name-canonical=true,push=true',
|
||||||
|
'--metadata-file',
|
||||||
|
str(metadata),
|
||||||
|
'--file',
|
||||||
|
dockerfile,
|
||||||
|
context,
|
||||||
|
env=env,
|
||||||
|
)
|
||||||
|
digest = json.loads(metadata.read_text())['containerimage.digest']
|
||||||
|
report['built'].append(name)
|
||||||
|
release['images'][image] = digest
|
||||||
|
release['inputs'][image] = inputs
|
||||||
|
if not DIGEST.fullmatch(digest):
|
||||||
|
raise ValueError('Image job returned an invalid digest')
|
||||||
|
output.write_text(json.dumps(release, indent=2) + '\n')
|
||||||
|
report['current'] = None
|
||||||
|
report['phase'] = 'Release file saved'
|
||||||
|
finally:
|
||||||
|
# Cleanup errors must neither leak credentials nor mask the original build error.
|
||||||
|
try:
|
||||||
|
subprocess.run( # noqa: S603
|
||||||
|
[
|
||||||
|
shutil.which('docker') or '/usr/bin/docker',
|
||||||
|
'buildx',
|
||||||
|
'prune',
|
||||||
|
'--builder',
|
||||||
|
'homelab-ci',
|
||||||
|
'--force',
|
||||||
|
'--max-used-space',
|
||||||
|
'1gb',
|
||||||
|
],
|
||||||
|
env=env,
|
||||||
|
timeout=60,
|
||||||
|
)
|
||||||
|
except (OSError, subprocess.TimeoutExpired):
|
||||||
|
print('CI builder cache cleanup deferred', flush=True)
|
||||||
|
finally:
|
||||||
|
shutil.rmtree(docker_config)
|
||||||
|
|
||||||
|
|
||||||
|
def write_summary(lines):
|
||||||
|
path = os.environ.get('GITHUB_STEP_SUMMARY')
|
||||||
|
if path:
|
||||||
|
try:
|
||||||
|
with Path(path).open('a') as stream:
|
||||||
|
stream.write('\n'.join(lines) + '\n\n')
|
||||||
|
except OSError:
|
||||||
|
print('WARNING: cannot write the job summary')
|
||||||
|
|
||||||
|
|
||||||
|
def check_summary():
|
||||||
|
lines = [
|
||||||
|
f'## {os.environ["SUMMARY_CHECK"]}',
|
||||||
|
'',
|
||||||
|
f'- Commit: `{os.environ.get("GITHUB_SHA", "unknown")}`',
|
||||||
|
f'- Result: **{os.environ["SUMMARY_RESULT"]}**',
|
||||||
|
]
|
||||||
|
if os.environ.get('SUMMARY_FAILED_STEP'):
|
||||||
|
lines.append(f'- Failed step: {os.environ["SUMMARY_FAILED_STEP"]}')
|
||||||
|
if os.environ['SUMMARY_RESULT'] != 'success':
|
||||||
|
lines.append('- Open the failed step log for the error details.')
|
||||||
|
write_summary(lines)
|
||||||
|
|
||||||
|
|
||||||
|
def build(output, name, plan):
|
||||||
|
report = {'phase': 'Check the source commit', 'current': None, 'built': [], 'reused': [], 'images': {}}
|
||||||
|
result = 'failure'
|
||||||
|
try:
|
||||||
|
build_images(output, report, name, plan)
|
||||||
|
result = 'success'
|
||||||
|
finally:
|
||||||
|
lines = [
|
||||||
|
f'## Image release `{os.environ.get("GITHUB_SHA", "unknown")}`',
|
||||||
|
'',
|
||||||
|
f'- Result: **{result}**',
|
||||||
|
f'- Last stage: {report["phase"]}',
|
||||||
|
]
|
||||||
|
if result == 'failure':
|
||||||
|
lines.append('- No release from this build can be deployed. Open the failed step log.')
|
||||||
|
if report['current']:
|
||||||
|
lines.append(f'- Image at the failure: `{report["current"]}`')
|
||||||
|
for title, key in (('Built', 'built'), ('Reused from successful CI', 'reused')):
|
||||||
|
lines.extend(['', f'### {title}'])
|
||||||
|
lines.extend(f'- `{name}`' for name in report[key])
|
||||||
|
if not report[key]:
|
||||||
|
lines.append('- None')
|
||||||
|
lines.extend(['', '### Completed image digests'])
|
||||||
|
lines.extend(f'- `{image}@{digest}`' for image, digest in report['images'].items())
|
||||||
|
if not report['images']:
|
||||||
|
lines.append('- None')
|
||||||
|
write_summary(lines)
|
||||||
|
|
||||||
|
|
||||||
|
def render(stream, destination):
|
||||||
|
release = validate_release(json.loads(Path(os.environ['RELEASE_FILE']).read_text()), os.environ['DEPLOY_SHA'])
|
||||||
|
image_line = re.compile(
|
||||||
|
r"^(\s*(?:-\s*)?image:\s*)(['\"]?)(gcr\.forust\.xyz/forust/[\w.-]+)(?::[\w.-]+|@sha256:[0-9a-f]{64})\2(\s*(?:#.*)?)$"
|
||||||
|
)
|
||||||
|
rendered = []
|
||||||
|
for line in stream:
|
||||||
|
match = image_line.fullmatch(line.rstrip('\n'))
|
||||||
|
if match:
|
||||||
|
prefix, quote, image, tail = match.groups()
|
||||||
|
if image not in release['images']:
|
||||||
|
raise ValueError(f'Owned image missing from checked release: {image}')
|
||||||
|
line = f'{prefix}{quote}{image}@{release["images"][image]}{quote}{tail}\n'
|
||||||
|
elif re.match(r'\s*(?:-\s*)?image:', line) and 'gcr.forust.xyz/forust/' in line:
|
||||||
|
raise ValueError('Unsupported owned image syntax; refusing to apply a mutable tag')
|
||||||
|
rendered.append(line)
|
||||||
|
destination.writelines(rendered)
|
||||||
|
|
||||||
|
|
||||||
|
def finalize_images(output, fragments, plan):
|
||||||
|
data = checked_plan(plan)
|
||||||
|
sha = data['sha']
|
||||||
|
release = {'version': 1, 'sha': sha, 'images': {}, 'inputs': {}}
|
||||||
|
for name in IMAGES:
|
||||||
|
fragment = json.loads((fragments / f'image-{name}' / 'image.json').read_text())
|
||||||
|
image = f'gcr.forust.xyz/forust/{name}'
|
||||||
|
if fragment.get('sha') != sha or fragment.get('version') != 1 or set(fragment.get('images', {})) != {image}:
|
||||||
|
raise ValueError('Image job artifact is missing or belongs to another commit')
|
||||||
|
target = next(t for t in data['targets'] if t['name'] == name)
|
||||||
|
if fragment.get('inputs') != {image: target['inputs']}:
|
||||||
|
raise ValueError('Image artifact does not match the build plan')
|
||||||
|
release['images'].update(fragment['images'])
|
||||||
|
release['inputs'].update(fragment['inputs'])
|
||||||
|
validate_release(release, sha)
|
||||||
|
# Only a complete set of successful image jobs can publish the release tags.
|
||||||
|
docker_config = tempfile.mkdtemp(prefix='homelab-registry-')
|
||||||
|
env = {**os.environ, 'DOCKER_CONFIG': docker_config}
|
||||||
|
try:
|
||||||
|
subprocess.run( # noqa: S603, S607
|
||||||
|
[
|
||||||
|
shutil.which('docker') or '/usr/bin/docker',
|
||||||
|
'login',
|
||||||
|
'gcr.forust.xyz',
|
||||||
|
'-u',
|
||||||
|
os.environ['REGISTRY_USERNAME'],
|
||||||
|
'--password-stdin',
|
||||||
|
],
|
||||||
|
input=os.environ['REGISTRY_PASSWORD'],
|
||||||
|
text=True,
|
||||||
|
check=True,
|
||||||
|
env=env,
|
||||||
|
)
|
||||||
|
for image, digest in release['images'].items():
|
||||||
|
command(
|
||||||
|
'docker',
|
||||||
|
'buildx',
|
||||||
|
'imagetools',
|
||||||
|
'create',
|
||||||
|
'--prefer-index=false',
|
||||||
|
'--tag',
|
||||||
|
f'{image}:sha-{sha}',
|
||||||
|
f'{image}@{digest}',
|
||||||
|
env=env,
|
||||||
|
timeout=90,
|
||||||
|
)
|
||||||
|
output.write_text(json.dumps(release, indent=2) + '\n')
|
||||||
|
finally:
|
||||||
|
shutil.rmtree(docker_config)
|
||||||
|
|
||||||
|
|
||||||
|
def main():
|
||||||
|
parser = argparse.ArgumentParser(description=__doc__)
|
||||||
|
parser.add_argument('action', choices=('prepare', 'image', 'finalize', 'gate', 'render', 'check-summary'))
|
||||||
|
parser.add_argument('--output', type=Path, default=Path('release.json'))
|
||||||
|
parser.add_argument('--ref', default='main')
|
||||||
|
parser.add_argument('--event-sha', default='')
|
||||||
|
parser.add_argument('--image', choices=IMAGES)
|
||||||
|
parser.add_argument('--plan', type=Path, default=Path('build-plan.json'))
|
||||||
|
parser.add_argument('--fragments', type=Path, default=Path('artifacts'))
|
||||||
|
args = parser.parse_args()
|
||||||
|
if args.action == 'check-summary':
|
||||||
|
check_summary()
|
||||||
|
elif args.action == 'render':
|
||||||
|
render(sys.stdin, sys.stdout)
|
||||||
|
elif args.action == 'gate':
|
||||||
|
gate(args.output, args.ref, args.event_sha)
|
||||||
|
elif args.action == 'prepare':
|
||||||
|
prepare_images(args.output)
|
||||||
|
elif args.action == 'image':
|
||||||
|
if not args.image:
|
||||||
|
parser.error('--image is required')
|
||||||
|
build(args.output, args.image, args.plan)
|
||||||
|
else:
|
||||||
|
finalize_images(args.output, args.fragments, args.plan)
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == '__main__':
|
||||||
|
main()
|
||||||
@@ -1,10 +1,26 @@
|
|||||||
name: renovate-ci
|
name: renovate-ci
|
||||||
|
|
||||||
on:
|
on:
|
||||||
pull_request:
|
# Read the workflow from the trusted base branch. PR code runs only on the
|
||||||
|
# unprivileged runner selected below.
|
||||||
|
pull_request_target:
|
||||||
|
paths:
|
||||||
|
- "renovate/**"
|
||||||
|
- ".gitea/workflows/renovate-ci.yaml"
|
||||||
|
- ".gitea/workflows/sync-renovate-configmap.sh"
|
||||||
|
- ".gitea/workflows/compose-lint.sh"
|
||||||
|
- ".gitea/workflows/install-ci-tools.sh"
|
||||||
|
- ".gitea/workflows/tool-versions.env"
|
||||||
push:
|
push:
|
||||||
branches:
|
branches:
|
||||||
- main
|
- main
|
||||||
|
paths:
|
||||||
|
- "renovate/**"
|
||||||
|
- ".gitea/workflows/renovate-ci.yaml"
|
||||||
|
- ".gitea/workflows/sync-renovate-configmap.sh"
|
||||||
|
- ".gitea/workflows/compose-lint.sh"
|
||||||
|
- ".gitea/workflows/install-ci-tools.sh"
|
||||||
|
- ".gitea/workflows/tool-versions.env"
|
||||||
workflow_dispatch:
|
workflow_dispatch:
|
||||||
|
|
||||||
permissions:
|
permissions:
|
||||||
@@ -12,37 +28,47 @@ permissions:
|
|||||||
|
|
||||||
jobs:
|
jobs:
|
||||||
validate-renovate:
|
validate-renovate:
|
||||||
runs-on: [self-hosted, linux, arch, homelab]
|
runs-on: ${{ github.event_name == 'push' && github.ref == 'refs/heads/main' && 'homelab' || 'homelab-pr' }}
|
||||||
timeout-minutes: 20
|
timeout-minutes: 20
|
||||||
steps:
|
steps:
|
||||||
- name: Checkout repository
|
- name: Checkout repository
|
||||||
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
|
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
|
||||||
|
with:
|
||||||
|
ref: ${{ github.event_name == 'pull_request_target' && github.event.pull_request.head.sha || github.sha }}
|
||||||
|
|
||||||
# renovate/k8s/cronjob.yaml is the single source of truth for the image tag,
|
# renovate/k8s/cronjob.yaml is the single source of truth for the version.
|
||||||
# so the same version that runs in the cluster is the one validated here.
|
- name: Resolve the deployed Renovate version
|
||||||
- name: Resolve the deployed Renovate image
|
|
||||||
id: image
|
id: image
|
||||||
shell: bash
|
shell: bash
|
||||||
run: |
|
run: |
|
||||||
set -euo pipefail
|
set -euo pipefail
|
||||||
image="$(sed -n 's|.*image:[[:space:]]*\(renovate/renovate:[^[:space:]]*\).*|\1|p' \
|
image="$(sed -n 's|.*image:[[:space:]]*\(renovate/renovate:[^[:space:]]*\).*|\1|p' \
|
||||||
renovate/k8s/cronjob.yaml | head -1)"
|
renovate/k8s/cronjob.yaml | head -1)"
|
||||||
if [ -z "$image" ]; then
|
if [[ ! "$image" =~ ^renovate/renovate:([0-9]+\.[0-9]+\.[0-9]+)$ ]]; then
|
||||||
echo "::error::no renovate/renovate image found in renovate/k8s/cronjob.yaml"
|
echo "::error::expected a pinned renovate/renovate semantic version in renovate/k8s/cronjob.yaml"
|
||||||
exit 1
|
exit 1
|
||||||
fi
|
fi
|
||||||
echo "using $image"
|
version="${BASH_REMATCH[1]}"
|
||||||
echo "image=$image" >> "$GITHUB_OUTPUT"
|
echo "using Renovate $version"
|
||||||
|
printf 'version=%s\n' "$version" >> "$GITHUB_OUTPUT"
|
||||||
|
|
||||||
- name: Validate Renovate repository config
|
- name: Prepare pinned validation tools
|
||||||
shell: bash
|
shell: bash
|
||||||
run: |
|
run: |
|
||||||
set -euo pipefail
|
set -euo pipefail
|
||||||
docker run --rm \
|
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh kubeconform node)"
|
||||||
-v "$PWD/renovate:/opt/renovate:ro" \
|
echo "$tools_dir" >> "$GITHUB_PATH"
|
||||||
-e RENOVATE_CONFIG_FILE=/opt/renovate/renovate.json \
|
|
||||||
"${{ steps.image.outputs.image }}" \
|
- name: Validate Renovate repository config
|
||||||
renovate-config-validator /opt/renovate/renovate.json
|
shell: bash
|
||||||
|
env:
|
||||||
|
RENOVATE_VERSION: ${{ steps.image.outputs.version }}
|
||||||
|
run: |
|
||||||
|
set -euo pipefail
|
||||||
|
npm_cache="$(mktemp -d "${RUNNER_TEMP:-/tmp}/renovate-npm-cache.XXXXXXXX")"
|
||||||
|
trap 'rm -rf "$npm_cache"' EXIT
|
||||||
|
NPM_CONFIG_CACHE="$npm_cache" RENOVATE_CONFIG_FILE="$PWD/renovate/renovate.json" \
|
||||||
|
npm exec --yes --package="renovate@${RENOVATE_VERSION}" -- renovate-config-validator
|
||||||
|
|
||||||
# The CronJob cannot read the repository, so renovate/k8s/configmap.yaml
|
# The CronJob cannot read the repository, so renovate/k8s/configmap.yaml
|
||||||
# carries an inlined copy of the config. Fail if it no longer matches.
|
# carries an inlined copy of the config. Fail if it no longer matches.
|
||||||
@@ -56,8 +82,6 @@ jobs:
|
|||||||
shell: bash
|
shell: bash
|
||||||
run: |
|
run: |
|
||||||
set -euo pipefail
|
set -euo pipefail
|
||||||
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh kubeconform)"
|
|
||||||
export PATH="$tools_dir:$PATH"
|
|
||||||
kubeconform \
|
kubeconform \
|
||||||
-strict \
|
-strict \
|
||||||
-ignore-missing-schemas \
|
-ignore-missing-schemas \
|
||||||
|
|||||||
@@ -32,11 +32,14 @@ concurrency:
|
|||||||
|
|
||||||
jobs:
|
jobs:
|
||||||
run-renovate:
|
run-renovate:
|
||||||
runs-on: [self-hosted, linux, arch, homelab]
|
if: github.ref == 'refs/heads/main'
|
||||||
|
runs-on: homelab
|
||||||
timeout-minutes: 60
|
timeout-minutes: 60
|
||||||
steps:
|
steps:
|
||||||
- name: Checkout repository
|
- name: Checkout repository
|
||||||
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
|
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
|
||||||
|
with:
|
||||||
|
ref: refs/heads/main
|
||||||
|
|
||||||
# renovate/k8s/cronjob.yaml is the single source of truth for the image tag.
|
# renovate/k8s/cronjob.yaml is the single source of truth for the image tag.
|
||||||
# Reading it here means this workflow validates and runs the exact version
|
# Reading it here means this workflow validates and runs the exact version
|
||||||
@@ -48,21 +51,23 @@ jobs:
|
|||||||
set -euo pipefail
|
set -euo pipefail
|
||||||
image="$(sed -n 's|.*image:[[:space:]]*\(renovate/renovate:[^[:space:]]*\).*|\1|p' \
|
image="$(sed -n 's|.*image:[[:space:]]*\(renovate/renovate:[^[:space:]]*\).*|\1|p' \
|
||||||
renovate/k8s/cronjob.yaml | head -1)"
|
renovate/k8s/cronjob.yaml | head -1)"
|
||||||
if [ -z "$image" ]; then
|
if [[ ! "$image" =~ ^renovate/renovate:[0-9]+\.[0-9]+\.[0-9]+$ ]]; then
|
||||||
echo "::error::no renovate/renovate image found in renovate/k8s/cronjob.yaml"
|
echo "::error::expected a pinned renovate/renovate semantic version in renovate/k8s/cronjob.yaml"
|
||||||
exit 1
|
exit 1
|
||||||
fi
|
fi
|
||||||
echo "using $image"
|
echo "using $image"
|
||||||
echo "image=$image" >> "$GITHUB_OUTPUT"
|
printf 'image=%s\n' "$image" >> "$GITHUB_OUTPUT"
|
||||||
|
|
||||||
- name: Validate Renovate config
|
- name: Validate Renovate config
|
||||||
shell: bash
|
shell: bash
|
||||||
|
env:
|
||||||
|
RENOVATE_IMAGE: ${{ steps.image.outputs.image }}
|
||||||
run: |
|
run: |
|
||||||
set -euo pipefail
|
set -euo pipefail
|
||||||
docker run --rm \
|
docker run --rm \
|
||||||
-v "$PWD/renovate/renovate.json:/opt/renovate/renovate.json:ro" \
|
-v "$PWD/renovate/renovate.json:/opt/renovate/renovate.json:ro" \
|
||||||
-e RENOVATE_CONFIG_FILE=/opt/renovate/renovate.json \
|
-e RENOVATE_CONFIG_FILE=/opt/renovate/renovate.json \
|
||||||
"${{ steps.image.outputs.image }}" \
|
"$RENOVATE_IMAGE" \
|
||||||
renovate-config-validator
|
renovate-config-validator
|
||||||
|
|
||||||
- name: Run Renovate
|
- name: Run Renovate
|
||||||
@@ -73,6 +78,7 @@ jobs:
|
|||||||
RENOVATE_REPOSITORIES: ${{ inputs.repositories }}
|
RENOVATE_REPOSITORIES: ${{ inputs.repositories }}
|
||||||
RENOVATE_DRY_RUN: ${{ inputs.dry_run && 'full' || '' }}
|
RENOVATE_DRY_RUN: ${{ inputs.dry_run && 'full' || '' }}
|
||||||
LOG_LEVEL: ${{ inputs.log_level }}
|
LOG_LEVEL: ${{ inputs.log_level }}
|
||||||
|
RENOVATE_IMAGE: ${{ steps.image.outputs.image }}
|
||||||
run: |
|
run: |
|
||||||
set -euo pipefail
|
set -euo pipefail
|
||||||
|
|
||||||
@@ -89,4 +95,4 @@ jobs:
|
|||||||
-e RENOVATE_CONFIG_FILE=/opt/renovate/renovate.json \
|
-e RENOVATE_CONFIG_FILE=/opt/renovate/renovate.json \
|
||||||
-e RENOVATE_BASE_DIR=/tmp/renovate \
|
-e RENOVATE_BASE_DIR=/tmp/renovate \
|
||||||
-e LOG_LEVEL="${LOG_LEVEL:-info}" \
|
-e LOG_LEVEL="${LOG_LEVEL:-info}" \
|
||||||
"${{ steps.image.outputs.image }}"
|
"$RENOVATE_IMAGE"
|
||||||
+62
-63
@@ -1,71 +1,70 @@
|
|||||||
#!/usr/bin/env bash
|
#!/usr/bin/env bash
|
||||||
# usage: ssh-run.sh <stage>
|
# The SSH client submits once and follows durable stages on workstation.
|
||||||
# Runs one deploy-lib.sh stage on the workstation over SSH.
|
|
||||||
set -euo pipefail
|
set -euo pipefail
|
||||||
|
|
||||||
: "${DEPLOY_HOST:?missing DEPLOY_HOST}"
|
: "${DEPLOY_HOST:?missing DEPLOY_HOST}"
|
||||||
: "${DEPLOY_USER:?missing DEPLOY_USER}"
|
: "${DEPLOY_USER:?missing DEPLOY_USER}"
|
||||||
: "${DEPLOY_KEY:?missing DEPLOY_SSH_KEY}"
|
: "${DEPLOY_KEY:?missing DEPLOY_SSH_KEY}"
|
||||||
|
: "${DEPLOY_KNOWN_HOSTS:?configure pinned DEPLOY_KNOWN_HOSTS}"
|
||||||
deploy_port="${DEPLOY_PORT:-22}"
|
: "${DEPLOY_RUN_ID:?missing DEPLOY_RUN_ID}"
|
||||||
deploy_path="${DEPLOY_PATH:-/srv/homelab}"
|
[[ "$DEPLOY_USER" =~ ^[A-Za-z_][A-Za-z0-9_.-]*$ ]] || exit 1
|
||||||
deploy_path="$(printf '%s' "$deploy_path" | tr -d '\"' | tr -d '\r' | xargs)"
|
[[ "$DEPLOY_HOST" =~ ^[A-Za-z0-9_.:-]+$ ]] || exit 1
|
||||||
|
[[ "$DEPLOY_RUN_ID" =~ ^[0-9]+-[0-9]+$ ]] || exit 1
|
||||||
# The private key is written to a per-run directory that is removed on exit, so a
|
[[ "${DEPLOY_PORT:-22}" =~ ^[0-9]+$ ]] || exit 1
|
||||||
# failed or cancelled job cannot leave deploy credentials in the runner's temp
|
|
||||||
# directory. Do not use a fixed path: apply-k8s and apply-compose run in parallel.
|
|
||||||
key_dir="$(mktemp -d "${RUNNER_TEMP:-/tmp}/homelab-deploy-key.XXXXXXXX")"
|
key_dir="$(mktemp -d "${RUNNER_TEMP:-/tmp}/homelab-deploy-key.XXXXXXXX")"
|
||||||
trap 'rm -rf "$key_dir"' EXIT INT TERM
|
trap 'rm -rf "$key_dir"' EXIT
|
||||||
|
chmod 700 "$key_dir"
|
||||||
ssh_key="$key_dir/deploy_key"
|
printf '%s\n' "$DEPLOY_KEY" >"$key_dir/key"
|
||||||
printf '%s\n' "$DEPLOY_KEY" > "$ssh_key"
|
printf '%s\n' "$DEPLOY_KNOWN_HOSTS" >"$key_dir/known_hosts"
|
||||||
chmod 600 "$ssh_key"
|
chmod 600 "$key_dir/key" "$key_dir/known_hosts"
|
||||||
|
ssh_opts=(-i "$key_dir/key" -p "${DEPLOY_PORT:-22}" -o BatchMode=yes -o StrictHostKeyChecking=yes
|
||||||
# A connection that died silently used to hang until the job timeout, and the
|
-o "UserKnownHostsFile=$key_dir/known_hosts" -o ConnectTimeout=15
|
||||||
# stage was never re-run: one flaky TCP session cost a whole 45-minute apply.
|
-o ServerAliveInterval=15 -o ServerAliveCountMax=4)
|
||||||
# ServerAlive* bounds how long a dead peer goes unnoticed, ConnectTimeout bounds
|
controller=.local/lib/homelab-deploy/controller.py
|
||||||
# setup. Only exit 255 - ssh's own transport failures - is retried. A stage that
|
case "${1:?start, apply, verify, smoke or summary required}" in
|
||||||
# fails on its own merits exits with the remote's status, so a real failure
|
start)
|
||||||
# still surfaces its own log instead of burning three attempts. The stages are
|
python3 - <<'PY' >"$key_dir/request.json"
|
||||||
# declarative applies, so re-running one that had already committed is harmless.
|
import json
|
||||||
ssh_opts=(
|
import os
|
||||||
-i "$ssh_key" -p "$deploy_port"
|
from pathlib import Path
|
||||||
-o BatchMode=yes -o StrictHostKeyChecking=accept-new
|
release = json.loads(Path('release.json').read_text())
|
||||||
-o ConnectTimeout=15
|
print(json.dumps({'release': release, 'mode': os.environ.get('DEPLOY_MODE', 'changed'),
|
||||||
-o ServerAliveInterval=15 -o ServerAliveCountMax=4
|
'refresh_images': os.environ.get('REFRESH_IMAGES', 'false') == 'true'}))
|
||||||
)
|
PY
|
||||||
|
for attempt in 1 2 3; do
|
||||||
rc=0
|
rc=0
|
||||||
# apply-k8s and apply-compose are separate workflow jobs so the graph stays
|
# shellcheck disable=SC2029 # The run ID and operation are validated local arguments, not remote variables.
|
||||||
# intact for the verify job, but on a single node they must not run at once:
|
ssh "${ssh_opts[@]}" "$DEPLOY_USER@$DEPLOY_HOST" python3 "$controller" start "$DEPLOY_RUN_ID" <"$key_dir/request.json" || rc=$?
|
||||||
# host docker churn on top of cluster churn is what melts the node (load 40+,
|
[ "$rc" -eq 0 ] && exit 0
|
||||||
# netbird/ssh die, helm is left pending-*). Serialize them on the workstation
|
[ "$rc" -eq 255 ] || exit "$rc"
|
||||||
# with a shared lock; whoever arrives second waits.
|
sleep 5
|
||||||
remote_cmd=(bash -se)
|
done
|
||||||
case "$1" in
|
exit "$rc"
|
||||||
apply-k8s | apply-compose)
|
|
||||||
remote_cmd=(flock -w 5400 /tmp/homelab-apply.lock bash -se)
|
|
||||||
;;
|
;;
|
||||||
|
apply|verify|smoke)
|
||||||
|
result=0
|
||||||
|
for attempt in 1 2 3; do
|
||||||
|
rc=0
|
||||||
|
# shellcheck disable=SC2029 # The run ID and operation are validated local arguments, not remote variables.
|
||||||
|
ssh "${ssh_opts[@]}" "$DEPLOY_USER@$DEPLOY_HOST" python3 "$controller" follow "$DEPLOY_RUN_ID" "$1" || rc=$?
|
||||||
|
[ "$rc" -eq 0 ] && break
|
||||||
|
[ "$rc" -eq 255 ] || { result="$rc"; break; }
|
||||||
|
echo "SSH disconnected; reconnecting to the existing deploy ($attempt/3)"
|
||||||
|
if [ "$attempt" -eq 3 ]; then result=255; break; fi
|
||||||
|
sleep 5
|
||||||
|
done
|
||||||
|
exit "$result"
|
||||||
|
;;
|
||||||
|
summary)
|
||||||
|
if [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
|
||||||
|
rc=0
|
||||||
|
# shellcheck disable=SC2029 # The run ID is validated above.
|
||||||
|
ssh "${ssh_opts[@]}" "$DEPLOY_USER@$DEPLOY_HOST" python3 "$controller" summary "$DEPLOY_RUN_ID" >"$key_dir/deploy-summary.md" || rc=$?
|
||||||
|
if [ "$rc" -eq 0 ]; then
|
||||||
|
cat "$key_dir/deploy-summary.md" >>"$GITHUB_STEP_SUMMARY" || echo "WARNING: cannot write the deploy summary"
|
||||||
|
else
|
||||||
|
echo 'Deploy summary is unavailable. The SSH connection failed or the controller did not respond. Check the job log.' >>"$GITHUB_STEP_SUMMARY" || true
|
||||||
|
fi
|
||||||
|
fi
|
||||||
|
;;
|
||||||
|
*) echo "Unknown SSH operation: $1" >&2; exit 1 ;;
|
||||||
esac
|
esac
|
||||||
for attempt in 1 2 3; do
|
|
||||||
if [ "$attempt" -gt 1 ]; then
|
|
||||||
echo ":: warning::ssh transport failed, retrying (${attempt}/3)"
|
|
||||||
sleep $((attempt * 5))
|
|
||||||
fi
|
|
||||||
rc=0
|
|
||||||
# shellcheck disable=SC2029 # remote_cmd/ssh_opts expand on the client on purpose: they select the local ssh invocation, only the heredoc runs remotely.
|
|
||||||
ssh "${ssh_opts[@]}" "${DEPLOY_USER}@${DEPLOY_HOST}" \
|
|
||||||
env "REPO=$deploy_path" "APPLY_PRUNE=${APPLY_PRUNE:-false}" \
|
|
||||||
"DEPLOY_SHA=${DEPLOY_SHA:-}" "DEPLOY_SNAPSHOT_DIR=${DEPLOY_SNAPSHOT_DIR:-}" \
|
|
||||||
"STAGE=$1" "${remote_cmd[@]}" <<'EOF' || rc=$?
|
|
||||||
source "$REPO/.gitea/workflows/deploy-lib.sh"
|
|
||||||
run_stage "$STAGE"
|
|
||||||
EOF
|
|
||||||
[ "$rc" -eq 0 ] && break
|
|
||||||
[ "$rc" -ne 255 ] && break
|
|
||||||
done
|
|
||||||
|
|
||||||
if [ "$rc" -ne 0 ]; then
|
|
||||||
echo ":: error::stage $1 failed over ssh (exit $rc)"
|
|
||||||
fi
|
|
||||||
exit "$rc"
|
|
||||||
@@ -34,3 +34,6 @@ NODE_VERSION="22.23.3"
|
|||||||
|
|
||||||
# Secret-reference regression tests parse rendered Kubernetes objects.
|
# Secret-reference regression tests parse rendered Kubernetes objects.
|
||||||
JQ_VERSION="1.8.1"
|
JQ_VERSION="1.8.1"
|
||||||
|
|
||||||
|
# BuildKit is the only auxiliary CI container; jobs themselves stay on the host.
|
||||||
|
BUILDKIT_IMAGE="moby/buildkit:v0.33.1"
|
||||||
@@ -94,7 +94,6 @@ replacements.txt
|
|||||||
.idea
|
.idea
|
||||||
|
|
||||||
# Temp files
|
# Temp files
|
||||||
edu_master/temp/
|
|
||||||
temp/*
|
temp/*
|
||||||
# Local-only tooling scratch space (pinned CI tools, verification scripts)
|
# Local-only tooling scratch space (pinned CI tools, verification scripts)
|
||||||
tmp/
|
tmp/
|
||||||
|
|||||||
@@ -31,7 +31,7 @@ services:
|
|||||||
- "traefik.http.routers.adguard-dev.entrypoints=websecure"
|
- "traefik.http.routers.adguard-dev.entrypoints=websecure"
|
||||||
- "traefik.http.routers.adguard-dev.tls=true"
|
- "traefik.http.routers.adguard-dev.tls=true"
|
||||||
# DoH Router
|
# DoH Router
|
||||||
- "traefik.http.routers.dns-over-https.rule=(Host(`dns.forust.xyz` || Host(`adguard.forust.xyz`)) && PathPrefix(`/dns-query`))"
|
- "traefik.http.routers.dns-over-https.rule=(Host(`dns.forust.xyz`) || Host(`adguard.forust.xyz`)) && PathPrefix(`/dns-query`)"
|
||||||
- "traefik.http.routers.dns-over-https.entrypoints=websecure"
|
- "traefik.http.routers.dns-over-https.entrypoints=websecure"
|
||||||
- "traefik.http.routers.dns-over-https.tls.certresolver=letsencrypt"
|
- "traefik.http.routers.dns-over-https.tls.certresolver=letsencrypt"
|
||||||
|
|
||||||
|
|||||||
@@ -20,7 +20,7 @@ spec:
|
|||||||
spec:
|
spec:
|
||||||
containers:
|
containers:
|
||||||
- name: cloudflared
|
- name: cloudflared
|
||||||
image: cloudflare/cloudflared:2026.9.3
|
image: cloudflare/cloudflared:2026.10.0
|
||||||
imagePullPolicy: IfNotPresent
|
imagePullPolicy: IfNotPresent
|
||||||
args:
|
args:
|
||||||
- tunnel
|
- tunnel
|
||||||
|
|||||||
@@ -1,14 +0,0 @@
|
|||||||
EDU_LOGIN=your_edu_login_here
|
|
||||||
EDU_PASSWORD=your_edu_password_here
|
|
||||||
EDU_URL_LOGIN=https://edu.edu.vn.ua/user/login
|
|
||||||
EDU_URL_VERIFY=https://edu.edu.vn.ua/course/userlist
|
|
||||||
PHPSESSID_INTERVAL=10
|
|
||||||
USER_AGENT="Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/142.0.0.0 Safari/537.36"
|
|
||||||
WEBINAR_URL=https://edu.edu.vn.ua/webinar/useractive
|
|
||||||
WEBINAR_CHECK_INTERVAL=60
|
|
||||||
REDIS_HOST=redis
|
|
||||||
REDIS_PORT=6379
|
|
||||||
PLAYWRIGHT_WS=ws://playwright-service:3000/ws
|
|
||||||
TZ=Europe/Kyiv
|
|
||||||
WEBINAR_TELEGRAM_TOKEN=your_telegram_bot_token_here
|
|
||||||
WEBINAR_ADMIN_ID=123456789
|
|
||||||
@@ -1 +0,0 @@
|
|||||||
1.56.0
|
|
||||||
@@ -1,49 +0,0 @@
|
|||||||
services:
|
|
||||||
redis:
|
|
||||||
image: redis:8.10.2-alpine
|
|
||||||
restart: unless-stopped
|
|
||||||
volumes:
|
|
||||||
- redis-data:/data
|
|
||||||
healthcheck:
|
|
||||||
test: ["CMD", "redis-cli", "ping"]
|
|
||||||
interval: 5s
|
|
||||||
timeout: 3s
|
|
||||||
retries: 5
|
|
||||||
|
|
||||||
playwright-service:
|
|
||||||
image: mcr.microsoft.com/playwright:v1.56.0-jammy
|
|
||||||
restart: unless-stopped
|
|
||||||
command: npx -y playwright@1.56.0 run-server --port 3000 --path /ws
|
|
||||||
|
|
||||||
session-keeper:
|
|
||||||
build: ./phpsessid-bot
|
|
||||||
image: gcr.forust.xyz/forust/session-keeper:prod
|
|
||||||
pull_policy: build
|
|
||||||
env_file: .env
|
|
||||||
restart: unless-stopped
|
|
||||||
depends_on:
|
|
||||||
redis:
|
|
||||||
condition: service_healthy
|
|
||||||
healthcheck:
|
|
||||||
test: ["CMD-SHELL", "redis-cli -h redis EXISTS EDU_PHPSESSID | grep -q 1"]
|
|
||||||
interval: 30s
|
|
||||||
timeout: 5s
|
|
||||||
retries: 10
|
|
||||||
start_period: 60s
|
|
||||||
|
|
||||||
webinar-checker:
|
|
||||||
build: ./webinar-checker
|
|
||||||
image: gcr.forust.xyz/forust/webinar-checker:prod
|
|
||||||
pull_policy: build
|
|
||||||
env_file: .env
|
|
||||||
restart: unless-stopped
|
|
||||||
depends_on:
|
|
||||||
redis:
|
|
||||||
condition: service_healthy
|
|
||||||
session-keeper:
|
|
||||||
condition: service_healthy
|
|
||||||
playwright-service:
|
|
||||||
condition: service_started
|
|
||||||
|
|
||||||
volumes:
|
|
||||||
redis-data:
|
|
||||||
Whitespace-only changes.
@@ -1,96 +0,0 @@
|
|||||||
apiVersion: monitoring.coreos.com/v1
|
|
||||||
kind: PrometheusRule
|
|
||||||
metadata:
|
|
||||||
name: edu-master-webinar
|
|
||||||
namespace: edu-master
|
|
||||||
labels:
|
|
||||||
release: prometheus-stack
|
|
||||||
spec:
|
|
||||||
groups:
|
|
||||||
- name: edu_master.webinar
|
|
||||||
rules:
|
|
||||||
# No successful webinar check for 5m (~2-3 missed 2-min checks).
|
|
||||||
# Catches: playwright hangs/timeouts, version skew, site changes, hung job.
|
|
||||||
# The last_success > 0 guard is mandatory: checker.py initialises
|
|
||||||
# last_success to 0, so without it `time() - 0` equals the current epoch
|
|
||||||
# and humanizeDuration renders ~20722d on every pod restart. Keep the
|
|
||||||
# duration expression on the left so $value stays the real gap.
|
|
||||||
- alert: WebinarCheckerNoSuccessfulCheck
|
|
||||||
expr: |
|
|
||||||
((time() - webinar_check_last_success_timestamp_seconds) > 300)
|
|
||||||
and (webinar_check_last_success_timestamp_seconds > 0)
|
|
||||||
and (webinar_check_last_run_timestamp_seconds > 0)
|
|
||||||
for: 2m
|
|
||||||
labels:
|
|
||||||
severity: critical
|
|
||||||
annotations:
|
|
||||||
summary: "Webinar checker has no successful check for 5m"
|
|
||||||
description: "edu-master/webinar-checker: last successful webinar check was {{ $value | humanizeDuration }} ago. Checks are failing or hanging (see consecutive failures alert). Notifications about new webinars are NOT being sent."
|
|
||||||
|
|
||||||
# Checks are running but none has ever succeeded since pod start.
|
|
||||||
# Split out from the rule above so a zeroed gauge never feeds
|
|
||||||
# humanizeDuration.
|
|
||||||
- alert: WebinarCheckerNeverSucceeded
|
|
||||||
expr: |
|
|
||||||
(webinar_check_last_success_timestamp_seconds == 0)
|
|
||||||
and (webinar_check_last_run_timestamp_seconds > 0)
|
|
||||||
for: 10m
|
|
||||||
labels:
|
|
||||||
severity: critical
|
|
||||||
annotations:
|
|
||||||
summary: "Webinar checker has never completed a successful check"
|
|
||||||
description: 'edu-master/webinar-checker: checks have been running for 10m but not one has ever succeeded since the pod started, so every check is failing. Check pod logs (Loki: {namespace="edu-master", container="webinar-checker"}).'
|
|
||||||
|
|
||||||
# Fast path: 3 consecutive failures (~6+ min at 2-min interval).
|
|
||||||
- alert: WebinarCheckerConsecutiveFailures
|
|
||||||
expr: |
|
|
||||||
webinar_check_consecutive_failures >= 3
|
|
||||||
for: 5m
|
|
||||||
labels:
|
|
||||||
severity: critical
|
|
||||||
annotations:
|
|
||||||
summary: "Webinar checker failing consecutively"
|
|
||||||
description: 'edu-master/webinar-checker: {{ $value }} consecutive webinar check failures (timeout / playwright error / page error). Check pod logs (Loki: {namespace="edu-master", container="webinar-checker"}).'
|
|
||||||
|
|
||||||
# Metrics endpoint not scraped for 10m: pod down, metrics server dead, or ServiceMonitor broken.
|
|
||||||
- alert: WebinarCheckerScrapeDown
|
|
||||||
expr: |
|
|
||||||
absent(webinar_check_last_run_timestamp_seconds) == 1
|
|
||||||
for: 10m
|
|
||||||
labels:
|
|
||||||
severity: critical
|
|
||||||
annotations:
|
|
||||||
summary: "Webinar checker metrics missing"
|
|
||||||
description: "edu-master/webinar-checker: no metrics series for 10m. Pod may be down, metrics server dead, or ServiceMonitor/Service broken. Webinar checks are unobserved."
|
|
||||||
|
|
||||||
# EDU session lost: session-keeper down or credentials expired. Without PHPSESSID every check is skipped.
|
|
||||||
- alert: EduPhpsessidMissing
|
|
||||||
expr: |
|
|
||||||
edu_phpsessid_present == 0
|
|
||||||
for: 10m
|
|
||||||
labels:
|
|
||||||
severity: critical
|
|
||||||
annotations:
|
|
||||||
summary: "EDU_PHPSESSID missing"
|
|
||||||
description: "edu-master: EDU_PHPSESSID absent from redis for 10m. Webinar/diari/schedule checks are all skipped. Check session-keeper logs and EDU credentials."
|
|
||||||
|
|
||||||
# Hard deps: checker and playwright deployments unavailable.
|
|
||||||
- alert: WebinarCheckerDeploymentDown
|
|
||||||
expr: |
|
|
||||||
kube_deployment_status_replicas_unavailable{deployment="webinar-checker", namespace="edu-master"} > 0
|
|
||||||
for: 10m
|
|
||||||
labels:
|
|
||||||
severity: critical
|
|
||||||
annotations:
|
|
||||||
summary: "Webinar checker deployment unavailable"
|
|
||||||
description: "edu-master/webinar-checker deployment has {{ $value }} unavailable replica(s) for 10m."
|
|
||||||
|
|
||||||
- alert: PlaywrightServiceDown
|
|
||||||
expr: |
|
|
||||||
kube_deployment_status_replicas_unavailable{deployment="playwright-service", namespace="edu-master"} > 0
|
|
||||||
for: 10m
|
|
||||||
labels:
|
|
||||||
severity: critical
|
|
||||||
annotations:
|
|
||||||
summary: "Playwright service unavailable"
|
|
||||||
description: "edu-master/playwright-service deployment has {{ $value }} unavailable replica(s) for 10m. All webinar/diari/schedule checks fail without it."
|
|
||||||
@@ -1,69 +0,0 @@
|
|||||||
apiVersion: apps/v1
|
|
||||||
kind: Deployment
|
|
||||||
metadata:
|
|
||||||
name: playwright-service
|
|
||||||
namespace: edu-master
|
|
||||||
labels:
|
|
||||||
app: edu-master-playwright
|
|
||||||
spec:
|
|
||||||
replicas: 1
|
|
||||||
selector:
|
|
||||||
matchLabels:
|
|
||||||
app: edu-master-playwright
|
|
||||||
strategy:
|
|
||||||
type: Recreate
|
|
||||||
template:
|
|
||||||
metadata:
|
|
||||||
labels:
|
|
||||||
app: edu-master-playwright
|
|
||||||
spec:
|
|
||||||
containers:
|
|
||||||
- name: playwright
|
|
||||||
# renovate: datasource=docker depName=mcr.microsoft.com/playwright versioning=docker
|
|
||||||
image: mcr.microsoft.com/playwright:v1.56.0-jammy
|
|
||||||
imagePullPolicy: IfNotPresent
|
|
||||||
# p95 412M, max 478M over 7 days, no limit before. Request is set at p95
|
|
||||||
# so the pod is not an eviction candidate; the limit stays above 2x the
|
|
||||||
# request because browser page lifetimes are unpredictable.
|
|
||||||
resources:
|
|
||||||
requests:
|
|
||||||
cpu: "200m"
|
|
||||||
memory: "416Mi"
|
|
||||||
limits:
|
|
||||||
memory: "1Gi"
|
|
||||||
command:
|
|
||||||
- npx
|
|
||||||
- -y
|
|
||||||
- playwright@1.56.0
|
|
||||||
- run-server
|
|
||||||
- --port
|
|
||||||
- "3000"
|
|
||||||
- --path
|
|
||||||
- /ws
|
|
||||||
ports:
|
|
||||||
- containerPort: 3000
|
|
||||||
readinessProbe:
|
|
||||||
tcpSocket:
|
|
||||||
port: 3000
|
|
||||||
initialDelaySeconds: 5
|
|
||||||
periodSeconds: 10
|
|
||||||
timeoutSeconds: 3
|
|
||||||
livenessProbe:
|
|
||||||
tcpSocket:
|
|
||||||
port: 3000
|
|
||||||
initialDelaySeconds: 15
|
|
||||||
periodSeconds: 20
|
|
||||||
timeoutSeconds: 3
|
|
||||||
---
|
|
||||||
apiVersion: v1
|
|
||||||
kind: Service
|
|
||||||
metadata:
|
|
||||||
name: playwright-service
|
|
||||||
namespace: edu-master
|
|
||||||
spec:
|
|
||||||
selector:
|
|
||||||
app: edu-master-playwright
|
|
||||||
ports:
|
|
||||||
- name: ws
|
|
||||||
port: 3000
|
|
||||||
targetPort: 3000
|
|
||||||
@@ -1,75 +0,0 @@
|
|||||||
apiVersion: apps/v1
|
|
||||||
kind: StatefulSet
|
|
||||||
metadata:
|
|
||||||
name: redis
|
|
||||||
namespace: edu-master
|
|
||||||
labels:
|
|
||||||
app: edu-master-redis
|
|
||||||
spec:
|
|
||||||
serviceName: redis
|
|
||||||
replicas: 1
|
|
||||||
selector:
|
|
||||||
matchLabels:
|
|
||||||
app: edu-master-redis
|
|
||||||
template:
|
|
||||||
metadata:
|
|
||||||
labels:
|
|
||||||
app: edu-master-redis
|
|
||||||
spec:
|
|
||||||
containers:
|
|
||||||
- name: redis
|
|
||||||
image: redis:8.10.2-alpine
|
|
||||||
imagePullPolicy: IfNotPresent
|
|
||||||
ports:
|
|
||||||
- containerPort: 6379
|
|
||||||
volumeMounts:
|
|
||||||
- name: redis-data
|
|
||||||
mountPath: /data
|
|
||||||
resources:
|
|
||||||
requests:
|
|
||||||
cpu: 25m
|
|
||||||
memory: 32Mi
|
|
||||||
limits:
|
|
||||||
cpu: 250m
|
|
||||||
memory: 128Mi
|
|
||||||
readinessProbe:
|
|
||||||
exec:
|
|
||||||
command: ["redis-cli", "ping"]
|
|
||||||
initialDelaySeconds: 5
|
|
||||||
periodSeconds: 5
|
|
||||||
timeoutSeconds: 3
|
|
||||||
livenessProbe:
|
|
||||||
exec:
|
|
||||||
command: ["redis-cli", "ping"]
|
|
||||||
initialDelaySeconds: 10
|
|
||||||
periodSeconds: 10
|
|
||||||
timeoutSeconds: 3
|
|
||||||
volumes:
|
|
||||||
- name: redis-data
|
|
||||||
persistentVolumeClaim:
|
|
||||||
claimName: redis-data-pvc
|
|
||||||
---
|
|
||||||
apiVersion: v1
|
|
||||||
kind: PersistentVolumeClaim
|
|
||||||
metadata:
|
|
||||||
name: redis-data-pvc
|
|
||||||
namespace: edu-master
|
|
||||||
spec:
|
|
||||||
accessModes:
|
|
||||||
- ReadWriteOnce
|
|
||||||
resources:
|
|
||||||
requests:
|
|
||||||
storage: 1Gi
|
|
||||||
---
|
|
||||||
apiVersion: v1
|
|
||||||
kind: Service
|
|
||||||
metadata:
|
|
||||||
name: redis
|
|
||||||
namespace: edu-master
|
|
||||||
spec:
|
|
||||||
selector:
|
|
||||||
app: edu-master-redis
|
|
||||||
ports:
|
|
||||||
- name: redis
|
|
||||||
port: 6379
|
|
||||||
targetPort: 6379
|
|
||||||
@@ -1,50 +0,0 @@
|
|||||||
# One-time Job to migrate redis state from docker compose to k8s (maintenance window).
|
|
||||||
# The .example file is not applied by the deploy pipeline (mask *.example.yaml).
|
|
||||||
#
|
|
||||||
# Runbook:
|
|
||||||
# 1. docker compose -f <repo>/edu_master/compose.yaml stop # SIGTERM -> redis will flush dump.rdb
|
|
||||||
# 2. docker run --rm -v edu_master_redis-data:/data \
|
|
||||||
# -v /tmp/edu-master-backup:/backup \
|
|
||||||
# redis:alpine sh -c "cp /data/dump.rdb /backup/ && ls -la /backup"
|
|
||||||
# 3. kubectl apply -f edu_master/k8s/namespace.yaml
|
|
||||||
# 4. kubectl apply -f <only the PVC from redis.yaml> # seed must come BEFORE redis pod starts
|
|
||||||
# 5. kubectl apply -f edu_master/k8s/restore-seed-job.yaml.example
|
|
||||||
# kubectl wait --for=condition=complete job/redis-restore-seed -n edu-master --timeout=120s
|
|
||||||
# 6. kubectl delete job redis-restore-seed -n edu-master
|
|
||||||
# 7. kubectl apply -f edu_master/k8s/ -R # apply remaining manifests
|
|
||||||
apiVersion: batch/v1
|
|
||||||
kind: Job
|
|
||||||
metadata:
|
|
||||||
name: redis-restore-seed
|
|
||||||
namespace: edu-master
|
|
||||||
spec:
|
|
||||||
backoffLimit: 2
|
|
||||||
ttlSecondsAfterFinished: 3600
|
|
||||||
template:
|
|
||||||
spec:
|
|
||||||
restartPolicy: Never
|
|
||||||
containers:
|
|
||||||
- name: seed
|
|
||||||
image: redis:alpine
|
|
||||||
command:
|
|
||||||
- /bin/sh
|
|
||||||
- -ec
|
|
||||||
- |
|
|
||||||
ls -la /backup
|
|
||||||
cp /backup/dump.rdb /data/dump.rdb
|
|
||||||
chmod 644 /data/dump.rdb
|
|
||||||
ls -la /data
|
|
||||||
volumeMounts:
|
|
||||||
- name: redis-data
|
|
||||||
mountPath: /data
|
|
||||||
- name: backup
|
|
||||||
mountPath: /backup
|
|
||||||
readOnly: true
|
|
||||||
volumes:
|
|
||||||
- name: redis-data
|
|
||||||
persistentVolumeClaim:
|
|
||||||
claimName: redis-data-pvc
|
|
||||||
- name: backup
|
|
||||||
hostPath:
|
|
||||||
path: /tmp/edu-master-backup
|
|
||||||
type: DirectoryOrCreate
|
|
||||||
@@ -1,29 +0,0 @@
|
|||||||
apiVersion: v1
|
|
||||||
kind: Secret
|
|
||||||
metadata:
|
|
||||||
name: edu-master-secrets
|
|
||||||
namespace: edu-master
|
|
||||||
type: Opaque
|
|
||||||
stringData:
|
|
||||||
# Session keeper credentials
|
|
||||||
KEEPER_LOGIN: ""
|
|
||||||
KEEPER_PASSWORD: ""
|
|
||||||
KEEPER_INTERVAL: "10"
|
|
||||||
# EDU links
|
|
||||||
EDU_URL_BASE: "https://edu.edu.vn.ua"
|
|
||||||
EDU_URL_LOGIN: "/user/login"
|
|
||||||
EDU_URL_COURSES: "/course/userlist"
|
|
||||||
EDU_URL_WEBINAR: "/webinar/useractive"
|
|
||||||
# Playwright
|
|
||||||
USER_AGENT: ""
|
|
||||||
PLAYWRIGHT_WS: "ws://playwright-service:3000/ws"
|
|
||||||
# Webinar-checker
|
|
||||||
WEBINAR_TELEGRAM_TOKEN: ""
|
|
||||||
WEBINAR_ADMIN_ID: ""
|
|
||||||
WEBINAR_CHECK_INTERVAL: "60"
|
|
||||||
# Prometheus metrics endpoint (scraped via ServiceMonitor, alerts in k8s/alerts.yaml)
|
|
||||||
METRICS_PORT: "8000"
|
|
||||||
# Database
|
|
||||||
REDIS_HOST: "redis"
|
|
||||||
REDIS_PORT: "6379"
|
|
||||||
TZ: "Europe/Kyiv"
|
|
||||||
@@ -1,15 +0,0 @@
|
|||||||
apiVersion: v1
|
|
||||||
kind: Service
|
|
||||||
metadata:
|
|
||||||
name: webinar-checker
|
|
||||||
namespace: edu-master
|
|
||||||
labels:
|
|
||||||
app: edu-master-webinar-checker
|
|
||||||
spec:
|
|
||||||
selector:
|
|
||||||
app: edu-master-webinar-checker
|
|
||||||
ports:
|
|
||||||
- name: metrics
|
|
||||||
port: 8000
|
|
||||||
targetPort: metrics
|
|
||||||
protocol: TCP
|
|
||||||
@@ -1,55 +0,0 @@
|
|||||||
apiVersion: apps/v1
|
|
||||||
kind: Deployment
|
|
||||||
metadata:
|
|
||||||
annotations:
|
|
||||||
reloader.stakater.com/auto: "true"
|
|
||||||
name: session-keeper
|
|
||||||
namespace: edu-master
|
|
||||||
labels:
|
|
||||||
app: edu-master-session-keeper
|
|
||||||
spec:
|
|
||||||
replicas: 1
|
|
||||||
selector:
|
|
||||||
matchLabels:
|
|
||||||
app: edu-master-session-keeper
|
|
||||||
strategy:
|
|
||||||
type: Recreate
|
|
||||||
template:
|
|
||||||
metadata:
|
|
||||||
labels:
|
|
||||||
app: edu-master-session-keeper
|
|
||||||
spec:
|
|
||||||
initContainers:
|
|
||||||
- name: wait-redis
|
|
||||||
image: redis:8.10.2-alpine
|
|
||||||
command:
|
|
||||||
- /bin/sh
|
|
||||||
- -ec
|
|
||||||
- |
|
|
||||||
i=0
|
|
||||||
until redis-cli -h redis ping | grep -q PONG; do
|
|
||||||
i=$((i+1))
|
|
||||||
[ "$i" -ge 300 ] && echo "TIMEOUT: redis not ready" && exit 1
|
|
||||||
sleep 2
|
|
||||||
done
|
|
||||||
echo "redis is ready"
|
|
||||||
containers:
|
|
||||||
- name: session-keeper
|
|
||||||
image: gcr.forust.xyz/forust/session-keeper:prod
|
|
||||||
envFrom:
|
|
||||||
- secretRef:
|
|
||||||
name: edu-master-secrets
|
|
||||||
resources:
|
|
||||||
requests:
|
|
||||||
cpu: 25m
|
|
||||||
memory: 32Mi
|
|
||||||
limits:
|
|
||||||
cpu: 250m
|
|
||||||
memory: 128Mi
|
|
||||||
readinessProbe:
|
|
||||||
exec:
|
|
||||||
command: ["/bin/sh", "-ec", "redis-cli -h redis EXISTS EDU_PHPSESSID | grep -q 1"]
|
|
||||||
initialDelaySeconds: 15
|
|
||||||
periodSeconds: 30
|
|
||||||
timeoutSeconds: 5
|
|
||||||
failureThreshold: 10
|
|
||||||
@@ -1,77 +0,0 @@
|
|||||||
apiVersion: apps/v1
|
|
||||||
kind: Deployment
|
|
||||||
metadata:
|
|
||||||
annotations:
|
|
||||||
reloader.stakater.com/auto: "true"
|
|
||||||
name: webinar-checker
|
|
||||||
namespace: edu-master
|
|
||||||
labels:
|
|
||||||
app: edu-master-webinar-checker
|
|
||||||
spec:
|
|
||||||
replicas: 1
|
|
||||||
selector:
|
|
||||||
matchLabels:
|
|
||||||
app: edu-master-webinar-checker
|
|
||||||
strategy:
|
|
||||||
type: Recreate
|
|
||||||
template:
|
|
||||||
metadata:
|
|
||||||
labels:
|
|
||||||
app: edu-master-webinar-checker
|
|
||||||
spec:
|
|
||||||
# Enforces dependency order like compose depends_on:
|
|
||||||
# redis healthy -> session-keeper healthy (EXISTS EDU_PHPSESSID) -> playwright started
|
|
||||||
initContainers:
|
|
||||||
- name: wait-deps
|
|
||||||
image: redis:8.10.2-alpine
|
|
||||||
command:
|
|
||||||
- /bin/sh
|
|
||||||
- -ec
|
|
||||||
- |
|
|
||||||
i=0
|
|
||||||
until redis-cli -h redis ping | grep -q PONG; do
|
|
||||||
i=$((i+1))
|
|
||||||
[ "$i" -ge 300 ] && echo "TIMEOUT: redis not ready" && exit 1
|
|
||||||
sleep 2
|
|
||||||
done
|
|
||||||
echo "redis ok"
|
|
||||||
until [ "$(redis-cli -h redis EXISTS EDU_PHPSESSID)" = "1" ]; do
|
|
||||||
i=$((i+1))
|
|
||||||
[ "$i" -ge 300 ] && echo "TIMEOUT: no PHPSESSID (session-keeper down?)" && exit 1
|
|
||||||
sleep 2
|
|
||||||
done
|
|
||||||
echo "PHPSESSID ok"
|
|
||||||
until nc -z playwright-service 3000; do
|
|
||||||
i=$((i+1))
|
|
||||||
[ "$i" -ge 300 ] && echo "TIMEOUT: playwright-service not reachable" && exit 1
|
|
||||||
sleep 2
|
|
||||||
done
|
|
||||||
echo "playwright ok"
|
|
||||||
containers:
|
|
||||||
- name: webinar-checker
|
|
||||||
image: gcr.forust.xyz/forust/webinar-checker:prod
|
|
||||||
ports:
|
|
||||||
- name: metrics
|
|
||||||
containerPort: 8000
|
|
||||||
protocol: TCP
|
|
||||||
readinessProbe:
|
|
||||||
httpGet:
|
|
||||||
path: /health
|
|
||||||
port: metrics
|
|
||||||
periodSeconds: 10
|
|
||||||
timeoutSeconds: 3
|
|
||||||
failureThreshold: 12
|
|
||||||
initialDelaySeconds: 10
|
|
||||||
envFrom:
|
|
||||||
- secretRef:
|
|
||||||
name: edu-master-secrets
|
|
||||||
env:
|
|
||||||
- name: TZ
|
|
||||||
value: "Europe/Kyiv"
|
|
||||||
resources:
|
|
||||||
requests:
|
|
||||||
cpu: "50m"
|
|
||||||
memory: "192Mi"
|
|
||||||
limits:
|
|
||||||
cpu: "600m"
|
|
||||||
memory: "384Mi"
|
|
||||||
@@ -1,15 +0,0 @@
|
|||||||
FROM python:3.11-slim
|
|
||||||
|
|
||||||
WORKDIR /app
|
|
||||||
|
|
||||||
# Install system dependencies
|
|
||||||
RUN apt-get update && apt-get install -y --no-install-recommends redis-tools && rm -rf /var/lib/apt/lists/*
|
|
||||||
|
|
||||||
# Install dependencies
|
|
||||||
RUN pip install --no-cache-dir requests==2.32.3 redis==5.2.1
|
|
||||||
|
|
||||||
# Copy application code
|
|
||||||
COPY . .
|
|
||||||
|
|
||||||
# Run the bot
|
|
||||||
CMD ["python", "bot.py"]
|
|
||||||
@@ -1,132 +0,0 @@
|
|||||||
import logging
|
|
||||||
import os
|
|
||||||
import time
|
|
||||||
from datetime import datetime
|
|
||||||
|
|
||||||
import redis
|
|
||||||
import requests
|
|
||||||
|
|
||||||
# Configure logging
|
|
||||||
logging.basicConfig(level=logging.INFO, format='%(asctime)s - %(levelname)s - %(message)s')
|
|
||||||
logger = logging.getLogger(__name__)
|
|
||||||
|
|
||||||
|
|
||||||
# Load configuration (adapted to .env keys)
|
|
||||||
def _env(key, default=None):
|
|
||||||
v = os.getenv(key, default)
|
|
||||||
if isinstance(v, str) and len(v) >= 2 and ((v[0] == '"' and v[-1] == '"') or (v[0] == "'" and v[-1] == "'")):
|
|
||||||
return v[1:-1]
|
|
||||||
return v
|
|
||||||
|
|
||||||
|
|
||||||
LOGIN = _env('KEEPER_LOGIN')
|
|
||||||
PASSWORD = _env('KEEPER_PASSWORD')
|
|
||||||
|
|
||||||
EDU_BASE = _env('EDU_URL_BASE', 'https://edu.edu.vn.ua')
|
|
||||||
EDU_LOGIN_PATH = _env('EDU_URL_LOGIN', '/user/login')
|
|
||||||
EDU_COURSES_PATH = _env('EDU_URL_COURSES', '/course/userlist')
|
|
||||||
URL_LOGIN = f'{EDU_BASE.rstrip("/")}/{EDU_LOGIN_PATH.lstrip("/")}'
|
|
||||||
URL_VERIFY = f'{EDU_BASE.rstrip("/")}/{EDU_COURSES_PATH.lstrip("/")}'
|
|
||||||
|
|
||||||
INTERVAL = int(_env('KEEPER_INTERVAL', 10))
|
|
||||||
USER_AGENT = _env(
|
|
||||||
'USER_AGENT',
|
|
||||||
'Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/142.0.0.0 Safari/537.36',
|
|
||||||
)
|
|
||||||
REDIS_HOST = _env('REDIS_HOST', 'redis')
|
|
||||||
REDIS_PORT = int(_env('REDIS_PORT', 6379))
|
|
||||||
|
|
||||||
SUCCESS_FILE = '/tmp/last_success' # noqa: S108
|
|
||||||
|
|
||||||
|
|
||||||
def touch_success_file():
|
|
||||||
"""Updates the timestamp of the success file for healthchecks."""
|
|
||||||
try:
|
|
||||||
with open(SUCCESS_FILE, 'w') as f:
|
|
||||||
f.write(str(datetime.now().timestamp()))
|
|
||||||
except Exception as e:
|
|
||||||
logger.error(f'Failed to touch success file: {e}')
|
|
||||||
|
|
||||||
|
|
||||||
def main():
|
|
||||||
logger.info('Starting Session Keeper Bot')
|
|
||||||
|
|
||||||
# Connect to Redis
|
|
||||||
try:
|
|
||||||
redis_client = redis.Redis(host=REDIS_HOST, port=REDIS_PORT, decode_responses=True)
|
|
||||||
redis_client.ping()
|
|
||||||
logger.info(f'Connected to Redis at {REDIS_HOST}:{REDIS_PORT}')
|
|
||||||
except Exception as e:
|
|
||||||
logger.error(f'Failed to connect to Redis: {e}')
|
|
||||||
return
|
|
||||||
|
|
||||||
session = requests.Session()
|
|
||||||
|
|
||||||
# Set headers
|
|
||||||
headers = {
|
|
||||||
'User-Agent': USER_AGENT,
|
|
||||||
'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,image/apng,*/*;q=0.8,application/signed-exchange;v=b3;q=0.7',
|
|
||||||
'Accept-Language': 'en-US,en;q=0.9',
|
|
||||||
'Cache-Control': 'max-age=0',
|
|
||||||
'Upgrade-Insecure-Requests': '1',
|
|
||||||
'Sec-Fetch-Site': 'same-origin',
|
|
||||||
'Sec-Fetch-Mode': 'navigate',
|
|
||||||
'Sec-Fetch-User': '?1',
|
|
||||||
'Sec-Fetch-Dest': 'document',
|
|
||||||
'Sec-Ch-Ua': '"Not_A Brand";v="99", "Chromium";v="142"',
|
|
||||||
'Sec-Ch-Ua-Mobile': '?0',
|
|
||||||
'Sec-Ch-Ua-Platform': '"Linux"',
|
|
||||||
'Accept-Encoding': 'gzip, deflate, br',
|
|
||||||
'Priority': 'u=0, i',
|
|
||||||
}
|
|
||||||
session.headers.update(headers)
|
|
||||||
|
|
||||||
while True:
|
|
||||||
try:
|
|
||||||
logger.info('Attempting login...')
|
|
||||||
|
|
||||||
# Login payload
|
|
||||||
payload = {'login': LOGIN, 'password': PASSWORD}
|
|
||||||
|
|
||||||
# Perform Login
|
|
||||||
# Note: The user request shows a POST to /user/login with form data
|
|
||||||
# We need to make sure we handle the PHPSESSID correctly.
|
|
||||||
# If we already have a PHPSESSID, requests will send it.
|
|
||||||
|
|
||||||
login_response = session.post(URL_LOGIN, data=payload, allow_redirects=True)
|
|
||||||
|
|
||||||
logger.info(f'Login Response Status: {login_response.status_code}')
|
|
||||||
logger.info(f'Cookies after login: {session.cookies.get_dict()}')
|
|
||||||
|
|
||||||
# Verify Session
|
|
||||||
logger.info('Verifying session...')
|
|
||||||
verify_response = session.get(URL_VERIFY, allow_redirects=False)
|
|
||||||
|
|
||||||
logger.info(f'Verify Response Status: {verify_response.status_code}')
|
|
||||||
|
|
||||||
if verify_response.status_code == 200:
|
|
||||||
logger.info('Session verification SUCCESS (200 OK).')
|
|
||||||
touch_success_file()
|
|
||||||
|
|
||||||
# Save PHPSESSID to Redis
|
|
||||||
phpsessid = session.cookies.get('PHPSESSID')
|
|
||||||
if phpsessid:
|
|
||||||
try:
|
|
||||||
redis_client.set('EDU_PHPSESSID', phpsessid)
|
|
||||||
logger.info(f'Saved PHPSESSID to Redis: {phpsessid}')
|
|
||||||
except Exception as e:
|
|
||||||
logger.error(f'Failed to save PHPSESSID to Redis: {e}')
|
|
||||||
elif verify_response.status_code == 302:
|
|
||||||
logger.warning('Session verification FAILED (302 Redirect). Session might be invalid.')
|
|
||||||
else:
|
|
||||||
logger.warning(f'Session verification returned unexpected status: {verify_response.status_code}')
|
|
||||||
|
|
||||||
except Exception as e:
|
|
||||||
logger.error(f'An error occurred: {e}')
|
|
||||||
|
|
||||||
logger.info(f'Sleeping for {INTERVAL} minutes...')
|
|
||||||
time.sleep(INTERVAL * 60)
|
|
||||||
|
|
||||||
|
|
||||||
if __name__ == '__main__':
|
|
||||||
main()
|
|
||||||
@@ -1,13 +0,0 @@
|
|||||||
FROM python:3.11-slim
|
|
||||||
|
|
||||||
WORKDIR /app
|
|
||||||
|
|
||||||
# renovate: datasource=pypi depName=playwright versioning=pep440
|
|
||||||
ARG PLAYWRIGHT_VERSION=1.56.0
|
|
||||||
|
|
||||||
# Install dependencies - PLAYWRIGHT_VERSION is single-source, renovate updates ARG above and all other places via regexManagers
|
|
||||||
RUN pip install --no-cache-dir pip==25.0.1 && pip install --no-cache-dir playwright==${PLAYWRIGHT_VERSION} redis==5.2.1 requests==2.32.3 "python-telegram-bot[job-queue]==21.10"
|
|
||||||
|
|
||||||
COPY checker.py .
|
|
||||||
|
|
||||||
CMD ["python", "checker.py"]
|
|
||||||
File diff suppressed because it is too large.
Load diff
@@ -20,6 +20,8 @@ data:
|
|||||||
|
|
||||||
GITEA__mailer__ENABLED: "false"
|
GITEA__mailer__ENABLED: "false"
|
||||||
|
|
||||||
|
GITEA__metrics__ENABLED: "true"
|
||||||
|
|
||||||
# No code/issue search needed: bleve reindexes the whole issue index on
|
# No code/issue search needed: bleve reindexes the whole issue index on
|
||||||
# every pod restart (cron.rebuild_issue_indexer RUN_AT_START) and hammers
|
# every pod restart (cron.rebuild_issue_indexer RUN_AT_START) and hammers
|
||||||
# the rotational disk for an hour. "db" serves issue search from postgres.
|
# the rotational disk for an hour. "db" serves issue search from postgres.
|
||||||
|
|||||||
@@ -3,6 +3,8 @@ kind: Service
|
|||||||
metadata:
|
metadata:
|
||||||
name: gitea-service
|
name: gitea-service
|
||||||
namespace: gitea
|
namespace: gitea
|
||||||
|
labels:
|
||||||
|
app: gitea
|
||||||
spec:
|
spec:
|
||||||
selector:
|
selector:
|
||||||
app: gitea
|
app: gitea
|
||||||
|
|||||||
@@ -7,7 +7,8 @@ spec:
|
|||||||
entryPoints:
|
entryPoints:
|
||||||
- websecure
|
- websecure
|
||||||
routes:
|
routes:
|
||||||
- match: Host(`gitea.forust.xyz`) || Host(`git.forust.xyz`)
|
# Metrics are scraped directly through the cluster Service.
|
||||||
|
- match: (Host(`gitea.forust.xyz`) || Host(`git.forust.xyz`)) && !PathPrefix(`/metrics`)
|
||||||
kind: Rule
|
kind: Rule
|
||||||
services:
|
services:
|
||||||
- name: gitea-service
|
- name: gitea-service
|
||||||
|
|||||||
@@ -0,0 +1,16 @@
|
|||||||
|
apiVersion: monitoring.coreos.com/v1
|
||||||
|
kind: ServiceMonitor
|
||||||
|
metadata:
|
||||||
|
name: gitea
|
||||||
|
namespace: gitea
|
||||||
|
labels:
|
||||||
|
release: prometheus-stack
|
||||||
|
spec:
|
||||||
|
selector:
|
||||||
|
matchLabels:
|
||||||
|
app: gitea
|
||||||
|
endpoints:
|
||||||
|
- port: http
|
||||||
|
path: /metrics
|
||||||
|
interval: 30s
|
||||||
|
scrapeTimeout: 10s
|
||||||
@@ -71,7 +71,7 @@ spec:
|
|||||||
name: glance-config
|
name: glance-config
|
||||||
- name: glance-assets
|
- name: glance-assets
|
||||||
configMap:
|
configMap:
|
||||||
name: glance-config
|
name: glance-assets
|
||||||
- name: docker-socket
|
- name: docker-socket
|
||||||
hostPath:
|
hostPath:
|
||||||
path: /var/run/docker.sock
|
path: /var/run/docker.sock
|
||||||
|
|||||||
@@ -7,7 +7,7 @@
|
|||||||
# # Dev server_url
|
# # Dev server_url
|
||||||
# server_url: https://hs.dev_internal_domain.internal
|
# server_url: https://hs.dev_internal_domain.internal
|
||||||
listen_addr: 0.0.0.0:8080
|
listen_addr: 0.0.0.0:8080
|
||||||
metrics_listen_addr: 127.0.0.1:9090
|
metrics_listen_addr: 0.0.0.0:9090
|
||||||
grpc_listen_addr: 127.0.0.1:50443
|
grpc_listen_addr: 127.0.0.1:50443
|
||||||
grpc_allow_insecure: false
|
grpc_allow_insecure: false
|
||||||
noise:
|
noise:
|
||||||
|
|||||||
@@ -3,6 +3,8 @@ kind: Service
|
|||||||
metadata:
|
metadata:
|
||||||
name: headscale-server-external
|
name: headscale-server-external
|
||||||
namespace: headscale
|
namespace: headscale
|
||||||
|
labels:
|
||||||
|
app: headscale
|
||||||
spec:
|
spec:
|
||||||
ports:
|
ports:
|
||||||
- port: 8080
|
- port: 8080
|
||||||
|
|||||||
@@ -0,0 +1,18 @@
|
|||||||
|
apiVersion: operator.victoriametrics.com/v1beta1
|
||||||
|
kind: VMServiceScrape
|
||||||
|
metadata:
|
||||||
|
name: headscale
|
||||||
|
namespace: headscale
|
||||||
|
labels:
|
||||||
|
release: prometheus-stack
|
||||||
|
spec:
|
||||||
|
# The external Service has a manually managed EndpointSlice, not Endpoints.
|
||||||
|
discoveryRole: endpointslice
|
||||||
|
selector:
|
||||||
|
matchLabels:
|
||||||
|
app: headscale
|
||||||
|
endpoints:
|
||||||
|
- port: metrics
|
||||||
|
path: /metrics
|
||||||
|
interval: 30s
|
||||||
|
scrapeTimeout: 10s
|
||||||
+1
-1
@@ -1,7 +1,7 @@
|
|||||||
services:
|
services:
|
||||||
homarr:
|
homarr:
|
||||||
container_name: homarr
|
container_name: homarr
|
||||||
image: ghcr.io/homarr-labs/homarr:v2.1.2
|
image: ghcr.io/homarr-labs/homarr:v2.2.0
|
||||||
restart: unless-stopped
|
restart: unless-stopped
|
||||||
volumes:
|
volumes:
|
||||||
- ./appdata:/appdata
|
- ./appdata:/appdata
|
||||||
|
|||||||
@@ -32,7 +32,7 @@ spec:
|
|||||||
serviceAccountName: homarr
|
serviceAccountName: homarr
|
||||||
containers:
|
containers:
|
||||||
- name: homarr
|
- name: homarr
|
||||||
image: ghcr.io/homarr-labs/homarr:v2.1.2
|
image: ghcr.io/homarr-labs/homarr:v2.2.0
|
||||||
envFrom:
|
envFrom:
|
||||||
- configMapRef:
|
- configMapRef:
|
||||||
name: homarr-config
|
name: homarr-config
|
||||||
|
|||||||
@@ -6,6 +6,10 @@ metadata:
|
|||||||
data:
|
data:
|
||||||
TZ: "Europe/Bratislava"
|
TZ: "Europe/Bratislava"
|
||||||
|
|
||||||
|
IMMICH_TELEMETRY_INCLUDE: "all"
|
||||||
|
IMMICH_API_METRICS_PORT: "8081"
|
||||||
|
IMMICH_MICROSERVICES_METRICS_PORT: "8082"
|
||||||
|
|
||||||
# The database in this namespace, not the shared one in the database
|
# The database in this namespace, not the shared one in the database
|
||||||
# namespace: v3 needs VectorChord, and only the dedicated image carries it.
|
# namespace: v3 needs VectorChord, and only the dedicated image carries it.
|
||||||
DB_HOSTNAME: "immich-postgres"
|
DB_HOSTNAME: "immich-postgres"
|
||||||
|
|||||||
@@ -3,6 +3,8 @@ kind: Service
|
|||||||
metadata:
|
metadata:
|
||||||
name: immich-service
|
name: immich-service
|
||||||
namespace: immich
|
namespace: immich
|
||||||
|
labels:
|
||||||
|
app: immich
|
||||||
spec:
|
spec:
|
||||||
selector:
|
selector:
|
||||||
app: immich
|
app: immich
|
||||||
@@ -10,6 +12,12 @@ spec:
|
|||||||
- name: http
|
- name: http
|
||||||
port: 2283
|
port: 2283
|
||||||
targetPort: 2283
|
targetPort: 2283
|
||||||
|
- name: api-metrics
|
||||||
|
port: 8081
|
||||||
|
targetPort: api-metrics
|
||||||
|
- name: worker-metrics
|
||||||
|
port: 8082
|
||||||
|
targetPort: worker-metrics
|
||||||
---
|
---
|
||||||
apiVersion: apps/v1
|
apiVersion: apps/v1
|
||||||
kind: Deployment
|
kind: Deployment
|
||||||
@@ -41,6 +49,10 @@ spec:
|
|||||||
ports:
|
ports:
|
||||||
- name: http
|
- name: http
|
||||||
containerPort: 2283
|
containerPort: 2283
|
||||||
|
- name: api-metrics
|
||||||
|
containerPort: 8081
|
||||||
|
- name: worker-metrics
|
||||||
|
containerPort: 8082
|
||||||
volumeMounts:
|
volumeMounts:
|
||||||
- name: immich-data
|
- name: immich-data
|
||||||
mountPath: /data
|
mountPath: /data
|
||||||
|
|||||||
@@ -0,0 +1,20 @@
|
|||||||
|
apiVersion: monitoring.coreos.com/v1
|
||||||
|
kind: ServiceMonitor
|
||||||
|
metadata:
|
||||||
|
name: immich
|
||||||
|
namespace: immich
|
||||||
|
labels:
|
||||||
|
release: prometheus-stack
|
||||||
|
spec:
|
||||||
|
selector:
|
||||||
|
matchLabels:
|
||||||
|
app: immich
|
||||||
|
endpoints:
|
||||||
|
- port: api-metrics
|
||||||
|
path: /metrics
|
||||||
|
interval: 30s
|
||||||
|
scrapeTimeout: 10s
|
||||||
|
- port: worker-metrics
|
||||||
|
path: /metrics
|
||||||
|
interval: 30s
|
||||||
|
scrapeTimeout: 10s
|
||||||
+1
-1
@@ -1,6 +1,6 @@
|
|||||||
services:
|
services:
|
||||||
n8n:
|
n8n:
|
||||||
image: docker.n8n.io/n8nio/n8n:2.42.3
|
image: docker.n8n.io/n8nio/n8n:2.43.0
|
||||||
container_name: n8n
|
container_name: n8n
|
||||||
restart: unless-stopped
|
restart: unless-stopped
|
||||||
environment:
|
environment:
|
||||||
|
|||||||
+1
-1
@@ -31,7 +31,7 @@ spec:
|
|||||||
spec:
|
spec:
|
||||||
containers:
|
containers:
|
||||||
- name: n8n
|
- name: n8n
|
||||||
image: docker.n8n.io/n8nio/n8n:2.42.3
|
image: docker.n8n.io/n8nio/n8n:2.43.0
|
||||||
envFrom:
|
envFrom:
|
||||||
- configMapRef:
|
- configMapRef:
|
||||||
name: n8n-config
|
name: n8n-config
|
||||||
|
|||||||
Executable
+109
@@ -0,0 +1,109 @@
|
|||||||
|
#!/bin/sh
|
||||||
|
set -eu
|
||||||
|
|
||||||
|
umask 077
|
||||||
|
|
||||||
|
TEMPLATE_PATH=/opt/netbird/config.template.yaml
|
||||||
|
RENDERED_PATH=/run/netbird/config.yaml
|
||||||
|
RELAY_SECRET_PATH=/run/secrets/relay_auth_secret
|
||||||
|
ENCRYPTION_KEY_PATH=/run/secrets/datastore_encryption_key
|
||||||
|
|
||||||
|
is_valid_proxy_subnet() {
|
||||||
|
candidate="$1"
|
||||||
|
case "$candidate" in
|
||||||
|
0.0.0.0/0)
|
||||||
|
return 1
|
||||||
|
;;
|
||||||
|
*/*)
|
||||||
|
address="${candidate%%/*}"
|
||||||
|
prefix="${candidate#*/}"
|
||||||
|
;;
|
||||||
|
*)
|
||||||
|
return 1
|
||||||
|
;;
|
||||||
|
esac
|
||||||
|
|
||||||
|
case "$prefix" in
|
||||||
|
0|[1-9]|[1-2][0-9]|3[0-2]) ;;
|
||||||
|
*)
|
||||||
|
return 1
|
||||||
|
;;
|
||||||
|
esac
|
||||||
|
|
||||||
|
old_ifs="$IFS"
|
||||||
|
IFS=.
|
||||||
|
# shellcheck disable=SC2086
|
||||||
|
set -- $address
|
||||||
|
IFS="$old_ifs"
|
||||||
|
[ "$#" -eq 4 ] || return 1
|
||||||
|
|
||||||
|
for octet do
|
||||||
|
case "$octet" in
|
||||||
|
0|[1-9]|[1-9][0-9]|1[0-9][0-9]|2[0-4][0-9]|25[0-5]) ;;
|
||||||
|
*)
|
||||||
|
return 1
|
||||||
|
;;
|
||||||
|
esac
|
||||||
|
done
|
||||||
|
}
|
||||||
|
|
||||||
|
read_secret() {
|
||||||
|
secret_path="$1"
|
||||||
|
|
||||||
|
if [ ! -r "$secret_path" ]; then
|
||||||
|
echo "Required secret is not readable: $secret_path" >&2
|
||||||
|
exit 1
|
||||||
|
fi
|
||||||
|
|
||||||
|
secret_value="$(cat "$secret_path")"
|
||||||
|
if [ -z "$secret_value" ]; then
|
||||||
|
echo "Required secret is empty: $secret_path" >&2
|
||||||
|
exit 1
|
||||||
|
fi
|
||||||
|
|
||||||
|
printf '%s' "$secret_value"
|
||||||
|
}
|
||||||
|
|
||||||
|
if [ -z "${NETBIRD_DOMAIN:-}" ]; then
|
||||||
|
echo "NETBIRD_DOMAIN must be set" >&2
|
||||||
|
exit 1
|
||||||
|
fi
|
||||||
|
|
||||||
|
case "$NETBIRD_DOMAIN" in
|
||||||
|
*[!A-Za-z0-9.-]*)
|
||||||
|
echo "NETBIRD_DOMAIN contains unsupported characters" >&2
|
||||||
|
exit 1
|
||||||
|
;;
|
||||||
|
esac
|
||||||
|
|
||||||
|
if [ -z "${NETBIRD_PROXY_SUBNET:-}" ] || [ "$NETBIRD_PROXY_SUBNET" = "auto" ]; then
|
||||||
|
echo "NETBIRD_PROXY_SUBNET must be an explicit IPv4 CIDR; run netbird/setup.sh first" >&2
|
||||||
|
exit 1
|
||||||
|
fi
|
||||||
|
if ! is_valid_proxy_subnet "$NETBIRD_PROXY_SUBNET"; then
|
||||||
|
echo "NETBIRD_PROXY_SUBNET must be a non-default IPv4 CIDR, for example 172.20.0.0/16" >&2
|
||||||
|
exit 1
|
||||||
|
fi
|
||||||
|
|
||||||
|
if [ "$#" -ne 2 ] || [ "$1" != "--config" ] || [ "$2" != "$RENDERED_PATH" ]; then
|
||||||
|
echo "Expected: --config $RENDERED_PATH" >&2
|
||||||
|
exit 1
|
||||||
|
fi
|
||||||
|
|
||||||
|
relay_secret="$(read_secret "$RELAY_SECRET_PATH")"
|
||||||
|
encryption_key="$(read_secret "$ENCRYPTION_KEY_PATH")"
|
||||||
|
|
||||||
|
mkdir -p "$(dirname "$RENDERED_PATH")"
|
||||||
|
sed \
|
||||||
|
-e "s|__NETBIRD_DOMAIN__|${NETBIRD_DOMAIN}|g" \
|
||||||
|
-e "s|__NETBIRD_AUTH_SECRET__|${relay_secret}|g" \
|
||||||
|
-e "s|__NETBIRD_ENCRYPTION_KEY__|${encryption_key}|g" \
|
||||||
|
-e "s|__NETBIRD_PROXY_SUBNET__|${NETBIRD_PROXY_SUBNET}|g" \
|
||||||
|
"$TEMPLATE_PATH" >"$RENDERED_PATH"
|
||||||
|
|
||||||
|
if grep -q '__NETBIRD_' "$RENDERED_PATH"; then
|
||||||
|
echo "Rendered NetBird configuration still contains unresolved placeholders" >&2
|
||||||
|
exit 1
|
||||||
|
fi
|
||||||
|
|
||||||
|
exec /go/bin/netbird-server "$@"
|
||||||
@@ -3,6 +3,8 @@ kind: Service
|
|||||||
metadata:
|
metadata:
|
||||||
name: netbird-server-service
|
name: netbird-server-service
|
||||||
namespace: netbird
|
namespace: netbird
|
||||||
|
labels:
|
||||||
|
app: netbird-server
|
||||||
spec:
|
spec:
|
||||||
selector:
|
selector:
|
||||||
app: netbird-server
|
app: netbird-server
|
||||||
@@ -11,6 +13,10 @@ spec:
|
|||||||
name: http
|
name: http
|
||||||
targetPort: 80
|
targetPort: 80
|
||||||
protocol: TCP
|
protocol: TCP
|
||||||
|
- port: 9090
|
||||||
|
name: metrics
|
||||||
|
targetPort: metrics
|
||||||
|
protocol: TCP
|
||||||
- port: 3478
|
- port: 3478
|
||||||
name: stun
|
name: stun
|
||||||
targetPort: 3478
|
targetPort: 3478
|
||||||
@@ -59,6 +65,9 @@ spec:
|
|||||||
- containerPort: 80
|
- containerPort: 80
|
||||||
name: http
|
name: http
|
||||||
protocol: TCP
|
protocol: TCP
|
||||||
|
- containerPort: 9090
|
||||||
|
name: metrics
|
||||||
|
protocol: TCP
|
||||||
- containerPort: 3478
|
- containerPort: 3478
|
||||||
name: stun
|
name: stun
|
||||||
protocol: UDP
|
protocol: UDP
|
||||||
|
|||||||
@@ -1,14 +1,14 @@
|
|||||||
apiVersion: monitoring.coreos.com/v1
|
apiVersion: monitoring.coreos.com/v1
|
||||||
kind: ServiceMonitor
|
kind: ServiceMonitor
|
||||||
metadata:
|
metadata:
|
||||||
name: webinar-checker
|
name: netbird-server
|
||||||
namespace: edu-master
|
namespace: netbird
|
||||||
labels:
|
labels:
|
||||||
release: prometheus-stack
|
release: prometheus-stack
|
||||||
spec:
|
spec:
|
||||||
selector:
|
selector:
|
||||||
matchLabels:
|
matchLabels:
|
||||||
app: edu-master-webinar-checker
|
app: netbird-server
|
||||||
endpoints:
|
endpoints:
|
||||||
- port: metrics
|
- port: metrics
|
||||||
path: /metrics
|
path: /metrics
|
||||||
Executable
+38
@@ -0,0 +1,38 @@
|
|||||||
|
#!/usr/bin/env bash
|
||||||
|
# Prepare local Compose configuration without replacing existing credentials.
|
||||||
|
set -euo pipefail
|
||||||
|
cd "$(dirname "${BASH_SOURCE[0]}")"
|
||||||
|
umask 077
|
||||||
|
if [ ! -f .env ]; then
|
||||||
|
cp .env.example .env
|
||||||
|
fi
|
||||||
|
|
||||||
|
if grep -q '^NETBIRD_PROXY_SUBNET=auto$' .env; then
|
||||||
|
subnet="$(docker network inspect proxy --format '{{range .IPAM.Config}}{{println .Subnet}}{{end}}' | awk '/^[0-9]+\./ { print; exit }')"
|
||||||
|
if [ -z "$subnet" ]; then
|
||||||
|
echo "No IPv4 subnet found on the Docker proxy network. Set NETBIRD_PROXY_SUBNET in .env." >&2
|
||||||
|
exit 1
|
||||||
|
fi
|
||||||
|
# The detected value must be safe to substitute into the env file.
|
||||||
|
if [[ ! "$subnet" =~ ^[0-9.]+/[0-9]+$ ]]; then
|
||||||
|
echo "Unexpected Docker network subnet: $subnet" >&2
|
||||||
|
exit 1
|
||||||
|
fi
|
||||||
|
sed -i "s|^NETBIRD_PROXY_SUBNET=auto$|NETBIRD_PROXY_SUBNET=$subnet|" .env
|
||||||
|
fi
|
||||||
|
|
||||||
|
mkdir -p secrets
|
||||||
|
chmod 700 secrets
|
||||||
|
for name in relay-auth-secret datastore-encryption-key; do
|
||||||
|
path="secrets/$name"
|
||||||
|
if [ -e "$path" ]; then
|
||||||
|
if [ ! -s "$path" ]; then
|
||||||
|
echo "Existing secret is empty: $path. Restore it before continuing." >&2
|
||||||
|
exit 1
|
||||||
|
fi
|
||||||
|
else
|
||||||
|
openssl rand -base64 32 >"$path"
|
||||||
|
fi
|
||||||
|
chmod 600 "$path"
|
||||||
|
done
|
||||||
|
printf '%s\n' 'Local files are ready. Review .env, then run docker compose config --quiet.'
|
||||||
@@ -1,6 +1,6 @@
|
|||||||
services:
|
services:
|
||||||
netronome:
|
netronome:
|
||||||
image: ghcr.io/autobrr/netronome:v0.15.0
|
image: ghcr.io/autobrr/netronome:v0.16.0
|
||||||
restart: unless-stopped
|
restart: unless-stopped
|
||||||
container_name: netronome
|
container_name: netronome
|
||||||
ports:
|
ports:
|
||||||
|
|||||||
@@ -34,7 +34,7 @@ spec:
|
|||||||
spec:
|
spec:
|
||||||
containers:
|
containers:
|
||||||
- name: netronome
|
- name: netronome
|
||||||
image: ghcr.io/autobrr/netronome:v0.15.0
|
image: ghcr.io/autobrr/netronome:v0.16.0
|
||||||
ports:
|
ports:
|
||||||
- name: netronome-port
|
- name: netronome-port
|
||||||
protocol: TCP
|
protocol: TCP
|
||||||
|
|||||||
@@ -0,0 +1,16 @@
|
|||||||
|
PAPERLESS_URL=https://papers.forust.xyz
|
||||||
|
PAPERLESS_ALLOWED_HOSTS=papers.forust.xyz,papers.workstation.internal
|
||||||
|
PAPERLESS_CSRF_TRUSTED_ORIGINS=https://papers.forust.xyz,https://papers.workstation.internal
|
||||||
|
PAPERLESS_TIME_ZONE=Europe/Bratislava
|
||||||
|
PAPERLESS_REDIS=redis://valkey:6379
|
||||||
|
PAPERLESS_DBENGINE=postgresql
|
||||||
|
PAPERLESS_DBHOST=homelab-postgres
|
||||||
|
PAPERLESS_DBNAME=paperless
|
||||||
|
PAPERLESS_DBUSER=paperless
|
||||||
|
PAPERLESS_DBPASS=<SET_THE_SAME_VALUE_AS_SHARED_POSTGRES_PAPERLESS_DB_PASSWORD>
|
||||||
|
PAPERLESS_OCR_LANGUAGE=rus+eng
|
||||||
|
PAPERLESS_OCR_LANGUAGES=rus
|
||||||
|
PAPERLESS_TASK_WORKERS=1
|
||||||
|
PAPERLESS_ADMIN_USER=admin
|
||||||
|
PAPERLESS_SECRET_KEY=<GENERATE_WITH_python3_-c_import_secrets;_print(secrets.token_urlsafe(64))>
|
||||||
|
PAPERLESS_ADMIN_PASSWORD=<SET_A_LONG_UNIQUE_PASSWORD>
|
||||||
@@ -0,0 +1,79 @@
|
|||||||
|
# Paperless-ngx
|
||||||
|
|
||||||
|
Paperless-ngx runs in the `paperless` namespace. It uses the shared PostgreSQL
|
||||||
|
service in the `database` namespace and Valkey for its task queue. The document
|
||||||
|
library, exports, and consume folder are stored on the `local-path-retain`
|
||||||
|
volume. The PVC size is fixed at 50 GiB because this storage class does not
|
||||||
|
support volume expansion.
|
||||||
|
|
||||||
|
The local route is `https://papers.workstation.internal`; the public route is
|
||||||
|
`https://papers.forust.xyz`. Both use TLS. Paperless keeps its own login and
|
||||||
|
password authentication. OCR is configured for Russian and English documents.
|
||||||
|
Keep `papers.forust.xyz` in the existing Cloudflare DDNS `DOMAINS` setting so
|
||||||
|
the public record follows the workstation address.
|
||||||
|
|
||||||
|
## Compose alternative
|
||||||
|
|
||||||
|
`compose.yaml` is an alternative to the active Kubernetes deployment. Do not
|
||||||
|
run both at the same time: they use the same Paperless database and route
|
||||||
|
names, but have separate document volumes.
|
||||||
|
|
||||||
|
The Compose variant uses the shared Compose PostgreSQL service on the
|
||||||
|
`homelab-database` Docker network. It does not start a PostgreSQL container.
|
||||||
|
The shared Compose database must be running and must have the `paperless`
|
||||||
|
database and role. Set `PAPERLESS_DBPASS` to the same password as
|
||||||
|
`PAPERLESS_DB_PASSWORD` in the shared PostgreSQL configuration.
|
||||||
|
|
||||||
|
To prepare and start the Compose variant:
|
||||||
|
|
||||||
|
```sh
|
||||||
|
cp paperless/.env.example paperless/.env
|
||||||
|
cd paperless
|
||||||
|
docker compose -f compose.yaml config --quiet
|
||||||
|
docker compose -f compose.yaml up -d
|
||||||
|
```
|
||||||
|
|
||||||
|
Create unique values for `PAPERLESS_SECRET_KEY` and
|
||||||
|
`PAPERLESS_ADMIN_PASSWORD` in `.env`. This directory has no Compose `active`
|
||||||
|
marker, so the repository deploy workflow does not start this alternative.
|
||||||
|
Stop the Kubernetes Paperless deployment before switching to Compose. Back up
|
||||||
|
and migrate the media files as well as the database; the Compose named volumes
|
||||||
|
are separate from the Kubernetes PVC.
|
||||||
|
|
||||||
|
## Prepare the secret
|
||||||
|
|
||||||
|
Create `k8s/secrets.yaml` on the workstation from
|
||||||
|
`k8s/secrets.yaml.example`. Set a unique random `PAPERLESS_SECRET_KEY`, a long
|
||||||
|
`PAPERLESS_ADMIN_PASSWORD`, and `PAPERLESS_DB_PASSWORD`.
|
||||||
|
|
||||||
|
Add the same `PAPERLESS_DB_PASSWORD` value to the local
|
||||||
|
`postgres/k8s/secrets.yaml` file. Keep both secret files out of Git. The
|
||||||
|
database bootstrap Job creates the `paperless` role and database from the
|
||||||
|
shared PostgreSQL secret. The job runs in the `database` namespace and needs
|
||||||
|
that namespace's existing `postgres-shared-secrets` Secret.
|
||||||
|
|
||||||
|
For example, generate a key with:
|
||||||
|
|
||||||
|
```sh
|
||||||
|
python3 -c 'import secrets; print(secrets.token_urlsafe(64))'
|
||||||
|
```
|
||||||
|
|
||||||
|
Then apply the secret before enabling the service:
|
||||||
|
|
||||||
|
```sh
|
||||||
|
kubectl apply -f paperless/k8s/namespace.yaml
|
||||||
|
kubectl apply -f postgres/k8s/secrets.yaml
|
||||||
|
kubectl apply -f paperless/k8s/secrets.yaml
|
||||||
|
```
|
||||||
|
|
||||||
|
The normal deploy workflow applies the remaining manifests when
|
||||||
|
`paperless/k8s/active` is present. Verify the rollout and ingress after deploy:
|
||||||
|
|
||||||
|
```sh
|
||||||
|
kubectl -n paperless rollout status deployment/paperless
|
||||||
|
kubectl -n paperless get pods,pvc,services
|
||||||
|
```
|
||||||
|
|
||||||
|
Back up the `paperless-data` PVC and the shared PostgreSQL database. The PVC
|
||||||
|
contains the originals, archived PDFs, and export/consume folders. Valkey has
|
||||||
|
no persistent volume; queued tasks are recreated after a restart.
|
||||||
@@ -0,0 +1,77 @@
|
|||||||
|
services:
|
||||||
|
paperless:
|
||||||
|
image: ghcr.io/paperless-ngx/paperless-ngx:3.2.1
|
||||||
|
container_name: paperless
|
||||||
|
restart: unless-stopped
|
||||||
|
env_file:
|
||||||
|
- .env
|
||||||
|
depends_on:
|
||||||
|
valkey:
|
||||||
|
condition: service_healthy
|
||||||
|
deploy:
|
||||||
|
resources:
|
||||||
|
limits:
|
||||||
|
cpus: "2.0"
|
||||||
|
memory: 2G
|
||||||
|
reservations:
|
||||||
|
cpus: "0.10"
|
||||||
|
memory: 512M
|
||||||
|
volumes:
|
||||||
|
- paperless-data:/usr/src/paperless/data
|
||||||
|
- paperless-media:/usr/src/paperless/media
|
||||||
|
- paperless-export:/usr/src/paperless/export
|
||||||
|
- paperless-consume:/usr/src/paperless/consume
|
||||||
|
networks:
|
||||||
|
- default
|
||||||
|
- proxy
|
||||||
|
- database
|
||||||
|
labels:
|
||||||
|
- "traefik.enable=true"
|
||||||
|
- "traefik.docker.network=proxy"
|
||||||
|
- "traefik.http.services.paperless-compose.loadbalancer.server.port=8000"
|
||||||
|
- "traefik.http.routers.paperless-compose.rule=Host(`papers.forust.xyz`)"
|
||||||
|
- "traefik.http.routers.paperless-compose.entrypoints=websecure"
|
||||||
|
- "traefik.http.routers.paperless-compose.tls.certresolver=letsencrypt"
|
||||||
|
- "traefik.http.routers.paperless-compose-local.rule=Host(`papers.workstation.internal`)"
|
||||||
|
- "traefik.http.routers.paperless-compose-local.entrypoints=websecure"
|
||||||
|
- "traefik.http.routers.paperless-compose-local.tls=true"
|
||||||
|
|
||||||
|
valkey:
|
||||||
|
image: valkey/valkey:9.0.3-alpine
|
||||||
|
container_name: paperless-valkey
|
||||||
|
restart: unless-stopped
|
||||||
|
command:
|
||||||
|
- valkey-server
|
||||||
|
- --save
|
||||||
|
- ""
|
||||||
|
- --appendonly
|
||||||
|
- "no"
|
||||||
|
healthcheck:
|
||||||
|
test: ["CMD", "valkey-cli", "ping"]
|
||||||
|
interval: 10s
|
||||||
|
timeout: 5s
|
||||||
|
retries: 5
|
||||||
|
deploy:
|
||||||
|
resources:
|
||||||
|
limits:
|
||||||
|
cpus: "0.25"
|
||||||
|
memory: 256M
|
||||||
|
reservations:
|
||||||
|
cpus: "0.025"
|
||||||
|
memory: 64M
|
||||||
|
networks:
|
||||||
|
- default
|
||||||
|
|
||||||
|
volumes:
|
||||||
|
paperless-data:
|
||||||
|
paperless-media:
|
||||||
|
paperless-export:
|
||||||
|
paperless-consume:
|
||||||
|
|
||||||
|
networks:
|
||||||
|
default:
|
||||||
|
proxy:
|
||||||
|
external: true
|
||||||
|
database:
|
||||||
|
name: homelab-database
|
||||||
|
external: true
|
||||||
@@ -0,0 +1 @@
|
|||||||
|
|
||||||
@@ -0,0 +1,28 @@
|
|||||||
|
apiVersion: cert-manager.io/v1
|
||||||
|
kind: Certificate
|
||||||
|
metadata:
|
||||||
|
name: paperless-prod-tls
|
||||||
|
namespace: paperless
|
||||||
|
spec:
|
||||||
|
secretName: paperless-prod-tls
|
||||||
|
dnsNames:
|
||||||
|
- papers.forust.xyz
|
||||||
|
issuerRef:
|
||||||
|
name: letsencrypt-prod
|
||||||
|
kind: ClusterIssuer
|
||||||
|
---
|
||||||
|
apiVersion: cert-manager.io/v1
|
||||||
|
kind: Certificate
|
||||||
|
metadata:
|
||||||
|
name: internal-wildcard-tls
|
||||||
|
namespace: paperless
|
||||||
|
spec:
|
||||||
|
secretName: internal-wildcard-tls
|
||||||
|
dnsNames:
|
||||||
|
- "*.workstation.internal"
|
||||||
|
- "*.gigaforust.internal"
|
||||||
|
- workstation.internal
|
||||||
|
- gigaforust.internal
|
||||||
|
issuerRef:
|
||||||
|
name: internal-ca
|
||||||
|
kind: ClusterIssuer
|
||||||
@@ -0,0 +1,58 @@
|
|||||||
|
apiVersion: batch/v1
|
||||||
|
kind: Job
|
||||||
|
metadata:
|
||||||
|
name: paperless-database-init
|
||||||
|
namespace: database
|
||||||
|
spec:
|
||||||
|
backoffLimit: 5
|
||||||
|
template:
|
||||||
|
metadata:
|
||||||
|
labels:
|
||||||
|
app.kubernetes.io/name: paperless-database-init
|
||||||
|
spec:
|
||||||
|
restartPolicy: OnFailure
|
||||||
|
containers:
|
||||||
|
- name: create-database
|
||||||
|
image: postgres:17.11-alpine
|
||||||
|
command:
|
||||||
|
- /bin/sh
|
||||||
|
- -ec
|
||||||
|
- |
|
||||||
|
PGPASSWORD="$POSTGRES_ADMIN_PASSWORD" psql \
|
||||||
|
--host postgres \
|
||||||
|
--username postgres \
|
||||||
|
--dbname postgres \
|
||||||
|
--set ON_ERROR_STOP=1 \
|
||||||
|
--set paperless_password="$PAPERLESS_DB_PASSWORD" <<'SQL'
|
||||||
|
SELECT format(
|
||||||
|
'CREATE ROLE paperless LOGIN PASSWORD %L',
|
||||||
|
:'paperless_password'
|
||||||
|
)
|
||||||
|
WHERE NOT EXISTS (
|
||||||
|
SELECT FROM pg_roles WHERE rolname = 'paperless'
|
||||||
|
)
|
||||||
|
\gexec
|
||||||
|
SELECT format('CREATE DATABASE paperless OWNER paperless')
|
||||||
|
WHERE NOT EXISTS (
|
||||||
|
SELECT FROM pg_database WHERE datname = 'paperless'
|
||||||
|
)
|
||||||
|
\gexec
|
||||||
|
SQL
|
||||||
|
env:
|
||||||
|
- name: POSTGRES_ADMIN_PASSWORD
|
||||||
|
valueFrom:
|
||||||
|
secretKeyRef:
|
||||||
|
name: postgres-shared-secrets
|
||||||
|
key: POSTGRES_ADMIN_PASSWORD
|
||||||
|
- name: PAPERLESS_DB_PASSWORD
|
||||||
|
valueFrom:
|
||||||
|
secretKeyRef:
|
||||||
|
name: postgres-shared-secrets
|
||||||
|
key: PAPERLESS_DB_PASSWORD
|
||||||
|
resources:
|
||||||
|
requests:
|
||||||
|
cpu: 10m
|
||||||
|
memory: 32Mi
|
||||||
|
limits:
|
||||||
|
cpu: 100m
|
||||||
|
memory: 128Mi
|
||||||
@@ -0,0 +1,33 @@
|
|||||||
|
apiVersion: traefik.io/v1alpha1
|
||||||
|
kind: IngressRoute
|
||||||
|
metadata:
|
||||||
|
name: paperless-prod
|
||||||
|
namespace: paperless
|
||||||
|
spec:
|
||||||
|
entryPoints:
|
||||||
|
- websecure
|
||||||
|
routes:
|
||||||
|
- match: Host(`papers.forust.xyz`)
|
||||||
|
kind: Rule
|
||||||
|
services:
|
||||||
|
- name: paperless
|
||||||
|
port: 8000
|
||||||
|
tls:
|
||||||
|
secretName: paperless-prod-tls
|
||||||
|
---
|
||||||
|
apiVersion: traefik.io/v1alpha1
|
||||||
|
kind: IngressRoute
|
||||||
|
metadata:
|
||||||
|
name: paperless-local
|
||||||
|
namespace: paperless
|
||||||
|
spec:
|
||||||
|
entryPoints:
|
||||||
|
- websecure
|
||||||
|
routes:
|
||||||
|
- match: Host(`papers.workstation.internal`)
|
||||||
|
kind: Rule
|
||||||
|
services:
|
||||||
|
- name: paperless
|
||||||
|
port: 8000
|
||||||
|
tls:
|
||||||
|
secretName: internal-wildcard-tls
|
||||||
@@ -1,4 +1,4 @@
|
|||||||
apiVersion: v1
|
apiVersion: v1
|
||||||
kind: Namespace
|
kind: Namespace
|
||||||
metadata:
|
metadata:
|
||||||
name: edu-master
|
name: paperless
|
||||||
@@ -0,0 +1,39 @@
|
|||||||
|
apiVersion: networking.k8s.io/v1
|
||||||
|
kind: NetworkPolicy
|
||||||
|
metadata:
|
||||||
|
name: paperless-ingress
|
||||||
|
namespace: paperless
|
||||||
|
spec:
|
||||||
|
podSelector:
|
||||||
|
matchLabels:
|
||||||
|
app.kubernetes.io/name: paperless
|
||||||
|
policyTypes:
|
||||||
|
- Ingress
|
||||||
|
ingress:
|
||||||
|
- from:
|
||||||
|
- namespaceSelector:
|
||||||
|
matchLabels:
|
||||||
|
kubernetes.io/metadata.name: traefik
|
||||||
|
ports:
|
||||||
|
- protocol: TCP
|
||||||
|
port: 8000
|
||||||
|
---
|
||||||
|
apiVersion: networking.k8s.io/v1
|
||||||
|
kind: NetworkPolicy
|
||||||
|
metadata:
|
||||||
|
name: paperless-valkey-ingress
|
||||||
|
namespace: paperless
|
||||||
|
spec:
|
||||||
|
podSelector:
|
||||||
|
matchLabels:
|
||||||
|
app.kubernetes.io/name: paperless-valkey
|
||||||
|
policyTypes:
|
||||||
|
- Ingress
|
||||||
|
ingress:
|
||||||
|
- from:
|
||||||
|
- podSelector:
|
||||||
|
matchLabels:
|
||||||
|
app.kubernetes.io/name: paperless
|
||||||
|
ports:
|
||||||
|
- protocol: TCP
|
||||||
|
port: 6379
|
||||||
@@ -0,0 +1,210 @@
|
|||||||
|
apiVersion: v1
|
||||||
|
kind: Service
|
||||||
|
metadata:
|
||||||
|
name: paperless
|
||||||
|
namespace: paperless
|
||||||
|
labels:
|
||||||
|
app.kubernetes.io/name: paperless
|
||||||
|
spec:
|
||||||
|
selector:
|
||||||
|
app.kubernetes.io/name: paperless
|
||||||
|
ports:
|
||||||
|
- name: http
|
||||||
|
port: 8000
|
||||||
|
targetPort: http
|
||||||
|
---
|
||||||
|
apiVersion: apps/v1
|
||||||
|
kind: Deployment
|
||||||
|
metadata:
|
||||||
|
name: paperless
|
||||||
|
namespace: paperless
|
||||||
|
labels:
|
||||||
|
app.kubernetes.io/name: paperless
|
||||||
|
spec:
|
||||||
|
replicas: 1
|
||||||
|
strategy:
|
||||||
|
type: Recreate
|
||||||
|
selector:
|
||||||
|
matchLabels:
|
||||||
|
app.kubernetes.io/name: paperless
|
||||||
|
template:
|
||||||
|
metadata:
|
||||||
|
labels:
|
||||||
|
app.kubernetes.io/name: paperless
|
||||||
|
spec:
|
||||||
|
enableServiceLinks: false
|
||||||
|
containers:
|
||||||
|
- name: paperless
|
||||||
|
image: ghcr.io/paperless-ngx/paperless-ngx:3.2.1
|
||||||
|
ports:
|
||||||
|
- name: http
|
||||||
|
containerPort: 8000
|
||||||
|
env:
|
||||||
|
- name: PAPERLESS_URL
|
||||||
|
value: https://papers.forust.xyz
|
||||||
|
- name: PAPERLESS_ALLOWED_HOSTS
|
||||||
|
value: papers.forust.xyz,papers.workstation.internal
|
||||||
|
- name: PAPERLESS_CSRF_TRUSTED_ORIGINS
|
||||||
|
value: https://papers.forust.xyz,https://papers.workstation.internal
|
||||||
|
- name: PAPERLESS_TIME_ZONE
|
||||||
|
value: Europe/Bratislava
|
||||||
|
- name: PAPERLESS_REDIS
|
||||||
|
value: redis://paperless-valkey:6379
|
||||||
|
- name: PAPERLESS_DBENGINE
|
||||||
|
value: postgresql
|
||||||
|
- name: PAPERLESS_DBHOST
|
||||||
|
value: postgres.database.svc.cluster.local
|
||||||
|
- name: PAPERLESS_DBNAME
|
||||||
|
value: paperless
|
||||||
|
- name: PAPERLESS_DBUSER
|
||||||
|
value: paperless
|
||||||
|
- name: PAPERLESS_DBPASS
|
||||||
|
valueFrom:
|
||||||
|
secretKeyRef:
|
||||||
|
name: paperless-secrets
|
||||||
|
key: PAPERLESS_DB_PASSWORD
|
||||||
|
- name: PAPERLESS_OCR_LANGUAGE
|
||||||
|
value: rus+eng
|
||||||
|
- name: PAPERLESS_OCR_LANGUAGES
|
||||||
|
value: rus
|
||||||
|
- name: PAPERLESS_TASK_WORKERS
|
||||||
|
value: "1"
|
||||||
|
- name: PAPERLESS_ADMIN_USER
|
||||||
|
value: admin
|
||||||
|
- name: PAPERLESS_SECRET_KEY
|
||||||
|
valueFrom:
|
||||||
|
secretKeyRef:
|
||||||
|
name: paperless-secrets
|
||||||
|
key: PAPERLESS_SECRET_KEY
|
||||||
|
- name: PAPERLESS_ADMIN_PASSWORD
|
||||||
|
valueFrom:
|
||||||
|
secretKeyRef:
|
||||||
|
name: paperless-secrets
|
||||||
|
key: PAPERLESS_ADMIN_PASSWORD
|
||||||
|
volumeMounts:
|
||||||
|
- name: documents
|
||||||
|
mountPath: /usr/src/paperless/data
|
||||||
|
subPath: data
|
||||||
|
- name: documents
|
||||||
|
mountPath: /usr/src/paperless/media
|
||||||
|
subPath: media
|
||||||
|
- name: documents
|
||||||
|
mountPath: /usr/src/paperless/export
|
||||||
|
subPath: export
|
||||||
|
- name: documents
|
||||||
|
mountPath: /usr/src/paperless/consume
|
||||||
|
subPath: consume
|
||||||
|
startupProbe:
|
||||||
|
httpGet:
|
||||||
|
path: /
|
||||||
|
port: http
|
||||||
|
httpHeaders:
|
||||||
|
- name: Host
|
||||||
|
value: papers.workstation.internal
|
||||||
|
failureThreshold: 60
|
||||||
|
periodSeconds: 10
|
||||||
|
timeoutSeconds: 5
|
||||||
|
readinessProbe:
|
||||||
|
httpGet:
|
||||||
|
path: /
|
||||||
|
port: http
|
||||||
|
httpHeaders:
|
||||||
|
- name: Host
|
||||||
|
value: papers.workstation.internal
|
||||||
|
periodSeconds: 10
|
||||||
|
timeoutSeconds: 5
|
||||||
|
livenessProbe:
|
||||||
|
httpGet:
|
||||||
|
path: /
|
||||||
|
port: http
|
||||||
|
httpHeaders:
|
||||||
|
- name: Host
|
||||||
|
value: papers.workstation.internal
|
||||||
|
periodSeconds: 30
|
||||||
|
timeoutSeconds: 5
|
||||||
|
resources:
|
||||||
|
requests:
|
||||||
|
cpu: 100m
|
||||||
|
memory: 512Mi
|
||||||
|
limits:
|
||||||
|
cpu: "2"
|
||||||
|
memory: 2Gi
|
||||||
|
volumes:
|
||||||
|
- name: documents
|
||||||
|
persistentVolumeClaim:
|
||||||
|
claimName: paperless-data
|
||||||
|
---
|
||||||
|
apiVersion: v1
|
||||||
|
kind: PersistentVolumeClaim
|
||||||
|
metadata:
|
||||||
|
name: paperless-data
|
||||||
|
namespace: paperless
|
||||||
|
spec:
|
||||||
|
accessModes:
|
||||||
|
- ReadWriteOnce
|
||||||
|
storageClassName: local-path-retain
|
||||||
|
resources:
|
||||||
|
requests:
|
||||||
|
storage: 50Gi
|
||||||
|
---
|
||||||
|
apiVersion: v1
|
||||||
|
kind: Service
|
||||||
|
metadata:
|
||||||
|
name: paperless-valkey
|
||||||
|
namespace: paperless
|
||||||
|
labels:
|
||||||
|
app.kubernetes.io/name: paperless-valkey
|
||||||
|
spec:
|
||||||
|
selector:
|
||||||
|
app.kubernetes.io/name: paperless-valkey
|
||||||
|
ports:
|
||||||
|
- name: redis
|
||||||
|
port: 6379
|
||||||
|
targetPort: redis
|
||||||
|
---
|
||||||
|
apiVersion: apps/v1
|
||||||
|
kind: Deployment
|
||||||
|
metadata:
|
||||||
|
name: paperless-valkey
|
||||||
|
namespace: paperless
|
||||||
|
labels:
|
||||||
|
app.kubernetes.io/name: paperless-valkey
|
||||||
|
spec:
|
||||||
|
replicas: 1
|
||||||
|
strategy:
|
||||||
|
type: Recreate
|
||||||
|
selector:
|
||||||
|
matchLabels:
|
||||||
|
app.kubernetes.io/name: paperless-valkey
|
||||||
|
template:
|
||||||
|
metadata:
|
||||||
|
labels:
|
||||||
|
app.kubernetes.io/name: paperless-valkey
|
||||||
|
spec:
|
||||||
|
containers:
|
||||||
|
- name: valkey
|
||||||
|
image: valkey/valkey:9.0.3-alpine
|
||||||
|
args:
|
||||||
|
- valkey-server
|
||||||
|
- --save
|
||||||
|
- ""
|
||||||
|
- --appendonly
|
||||||
|
- "no"
|
||||||
|
ports:
|
||||||
|
- name: redis
|
||||||
|
containerPort: 6379
|
||||||
|
readinessProbe:
|
||||||
|
exec:
|
||||||
|
command: ["valkey-cli", "ping"]
|
||||||
|
periodSeconds: 10
|
||||||
|
livenessProbe:
|
||||||
|
exec:
|
||||||
|
command: ["valkey-cli", "ping"]
|
||||||
|
periodSeconds: 30
|
||||||
|
resources:
|
||||||
|
requests:
|
||||||
|
cpu: 25m
|
||||||
|
memory: 64Mi
|
||||||
|
limits:
|
||||||
|
cpu: 250m
|
||||||
|
memory: 256Mi
|
||||||
@@ -0,0 +1,10 @@
|
|||||||
|
apiVersion: v1
|
||||||
|
kind: Secret
|
||||||
|
metadata:
|
||||||
|
name: paperless-secrets
|
||||||
|
namespace: paperless
|
||||||
|
type: Opaque
|
||||||
|
stringData:
|
||||||
|
PAPERLESS_SECRET_KEY: "<GENERATE_WITH_python3_-c_import_secrets;_print(secrets.token_urlsafe(64))>"
|
||||||
|
PAPERLESS_ADMIN_PASSWORD: "<SET_A_LONG_UNIQUE_PASSWORD>"
|
||||||
|
PAPERLESS_DB_PASSWORD: "<SET_THE_SAME_VALUE_AS_database_PAPERLESS_DB_PASSWORD>"
|
||||||
@@ -1,6 +1,8 @@
|
|||||||
POSTGRES_ADMIN_PASSWORD=
|
POSTGRES_ADMIN_PASSWORD=
|
||||||
AUTHENTIK_DB_PASSWORD=
|
AUTHENTIK_DB_PASSWORD=
|
||||||
GITEA_DB_PASSWORD=
|
GITEA_DB_PASSWORD=
|
||||||
|
NETBOX_DB_PASSWORD=
|
||||||
NETRONOME_DB_PASSWORD=
|
NETRONOME_DB_PASSWORD=
|
||||||
PENPOT_DB_PASSWORD=
|
PENPOT_DB_PASSWORD=
|
||||||
|
PAPERLESS_DB_PASSWORD=
|
||||||
STATUSPAGE_DB_PASSWORD=
|
STATUSPAGE_DB_PASSWORD=
|
||||||
+12
-11
@@ -1,20 +1,21 @@
|
|||||||
# Shared PostgreSQL
|
# Shared PostgreSQL
|
||||||
|
|
||||||
This directory contains the shared PostgreSQL 17 deployment for Authentik,
|
This directory contains the shared PostgreSQL 17 deployment for Authentik,
|
||||||
Gitea, NetBox, Netronome, and Statuspage. It creates one database and one login role
|
Gitea, NetBox, Netronome, Paperless-ngx, and Statuspage. It creates one database
|
||||||
per service. Per-service standalone databases were removed after the
|
and one login role per service. Per-service standalone databases were removed
|
||||||
migration (Sep 2026); Penpot stays on its own compose PostgreSQL (archived,
|
after the migration (Sep 2026); Penpot stays on its own compose PostgreSQL
|
||||||
not part of the shared instance).
|
(archived, not part of the shared instance).
|
||||||
|
|
||||||
## Compatibility baseline
|
## Compatibility baseline
|
||||||
|
|
||||||
| Service | Current application | Shared PostgreSQL 17 |
|
| Service | Current application | Shared PostgreSQL 17 |
|
||||||
| ---------- | ------------------- | -------------------------------------- |
|
| ------------- | ------------------- | -------------------------------------- |
|
||||||
| Authentik | 2025.10.x | Supported (Authentik requires 14+) |
|
| Authentik | 2025.10.x | Supported (Authentik requires 14+) |
|
||||||
| Gitea | 1.27.3 | Supported (Gitea requires 12+) |
|
| Gitea | 1.27.3 | Supported (Gitea requires 12+) |
|
||||||
| NetBox | 4.7.x | Supported (NetBox 4.x requires 13+) |
|
| NetBox | 4.7.x | Supported (NetBox 4.x requires 13+) |
|
||||||
| Netronome | 0.14.0 | Supported (upstream's example uses 17) |
|
| Netronome | 0.14.0 | Supported (upstream's example uses 17) |
|
||||||
| Statuspage | custom | Supported |
|
| Paperless-ngx | 3.2.1 | Supported |
|
||||||
|
| Statuspage | custom | Supported |
|
||||||
|
|
||||||
A major-version change must use a logical dump/restore; changing only the
|
A major-version change must use a logical dump/restore; changing only the
|
||||||
image tag while keeping a data directory is not supported.
|
image tag while keeping a data directory is not supported.
|
||||||
|
|||||||
@@ -6,6 +6,7 @@ set -euo pipefail
|
|||||||
: "${NETBOX_DB_PASSWORD:?NETBOX_DB_PASSWORD is required}"
|
: "${NETBOX_DB_PASSWORD:?NETBOX_DB_PASSWORD is required}"
|
||||||
: "${NETRONOME_DB_PASSWORD:?NETRONOME_DB_PASSWORD is required}"
|
: "${NETRONOME_DB_PASSWORD:?NETRONOME_DB_PASSWORD is required}"
|
||||||
: "${PENPOT_DB_PASSWORD:?PENPOT_DB_PASSWORD is required}"
|
: "${PENPOT_DB_PASSWORD:?PENPOT_DB_PASSWORD is required}"
|
||||||
|
: "${PAPERLESS_DB_PASSWORD:?PAPERLESS_DB_PASSWORD is required}"
|
||||||
: "${STATUSPAGE_DB_PASSWORD:?STATUSPAGE_DB_PASSWORD is required}"
|
: "${STATUSPAGE_DB_PASSWORD:?STATUSPAGE_DB_PASSWORD is required}"
|
||||||
|
|
||||||
create_role_and_database() {
|
create_role_and_database() {
|
||||||
@@ -27,4 +28,5 @@ create_role_and_database gitea gitea "$GITEA_DB_PASSWORD"
|
|||||||
create_role_and_database netbox netbox "$NETBOX_DB_PASSWORD"
|
create_role_and_database netbox netbox "$NETBOX_DB_PASSWORD"
|
||||||
create_role_and_database netronome netronome "$NETRONOME_DB_PASSWORD"
|
create_role_and_database netronome netronome "$NETRONOME_DB_PASSWORD"
|
||||||
create_role_and_database penpot penpot "$PENPOT_DB_PASSWORD"
|
create_role_and_database penpot penpot "$PENPOT_DB_PASSWORD"
|
||||||
|
create_role_and_database paperless paperless "$PAPERLESS_DB_PASSWORD"
|
||||||
create_role_and_database statuspage statuspage "$STATUSPAGE_DB_PASSWORD"
|
create_role_and_database statuspage statuspage "$STATUSPAGE_DB_PASSWORD"
|
||||||
@@ -26,6 +26,18 @@ spec:
|
|||||||
- namespaceSelector:
|
- namespaceSelector:
|
||||||
matchLabels:
|
matchLabels:
|
||||||
kubernetes.io/metadata.name: penpot
|
kubernetes.io/metadata.name: penpot
|
||||||
|
- namespaceSelector:
|
||||||
|
matchLabels:
|
||||||
|
kubernetes.io/metadata.name: paperless
|
||||||
|
podSelector:
|
||||||
|
matchLabels:
|
||||||
|
app.kubernetes.io/name: paperless
|
||||||
|
- namespaceSelector:
|
||||||
|
matchLabels:
|
||||||
|
kubernetes.io/metadata.name: database
|
||||||
|
podSelector:
|
||||||
|
matchLabels:
|
||||||
|
app.kubernetes.io/name: paperless-database-init
|
||||||
- namespaceSelector:
|
- namespaceSelector:
|
||||||
matchLabels:
|
matchLabels:
|
||||||
kubernetes.io/metadata.name: statuspage
|
kubernetes.io/metadata.name: statuspage
|
||||||
|
|||||||
@@ -126,6 +126,7 @@ data:
|
|||||||
: "${NETBOX_DB_PASSWORD:?NETBOX_DB_PASSWORD is required}"
|
: "${NETBOX_DB_PASSWORD:?NETBOX_DB_PASSWORD is required}"
|
||||||
: "${NETRONOME_DB_PASSWORD:?NETRONOME_DB_PASSWORD is required}"
|
: "${NETRONOME_DB_PASSWORD:?NETRONOME_DB_PASSWORD is required}"
|
||||||
: "${PENPOT_DB_PASSWORD:?PENPOT_DB_PASSWORD is required}"
|
: "${PENPOT_DB_PASSWORD:?PENPOT_DB_PASSWORD is required}"
|
||||||
|
: "${PAPERLESS_DB_PASSWORD:?PAPERLESS_DB_PASSWORD is required}"
|
||||||
: "${STATUSPAGE_DB_PASSWORD:?STATUSPAGE_DB_PASSWORD is required}"
|
: "${STATUSPAGE_DB_PASSWORD:?STATUSPAGE_DB_PASSWORD is required}"
|
||||||
|
|
||||||
create_role_and_database() {
|
create_role_and_database() {
|
||||||
@@ -147,4 +148,5 @@ data:
|
|||||||
create_role_and_database netbox netbox "$NETBOX_DB_PASSWORD"
|
create_role_and_database netbox netbox "$NETBOX_DB_PASSWORD"
|
||||||
create_role_and_database netronome netronome "$NETRONOME_DB_PASSWORD"
|
create_role_and_database netronome netronome "$NETRONOME_DB_PASSWORD"
|
||||||
create_role_and_database penpot penpot "$PENPOT_DB_PASSWORD"
|
create_role_and_database penpot penpot "$PENPOT_DB_PASSWORD"
|
||||||
|
create_role_and_database paperless paperless "$PAPERLESS_DB_PASSWORD"
|
||||||
create_role_and_database statuspage statuspage "$STATUSPAGE_DB_PASSWORD"
|
create_role_and_database statuspage statuspage "$STATUSPAGE_DB_PASSWORD"
|
||||||
@@ -11,4 +11,5 @@ stringData:
|
|||||||
NETBOX_DB_PASSWORD: ""
|
NETBOX_DB_PASSWORD: ""
|
||||||
NETRONOME_DB_PASSWORD: ""
|
NETRONOME_DB_PASSWORD: ""
|
||||||
PENPOT_DB_PASSWORD: ""
|
PENPOT_DB_PASSWORD: ""
|
||||||
|
PAPERLESS_DB_PASSWORD: ""
|
||||||
STATUSPAGE_DB_PASSWORD: ""
|
STATUSPAGE_DB_PASSWORD: ""
|
||||||
@@ -0,0 +1,43 @@
|
|||||||
|
# VictoriaMetrics
|
||||||
|
|
||||||
|
The `victoria-operator` Helm release converts Prometheus Operator
|
||||||
|
`ServiceMonitor` resources into owned `VMServiceScrape` resources. The
|
||||||
|
`VMAgent` selects converted scrapes labeled `release: prometheus-stack` in all
|
||||||
|
namespaces and writes them to the existing single-node VictoriaMetrics
|
||||||
|
instance. Changes to selected `ServiceMonitor` resources are reconciled
|
||||||
|
automatically; there is no copied Prometheus scrape-config blob to regenerate.
|
||||||
|
|
||||||
|
The agent drops targets for the Prometheus server service to avoid duplicating
|
||||||
|
its self-scrape. `scraper: victoria` identifies the samples ingested by this
|
||||||
|
VMAgent.
|
||||||
|
|
||||||
|
The VictoriaMetrics Operator chart and its CRDs are installed before the
|
||||||
|
Kubernetes manifests by the normal deploy workflow. On a cluster where the
|
||||||
|
operator CRDs are not installed yet, CI skips the server-side dry-run of the
|
||||||
|
`VMAgent` resource; the deploy installs the chart before applying that resource.
|
||||||
|
|
||||||
|
## Application metrics
|
||||||
|
|
||||||
|
The application ServiceMonitors use a 30s interval and a 10s timeout:
|
||||||
|
|
||||||
|
- Headscale: the external Service points to the Compose host on port 19090.
|
||||||
|
A VMServiceScrape uses EndpointSlice discovery for this manually managed target.
|
||||||
|
The Compose configuration must bind metrics to `0.0.0.0:9090`.
|
||||||
|
- NetBird: the combined server exports `/metrics` on port 9090. The existing
|
||||||
|
`server.metricsPort` setting enables the listener.
|
||||||
|
- Gitea: `GITEA__metrics__ENABLED` enables `/metrics` on the HTTP port. The public
|
||||||
|
ingress excludes this path. The monitor uses the internal Service directly.
|
||||||
|
- Immich: `IMMICH_TELEMETRY_INCLUDE=all` enables API and worker metrics on ports
|
||||||
|
8081 and 8082. The monitor scrapes both ports on each server replica.
|
||||||
|
|
||||||
|
Deploy through the existing CI and deploy workflow. Gitea and Immich reload their
|
||||||
|
ConfigMap changes through Reloader. Check the VMAgent targets after deployment
|
||||||
|
and query `up{scraper="victoria",namespace=~"netbird|gitea|immich|headscale"}` in
|
||||||
|
VictoriaMetrics. All targets should report 1.
|
||||||
|
|
||||||
|
For rollback, revert the application metrics changes, run CI, and deploy the
|
||||||
|
revert. Remove the three application ServiceMonitors and the Headscale VMServiceScrape explicitly: the deployment
|
||||||
|
workflow applies manifests and does not prune removed resources.
|
||||||
|
|
||||||
|
For Headscale rollback, remove its VMServiceScrape and Service label, restore the
|
||||||
|
previous Compose metrics bind address, and restart only the Headscale service.
|
||||||
@@ -38,6 +38,8 @@ grafana:
|
|||||||
|
|
||||||
# One block covers both the dashboards and datasources sidecars (p95 91M / 80M).
|
# One block covers both the dashboards and datasources sidecars (p95 91M / 80M).
|
||||||
sidecar:
|
sidecar:
|
||||||
|
datasources:
|
||||||
|
defaultDatasourceEnabled: false
|
||||||
resources:
|
resources:
|
||||||
requests:
|
requests:
|
||||||
memory: "96Mi"
|
memory: "96Mi"
|
||||||
@@ -50,9 +52,18 @@ grafana:
|
|||||||
type: loki
|
type: loki
|
||||||
url: http://loki-gateway.prometheus.svc.cluster.local
|
url: http://loki-gateway.prometheus.svc.cluster.local
|
||||||
access: proxy
|
access: proxy
|
||||||
|
- name: VictoriaMetrics
|
||||||
|
type: prometheus
|
||||||
|
url: http://victoria-metrics.prometheus.svc.cluster.local:8428
|
||||||
|
access: proxy
|
||||||
|
isDefault: true
|
||||||
|
|
||||||
prometheus:
|
prometheus:
|
||||||
prometheusSpec:
|
prometheusSpec:
|
||||||
|
# VM trial: vmagent scrapes and remote-writes to VictoriaMetrics, so the
|
||||||
|
# Prometheus server itself stands down. Encoded here (not a kubectl patch)
|
||||||
|
# so helm keeps owning spec.replicas and upgrades do not conflict on it.
|
||||||
|
replicas: 0
|
||||||
retention: 60d
|
retention: 60d
|
||||||
retentionSize: 32GB
|
retentionSize: 32GB
|
||||||
storageSpec:
|
storageSpec:
|
||||||
|
|||||||
@@ -31,3 +31,105 @@ spec:
|
|||||||
port: 80
|
port: 80
|
||||||
tls:
|
tls:
|
||||||
secretName: internal-wildcard-tls
|
secretName: internal-wildcard-tls
|
||||||
|
---
|
||||||
|
apiVersion: traefik.io/v1alpha1
|
||||||
|
kind: IngressRoute
|
||||||
|
metadata:
|
||||||
|
name: prometheus-local
|
||||||
|
namespace: prometheus
|
||||||
|
spec:
|
||||||
|
entryPoints:
|
||||||
|
- websecure
|
||||||
|
routes:
|
||||||
|
- match: Host(`prom.workstation.internal`) || Host(`prom.gigaforust.internal`)
|
||||||
|
kind: Rule
|
||||||
|
services:
|
||||||
|
- name: prometheus-stack-kube-prom-prometheus
|
||||||
|
port: 9090
|
||||||
|
tls:
|
||||||
|
secretName: internal-wildcard-tls
|
||||||
|
---
|
||||||
|
apiVersion: traefik.io/v1alpha1
|
||||||
|
kind: IngressRoute
|
||||||
|
metadata:
|
||||||
|
name: alertmanager-local
|
||||||
|
namespace: prometheus
|
||||||
|
spec:
|
||||||
|
entryPoints:
|
||||||
|
- websecure
|
||||||
|
routes:
|
||||||
|
- match: Host(`am.workstation.internal`) || Host(`am.gigaforust.internal`)
|
||||||
|
kind: Rule
|
||||||
|
services:
|
||||||
|
- name: prometheus-stack-kube-prom-alertmanager
|
||||||
|
port: 9093
|
||||||
|
tls:
|
||||||
|
secretName: internal-wildcard-tls
|
||||||
|
---
|
||||||
|
apiVersion: traefik.io/v1alpha1
|
||||||
|
kind: IngressRoute
|
||||||
|
metadata:
|
||||||
|
name: loki-local
|
||||||
|
namespace: prometheus
|
||||||
|
spec:
|
||||||
|
entryPoints:
|
||||||
|
- websecure
|
||||||
|
routes:
|
||||||
|
- match: Host(`loki.workstation.internal`) || Host(`loki.gigaforust.internal`)
|
||||||
|
kind: Rule
|
||||||
|
services:
|
||||||
|
- name: loki-gateway
|
||||||
|
port: 80
|
||||||
|
tls:
|
||||||
|
secretName: internal-wildcard-tls
|
||||||
|
---
|
||||||
|
apiVersion: traefik.io/v1alpha1
|
||||||
|
kind: IngressRoute
|
||||||
|
metadata:
|
||||||
|
name: alloy-local
|
||||||
|
namespace: prometheus
|
||||||
|
spec:
|
||||||
|
entryPoints:
|
||||||
|
- websecure
|
||||||
|
routes:
|
||||||
|
- match: Host(`alloy.workstation.internal`) || Host(`alloy.gigaforust.internal`)
|
||||||
|
kind: Rule
|
||||||
|
services:
|
||||||
|
- name: alloy
|
||||||
|
port: 12345
|
||||||
|
tls:
|
||||||
|
secretName: internal-wildcard-tls
|
||||||
|
---
|
||||||
|
apiVersion: traefik.io/v1alpha1
|
||||||
|
kind: IngressRoute
|
||||||
|
metadata:
|
||||||
|
name: victoria-local
|
||||||
|
namespace: prometheus
|
||||||
|
spec:
|
||||||
|
entryPoints:
|
||||||
|
- websecure
|
||||||
|
routes:
|
||||||
|
- match: Host(`victoria.workstation.internal`) || Host(`victoria.gigaforust.internal`)
|
||||||
|
kind: Rule
|
||||||
|
services:
|
||||||
|
- name: victoria-metrics
|
||||||
|
port: 8428
|
||||||
|
tls:
|
||||||
|
secretName: internal-wildcard-tls
|
||||||
|
---
|
||||||
|
apiVersion: traefik.io/v1alpha1
|
||||||
|
kind: IngressRoute
|
||||||
|
metadata:
|
||||||
|
name: vmalert-local
|
||||||
|
namespace: prometheus
|
||||||
|
spec:
|
||||||
|
entryPoints:
|
||||||
|
- websecure
|
||||||
|
routes:
|
||||||
|
- match: Host(`vmalert.workstation.internal`) || Host(`vmalert.gigaforust.internal`)
|
||||||
|
kind: Rule
|
||||||
|
services:
|
||||||
|
- name: vmalert
|
||||||
|
port: 8880
|
||||||
|
tls:
|
||||||
|
secretName: internal-wildcard-tls
|
||||||
@@ -0,0 +1,12 @@
|
|||||||
|
nameOverride: victoria-operator
|
||||||
|
|
||||||
|
operator:
|
||||||
|
enable_converter_ownership: true
|
||||||
|
|
||||||
|
resources:
|
||||||
|
requests:
|
||||||
|
cpu: 50m
|
||||||
|
memory: 96Mi
|
||||||
|
limits:
|
||||||
|
cpu: 200m
|
||||||
|
memory: 256Mi
|
||||||
@@ -0,0 +1,79 @@
|
|||||||
|
apiVersion: v1
|
||||||
|
kind: Service
|
||||||
|
metadata:
|
||||||
|
name: victoria-metrics
|
||||||
|
namespace: prometheus
|
||||||
|
spec:
|
||||||
|
selector:
|
||||||
|
app: victoria-metrics
|
||||||
|
ports:
|
||||||
|
- port: 8428
|
||||||
|
targetPort: 8428
|
||||||
|
---
|
||||||
|
apiVersion: v1
|
||||||
|
kind: PersistentVolumeClaim
|
||||||
|
metadata:
|
||||||
|
name: victoria-pvc
|
||||||
|
namespace: prometheus
|
||||||
|
spec:
|
||||||
|
resources:
|
||||||
|
requests:
|
||||||
|
storage: 10Gi
|
||||||
|
volumeMode: Filesystem
|
||||||
|
accessModes:
|
||||||
|
- ReadWriteOnce
|
||||||
|
---
|
||||||
|
apiVersion: apps/v1
|
||||||
|
kind: Deployment
|
||||||
|
metadata:
|
||||||
|
name: victoria-deployment
|
||||||
|
namespace: prometheus
|
||||||
|
spec:
|
||||||
|
replicas: 1
|
||||||
|
selector:
|
||||||
|
matchLabels:
|
||||||
|
app: victoria-metrics
|
||||||
|
strategy:
|
||||||
|
type: Recreate
|
||||||
|
template:
|
||||||
|
metadata:
|
||||||
|
labels:
|
||||||
|
app: victoria-metrics
|
||||||
|
spec:
|
||||||
|
containers:
|
||||||
|
- name: victoria
|
||||||
|
image: victoriametrics/victoria-metrics:v1.153.0-scratch
|
||||||
|
args:
|
||||||
|
- -storageDataPath=/vmdata
|
||||||
|
- -retentionPeriod=30d
|
||||||
|
- -httpListenAddr=:8428
|
||||||
|
ports:
|
||||||
|
- containerPort: 8428
|
||||||
|
readinessProbe:
|
||||||
|
httpGet:
|
||||||
|
path: /health
|
||||||
|
port: 8428
|
||||||
|
initialDelaySeconds: 15
|
||||||
|
periodSeconds: 10
|
||||||
|
failureThreshold: 6
|
||||||
|
livenessProbe:
|
||||||
|
httpGet:
|
||||||
|
path: /health
|
||||||
|
port: 8428
|
||||||
|
initialDelaySeconds: 60
|
||||||
|
periodSeconds: 30
|
||||||
|
failureThreshold: 3
|
||||||
|
volumeMounts:
|
||||||
|
- name: vmdata
|
||||||
|
mountPath: /vmdata
|
||||||
|
resources:
|
||||||
|
requests:
|
||||||
|
cpu: "100m"
|
||||||
|
memory: "256Mi"
|
||||||
|
limits:
|
||||||
|
cpu: "1000m"
|
||||||
|
memory: "1Gi"
|
||||||
|
volumes:
|
||||||
|
- name: vmdata
|
||||||
|
persistentVolumeClaim:
|
||||||
|
claimName: victoria-pvc
|
||||||
@@ -0,0 +1,29 @@
|
|||||||
|
apiVersion: operator.victoriametrics.com/v1beta1
|
||||||
|
kind: VMAgent
|
||||||
|
metadata:
|
||||||
|
name: vmagent
|
||||||
|
namespace: prometheus
|
||||||
|
spec:
|
||||||
|
image:
|
||||||
|
tag: v1.153.0
|
||||||
|
scrapeInterval: 30s
|
||||||
|
externalLabels:
|
||||||
|
scraper: victoria
|
||||||
|
serviceScrapeNamespaceSelector: {}
|
||||||
|
serviceScrapeSelector:
|
||||||
|
matchLabels:
|
||||||
|
release: prometheus-stack
|
||||||
|
globalScrapeRelabelConfigs:
|
||||||
|
- action: drop
|
||||||
|
source_labels:
|
||||||
|
- __meta_kubernetes_service_name
|
||||||
|
regex: prometheus-stack-kube-prom-prometheus
|
||||||
|
remoteWrite:
|
||||||
|
- url: http://victoria-metrics.prometheus.svc.cluster.local:8428/api/v1/write
|
||||||
|
resources:
|
||||||
|
requests:
|
||||||
|
cpu: 100m
|
||||||
|
memory: 256Mi
|
||||||
|
limits:
|
||||||
|
cpu: "1000m"
|
||||||
|
memory: 1Gi
|
||||||
@@ -0,0 +1,70 @@
|
|||||||
|
apiVersion: v1
|
||||||
|
kind: Service
|
||||||
|
metadata:
|
||||||
|
name: vmalert
|
||||||
|
namespace: prometheus
|
||||||
|
spec:
|
||||||
|
selector:
|
||||||
|
app: vmalert
|
||||||
|
ports:
|
||||||
|
- port: 8880
|
||||||
|
targetPort: 8880
|
||||||
|
---
|
||||||
|
apiVersion: apps/v1
|
||||||
|
kind: Deployment
|
||||||
|
metadata:
|
||||||
|
name: vmalert-deployment
|
||||||
|
namespace: prometheus
|
||||||
|
spec:
|
||||||
|
replicas: 1
|
||||||
|
selector:
|
||||||
|
matchLabels:
|
||||||
|
app: vmalert
|
||||||
|
strategy:
|
||||||
|
type: Recreate
|
||||||
|
template:
|
||||||
|
metadata:
|
||||||
|
labels:
|
||||||
|
app: vmalert
|
||||||
|
spec:
|
||||||
|
containers:
|
||||||
|
- name: vmalert
|
||||||
|
image: victoriametrics/vmalert:v1.153.0
|
||||||
|
args:
|
||||||
|
- -datasource.url=http://victoria-metrics.prometheus.svc.cluster.local:8428
|
||||||
|
- -remoteWrite.url=http://victoria-metrics.prometheus.svc.cluster.local:8428
|
||||||
|
- -notifier.url=http://prometheus-stack-kube-prom-alertmanager.prometheus.svc.cluster.local:9093
|
||||||
|
- -rule=/etc/vm/rules/*.yaml
|
||||||
|
- -evaluationInterval=60s
|
||||||
|
- -httpListenAddr=:8880
|
||||||
|
ports:
|
||||||
|
- containerPort: 8880
|
||||||
|
readinessProbe:
|
||||||
|
httpGet:
|
||||||
|
path: /metrics
|
||||||
|
port: 8880
|
||||||
|
initialDelaySeconds: 15
|
||||||
|
periodSeconds: 10
|
||||||
|
failureThreshold: 6
|
||||||
|
livenessProbe:
|
||||||
|
httpGet:
|
||||||
|
path: /metrics
|
||||||
|
port: 8880
|
||||||
|
initialDelaySeconds: 60
|
||||||
|
periodSeconds: 30
|
||||||
|
failureThreshold: 3
|
||||||
|
volumeMounts:
|
||||||
|
- name: rules
|
||||||
|
mountPath: /etc/vm/rules
|
||||||
|
readOnly: true
|
||||||
|
resources:
|
||||||
|
requests:
|
||||||
|
cpu: "50m"
|
||||||
|
memory: "64Mi"
|
||||||
|
limits:
|
||||||
|
cpu: "200m"
|
||||||
|
memory: "256Mi"
|
||||||
|
volumes:
|
||||||
|
- name: rules
|
||||||
|
configMap:
|
||||||
|
name: prometheus-prometheus-stack-kube-prom-prometheus-rulefiles-0
|
||||||
+43
-49
@@ -19,6 +19,9 @@ data:
|
|||||||
"dependencyDashboard": true,
|
"dependencyDashboard": true,
|
||||||
"prCreation": "immediate",
|
"prCreation": "immediate",
|
||||||
"labels": ["dependencies", "automated"],
|
"labels": ["dependencies", "automated"],
|
||||||
|
"docker-compose": {
|
||||||
|
"managerFilePatterns": ["renovate/renovate-compose.yaml"]
|
||||||
|
},
|
||||||
"helm-values": {
|
"helm-values": {
|
||||||
"managerFilePatterns": ["/k8s/.+values\\.ya?ml$/"]
|
"managerFilePatterns": ["/k8s/.+values\\.ya?ml$/"]
|
||||||
},
|
},
|
||||||
@@ -26,35 +29,28 @@ data:
|
|||||||
"managerFilePatterns": ["/k8s/.+\\.ya?ml$/"]
|
"managerFilePatterns": ["/k8s/.+\\.ya?ml$/"]
|
||||||
},
|
},
|
||||||
"customManagers": [
|
"customManagers": [
|
||||||
{
|
|
||||||
"customType": "regex",
|
|
||||||
"description": "singlesource: playwright npm version pinned in npx command (k8s + compose)",
|
|
||||||
"managerFilePatterns": ["^edu_master/k8s/playwright\\.yaml$", "^edu_master/compose\\.yaml$"],
|
|
||||||
"matchStrings": ["playwright@(?<currentValue>\\d+\\.\\d+\\.\\d+)"],
|
|
||||||
"datasourceTemplate": "npm",
|
|
||||||
"depNameTemplate": "playwright"
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"customType": "regex",
|
|
||||||
"description": "singlesource: PLAYWRIGHT_VERSION file",
|
|
||||||
"managerFilePatterns": ["^edu_master/PLAYWRIGHT_VERSION$"],
|
|
||||||
"matchStrings": ["^(?<currentValue>\\d+\\.\\d+\\.\\d+)$"],
|
|
||||||
"datasourceTemplate": "pypi",
|
|
||||||
"depNameTemplate": "playwright"
|
|
||||||
},
|
|
||||||
{
|
{
|
||||||
"customType": "regex",
|
"customType": "regex",
|
||||||
"description": "kube-prometheus-stack chart version pinned in the deploy workflow",
|
"description": "kube-prometheus-stack chart version pinned in the deploy workflow",
|
||||||
"managerFilePatterns": ["^\\.gitea/workflows/deploy-lib\\.sh$"],
|
"managerFilePatterns": [".gitea/workflows/deploy-lib.sh"],
|
||||||
"matchStrings": ["\\|prometheus-community/kube-prometheus-stack\\|prometheus\\|(?<currentValue>[0-9.]+)\\|"],
|
"matchStrings": ["\\|prometheus-community/kube-prometheus-stack\\|prometheus\\|(?<currentValue>[0-9.]+)\\|"],
|
||||||
"datasourceTemplate": "helm",
|
"datasourceTemplate": "helm",
|
||||||
"depNameTemplate": "kube-prometheus-stack",
|
"depNameTemplate": "kube-prometheus-stack",
|
||||||
"registryUrlTemplate": "https://prometheus-community.github.io/helm-charts"
|
"registryUrlTemplate": "https://prometheus-community.github.io/helm-charts"
|
||||||
},
|
},
|
||||||
|
{
|
||||||
|
"customType": "regex",
|
||||||
|
"description": "VictoriaMetrics Operator chart version pinned in the deploy workflow",
|
||||||
|
"managerFilePatterns": [".gitea/workflows/deploy-lib.sh"],
|
||||||
|
"matchStrings": ["\\|victoriametrics/victoria-metrics-operator\\|prometheus\\|(?<currentValue>[0-9.]+)\\|"],
|
||||||
|
"datasourceTemplate": "helm",
|
||||||
|
"depNameTemplate": "victoria-metrics-operator",
|
||||||
|
"registryUrlTemplate": "https://victoriametrics.github.io/helm-charts"
|
||||||
|
},
|
||||||
{
|
{
|
||||||
"customType": "regex",
|
"customType": "regex",
|
||||||
"description": "grafana/loki chart version pinned in the deploy workflow",
|
"description": "grafana/loki chart version pinned in the deploy workflow",
|
||||||
"managerFilePatterns": ["^\\.gitea/workflows/deploy-lib\\.sh$"],
|
"managerFilePatterns": [".gitea/workflows/deploy-lib.sh"],
|
||||||
"matchStrings": ["\\|grafana/loki\\|prometheus\\|(?<currentValue>[0-9.]+)\\|"],
|
"matchStrings": ["\\|grafana/loki\\|prometheus\\|(?<currentValue>[0-9.]+)\\|"],
|
||||||
"datasourceTemplate": "helm",
|
"datasourceTemplate": "helm",
|
||||||
"depNameTemplate": "loki",
|
"depNameTemplate": "loki",
|
||||||
@@ -63,7 +59,7 @@ data:
|
|||||||
{
|
{
|
||||||
"customType": "regex",
|
"customType": "regex",
|
||||||
"description": "grafana/alloy chart version pinned in the deploy workflow",
|
"description": "grafana/alloy chart version pinned in the deploy workflow",
|
||||||
"managerFilePatterns": ["^\\.gitea/workflows/deploy-lib\\.sh$"],
|
"managerFilePatterns": [".gitea/workflows/deploy-lib.sh"],
|
||||||
"matchStrings": ["\\|grafana/alloy\\|prometheus\\|(?<currentValue>[0-9.]+)\\|"],
|
"matchStrings": ["\\|grafana/alloy\\|prometheus\\|(?<currentValue>[0-9.]+)\\|"],
|
||||||
"datasourceTemplate": "helm",
|
"datasourceTemplate": "helm",
|
||||||
"depNameTemplate": "alloy",
|
"depNameTemplate": "alloy",
|
||||||
@@ -72,7 +68,7 @@ data:
|
|||||||
{
|
{
|
||||||
"customType": "regex",
|
"customType": "regex",
|
||||||
"description": "actionlint version used by the ci workflow",
|
"description": "actionlint version used by the ci workflow",
|
||||||
"managerFilePatterns": ["^\\.gitea/workflows/tool-versions\\.env$"],
|
"managerFilePatterns": [".gitea/workflows/tool-versions.env"],
|
||||||
"matchStrings": ["(?:^|\\n)ACTIONLINT_VERSION=\"(?<currentValue>[0-9.]+)\""],
|
"matchStrings": ["(?:^|\\n)ACTIONLINT_VERSION=\"(?<currentValue>[0-9.]+)\""],
|
||||||
"datasourceTemplate": "github-tags",
|
"datasourceTemplate": "github-tags",
|
||||||
"depNameTemplate": "rhysd/actionlint"
|
"depNameTemplate": "rhysd/actionlint"
|
||||||
@@ -80,7 +76,7 @@ data:
|
|||||||
{
|
{
|
||||||
"customType": "regex",
|
"customType": "regex",
|
||||||
"description": "shellcheck version used by the ci workflow",
|
"description": "shellcheck version used by the ci workflow",
|
||||||
"managerFilePatterns": ["^\\.gitea/workflows/tool-versions\\.env$"],
|
"managerFilePatterns": [".gitea/workflows/tool-versions.env"],
|
||||||
"matchStrings": ["(?:^|\\n)SHELLCHECK_VERSION=\"(?<currentValue>[0-9.]+)\""],
|
"matchStrings": ["(?:^|\\n)SHELLCHECK_VERSION=\"(?<currentValue>[0-9.]+)\""],
|
||||||
"datasourceTemplate": "github-tags",
|
"datasourceTemplate": "github-tags",
|
||||||
"depNameTemplate": "koalaman/shellcheck"
|
"depNameTemplate": "koalaman/shellcheck"
|
||||||
@@ -88,7 +84,7 @@ data:
|
|||||||
{
|
{
|
||||||
"customType": "regex",
|
"customType": "regex",
|
||||||
"description": "kubeconform version used by the ci workflow",
|
"description": "kubeconform version used by the ci workflow",
|
||||||
"managerFilePatterns": ["^\\.gitea/workflows/tool-versions\\.env$"],
|
"managerFilePatterns": [".gitea/workflows/tool-versions.env"],
|
||||||
"matchStrings": ["(?:^|\\n)KUBECONFORM_VERSION=\"(?<currentValue>[0-9.]+)\""],
|
"matchStrings": ["(?:^|\\n)KUBECONFORM_VERSION=\"(?<currentValue>[0-9.]+)\""],
|
||||||
"datasourceTemplate": "github-tags",
|
"datasourceTemplate": "github-tags",
|
||||||
"depNameTemplate": "yannh/kubeconform"
|
"depNameTemplate": "yannh/kubeconform"
|
||||||
@@ -96,7 +92,7 @@ data:
|
|||||||
{
|
{
|
||||||
"customType": "regex",
|
"customType": "regex",
|
||||||
"description": "uv version used to build the pytest venv",
|
"description": "uv version used to build the pytest venv",
|
||||||
"managerFilePatterns": ["^\\.gitea/workflows/tool-versions\\.env$"],
|
"managerFilePatterns": [".gitea/workflows/tool-versions.env"],
|
||||||
"matchStrings": ["(?:^|\\n)UV_VERSION=\"(?<currentValue>[0-9.]+)\""],
|
"matchStrings": ["(?:^|\\n)UV_VERSION=\"(?<currentValue>[0-9.]+)\""],
|
||||||
"datasourceTemplate": "github-tags",
|
"datasourceTemplate": "github-tags",
|
||||||
"depNameTemplate": "astral-sh/uv"
|
"depNameTemplate": "astral-sh/uv"
|
||||||
@@ -104,7 +100,7 @@ data:
|
|||||||
{
|
{
|
||||||
"customType": "regex",
|
"customType": "regex",
|
||||||
"description": "prettier version used by the ci workflow",
|
"description": "prettier version used by the ci workflow",
|
||||||
"managerFilePatterns": ["^\\.gitea/workflows/tool-versions\\.env$"],
|
"managerFilePatterns": [".gitea/workflows/tool-versions.env"],
|
||||||
"matchStrings": ["(?:^|\\n)PRETTIER_VERSION=\"(?<currentValue>[0-9.]+)\""],
|
"matchStrings": ["(?:^|\\n)PRETTIER_VERSION=\"(?<currentValue>[0-9.]+)\""],
|
||||||
"datasourceTemplate": "npm",
|
"datasourceTemplate": "npm",
|
||||||
"depNameTemplate": "prettier"
|
"depNameTemplate": "prettier"
|
||||||
@@ -112,7 +108,7 @@ data:
|
|||||||
{
|
{
|
||||||
"customType": "regex",
|
"customType": "regex",
|
||||||
"description": "ruff version used by the ci workflow",
|
"description": "ruff version used by the ci workflow",
|
||||||
"managerFilePatterns": ["^\\.gitea/workflows/tool-versions\\.env$"],
|
"managerFilePatterns": [".gitea/workflows/tool-versions.env"],
|
||||||
"matchStrings": ["(?:^|\\n)RUFF_VERSION=\"(?<currentValue>[0-9.]+)\""],
|
"matchStrings": ["(?:^|\\n)RUFF_VERSION=\"(?<currentValue>[0-9.]+)\""],
|
||||||
"datasourceTemplate": "pypi",
|
"datasourceTemplate": "pypi",
|
||||||
"depNameTemplate": "ruff"
|
"depNameTemplate": "ruff"
|
||||||
@@ -120,7 +116,7 @@ data:
|
|||||||
{
|
{
|
||||||
"customType": "regex",
|
"customType": "regex",
|
||||||
"description": "pip-audit version used by the ci workflow",
|
"description": "pip-audit version used by the ci workflow",
|
||||||
"managerFilePatterns": ["^\\.gitea/workflows/tool-versions\\.env$"],
|
"managerFilePatterns": [".gitea/workflows/tool-versions.env"],
|
||||||
"matchStrings": ["(?:^|\\n)PIP_AUDIT_VERSION=\"(?<currentValue>[0-9.]+)\""],
|
"matchStrings": ["(?:^|\\n)PIP_AUDIT_VERSION=\"(?<currentValue>[0-9.]+)\""],
|
||||||
"datasourceTemplate": "pypi",
|
"datasourceTemplate": "pypi",
|
||||||
"depNameTemplate": "pip-audit"
|
"depNameTemplate": "pip-audit"
|
||||||
@@ -128,7 +124,7 @@ data:
|
|||||||
{
|
{
|
||||||
"customType": "regex",
|
"customType": "regex",
|
||||||
"description": "yamllint version used by the ci workflow",
|
"description": "yamllint version used by the ci workflow",
|
||||||
"managerFilePatterns": ["^\\.gitea/workflows/tool-versions\\.env$"],
|
"managerFilePatterns": [".gitea/workflows/tool-versions.env"],
|
||||||
"matchStrings": ["(?:^|\\n)YAMLLINT_VERSION=\"(?<currentValue>[0-9.]+)\""],
|
"matchStrings": ["(?:^|\\n)YAMLLINT_VERSION=\"(?<currentValue>[0-9.]+)\""],
|
||||||
"datasourceTemplate": "pypi",
|
"datasourceTemplate": "pypi",
|
||||||
"depNameTemplate": "yamllint"
|
"depNameTemplate": "yamllint"
|
||||||
@@ -136,7 +132,7 @@ data:
|
|||||||
{
|
{
|
||||||
"customType": "regex",
|
"customType": "regex",
|
||||||
"description": "hadolint version used by the ci workflow",
|
"description": "hadolint version used by the ci workflow",
|
||||||
"managerFilePatterns": ["^\\.gitea/workflows/tool-versions\\.env$"],
|
"managerFilePatterns": [".gitea/workflows/tool-versions.env"],
|
||||||
"matchStrings": ["(?:^|\\n)HADOLINT_VERSION=\"(?<currentValue>[0-9.]+)\""],
|
"matchStrings": ["(?:^|\\n)HADOLINT_VERSION=\"(?<currentValue>[0-9.]+)\""],
|
||||||
"datasourceTemplate": "github-tags",
|
"datasourceTemplate": "github-tags",
|
||||||
"depNameTemplate": "hadolint/hadolint"
|
"depNameTemplate": "hadolint/hadolint"
|
||||||
@@ -144,7 +140,7 @@ data:
|
|||||||
{
|
{
|
||||||
"customType": "regex",
|
"customType": "regex",
|
||||||
"description": "node version the ci workflow runs npm with",
|
"description": "node version the ci workflow runs npm with",
|
||||||
"managerFilePatterns": ["^\\.gitea/workflows/tool-versions\\.env$"],
|
"managerFilePatterns": [".gitea/workflows/tool-versions.env"],
|
||||||
"matchStrings": ["(?:^|\\n)NODE_VERSION=\"(?<currentValue>[0-9.]+)\""],
|
"matchStrings": ["(?:^|\\n)NODE_VERSION=\"(?<currentValue>[0-9.]+)\""],
|
||||||
"datasourceTemplate": "node",
|
"datasourceTemplate": "node",
|
||||||
"depNameTemplate": "node"
|
"depNameTemplate": "node"
|
||||||
@@ -152,16 +148,23 @@ data:
|
|||||||
{
|
{
|
||||||
"customType": "regex",
|
"customType": "regex",
|
||||||
"description": "stakater/reloader chart version pinned in the deploy workflow",
|
"description": "stakater/reloader chart version pinned in the deploy workflow",
|
||||||
"managerFilePatterns": ["^\\.gitea/workflows/deploy-lib\\.sh$"],
|
"managerFilePatterns": [".gitea/workflows/deploy-lib.sh"],
|
||||||
"matchStrings": ["\\|stakater/reloader\\|reloader\\|(?<currentValue>[0-9.]+)\\|"],
|
"matchStrings": ["\\|stakater/reloader\\|reloader\\|(?<currentValue>[0-9.]+)\\|"],
|
||||||
"datasourceTemplate": "helm",
|
"datasourceTemplate": "helm",
|
||||||
"depNameTemplate": "reloader",
|
"depNameTemplate": "reloader",
|
||||||
"registryUrlTemplate": "https://stakater.github.io/stakater-charts"
|
"registryUrlTemplate": "https://stakater.github.io/stakater-charts"
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"customType": "regex",
|
||||||
|
"description": "Pinned CI BuildKit helper image",
|
||||||
|
"managerFilePatterns": [".gitea/workflows/tool-versions.env"],
|
||||||
|
"matchStrings": ["BUILDKIT_IMAGE=\"(?<depName>moby/buildkit):(?<currentValue>[^\"\\n]+)\""],
|
||||||
|
"datasourceTemplate": "docker"
|
||||||
}
|
}
|
||||||
],
|
],
|
||||||
"packageRules": [
|
"packageRules": [
|
||||||
{
|
{
|
||||||
"description": "Automerge digest and patch updates - safe by definition, review adds nothing, keeps the renovate queue and the deploy line short. Specific no-automerge rules below still override this for playwright, helm and majors.",
|
"description": "Automerge ordinary digest and patch updates after successful checks; specific manual-review rules below override this.",
|
||||||
"matchUpdateTypes": ["digest", "patch"],
|
"matchUpdateTypes": ["digest", "patch"],
|
||||||
"automerge": true
|
"automerge": true
|
||||||
},
|
},
|
||||||
@@ -172,23 +175,18 @@ data:
|
|||||||
"groupSlug": "all-minor",
|
"groupSlug": "all-minor",
|
||||||
"automerge": false
|
"automerge": false
|
||||||
},
|
},
|
||||||
|
{
|
||||||
|
"description": "Group ordinary patch updates; the specific groups and manual-review rules below take precedence",
|
||||||
|
"matchUpdateTypes": ["patch"],
|
||||||
|
"groupName": "all patch updates",
|
||||||
|
"groupSlug": "all-patch"
|
||||||
|
},
|
||||||
{
|
{
|
||||||
"description": "Keep private homelab images unchanged",
|
"description": "Keep private homelab images unchanged",
|
||||||
"matchDatasources": ["docker"],
|
"matchDatasources": ["docker"],
|
||||||
"matchPackageNames": ["/gcr\\.forust\\.xyz\\/forust\\/.+/"],
|
"matchPackageNames": ["/gcr\\.forust\\.xyz\\/forust\\/.+/"],
|
||||||
"enabled": false
|
"enabled": false
|
||||||
},
|
},
|
||||||
{
|
|
||||||
"description": "singlesource playwright - use whichever version is found, keep docker+pypi+npm in sync",
|
|
||||||
"matchPackageNames": ["playwright", "mcr.microsoft.com/playwright"],
|
|
||||||
"groupName": "playwright singlesource",
|
|
||||||
"groupSlug": "playwright"
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"description": "playwright must not automerge - version skew breaks the WS handshake (checker.py:1523 vs playwright.yaml:20)",
|
|
||||||
"matchPackageNames": ["playwright", "mcr.microsoft.com/playwright"],
|
|
||||||
"automerge": false
|
|
||||||
},
|
|
||||||
{
|
{
|
||||||
"description": "Renovate updates itself in lockstep across the CronJob and the Compose file",
|
"description": "Renovate updates itself in lockstep across the CronJob and the Compose file",
|
||||||
"matchPackageNames": ["renovate/renovate"],
|
"matchPackageNames": ["renovate/renovate"],
|
||||||
@@ -196,7 +194,7 @@ data:
|
|||||||
"automerge": false
|
"automerge": false
|
||||||
},
|
},
|
||||||
{
|
{
|
||||||
"description": "CI runs npm on the node the panel image is built from - the NODE_VERSION pin in tool-versions.env and node:22-alpine in the Dockerfile are the same dependency and move as one",
|
"description": "Keep CI Node runtime updates in a separate, manually reviewed group",
|
||||||
"matchPackageNames": ["node"],
|
"matchPackageNames": ["node"],
|
||||||
"groupName": "node runtime",
|
"groupName": "node runtime",
|
||||||
"groupSlug": "node",
|
"groupSlug": "node",
|
||||||
@@ -205,6 +203,8 @@ data:
|
|||||||
{
|
{
|
||||||
"description": "Helm chart bumps change PVC fields and admission behaviour, keep them reviewable",
|
"description": "Helm chart bumps change PVC fields and admission behaviour, keep them reviewable",
|
||||||
"matchDatasources": ["helm"],
|
"matchDatasources": ["helm"],
|
||||||
|
"groupName": "Helm chart {{depName}}",
|
||||||
|
"groupSlug": "helm-{{depName}}",
|
||||||
"automerge": false
|
"automerge": false
|
||||||
},
|
},
|
||||||
{
|
{
|
||||||
@@ -213,12 +213,6 @@ data:
|
|||||||
"dependencyDashboardApproval": true,
|
"dependencyDashboardApproval": true,
|
||||||
"automerge": false
|
"automerge": false
|
||||||
},
|
},
|
||||||
{
|
|
||||||
"description": "Group patch updates from all sources - automerge still applies via the digest/patch rule above (helm/playwright stay manual via their own rules)",
|
|
||||||
"matchUpdateTypes": ["patch"],
|
|
||||||
"groupName": "all patch updates",
|
|
||||||
"groupSlug": "all-patch"
|
|
||||||
},
|
|
||||||
{
|
{
|
||||||
"description": "Python Y-bumps break compat (3.11->3.12->3.13->3.14) - keep the base image out of the shared minor/patch groups, review every bump separately. Placed last so its groupName wins.",
|
"description": "Python Y-bumps break compat (3.11->3.12->3.13->3.14) - keep the base image out of the shared minor/patch groups, review every bump separately. Placed last so its groupName wins.",
|
||||||
"matchDatasources": ["docker"],
|
"matchDatasources": ["docker"],
|
||||||
|
|||||||
@@ -19,7 +19,7 @@ spec:
|
|||||||
restartPolicy: Never
|
restartPolicy: Never
|
||||||
containers:
|
containers:
|
||||||
- name: renovate
|
- name: renovate
|
||||||
image: renovate/renovate:44.136.0
|
image: renovate/renovate:44.140.0
|
||||||
env:
|
env:
|
||||||
- name: RENOVATE_PLATFORM
|
- name: RENOVATE_PLATFORM
|
||||||
value: gitea
|
value: gitea
|
||||||
|
|||||||
@@ -2,7 +2,7 @@ services:
|
|||||||
renovate:
|
renovate:
|
||||||
# Kept in step with renovate/k8s/cronjob.yaml by the "renovate self-update"
|
# Kept in step with renovate/k8s/cronjob.yaml by the "renovate self-update"
|
||||||
# package rule in renovate/renovate.json.
|
# package rule in renovate/renovate.json.
|
||||||
image: renovate/renovate:44.115.9
|
image: renovate/renovate:44.136.0
|
||||||
container_name: renovate
|
container_name: renovate
|
||||||
restart: "no"
|
restart: "no"
|
||||||
env_file:
|
env_file:
|
||||||
|
|||||||
+43
-49
@@ -8,6 +8,9 @@
|
|||||||
"dependencyDashboard": true,
|
"dependencyDashboard": true,
|
||||||
"prCreation": "immediate",
|
"prCreation": "immediate",
|
||||||
"labels": ["dependencies", "automated"],
|
"labels": ["dependencies", "automated"],
|
||||||
|
"docker-compose": {
|
||||||
|
"managerFilePatterns": ["renovate/renovate-compose.yaml"]
|
||||||
|
},
|
||||||
"helm-values": {
|
"helm-values": {
|
||||||
"managerFilePatterns": ["/k8s/.+values\\.ya?ml$/"]
|
"managerFilePatterns": ["/k8s/.+values\\.ya?ml$/"]
|
||||||
},
|
},
|
||||||
@@ -15,35 +18,28 @@
|
|||||||
"managerFilePatterns": ["/k8s/.+\\.ya?ml$/"]
|
"managerFilePatterns": ["/k8s/.+\\.ya?ml$/"]
|
||||||
},
|
},
|
||||||
"customManagers": [
|
"customManagers": [
|
||||||
{
|
|
||||||
"customType": "regex",
|
|
||||||
"description": "singlesource: playwright npm version pinned in npx command (k8s + compose)",
|
|
||||||
"managerFilePatterns": ["^edu_master/k8s/playwright\\.yaml$", "^edu_master/compose\\.yaml$"],
|
|
||||||
"matchStrings": ["playwright@(?<currentValue>\\d+\\.\\d+\\.\\d+)"],
|
|
||||||
"datasourceTemplate": "npm",
|
|
||||||
"depNameTemplate": "playwright"
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"customType": "regex",
|
|
||||||
"description": "singlesource: PLAYWRIGHT_VERSION file",
|
|
||||||
"managerFilePatterns": ["^edu_master/PLAYWRIGHT_VERSION$"],
|
|
||||||
"matchStrings": ["^(?<currentValue>\\d+\\.\\d+\\.\\d+)$"],
|
|
||||||
"datasourceTemplate": "pypi",
|
|
||||||
"depNameTemplate": "playwright"
|
|
||||||
},
|
|
||||||
{
|
{
|
||||||
"customType": "regex",
|
"customType": "regex",
|
||||||
"description": "kube-prometheus-stack chart version pinned in the deploy workflow",
|
"description": "kube-prometheus-stack chart version pinned in the deploy workflow",
|
||||||
"managerFilePatterns": ["^\\.gitea/workflows/deploy-lib\\.sh$"],
|
"managerFilePatterns": [".gitea/workflows/deploy-lib.sh"],
|
||||||
"matchStrings": ["\\|prometheus-community/kube-prometheus-stack\\|prometheus\\|(?<currentValue>[0-9.]+)\\|"],
|
"matchStrings": ["\\|prometheus-community/kube-prometheus-stack\\|prometheus\\|(?<currentValue>[0-9.]+)\\|"],
|
||||||
"datasourceTemplate": "helm",
|
"datasourceTemplate": "helm",
|
||||||
"depNameTemplate": "kube-prometheus-stack",
|
"depNameTemplate": "kube-prometheus-stack",
|
||||||
"registryUrlTemplate": "https://prometheus-community.github.io/helm-charts"
|
"registryUrlTemplate": "https://prometheus-community.github.io/helm-charts"
|
||||||
},
|
},
|
||||||
|
{
|
||||||
|
"customType": "regex",
|
||||||
|
"description": "VictoriaMetrics Operator chart version pinned in the deploy workflow",
|
||||||
|
"managerFilePatterns": [".gitea/workflows/deploy-lib.sh"],
|
||||||
|
"matchStrings": ["\\|victoriametrics/victoria-metrics-operator\\|prometheus\\|(?<currentValue>[0-9.]+)\\|"],
|
||||||
|
"datasourceTemplate": "helm",
|
||||||
|
"depNameTemplate": "victoria-metrics-operator",
|
||||||
|
"registryUrlTemplate": "https://victoriametrics.github.io/helm-charts"
|
||||||
|
},
|
||||||
{
|
{
|
||||||
"customType": "regex",
|
"customType": "regex",
|
||||||
"description": "grafana/loki chart version pinned in the deploy workflow",
|
"description": "grafana/loki chart version pinned in the deploy workflow",
|
||||||
"managerFilePatterns": ["^\\.gitea/workflows/deploy-lib\\.sh$"],
|
"managerFilePatterns": [".gitea/workflows/deploy-lib.sh"],
|
||||||
"matchStrings": ["\\|grafana/loki\\|prometheus\\|(?<currentValue>[0-9.]+)\\|"],
|
"matchStrings": ["\\|grafana/loki\\|prometheus\\|(?<currentValue>[0-9.]+)\\|"],
|
||||||
"datasourceTemplate": "helm",
|
"datasourceTemplate": "helm",
|
||||||
"depNameTemplate": "loki",
|
"depNameTemplate": "loki",
|
||||||
@@ -52,7 +48,7 @@
|
|||||||
{
|
{
|
||||||
"customType": "regex",
|
"customType": "regex",
|
||||||
"description": "grafana/alloy chart version pinned in the deploy workflow",
|
"description": "grafana/alloy chart version pinned in the deploy workflow",
|
||||||
"managerFilePatterns": ["^\\.gitea/workflows/deploy-lib\\.sh$"],
|
"managerFilePatterns": [".gitea/workflows/deploy-lib.sh"],
|
||||||
"matchStrings": ["\\|grafana/alloy\\|prometheus\\|(?<currentValue>[0-9.]+)\\|"],
|
"matchStrings": ["\\|grafana/alloy\\|prometheus\\|(?<currentValue>[0-9.]+)\\|"],
|
||||||
"datasourceTemplate": "helm",
|
"datasourceTemplate": "helm",
|
||||||
"depNameTemplate": "alloy",
|
"depNameTemplate": "alloy",
|
||||||
@@ -61,7 +57,7 @@
|
|||||||
{
|
{
|
||||||
"customType": "regex",
|
"customType": "regex",
|
||||||
"description": "actionlint version used by the ci workflow",
|
"description": "actionlint version used by the ci workflow",
|
||||||
"managerFilePatterns": ["^\\.gitea/workflows/tool-versions\\.env$"],
|
"managerFilePatterns": [".gitea/workflows/tool-versions.env"],
|
||||||
"matchStrings": ["(?:^|\\n)ACTIONLINT_VERSION=\"(?<currentValue>[0-9.]+)\""],
|
"matchStrings": ["(?:^|\\n)ACTIONLINT_VERSION=\"(?<currentValue>[0-9.]+)\""],
|
||||||
"datasourceTemplate": "github-tags",
|
"datasourceTemplate": "github-tags",
|
||||||
"depNameTemplate": "rhysd/actionlint"
|
"depNameTemplate": "rhysd/actionlint"
|
||||||
@@ -69,7 +65,7 @@
|
|||||||
{
|
{
|
||||||
"customType": "regex",
|
"customType": "regex",
|
||||||
"description": "shellcheck version used by the ci workflow",
|
"description": "shellcheck version used by the ci workflow",
|
||||||
"managerFilePatterns": ["^\\.gitea/workflows/tool-versions\\.env$"],
|
"managerFilePatterns": [".gitea/workflows/tool-versions.env"],
|
||||||
"matchStrings": ["(?:^|\\n)SHELLCHECK_VERSION=\"(?<currentValue>[0-9.]+)\""],
|
"matchStrings": ["(?:^|\\n)SHELLCHECK_VERSION=\"(?<currentValue>[0-9.]+)\""],
|
||||||
"datasourceTemplate": "github-tags",
|
"datasourceTemplate": "github-tags",
|
||||||
"depNameTemplate": "koalaman/shellcheck"
|
"depNameTemplate": "koalaman/shellcheck"
|
||||||
@@ -77,7 +73,7 @@
|
|||||||
{
|
{
|
||||||
"customType": "regex",
|
"customType": "regex",
|
||||||
"description": "kubeconform version used by the ci workflow",
|
"description": "kubeconform version used by the ci workflow",
|
||||||
"managerFilePatterns": ["^\\.gitea/workflows/tool-versions\\.env$"],
|
"managerFilePatterns": [".gitea/workflows/tool-versions.env"],
|
||||||
"matchStrings": ["(?:^|\\n)KUBECONFORM_VERSION=\"(?<currentValue>[0-9.]+)\""],
|
"matchStrings": ["(?:^|\\n)KUBECONFORM_VERSION=\"(?<currentValue>[0-9.]+)\""],
|
||||||
"datasourceTemplate": "github-tags",
|
"datasourceTemplate": "github-tags",
|
||||||
"depNameTemplate": "yannh/kubeconform"
|
"depNameTemplate": "yannh/kubeconform"
|
||||||
@@ -85,7 +81,7 @@
|
|||||||
{
|
{
|
||||||
"customType": "regex",
|
"customType": "regex",
|
||||||
"description": "uv version used to build the pytest venv",
|
"description": "uv version used to build the pytest venv",
|
||||||
"managerFilePatterns": ["^\\.gitea/workflows/tool-versions\\.env$"],
|
"managerFilePatterns": [".gitea/workflows/tool-versions.env"],
|
||||||
"matchStrings": ["(?:^|\\n)UV_VERSION=\"(?<currentValue>[0-9.]+)\""],
|
"matchStrings": ["(?:^|\\n)UV_VERSION=\"(?<currentValue>[0-9.]+)\""],
|
||||||
"datasourceTemplate": "github-tags",
|
"datasourceTemplate": "github-tags",
|
||||||
"depNameTemplate": "astral-sh/uv"
|
"depNameTemplate": "astral-sh/uv"
|
||||||
@@ -93,7 +89,7 @@
|
|||||||
{
|
{
|
||||||
"customType": "regex",
|
"customType": "regex",
|
||||||
"description": "prettier version used by the ci workflow",
|
"description": "prettier version used by the ci workflow",
|
||||||
"managerFilePatterns": ["^\\.gitea/workflows/tool-versions\\.env$"],
|
"managerFilePatterns": [".gitea/workflows/tool-versions.env"],
|
||||||
"matchStrings": ["(?:^|\\n)PRETTIER_VERSION=\"(?<currentValue>[0-9.]+)\""],
|
"matchStrings": ["(?:^|\\n)PRETTIER_VERSION=\"(?<currentValue>[0-9.]+)\""],
|
||||||
"datasourceTemplate": "npm",
|
"datasourceTemplate": "npm",
|
||||||
"depNameTemplate": "prettier"
|
"depNameTemplate": "prettier"
|
||||||
@@ -101,7 +97,7 @@
|
|||||||
{
|
{
|
||||||
"customType": "regex",
|
"customType": "regex",
|
||||||
"description": "ruff version used by the ci workflow",
|
"description": "ruff version used by the ci workflow",
|
||||||
"managerFilePatterns": ["^\\.gitea/workflows/tool-versions\\.env$"],
|
"managerFilePatterns": [".gitea/workflows/tool-versions.env"],
|
||||||
"matchStrings": ["(?:^|\\n)RUFF_VERSION=\"(?<currentValue>[0-9.]+)\""],
|
"matchStrings": ["(?:^|\\n)RUFF_VERSION=\"(?<currentValue>[0-9.]+)\""],
|
||||||
"datasourceTemplate": "pypi",
|
"datasourceTemplate": "pypi",
|
||||||
"depNameTemplate": "ruff"
|
"depNameTemplate": "ruff"
|
||||||
@@ -109,7 +105,7 @@
|
|||||||
{
|
{
|
||||||
"customType": "regex",
|
"customType": "regex",
|
||||||
"description": "pip-audit version used by the ci workflow",
|
"description": "pip-audit version used by the ci workflow",
|
||||||
"managerFilePatterns": ["^\\.gitea/workflows/tool-versions\\.env$"],
|
"managerFilePatterns": [".gitea/workflows/tool-versions.env"],
|
||||||
"matchStrings": ["(?:^|\\n)PIP_AUDIT_VERSION=\"(?<currentValue>[0-9.]+)\""],
|
"matchStrings": ["(?:^|\\n)PIP_AUDIT_VERSION=\"(?<currentValue>[0-9.]+)\""],
|
||||||
"datasourceTemplate": "pypi",
|
"datasourceTemplate": "pypi",
|
||||||
"depNameTemplate": "pip-audit"
|
"depNameTemplate": "pip-audit"
|
||||||
@@ -117,7 +113,7 @@
|
|||||||
{
|
{
|
||||||
"customType": "regex",
|
"customType": "regex",
|
||||||
"description": "yamllint version used by the ci workflow",
|
"description": "yamllint version used by the ci workflow",
|
||||||
"managerFilePatterns": ["^\\.gitea/workflows/tool-versions\\.env$"],
|
"managerFilePatterns": [".gitea/workflows/tool-versions.env"],
|
||||||
"matchStrings": ["(?:^|\\n)YAMLLINT_VERSION=\"(?<currentValue>[0-9.]+)\""],
|
"matchStrings": ["(?:^|\\n)YAMLLINT_VERSION=\"(?<currentValue>[0-9.]+)\""],
|
||||||
"datasourceTemplate": "pypi",
|
"datasourceTemplate": "pypi",
|
||||||
"depNameTemplate": "yamllint"
|
"depNameTemplate": "yamllint"
|
||||||
@@ -125,7 +121,7 @@
|
|||||||
{
|
{
|
||||||
"customType": "regex",
|
"customType": "regex",
|
||||||
"description": "hadolint version used by the ci workflow",
|
"description": "hadolint version used by the ci workflow",
|
||||||
"managerFilePatterns": ["^\\.gitea/workflows/tool-versions\\.env$"],
|
"managerFilePatterns": [".gitea/workflows/tool-versions.env"],
|
||||||
"matchStrings": ["(?:^|\\n)HADOLINT_VERSION=\"(?<currentValue>[0-9.]+)\""],
|
"matchStrings": ["(?:^|\\n)HADOLINT_VERSION=\"(?<currentValue>[0-9.]+)\""],
|
||||||
"datasourceTemplate": "github-tags",
|
"datasourceTemplate": "github-tags",
|
||||||
"depNameTemplate": "hadolint/hadolint"
|
"depNameTemplate": "hadolint/hadolint"
|
||||||
@@ -133,7 +129,7 @@
|
|||||||
{
|
{
|
||||||
"customType": "regex",
|
"customType": "regex",
|
||||||
"description": "node version the ci workflow runs npm with",
|
"description": "node version the ci workflow runs npm with",
|
||||||
"managerFilePatterns": ["^\\.gitea/workflows/tool-versions\\.env$"],
|
"managerFilePatterns": [".gitea/workflows/tool-versions.env"],
|
||||||
"matchStrings": ["(?:^|\\n)NODE_VERSION=\"(?<currentValue>[0-9.]+)\""],
|
"matchStrings": ["(?:^|\\n)NODE_VERSION=\"(?<currentValue>[0-9.]+)\""],
|
||||||
"datasourceTemplate": "node",
|
"datasourceTemplate": "node",
|
||||||
"depNameTemplate": "node"
|
"depNameTemplate": "node"
|
||||||
@@ -141,16 +137,23 @@
|
|||||||
{
|
{
|
||||||
"customType": "regex",
|
"customType": "regex",
|
||||||
"description": "stakater/reloader chart version pinned in the deploy workflow",
|
"description": "stakater/reloader chart version pinned in the deploy workflow",
|
||||||
"managerFilePatterns": ["^\\.gitea/workflows/deploy-lib\\.sh$"],
|
"managerFilePatterns": [".gitea/workflows/deploy-lib.sh"],
|
||||||
"matchStrings": ["\\|stakater/reloader\\|reloader\\|(?<currentValue>[0-9.]+)\\|"],
|
"matchStrings": ["\\|stakater/reloader\\|reloader\\|(?<currentValue>[0-9.]+)\\|"],
|
||||||
"datasourceTemplate": "helm",
|
"datasourceTemplate": "helm",
|
||||||
"depNameTemplate": "reloader",
|
"depNameTemplate": "reloader",
|
||||||
"registryUrlTemplate": "https://stakater.github.io/stakater-charts"
|
"registryUrlTemplate": "https://stakater.github.io/stakater-charts"
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"customType": "regex",
|
||||||
|
"description": "Pinned CI BuildKit helper image",
|
||||||
|
"managerFilePatterns": [".gitea/workflows/tool-versions.env"],
|
||||||
|
"matchStrings": ["BUILDKIT_IMAGE=\"(?<depName>moby/buildkit):(?<currentValue>[^\"\\n]+)\""],
|
||||||
|
"datasourceTemplate": "docker"
|
||||||
}
|
}
|
||||||
],
|
],
|
||||||
"packageRules": [
|
"packageRules": [
|
||||||
{
|
{
|
||||||
"description": "Automerge digest and patch updates - safe by definition, review adds nothing, keeps the renovate queue and the deploy line short. Specific no-automerge rules below still override this for playwright, helm and majors.",
|
"description": "Automerge ordinary digest and patch updates after successful checks; specific manual-review rules below override this.",
|
||||||
"matchUpdateTypes": ["digest", "patch"],
|
"matchUpdateTypes": ["digest", "patch"],
|
||||||
"automerge": true
|
"automerge": true
|
||||||
},
|
},
|
||||||
@@ -161,23 +164,18 @@
|
|||||||
"groupSlug": "all-minor",
|
"groupSlug": "all-minor",
|
||||||
"automerge": false
|
"automerge": false
|
||||||
},
|
},
|
||||||
|
{
|
||||||
|
"description": "Group ordinary patch updates; the specific groups and manual-review rules below take precedence",
|
||||||
|
"matchUpdateTypes": ["patch"],
|
||||||
|
"groupName": "all patch updates",
|
||||||
|
"groupSlug": "all-patch"
|
||||||
|
},
|
||||||
{
|
{
|
||||||
"description": "Keep private homelab images unchanged",
|
"description": "Keep private homelab images unchanged",
|
||||||
"matchDatasources": ["docker"],
|
"matchDatasources": ["docker"],
|
||||||
"matchPackageNames": ["/gcr\\.forust\\.xyz\\/forust\\/.+/"],
|
"matchPackageNames": ["/gcr\\.forust\\.xyz\\/forust\\/.+/"],
|
||||||
"enabled": false
|
"enabled": false
|
||||||
},
|
},
|
||||||
{
|
|
||||||
"description": "singlesource playwright - use whichever version is found, keep docker+pypi+npm in sync",
|
|
||||||
"matchPackageNames": ["playwright", "mcr.microsoft.com/playwright"],
|
|
||||||
"groupName": "playwright singlesource",
|
|
||||||
"groupSlug": "playwright"
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"description": "playwright must not automerge - version skew breaks the WS handshake (checker.py:1523 vs playwright.yaml:20)",
|
|
||||||
"matchPackageNames": ["playwright", "mcr.microsoft.com/playwright"],
|
|
||||||
"automerge": false
|
|
||||||
},
|
|
||||||
{
|
{
|
||||||
"description": "Renovate updates itself in lockstep across the CronJob and the Compose file",
|
"description": "Renovate updates itself in lockstep across the CronJob and the Compose file",
|
||||||
"matchPackageNames": ["renovate/renovate"],
|
"matchPackageNames": ["renovate/renovate"],
|
||||||
@@ -185,7 +183,7 @@
|
|||||||
"automerge": false
|
"automerge": false
|
||||||
},
|
},
|
||||||
{
|
{
|
||||||
"description": "CI runs npm on the node the panel image is built from - the NODE_VERSION pin in tool-versions.env and node:22-alpine in the Dockerfile are the same dependency and move as one",
|
"description": "Keep CI Node runtime updates in a separate, manually reviewed group",
|
||||||
"matchPackageNames": ["node"],
|
"matchPackageNames": ["node"],
|
||||||
"groupName": "node runtime",
|
"groupName": "node runtime",
|
||||||
"groupSlug": "node",
|
"groupSlug": "node",
|
||||||
@@ -194,6 +192,8 @@
|
|||||||
{
|
{
|
||||||
"description": "Helm chart bumps change PVC fields and admission behaviour, keep them reviewable",
|
"description": "Helm chart bumps change PVC fields and admission behaviour, keep them reviewable",
|
||||||
"matchDatasources": ["helm"],
|
"matchDatasources": ["helm"],
|
||||||
|
"groupName": "Helm chart {{depName}}",
|
||||||
|
"groupSlug": "helm-{{depName}}",
|
||||||
"automerge": false
|
"automerge": false
|
||||||
},
|
},
|
||||||
{
|
{
|
||||||
@@ -202,12 +202,6 @@
|
|||||||
"dependencyDashboardApproval": true,
|
"dependencyDashboardApproval": true,
|
||||||
"automerge": false
|
"automerge": false
|
||||||
},
|
},
|
||||||
{
|
|
||||||
"description": "Group patch updates from all sources - automerge still applies via the digest/patch rule above (helm/playwright stay manual via their own rules)",
|
|
||||||
"matchUpdateTypes": ["patch"],
|
|
||||||
"groupName": "all patch updates",
|
|
||||||
"groupSlug": "all-patch"
|
|
||||||
},
|
|
||||||
{
|
{
|
||||||
"description": "Python Y-bumps break compat (3.11->3.12->3.13->3.14) - keep the base image out of the shared minor/patch groups, review every bump separately. Placed last so its groupName wins.",
|
"description": "Python Y-bumps break compat (3.11->3.12->3.13->3.14) - keep the base image out of the shared minor/patch groups, review every bump separately. Placed last so its groupName wins.",
|
||||||
"matchDatasources": ["docker"],
|
"matchDatasources": ["docker"],
|
||||||
|
|||||||
@@ -12,11 +12,11 @@ services:
|
|||||||
- "traefik.enable=true"
|
- "traefik.enable=true"
|
||||||
- "traefik.http.services.searxng.loadbalancer.server.port=8080"
|
- "traefik.http.services.searxng.loadbalancer.server.port=8080"
|
||||||
# Prod Router
|
# Prod Router
|
||||||
- "traefik.http.routers.searxng.rule=Host(`s.forust.xyz` || `search.forust.xyz`)"
|
- "traefik.http.routers.searxng.rule=Host(`s.forust.xyz`) || Host(`search.forust.xyz`)"
|
||||||
- "traefik.http.routers.searxng.entrypoints=websecure"
|
- "traefik.http.routers.searxng.entrypoints=websecure"
|
||||||
- "traefik.http.routers.searxng.tls.certresolver=letsencrypt"
|
- "traefik.http.routers.searxng.tls.certresolver=letsencrypt"
|
||||||
# Local Router
|
# Local Router
|
||||||
- "traefik.http.routers.searxng-local.rule=Host(`s.workstation.internal` || `searxng.workstation.internal`)"
|
- "traefik.http.routers.searxng-local.rule=Host(`s.workstation.internal`) || Host(`searxng.workstation.internal`)"
|
||||||
- "traefik.http.routers.searxng-local.entrypoints=websecure"
|
- "traefik.http.routers.searxng-local.entrypoints=websecure"
|
||||||
- "traefik.http.routers.searxng-local.tls=true"
|
- "traefik.http.routers.searxng-local.tls=true"
|
||||||
# Dev Router
|
# Dev Router
|
||||||
|
|||||||
Whitespace-only changes.
@@ -19,7 +19,7 @@ services:
|
|||||||
- streaming
|
- streaming
|
||||||
|
|
||||||
qbittorrent:
|
qbittorrent:
|
||||||
image: lscr.io/linuxserver/qbittorrent:5.2.4
|
image: lscr.io/linuxserver/qbittorrent:20.04.1
|
||||||
container_name: qbittorrent
|
container_name: qbittorrent
|
||||||
restart: unless-stopped
|
restart: unless-stopped
|
||||||
environment:
|
environment:
|
||||||
|
|||||||
Whitespace-only changes.
+1
-1
@@ -1,6 +1,6 @@
|
|||||||
services:
|
services:
|
||||||
termix:
|
termix:
|
||||||
image: ghcr.io/lukegus/termix:2.9.1
|
image: ghcr.io/lukegus/termix:2.9.2
|
||||||
container_name: termix
|
container_name: termix
|
||||||
restart: unless-stopped
|
restart: unless-stopped
|
||||||
# ports:
|
# ports:
|
||||||
|
|||||||
Loaded 100 of 108 files, more files were not shown because too many files have changed in this diff.
Show more
Reference in new issue
Block a user