f0b040a4f13275e2b5715af48405b51d5f293a91
97
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
f0b040a4f1 |
chore(deps): update all minor updates
ci / Compose (pull_request) Successful in 14s
ci / Workflows (pull_request) Successful in 7s
ci / Formatting (pull_request) Successful in 22s
ci / Kubernetes (pull_request) Successful in 8s
ci / Shell (pull_request) Successful in 24s
ci / Python and tests (pull_request) Successful in 9s
ci / YAML (pull_request) Successful in 11s
ci / Dockerfiles (pull_request) Successful in 6s
ci / image-plan (pull_request) Skipped
ci / Image (${{ matrix.name }}) (pull_request) Skipped
ci / build (pull_request) Skipped
renovate-ci / validate-renovate (pull_request_target) Successful in 3m13s
|
||
|
|
4c7c53e0f2 |
fix(cicd): align apply timeout with stage budgets
ci / Compose (push) Successful in 11s
ci / Workflows (push) Successful in 6s
ci / Shell (push) Successful in 20s
ci / Formatting (push) Successful in 17s
ci / Python and tests (push) Successful in 7s
ci / YAML (push) Successful in 10s
ci / Dockerfiles (push) Successful in 4s
ci / image-plan (push) Successful in 13s
ci / Image (error-pages) (push) Successful in 15s
ci / Image (forust-homepage) (push) Successful in 15s
ci / Kubernetes (push) Successful in 7s
ci / Image (xdfnx-homepage) (push) Successful in 15s
ci / build (push) Successful in 19s
|
||
|
|
64962d1a63 |
fix(cicd): preserve AIO tag in recovery snapshot
ci / Workflows (pull_request) Successful in 6s
ci / Kubernetes (pull_request) Successful in 6s
ci / image-plan (pull_request) Skipped
ci / Image (${{ matrix.name }}) (pull_request) Skipped
ci / build (pull_request) Skipped
ci / Compose (pull_request) Successful in 15s
ci / Shell (pull_request) Successful in 19s
ci / Formatting (pull_request) Successful in 30s
ci / Python and tests (pull_request) Successful in 8s
ci / YAML (pull_request) Successful in 10s
ci / Dockerfiles (pull_request) Successful in 4s
|
||
|
|
8203ba1b0b |
fix(cicd): preserve Nextcloud AIO image tag
ci / Workflows (pull_request) Successful in 9s
ci / Compose (pull_request) Successful in 14s
ci / Shell (pull_request) Successful in 19s
ci / Formatting (pull_request) Successful in 33s
ci / Python and tests (pull_request) Successful in 16s
ci / Dockerfiles (pull_request) Successful in 14s
ci / Kubernetes (pull_request) Successful in 17s
ci / Image (${{ matrix.name }}) (pull_request) Skipped
ci / build (pull_request) Skipped
ci / YAML (pull_request) Successful in 19s
ci / image-plan (pull_request) Skipped
|
||
|
|
73d2af73e5 |
fix(cicd): support Helm 4 release listing
ci / build (pull_request) Skipped
ci / Workflows (pull_request) Successful in 7s
ci / Python and tests (pull_request) Successful in 5s
ci / Compose (pull_request) Successful in 11s
ci / Shell (pull_request) Successful in 17s
ci / Formatting (pull_request) Successful in 17s
ci / YAML (pull_request) Successful in 8s
ci / Dockerfiles (pull_request) Successful in 5s
ci / Kubernetes (pull_request) Successful in 7s
ci / image-plan (pull_request) Skipped
ci / Image (${{ matrix.name }}) (pull_request) Skipped
|
||
|
|
69accd1752 |
fix(cicd): use Gitea artifact v4 backend
ci / Shell (push) Skipped
ci / Formatting (push) Skipped
ci / Python and tests (push) Skipped
ci / YAML (push) Skipped
ci / Dockerfiles (push) Skipped
ci / Kubernetes (push) Skipped
ci / Compose (pull_request) Successful in 11s
ci / Kubernetes (pull_request) Successful in 7s
ci / Compose (push) Skipped
ci / Workflows (push) Skipped
ci / Workflows (pull_request) Successful in 7s
ci / Shell (pull_request) Successful in 18s
ci / Formatting (pull_request) Successful in 20s
ci / Python and tests (pull_request) Successful in 7s
ci / YAML (pull_request) Successful in 8s
ci / Dockerfiles (pull_request) Successful in 4s
ci / image-plan (pull_request) Skipped
ci / Image (${{ matrix.name }}) (pull_request) Skipped
ci / build (pull_request) Skipped
|
||
|
|
d7441bbbc2 |
fix(cicd): isolate pull request runner jobs
ci / Compose (push) Skipped
ci / Workflows (push) Skipped
ci / Shell (push) Skipped
ci / Formatting (push) Skipped
ci / Python and tests (push) Skipped
ci / Compose (pull_request) Successful in 37s
ci / Shell (pull_request) Successful in 18s
ci / YAML (push) Skipped
ci / Dockerfiles (push) Skipped
ci / Kubernetes (push) Skipped
ci / Workflows (pull_request) Successful in 8s
ci / Python and tests (pull_request) Successful in 11s
ci / Kubernetes (pull_request) Successful in 8s
ci / Formatting (pull_request) Successful in 18s
ci / YAML (pull_request) Successful in 11s
ci / Dockerfiles (pull_request) Successful in 7s
ci / image-plan (pull_request) Skipped
ci / Image (${{ matrix.name }}) (pull_request) Skipped
ci / build (pull_request) Skipped
|
||
|
|
95e4d8f875 |
ci: add per-image jobs to homelab CI
ci / Compose (push) Skipped
ci / Workflows (push) Skipped
ci / Shell (push) Skipped
ci / Formatting (push) Skipped
ci / Python and tests (push) Skipped
ci / YAML (push) Skipped
ci / Dockerfiles (push) Skipped
ci / Kubernetes (push) Skipped
ci / Compose (pull_request) Successful in 11s
ci / Workflows (pull_request) Successful in 7s
ci / Shell (pull_request) Successful in 15s
ci / Formatting (pull_request) Successful in 15s
ci / Python and tests (pull_request) Successful in 5s
ci / YAML (pull_request) Successful in 7s
ci / Dockerfiles (pull_request) Successful in 5s
ci / Kubernetes (pull_request) Successful in 6s
ci / image-plan (pull_request) Skipped
ci / Image (${{ matrix.name }}) (pull_request) Skipped
ci / build (pull_request) Skipped
renovate-ci / validate-renovate (pull_request) Successful in 11s
|
||
|
|
fb400eea6e |
Use plain text for the deploy request details
ci / Compose (push) Skipped
ci / Shell (push) Skipped
ci / Formatting (push) Skipped
ci / YAML (push) Skipped
ci / Dockerfiles (push) Skipped
ci / Kubernetes (push) Skipped
ci / YAML (pull_request) Successful in 11s
ci / build (pull_request) Skipped
ci / Workflows (push) Skipped
ci / Python and tests (push) Skipped
ci / Compose (pull_request) Successful in 13s
ci / Workflows (pull_request) Successful in 9s
ci / Shell (pull_request) Successful in 21s
ci / Formatting (pull_request) Successful in 22s
ci / Python and tests (pull_request) Successful in 6s
ci / Dockerfiles (pull_request) Successful in 5s
ci / Kubernetes (pull_request) Successful in 7s
|
||
|
|
0c76426c17 |
Show failure details in CI and deploy summaries
ci / Formatting (push) Skipped
ci / Python and tests (push) Skipped
ci / Kubernetes (push) Skipped
ci / Compose (push) Skipped
ci / Shell (push) Skipped
ci / YAML (push) Skipped
ci / Dockerfiles (push) Skipped
ci / Workflows (push) Skipped
ci / Compose (pull_request) Successful in 11s
ci / Workflows (pull_request) Failing after 8s
ci / Formatting (pull_request) Canceled after 0s
ci / Python and tests (pull_request) Canceled after 0s
ci / YAML (pull_request) Canceled after 0s
ci / Dockerfiles (pull_request) Canceled after 0s
ci / Kubernetes (pull_request) Canceled after 0s
ci / build (pull_request) Canceled after 0s
ci / Shell (pull_request) Canceled after 7s
|
||
|
|
d74822cd27 |
Add CI and deploy summaries
ci / Compose (push) Skipped
ci / Workflows (push) Skipped
ci / Shell (push) Skipped
ci / Formatting (push) Skipped
ci / Python and tests (push) Skipped
ci / YAML (push) Skipped
ci / Kubernetes (push) Skipped
ci / Dockerfiles (push) Skipped
ci / Compose (pull_request) Successful in 14s
ci / Workflows (pull_request) Successful in 7s
ci / Shell (pull_request) Successful in 23s
ci / Formatting (pull_request) Successful in 26s
ci / Python and tests (pull_request) Successful in 8s
ci / YAML (pull_request) Successful in 14s
ci / Dockerfiles (pull_request) Successful in 7s
ci / Kubernetes (pull_request) Successful in 10s
ci / build (pull_request) Skipped
|
||
|
|
b677d553b4 |
refactor(ci): split checks into visible jobs
ci / YAML (pull_request) Successful in 11s
ci / Dockerfiles (pull_request) Successful in 6s
ci / Kubernetes (pull_request) Successful in 6s
ci / build (pull_request) Skipped
renovate-ci / validate-renovate (push) Skipped
renovate-ci / validate-renovate (pull_request) Skipped
ci / Compose (pull_request) Successful in 10s
ci / Workflows (pull_request) Successful in 5s
ci / Shell (pull_request) Successful in 22s
ci / Formatting (pull_request) Successful in 21s
ci / Python and tests (pull_request) Successful in 8s
|
||
|
|
9a76529be8 | refactor(ci): use native runner and durable incremental deploys | ||
|
|
5f9354b9a8 |
fix(edu): deploy reviewed application release with Redis authentication
renovate-ci / validate-renovate (push) Skipped
ci / lint-compose (push) Successful in 14s
ci / lint-actionlint (push) Successful in 6s
ci / lint-shellcheck (push) Successful in 13s
ci / lint-prettier (push) Successful in 19s
ci / lint-ruff (push) Successful in 9s
ci / lint-yaml (push) Successful in 10s
ci / lint-dockerfiles (push) Successful in 6s
ci / validate (push) Successful in 10s
ci / build (push) Successful in 30s
|
||
|
|
cef499de73 |
fix(deploy): skip VMAgent in secrets check before CRD install
renovate-ci / validate-renovate (push) Skipped
ci / lint-compose (push) Successful in 12s
ci / lint-actionlint (push) Successful in 6s
ci / lint-shellcheck (push) Successful in 15s
ci / lint-ruff (push) Successful in 6s
ci / lint-dockerfiles (push) Successful in 6s
ci / lint-prettier (push) Successful in 21s
ci / lint-yaml (push) Successful in 10s
ci / validate (push) Successful in 8s
ci / build (push) Successful in 21s
check_referenced_secrets ran kubectl create on vmagent.yaml even when the VMAgent CRD is not installed yet, failing validate with 'no matches for kind VMAgent'. Apply the same skip_uninstalled_vmagent_crd guard used by both dry-run loops. |
||
|
|
1b70a55300 |
fix(ci): handle malformed push before SHA
renovate-ci / validate-renovate (push) Skipped
ci / lint-compose (push) Successful in 10s
ci / lint-actionlint (push) Successful in 7s
ci / lint-shellcheck (push) Successful in 15s
ci / lint-prettier (push) Failing after 22s
ci / lint-ruff (push) Failing after 2s
ci / lint-yaml (push) Failing after 2s
ci / lint-dockerfiles (push) Failing after 3s
ci / validate (push) Failing after 2s
ci / build (push) Skipped
|
||
|
|
3f2b4e9acf |
fix(deploy): skip VMAgent preflight before CRD install
renovate-ci / validate-renovate (push) Skipped
ci / lint-compose (push) Successful in 11s
ci / lint-actionlint (push) Successful in 5s
ci / lint-shellcheck (push) Successful in 10s
ci / lint-prettier (push) Successful in 15s
ci / lint-ruff (push) Successful in 8s
ci / lint-yaml (push) Successful in 10s
ci / lint-dockerfiles (push) Successful in 7s
ci / validate (push) Successful in 10s
ci / build (push) Successful in 25s
|
||
|
|
8c0e36a5c0 |
feat(monitoring): replace scrape dump with vmagent
ci / lint-prettier (push) Skipped
ci / lint-ruff (push) Skipped
ci / lint-yaml (push) Skipped
ci / lint-dockerfiles (push) Skipped
ci / validate (push) Skipped
renovate-ci / validate-renovate (push) Skipped
ci / lint-compose (pull_request) Successful in 11s
ci / lint-actionlint (pull_request) Successful in 6s
ci / lint-shellcheck (pull_request) Successful in 16s
ci / lint-dockerfiles (pull_request) Successful in 6s
ci / lint-prettier (pull_request) Successful in 16s
ci / lint-ruff (pull_request) Successful in 7s
ci / lint-yaml (pull_request) Successful in 10s
ci / validate (pull_request) Successful in 7s
ci / build (pull_request) Skipped
renovate-ci / validate-renovate (pull_request) Failing after 1m40s
|
||
|
|
24f84bab2f |
fix(deploy): skip verification when deployment is disabled
ci / lint-yaml (push) Skipped
ci / lint-dockerfiles (push) Skipped
ci / validate (push) Skipped
renovate-ci / validate-renovate (push) Skipped
ci / lint-prettier (push) Skipped
ci / lint-ruff (push) Skipped
renovate-ci / validate-renovate (pull_request) Skipped
ci / lint-compose (pull_request) Canceled after 0s
ci / lint-actionlint (pull_request) Canceled after 0s
ci / lint-shellcheck (pull_request) Canceled after 0s
ci / lint-prettier (pull_request) Canceled after 0s
ci / lint-ruff (pull_request) Canceled after 0s
ci / lint-yaml (pull_request) Canceled after 0s
ci / lint-dockerfiles (pull_request) Canceled after 0s
ci / validate (pull_request) Canceled after 0s
ci / build (pull_request) Canceled after 0s
|
||
|
|
8729cb5062 |
fix(netbird): restore Compose setup and server entrypoint
ci / validate (push) Skipped
ci / lint-prettier (push) Skipped
ci / lint-ruff (push) Skipped
ci / lint-yaml (push) Skipped
ci / lint-dockerfiles (push) Skipped
renovate-ci / validate-renovate (push) Skipped
renovate-ci / validate-renovate (pull_request) Skipped
ci / lint-compose (pull_request) Successful in 12s
ci / lint-actionlint (pull_request) Successful in 8s
ci / lint-shellcheck (pull_request) Successful in 21s
ci / lint-prettier (pull_request) Successful in 17s
ci / lint-ruff (pull_request) Successful in 6s
ci / lint-yaml (pull_request) Successful in 9s
ci / lint-dockerfiles (pull_request) Successful in 5s
ci / validate (pull_request) Successful in 6s
ci / build (pull_request) Skipped
|
||
|
|
b357ef95d8 |
fix(deploy): only create automatic runs for main CI
ci / lint-ruff (push) Skipped
ci / lint-yaml (push) Skipped
ci / lint-dockerfiles (push) Skipped
ci / validate (push) Skipped
ci / lint-prettier (push) Skipped
renovate-ci / validate-renovate (push) Skipped
ci / lint-compose (pull_request) Successful in 13s
ci / lint-actionlint (pull_request) Successful in 6s
ci / lint-shellcheck (pull_request) Successful in 11s
ci / lint-prettier (pull_request) Successful in 18s
ci / lint-ruff (pull_request) Successful in 8s
ci / lint-yaml (pull_request) Successful in 13s
ci / lint-dockerfiles (pull_request) Successful in 5s
ci / validate (pull_request) Successful in 9s
ci / build (pull_request) Skipped
renovate-ci / validate-renovate (pull_request) Successful in 12s
|
||
|
|
67422663b7 |
fix(ci): avoid duplicate branch and PR runs
ci / validate (push) Skipped
ci / lint-prettier (push) Skipped
ci / lint-ruff (push) Skipped
ci / lint-yaml (push) Skipped
ci / lint-dockerfiles (push) Skipped
renovate-ci / validate-renovate (push) Skipped
ci / lint-compose (pull_request) Canceled after 0s
ci / lint-actionlint (pull_request) Canceled after 0s
ci / lint-shellcheck (pull_request) Canceled after 0s
ci / lint-prettier (pull_request) Canceled after 0s
ci / lint-ruff (pull_request) Canceled after 0s
ci / lint-yaml (pull_request) Canceled after 0s
ci / lint-dockerfiles (pull_request) Canceled after 0s
ci / validate (pull_request) Canceled after 0s
ci / build (pull_request) Canceled after 0s
renovate-ci / validate-renovate (pull_request) Successful in 10s
|
||
|
|
c098807aa4 |
fix(deploy): reject destructive per-file pruning before apply
renovate-ci / validate-renovate (push) Skipped
ci / lint-compose (push) Successful in 14s
ci / lint-actionlint (push) Successful in 8s
ci / lint-shellcheck (push) Successful in 13s
ci / lint-prettier (push) Successful in 19s
ci / lint-ruff (push) Successful in 8s
ci / lint-yaml (push) Successful in 12s
ci / lint-dockerfiles (push) Successful in 8s
ci / validate (push) Successful in 10s
ci / build (push) Skipped
ci / lint-compose (pull_request) Successful in 12s
ci / lint-actionlint (pull_request) Successful in 5s
ci / lint-shellcheck (pull_request) Successful in 11s
ci / lint-prettier (pull_request) Successful in 16s
ci / lint-ruff (pull_request) Successful in 8s
ci / lint-yaml (pull_request) Successful in 11s
ci / lint-dockerfiles (pull_request) Successful in 7s
ci / validate (pull_request) Successful in 7s
ci / build (pull_request) Skipped
renovate-ci / validate-renovate (pull_request) Successful in 10s
|
||
|
|
8d3185f8ab |
fix(ci): install jq for deploy validation regressions
renovate-ci / validate-renovate (push) Skipped
ci / lint-compose (push) Successful in 8s
ci / lint-actionlint (push) Successful in 5s
ci / lint-shellcheck (push) Successful in 14s
ci / lint-prettier (push) Successful in 15s
ci / lint-ruff (push) Successful in 5s
ci / lint-yaml (push) Successful in 9s
ci / lint-dockerfiles (push) Successful in 7s
ci / validate (push) Successful in 6s
ci / build (push) Skipped
ci / lint-compose (pull_request) Successful in 11s
ci / lint-actionlint (pull_request) Successful in 7s
ci / lint-shellcheck (pull_request) Successful in 12s
ci / lint-prettier (pull_request) Successful in 14s
ci / lint-ruff (pull_request) Successful in 5s
ci / lint-yaml (pull_request) Successful in 11s
ci / lint-dockerfiles (pull_request) Successful in 6s
ci / validate (pull_request) Successful in 6s
ci / build (pull_request) Skipped
renovate-ci / validate-renovate (pull_request) Successful in 8s
|
||
|
|
ff40a71145 |
fix(deploy): resolve Compose config and check namespaced pod Secrets
renovate-ci / validate-renovate (push) Skipped
ci / lint-compose (push) Successful in 10s
ci / lint-actionlint (push) Successful in 5s
ci / lint-shellcheck (push) Failing after 10s
ci / lint-prettier (push) Successful in 13s
ci / lint-ruff (push) Successful in 5s
ci / lint-yaml (push) Successful in 11s
ci / lint-dockerfiles (push) Successful in 7s
ci / validate (push) Successful in 8s
ci / build (push) Skipped
ci / lint-compose (pull_request) Canceled after 0s
ci / lint-actionlint (pull_request) Canceled after 0s
ci / lint-shellcheck (pull_request) Canceled after 0s
ci / lint-prettier (pull_request) Canceled after 0s
ci / lint-ruff (pull_request) Canceled after 0s
ci / lint-yaml (pull_request) Canceled after 0s
ci / lint-dockerfiles (pull_request) Canceled after 0s
ci / validate (pull_request) Canceled after 0s
ci / build (pull_request) Canceled after 0s
renovate-ci / validate-renovate (pull_request) Successful in 8s
|
||
|
|
51a73fb213 | chore(gitea): switch domain to git.forust.xyz (28.0.0 update) | ||
|
|
22c2f1e108 |
extract dtek_notif subtree to its own repo
ci / lint-compose (push) Canceled after 0s
ci / lint-actionlint (push) Canceled after 0s
ci / lint-shellcheck (push) Canceled after 0s
ci / lint-prettier (push) Canceled after 0s
ci / lint-ruff (push) Canceled after 0s
ci / lint-yaml (push) Canceled after 0s
ci / lint-dockerfiles (push) Canceled after 0s
ci / validate (push) Canceled after 0s
ci / build (push) Canceled after 0s
renovate-ci / validate-renovate (push) Successful in 10s
|
||
|
|
0859479c0f |
feat(ingress): replace traefik crowdsec plugin with firewall bouncer
ci / lint-compose (push) Successful in 9s
ci / lint-actionlint (push) Successful in 4s
ci / lint-shellcheck (push) Successful in 7s
ci / lint-prettier (push) Successful in 12s
ci / lint-ruff (push) Successful in 6s
ci / lint-yaml (push) Successful in 9s
ci / lint-dockerfiles (push) Successful in 5s
ci / validate (push) Successful in 5s
renovate-ci / validate-renovate (push) Successful in 7s
ci / build (push) Failing after 14m22s
Move L3 enforcement to the host firewall-bouncer (systemd, nftables): drop the Traefik plugin, its secrets volume and the crowdsec Middleware, remove bouncer refs from all IngressRoutes. Disable the http-generic-bf scenario (403-burst bans hurt legit automation under L3 enforcement). Add a Gateway API PoC for homepages prod and CrowdSec PrometheusRule alerts. |
||
|
|
872f64b887 |
fix(ci): silence intentional SC2029 in ssh-run.sh
ci / lint-compose (push) Successful in 11s
ci / lint-actionlint (push) Successful in 7s
ci / lint-shellcheck (push) Successful in 9s
ci / lint-prettier (push) Successful in 16s
ci / lint-ruff (push) Successful in 7s
ci / lint-yaml (push) Successful in 12s
ci / lint-dockerfiles (push) Successful in 8s
ci / validate (push) Successful in 8s
renovate-ci / validate-renovate (push) Successful in 22s
ci / build (push) Successful in 38s
|
||
|
|
6f4cd03f4b |
feat(deploy): serialize apply stages and wait for calm node
apply-k8s and apply-compose share a workstation flock so host docker churn never overlaps cluster churn. Each helm upgrade and the apply loop wait up to 10m for load <28 first, so a deploy never piles onto an already-hot node (the load-40/netbird-death/pending-helm spiral). |
||
|
|
e56662194b |
fix(ci): log in to registry for manifest-only main pushes
ci / lint-compose (push) Successful in 11s
ci / lint-actionlint (push) Successful in 6s
ci / lint-shellcheck (push) Successful in 9s
ci / lint-prettier (push) Successful in 15s
ci / lint-ruff (push) Successful in 8s
ci / lint-yaml (push) Successful in 13s
ci / lint-dockerfiles (push) Successful in 8s
ci / validate (push) Successful in 9s
renovate-ci / validate-renovate (push) Successful in 10s
ci / build (push) Successful in 27s
The pin step writes manifest PUTs on every main push, but login was gated on services != ''. Manifest-only pushes skipped login and pushed anonymously (401); this only worked before via a stale persistent login on the old runner. |
||
|
|
5dcad7eb38 |
fix(ci): resolve uv from BIN_DIR in install-ci-tools
ci / lint-compose (push) Successful in 14s
ci / lint-actionlint (push) Successful in 10s
ci / lint-shellcheck (push) Successful in 11s
ci / lint-prettier (push) Successful in 17s
ci / lint-ruff (push) Successful in 9s
ci / lint-yaml (push) Successful in 12s
ci / lint-dockerfiles (push) Successful in 9s
ci / validate (push) Successful in 9s
renovate-ci / validate-renovate (push) Successful in 1m34s
ci / build (push) Failing after 3m3s
Callers prepend BIN_DIR to PATH only after the script exits, so the bare uv invocation in install_uv_tool died with 127 on clean runners. Export BIN_DIR to PATH inside the script and invoke the just-installed binary by absolute path. |
||
|
|
d0872bc918 |
feat(deploy): autodeploy off unless AUTODEPLOY=true
renovate-ci / validate-renovate (push) Successful in 2m1s
ci / lint-shellcheck (push) Successful in 24s
ci / lint-prettier (push) Successful in 6s
ci / lint-ruff (push) Successful in 2s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 2s
ci / validate (push) Successful in 3s
ci / lint-compose (push) Successful in 16s
ci / lint-actionlint (push) Successful in 8s
ci / build (push) Failing after 18s
Pushes no longer reach the cluster by default; set the AUTODEPLOY repo variable to 'true' to re-enable, or dispatch manually. Skips propagate through the existing needs chain. |
||
|
|
91dd749a29 |
feat(deploy): AUTODEPLOY kill-switch via repo variable
ci / lint-compose (push) Successful in 3s
ci / lint-actionlint (push) Successful in 1s
ci / lint-shellcheck (push) Successful in 2s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 2s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 3s
ci / validate (push) Successful in 2s
renovate-ci / validate-renovate (push) Successful in 15s
ci / build (push) Canceled after 0s
Set AUTODEPLOY=false under Settings -> Actions -> Variables and pushes stop deploying with no commit; unset means on. Manual Run workflow always bypasses the switch. The existing needs/skipped chaining propagates the skip through verify and smoke untouched. |
||
|
|
436fd1ecae |
chore: extract userbot subtree to its own repo
ci / lint-compose (push) Successful in 3s
ci / lint-actionlint (push) Successful in 1s
ci / lint-shellcheck (push) Successful in 2s
ci / lint-prettier (push) Successful in 2s
ci / lint-ruff (push) Successful in 2s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 1s
ci / validate (push) Successful in 2s
renovate-ci / validate-renovate (push) Successful in 17s
ci / build (push) Successful in 8s
userbot/ (bot plus panel) now lives at /home/forust/userbot as a clone of forust/userbot instead of a subtree in homelab. Cleans up the pipeline references that only existed for it: scan-deps, test-backend and test-frontend jobs, the userbot build matrix entries, the deploy panel hook, and the userbot-only pyrightconfig. Running cluster workloads are untouched; homelab just stops building and testing upstream's code. |
||
|
|
2cb06debc5 |
fix(deploy): retry registry lookups with a timeout
ci / lint-compose (push) Successful in 3s
ci / lint-actionlint (push) Successful in 1s
ci / lint-shellcheck (push) Successful in 2s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 1s
ci / scan-deps (push) Successful in 19s
ci / test-backend (push) Successful in 8s
ci / test-frontend (push) Successful in 10s
ci / validate (push) Successful in 2s
renovate-ci / validate-renovate (push) Successful in 14s
ci / build (push) Successful in 7s
A single blink of the registry failed render_pinned for the whole file and redded the apply stage. registry_digest now retries 3 times under a 25s timeout with a warning per attempt; empty still means unresolvable and callers report it by name as before. |
||
|
|
76f39da90c |
fix(ci): skip heavy jobs on renovate branches, automerge digest and patch
ci / lint-compose (push) Canceled after 0s
ci / lint-actionlint (push) Canceled after 0s
ci / lint-shellcheck (push) Canceled after 0s
ci / lint-prettier (push) Canceled after 0s
ci / lint-ruff (push) Canceled after 0s
ci / lint-yaml (push) Canceled after 0s
ci / lint-dockerfiles (push) Canceled after 0s
ci / scan-deps (push) Canceled after 0s
ci / test-backend (push) Canceled after 0s
ci / test-frontend (push) Canceled after 0s
ci / validate (push) Canceled after 0s
ci / build (push) Canceled after 0s
renovate-ci / validate-renovate (push) Canceled after 0s
Renovate branches only carry version/digest bumps, so scan-deps, test-backend, test-frontend and build just burn runner time on the box that also serves prod. Static checks and validate still run. Digest and patch updates automerge (playwright, helm and major rules below still override to no-automerge). Also replaces deprecated helm --atomic with --wait --rollback-on-failure. |
||
|
|
dde6b1c743 |
fix(deploy): recover helm releases from pending-* and skip helm-owned rollbacks
An --atomic upgrade whose own rollback never finishes leaves the release in pending-*, blocking every future run until a human rolls back (loki rev 18/21). Recover automatically before and after each upgrade, and fail loud when recovery does not land on deployed. Also skip helm-managed workloads in rollback_workloads: rollout undo there would step back to the revision --atomic just escaped. |
||
|
|
a5409edbf2 |
fix(deploy): fail the smoke stage when Traefik has no route for a host
ci / lint-compose (push) Successful in 5s
ci / lint-actionlint (push) Successful in 2s
ci / lint-shellcheck (push) Successful in 3s
ci / lint-prettier (push) Successful in 2s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 3s
ci / lint-dockerfiles (push) Successful in 3s
ci / scan-deps (push) Successful in 15s
renovate-ci / validate-renovate (push) Successful in 1m15s
ci / test-frontend (push) Successful in 14s
ci / validate (push) Successful in 4s
ci / test-backend (push) Failing after 13m28s
ci / build (push) Skipped
The smoke stage treats any HTTP response as proof the service is serving, which is right -- a 302 to a login or a 404 from a path the app does not serve still means the chain is intact. But a 404 is not evidence of that on its own: a router Traefik refused to build answers with exactly the same 404 and nothing behind it. That is not hypothetical. The crowdsec bouncer is a plugin, and when Traefik cannot fetch it at startup it disables the plugin without failing, then drops every router whose chain referenced it. Sixteen routes answered 404 and the stage printed `ok` for all sixteen, because a dropped router and an unserved path are indistinguishable from outside. The Kubernetes objects cannot tell us either: the IngressRoute is still sitting there looking healthy, the router Traefik built from it is simply not there. So ask Traefik. api.insecure is already on for the internal entrypoint and the router list says which hosts it matches right now. Every probed host has to appear in that list. HTTP routers only -- the TCP ones match on a HostSNI wildcard and the UDP ones carry no rule at all, both selected by entrypoint and port, so neither can answer the question. A router mid-rollout is legitimately absent for a moment, so the list is re-read twice over 20s; a plugin that failed to load stays absent and waiting cannot rescue it. An unreadable router list fails the stage rather than skipping the check, since a check that cannot run is not a passing check. Verified against the live cluster: all 23 public routes have a router and the stage passes. With gitea, grafana and uptime removed from that list the probes still answer and the stage fails on exactly those three. |
||
|
|
c70d2db3a1 |
ci: deploy the image the commit built, not whatever the tag points at
Every service tracked the mutable `:prod` tag, so a deploy applied whatever that tag happened to name at the time rather than the commit it was deploying. A rollback had no way to state what it was rolling back to, and two deploys of one commit could land different images. CI now publishes an immutable `sha-<commit12>` tag beside `:prod` on main, and re-tags it for every image a push did not rebuild. That re-tag copies the manifest list, so no layer moves. The deploy resolves the immutable tag to a digest and pins the workload to it, and only falls back to the moving tag when the immutable one cannot be resolved -- which it says out loud, because that fallback is the deploy quietly ceasing to be reproducible from its own commit. The image list comes out of the tree with git grep rather than being written out a second time, so adding a service no longer means keeping two lists in step. build also gains the three jobs it was skipping -- scan-deps, test-backend, test-frontend -- so a change that breaks them cannot be tagged at all. The two run blocks where a mid-loop failure was survivable now run under set -euo pipefail: the build loop and the service detector both carried on past an error and could report a green build having produced nothing. The registry password moves from run: substitution into an env: block. A quote, a backtick or a $(...) in the password is parsed as shell before the command ever runs, and a login that failed that way looked exactly like a build that failed. The apply and verify timeouts stay at 45 and 30 minutes. The comments now record the arithmetic that says so rather than leaving the numbers to be raised on the next scare: three no-op helm upgrades run 3-5 minutes, one broken release is a single 10 minute rollback because the loop aborts on the first failure, and the apply loop itself is about a minute. That is roughly 15 minutes of work against a 45 minute budget. verify is 32 workloads at 8 wide -- four waves of 300 seconds, 20 minutes -- which leaves room for two serial rollbacks, and only becomes derivable at 45 once rollback_workloads is parallelised. |
||
|
|
11e92fdf4e |
fix(deploy): bound ssh hangs and retry the stage on transport loss
A connection that died silently used to hang until the job timeout, and the stage was never re-run. One flaky TCP session cost a whole 45-minute apply, and the symptom - a job that stops mid-output with no error - is what made the last few deploy failures expensive to read. ServerAliveInterval/CountMax cap how long a dead peer goes unnoticed at ~60s, ConnectTimeout caps setup. Only exit 255 - ssh's own transport failures - is retried, up to three attempts with a growing gap. A stage that fails on its own merits exits with the remote's status, so a real failure surfaces its own log immediately instead of being repeated three times over 45 minutes. The stages are declarative applies, so re-running one that had already committed is harmless. The stage environment now goes through `env` as separate argv entries rather than one interpolated string, so nothing in REPO, DEPLOY_SHA or DEPLOY_SNAPSHOT_DIR is re-split by the remote shell. Verified against a stubbed ssh: clean run attempts once, a single transport failure recovers on attempt 2 and exits 0, three failures give up preserving 255, and a stage failing with 1 or 7 attempts once and passes the code through unchanged. Also records why USERBOT_IMAGE stays on the prod tag: render_pinned rewrites only plain `image:` lines, and this ref is what the panel injects into the per-instance Deployments it creates, so those instances track the tag rather than the panel's own resolved digest. The two panel-created instances currently in the cluster are digest-pinned, so the panel does accept one either way; the tag is the choice, not a limitation. |
||
|
|
6a9a460769 |
fix(deploy): let a pinning failure explain itself
ci / lint-compose (push) Successful in 3s
ci / lint-actionlint (push) Successful in 1s
ci / lint-shellcheck (push) Successful in 2s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 3s
ci / lint-dockerfiles (push) Successful in 2s
ci / scan-deps (push) Successful in 15s
ci / test-backend (push) Successful in 6s
ci / test-frontend (push) Successful in 10s
ci / validate (push) Successful in 2s
renovate-ci / validate-renovate (push) Successful in 26s
ci / build (push) Successful in 1s
registry_digest was written to return an empty string for a ref the registry does not have, so render_pinned could print "cannot resolve <ref>" and stop. It could not do that. Every caller runs under set -euo pipefail, pipefail reports the rightmost non-zero stage, and the failed docker manifest inspect made the assignment itself fail, which set -e turns into an immediate exit. render_pinned therefore died silently on the first unresolvable ref: nothing on stderr, nothing on stdout, exit 1. The apply loop piped that empty stream into kubectl, so the whole deploy stopped with "error: no objects passed to apply" - kubectl guessing at a cause, with the actual reason nowhere in the log. The missing message is the reason the |
||
|
|
f54589a05c |
fix(deploy): put the snapshot where the deploy user can write it
ci / lint-compose (push) Successful in 4s
ci / lint-actionlint (push) Successful in 1s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 2s
ci / lint-shellcheck (push) Successful in 2s
ci / lint-prettier (push) Successful in 3s
ci / test-backend (push) Successful in 7s
ci / lint-dockerfiles (push) Successful in 3s
ci / scan-deps (push) Successful in 15s
ci / test-frontend (push) Successful in 11s
ci / validate (push) Successful in 7s
renovate-ci / validate-renovate (push) Successful in 29s
ci / build (push) Successful in 1s
The first deploy to actually run died on its very first action, and the error the other job reported was only the consequence. DEPLOY_SNAPSHOT_DIR defaulted to /var/backups/homelab-deploy. The deploy is unprivileged, and this Arch host has no /var/backups at all, so snapshot_dir's mkdir -p had to create it under root-owned /var and got Permission denied. It refused to go on, which is exactly what the guard is for, so no workload was touched - but the verify job then found no pointer and could only say to go look by hand. Defaulting to the deploy user's own XDG state directory fixes it with no root and no setup step, and keeps the guard: an unwritable snapshot dir still stops the deploy before the first apply. ssh-run.sh now forwards DEPLOY_SNAPSHOT_DIR too, so the path is overridable without editing the library. Verified on the workstation as the unprivileged user: pointer published, commit recorded, 71 workload generations and three helm releases captured, and the stale-pointer refusal still works. |
||
|
|
f49d91b63d |
Update deploy-lib.sh
ci / lint-compose (push) Successful in 3s
ci / lint-actionlint (push) Successful in 2s
ci / lint-shellcheck (push) Successful in 2s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 3s
ci / lint-dockerfiles (push) Successful in 2s
ci / scan-deps (push) Successful in 17s
ci / test-backend (push) Successful in 7s
ci / test-frontend (push) Successful in 10s
ci / validate (push) Successful in 2s
renovate-ci / validate-renovate (push) Successful in 13s
ci / build (push) Successful in 0s
|
||
|
|
4f74fe1778 |
ci: stop inheriting the runner's python and node
The runner executes jobs on the host rather than in a container, so a
workflow that says 'python3' or 'npm' is really saying 'whatever this
machine happens to have today'. Both of the test jobs added in
|
||
|
|
a5d384a4d8 |
feat(deploy): probe every active service after a deploy, rollouts included
ci / lint-compose (push) Successful in 4s
ci / lint-actionlint (push) Successful in 1s
ci / lint-shellcheck (push) Successful in 2s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 2s
ci / scan-deps (push) Failing after 13s
ci / test-backend (push) Failing after 10s
ci / test-frontend (push) Failing after 1s
ci / validate (push) Successful in 2s
renovate-ci / validate-renovate (push) Successful in 59s
ci / build (push) Successful in 1m50s
verify-k8s watches rollouts, which reports that pods converged. It cannot tell
a converged pod from a serving one. A Service selector pointing at a port
nothing listens on, a 500 from the app itself, a Traefik route that stopped
matching, a pod that OOMKilled early enough to still count as Available for the
duration of the check -- all of those are green at the rollout level and broken
for whoever opens the URL.
So ask what users ask. A smoke stage probes the public route of every active
service and fails on a transport error, a 5xx, or a 000, which curl reports
when it exits cleanly and nothing replied. Everything else passes, including 4xx:
a 404 from a path the service does not serve and a 302 to a login both prove
Traefik matched the host, the Service resolved to a pod and the pod answered,
which is the whole claim being tested.
An empty host list is an error, not a pass. Zero names means the extraction
broke, and reporting a clean deploy off a broken grep is the failure mode this
job exists to catch.
It runs on always() and after verify-k8s rather than before it, because a
rollback is when a route most needs re-checking. It only skips when verify-k8s
did, which is when nothing was deployed at all.
Two things worth writing down, because both were wrong on the first pass:
Stripping comments before reading the routes is not optional. naio and xui are
still in the tree commented out, and a plain grep picks both up and then reports
two services as unreachable when nobody ever deployed them. The apex
forust.xyz also needs a filter that admits it, so /\.forust\.xyz$/ quietly
dropped the site root.
And the 5xx test was written as ${code%%[0-9]*} != 5, which is empty for every
three-digit code, so a 500 was reported as ok. A case glob on the leading digit
is what actually works.
23 routes answer today, in 1.6s. The 5xx branch is the one part no live service
here exercises, so it was checked by running the block over 200 through 599 and
000 rather than against a real response.
|
||
|
|
baedea504d |
ci: scan dependencies for new advisories and stop handing out a write token
Two things, both about not finding out late. No workflow declared `permissions`, so all eighteen jobs across the four workflows ran on a token with the default full repository scope. Every one of them only checks out code, and deploy reaches the cluster over SSH with the deploy key, and Renovate writes through its own bot PAT rather than the Actions token. So `contents: read` is all any of them needed. The panel image ships 15 known advisories and nothing was looking. Add a scan-deps job that fails on anything new, and record the eight current ones by ID in the workflow. It is a list rather than a baseline count so that the diff that accepts an advisory says so in words, and it lives in our workflow instead of the package manifest so a subtree sync from forust/userbot cannot quietly widen the exemption. Both halves were checked to fail on a regression, not just to pass today: removing one --ignore-vuln turns the Python step red, and dropping --audit-level to moderate turns the npm one red on the devalue advisory. npm audits production dependencies only. All seven findings in the full tree are build- or test-time: the esbuild advisory needs a vite dev server exposed to the internet, and nanoid's infinite loop needs a custom generator called with size 0, which postcss does not do. None are in the 91 kB bundle the panel serves, so failing on them would be noise that trains people to ignore the job. The starlette entries are the reason the job is not "fail on everything": fastapi 0.115.12 pins starlette<0.47.0 and the last four fixes need 0.49.1 through 1.3.1, so clearing them is a jump to fastapi 0.141.x and is upstream's call, not a drive-by. Four of the seven are reachable in principle, which the comment on the job sets out. The panel answers only on userbot.workstation.internal with no public route, which is what keeps those four from being an internet-facing DoS. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
30995ee009 |
ci: pin the last four linters instead of trusting the runner
prettier, ruff, yamllint and hadolint were the only CI tools still called bare, straight off whatever the runner happened to have installed. Pin them in tool-versions.env like the other three and install them the same way, so the versions Renovate moves are the versions CI runs. Each pinned version equals what is already on the runner, so this changes what CI does not at all today. It changes what CI does on a rebuilt runner: the pinned one gets installed over the drift. The four need four different mechanisms, which is why this is not one pattern: hadolint a bare binary per platform, like actionlint ruff, yamllint PyPI wheels, unpacked by uv prettier an npm tarball, unpacked by tar prettier is the awkward one. Its entry point requires ../package.json relative to its own real path, so copying the single file out -- which is what every other installer here does -- yields a module-not-found at the first run. It keeps its package directory in a versioned one next to a relative symlink, and the tarball ships bin/ without the exec bit, so that needs chmod too. hadolint's release names one platform uname-style and the other Go-style (x86_64 but arm64), which 404s on the first architecture if you assume otherwise. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
3a05d86e3e |
ci: run svelte-check, which was already a dependency with no script
svelte-check sits in devDependencies at ^4.7.3 and nothing in the repository ever invoked it, so the type errors it reports had no path to a human. Point a script at it and run it in the frontend job, and it is clean: 0 errors, 0 warnings. It shares the one `npm ci` with the test step. A second install would have doubled the slowest part of the job to learn exactly the same thing. The lockfile is untouched, because scripts are not part of what it pins. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
c00a4724f5 |
ci: actually run the test suites that exist in the tree
The panel ships 25 pytest tests and 2 vitest tests. Nothing executed them:
there was no job, no local dev loop, and nothing that would have noticed when
one of them rotted. They pass, and they are 8 seconds of work, which is the
argument for having them.
Both jobs mirror how the image is built rather than how a developer would run
them by hand: `npm ci` because that is what the Dockerfile does, so the tree
under test is the tree that ships, and requirements-dev.txt through uv, which
is now pinned like the other CI tools.
The backend job runs `python -m pytest`, not bare `pytest`. The tests import
`app.*` relative to the backend directory, and only the `-m` form puts the
working directory on sys.path.
ruff format --check joins ruff check in the lint job. It needed
|