5.7 KiB
Homelab CI/CD
The native Gitea runner runs on vps; production runs on workstation.
Jobs run on homelab:host, one at a time. No job images or Kubernetes credentials
are needed on the VPS. Builds use one pinned BuildKit helper container. CI and deploy are separate workflows.
Runner installation
Install Docker Engine with Compose and Buildx, Git, Python 3.11+, Bash, curl,
GNU tar/xz, flock and systemd using the host's package manager. Keep the existing
Gitea runner 3.0.2 binary at /usr/local/bin/gitea-runner.
From this checkout on the VPS:
sudo bash .gitea/runner/setup-runner.sh
The installer reuses /var/lib/gitea-runner/.runner and the existing service.
For a new host, install the same runner binary and register as gitea-runner
using the registration token interactively, label homelab:host, and working
directory /var/lib/gitea-runner; then rerun the installer. Tokens never belong
in this repository or command-line examples.
Pinned tools live in the runner user's ~/.cache/homelab-ci; CI repairs version
drift there. Installations are locked. Buildx uses only the homelab-ci builder,
pushes directly to the registry, and caps retained local cache at 1 GiB with a
2 GiB free-space target. This is not a hard limit on peak build disk usage.
Nothing runs docker system prune, removes unrelated images, or deletes volumes.
Workstation setup
As the existing SSH deploy user on workstation:
sudo loginctl enable-linger forust
bash .gitea/runner/setup-workstation.sh
The controller uses /srv/homelab as the persistent configuration tree and makes
a detached source worktree for each SHA. It never resets /srv/homelab, moves
local configuration, renames Compose projects, or changes volume names.
The installer records the current Kubernetes context and cluster UID in
~/.config/homelab-deploy/environment. Check these before installing.
Configure Gitea Actions Variables:
DEPLOY_HOST,DEPLOY_USER,DEPLOY_PORT: the existing VPS-to-workstation SSH endpoint.DEPLOY_KNOWN_HOSTS: workstation's verified SSH host key entry for that endpoint.AUTODEPLOY:falseinitially;trueenables deployment after successful main CI.
Keep DEPLOY_SSH_KEY, REGISTRY_USERNAME and REGISTRY_PASSWORD in Actions
Secrets. Legacy endpoint secrets remain accepted during migration. The Actions
token must have repository read and Actions read access for release downloads.
The deploy user's existing Docker registry authentication remains necessary.
Releases and deployment
CI publishes release-<full SHA> as a Gitea artifact with all three owned image
digests and build input fingerprints. Unchanged images are reused only from a
successful main CI artifact, never from :prod. EDU images remain pinned to the
digests released by their application repository. Expired artifacts cause CI to
rebuild images; they block deployment until CI is rerun.
Run deploy from main with deploy_ref=main or a checked SHA:
full: required for the first baseline; reconcile all active components.changed: compare with the last fully successful production deploy.plan: validate configuration and show selection without changing production resources.refresh_images=true: explicitly refresh mutable third-party Compose tags.
The manual and automatic paths both require successful CI, a successful build
job and the exact SHA's release artifact. PRs cannot publish images or deploy.
Removed resources are reported and require explicit removal; no automatic prune.
Service dependencies are listed in .gitea/deploy-dependencies.json.
A workstation user systemd service holds the deploy lock across validation,
sequential apply, verification and smoke checks. SSH clients only submit/follow:
disconnecting or cancelling the Actions client does not kill production apply.
Retrying the same run ID does not start another apply. ExecStopPost recovers
interrupted runs before the unit finishes. Kubernetes rolls back to captured
revisions; configuration and persistent data are not reverted.
Status and recovery
--retry repeats failed verification and smoke checks, never apply. Recovery
keeps a failed deploy out of the successful baseline, even after rollback.
On workstation (replace the numeric ID with Actions run ID and attempt):
python3 ~/.local/lib/homelab-deploy/controller.py status 123-1
python3 ~/.local/lib/homelab-deploy/controller.py recover 123-1 --retry
journalctl --user -u homelab-deploy@123-1
Runs live in ~/.local/state/homelab-deploy/runs. Compose stores resolved configs
with restricted permissions; these may contain credentials and must never be
uploaded as CI artifacts. Stage logs print the exact manual recovery command
using compose-before/<stack>.json, the original project directory and project
name. Compose does not automatically roll back, and Nextcloud AIO's child
containers remain managed by AIO. Preserve its own backups for data recovery.
The controller retains twenty successful/planned runs and preserves failures. Update the workstation dispatcher only when no deploy is running.
Validation and migration rollback
python3 -m unittest discover -s tests -v
bash .gitea/tests/deploy-validation.sh
Test on a separate namespace before the initial production full run. Check a
failed rollout, interrupted SSH and repeated run ID, and verify that an isolated
service change does not upgrade unrelated Helm releases or Compose stacks.
To roll back the migration, disable autodeploy and finish or recover the remote
run first. Restore the runner config/unit from .before-<timestamp> backups,
reload systemd and restart the runner. Restore the prior workflows from Git.
Production data and persistent volumes stay where they were. Do not remove run
state or Compose recovery files until recovery is confirmed.