fix(deploy): bound ssh hangs and retry the stage on transport loss
A connection that died silently used to hang until the job timeout, and the stage was never re-run. One flaky TCP session cost a whole 45-minute apply, and the symptom - a job that stops mid-output with no error - is what made the last few deploy failures expensive to read. ServerAliveInterval/CountMax cap how long a dead peer goes unnoticed at ~60s, ConnectTimeout caps setup. Only exit 255 - ssh's own transport failures - is retried, up to three attempts with a growing gap. A stage that fails on its own merits exits with the remote's status, so a real failure surfaces its own log immediately instead of being repeated three times over 45 minutes. The stages are declarative applies, so re-running one that had already committed is harmless. The stage environment now goes through `env` as separate argv entries rather than one interpolated string, so nothing in REPO, DEPLOY_SHA or DEPLOY_SNAPSHOT_DIR is re-split by the remote shell. Verified against a stubbed ssh: clean run attempts once, a single transport failure recovers on attempt 2 and exits 0, three failures give up preserving 255, and a stage failing with 1 or 7 attempts once and passes the code through unchanged. Also records why USERBOT_IMAGE stays on the prod tag: render_pinned rewrites only plain `image:` lines, and this ref is what the panel injects into the per-instance Deployments it creates, so those instances track the tag rather than the panel's own resolved digest. The two panel-created instances currently in the cluster are digest-pinned, so the panel does accept one either way; the tag is the choice, not a limitation.
This commit is contained in:
1 parent
b1f98fc148
commit
11e92fdf4e
2 files changed
+37
-4
No files matched your search
@@ -21,10 +21,39 @@ ssh_key="$key_dir/deploy_key"
|
|||||||
printf '%s\n' "$DEPLOY_KEY" > "$ssh_key"
|
printf '%s\n' "$DEPLOY_KEY" > "$ssh_key"
|
||||||
chmod 600 "$ssh_key"
|
chmod 600 "$ssh_key"
|
||||||
|
|
||||||
ssh -i "$ssh_key" -p "$deploy_port" \
|
# A connection that died silently used to hang until the job timeout, and the
|
||||||
-o BatchMode=yes -o StrictHostKeyChecking=accept-new \
|
# stage was never re-run: one flaky TCP session cost a whole 45-minute apply.
|
||||||
"${DEPLOY_USER}@${DEPLOY_HOST}" \
|
# ServerAlive* bounds how long a dead peer goes unnoticed, ConnectTimeout bounds
|
||||||
"REPO=$deploy_path APPLY_PRUNE=${APPLY_PRUNE:-false} DEPLOY_SHA=${DEPLOY_SHA:-} DEPLOY_SNAPSHOT_DIR=${DEPLOY_SNAPSHOT_DIR:-} STAGE=$1 bash -se" <<'EOF'
|
# setup. Only exit 255 - ssh's own transport failures - is retried. A stage that
|
||||||
|
# fails on its own merits exits with the remote's status, so a real failure
|
||||||
|
# still surfaces its own log instead of burning three attempts. The stages are
|
||||||
|
# declarative applies, so re-running one that had already committed is harmless.
|
||||||
|
ssh_opts=(
|
||||||
|
-i "$ssh_key" -p "$deploy_port"
|
||||||
|
-o BatchMode=yes -o StrictHostKeyChecking=accept-new
|
||||||
|
-o ConnectTimeout=15
|
||||||
|
-o ServerAliveInterval=15 -o ServerAliveCountMax=4
|
||||||
|
)
|
||||||
|
|
||||||
|
rc=0
|
||||||
|
for attempt in 1 2 3; do
|
||||||
|
if [ "$attempt" -gt 1 ]; then
|
||||||
|
echo ":: warning::ssh transport failed, retrying (${attempt}/3)"
|
||||||
|
sleep $((attempt * 5))
|
||||||
|
fi
|
||||||
|
rc=0
|
||||||
|
ssh "${ssh_opts[@]}" "${DEPLOY_USER}@${DEPLOY_HOST}" \
|
||||||
|
env "REPO=$deploy_path" "APPLY_PRUNE=${APPLY_PRUNE:-false}" \
|
||||||
|
"DEPLOY_SHA=${DEPLOY_SHA:-}" "DEPLOY_SNAPSHOT_DIR=${DEPLOY_SNAPSHOT_DIR:-}" \
|
||||||
|
"STAGE=$1" bash -se <<'EOF' || rc=$?
|
||||||
source "$REPO/.gitea/workflows/deploy-lib.sh"
|
source "$REPO/.gitea/workflows/deploy-lib.sh"
|
||||||
run_stage "$STAGE"
|
run_stage "$STAGE"
|
||||||
EOF
|
EOF
|
||||||
|
[ "$rc" -eq 0 ] && break
|
||||||
|
[ "$rc" -ne 255 ] && break
|
||||||
|
done
|
||||||
|
|
||||||
|
if [ "$rc" -ne 0 ]; then
|
||||||
|
echo ":: error::stage $1 failed over ssh (exit $rc)"
|
||||||
|
fi
|
||||||
|
exit "$rc"
|
||||||
@@ -210,6 +210,10 @@ spec:
|
|||||||
- name: USERBOT_IMAGE
|
- name: USERBOT_IMAGE
|
||||||
# The build pushes main/prod only. The deploy resolves every gcr ref in
|
# The build pushes main/prod only. The deploy resolves every gcr ref in
|
||||||
# this file, so a :latest here aborts the whole apply as unresolvable.
|
# this file, so a :latest here aborts the whole apply as unresolvable.
|
||||||
|
# Deliberately left on the tag: render_pinned only rewrites plain
|
||||||
|
# `image:` lines to a digest, and this ref is what the panel injects
|
||||||
|
# into the per-instance Deployments it creates. Instances track prod
|
||||||
|
# rather than the panel's own resolved digest.
|
||||||
value: gcr.forust.xyz/forust/userbot:prod
|
value: gcr.forust.xyz/forust/userbot:prod
|
||||||
- name: USERBOT_STORAGE_CLASS
|
- name: USERBOT_STORAGE_CLASS
|
||||||
value: local-path-retain
|
value: local-path-retain
|
||||||
|
|||||||
Reference in new issue
Block a user