fix(edu-master): stop 20722d false alert on zeroed last_success gauge
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 1s
ci / validate (push) Successful in 2s
deploy / preflight (push) Successful in 2s
renovate-ci / validate-renovate (push) Successful in 14s
ci / build (push) Successful in 2s
deploy / validate (push) Successful in 1m52s
deploy / apply-k8s (push) Successful in 1m48s
deploy / apply-compose (push) Successful in 13s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 1s
ci / validate (push) Successful in 2s
deploy / preflight (push) Successful in 2s
renovate-ci / validate-renovate (push) Successful in 14s
ci / build (push) Successful in 2s
deploy / validate (push) Successful in 1m52s
deploy / apply-k8s (push) Successful in 1m48s
deploy / apply-compose (push) Successful in 13s
checker.py initialises last_success to 0, so right after a pod restart `time() - last_success` equals the current epoch. The rule compared that against 300, went firing instantly, and humanizeDuration rendered the raw epoch delta as ~20722d. The last_run > 0 guard did not help because a run happens long before the first success. Guard the duration rule on last_success > 0, keeping the duration expression on the left of `and` so $value stays the real gap, and add a separate WebinarCheckerNeverSucceeded rule for the zeroed-gauge case so a checker that has never succeeded is still caught.
This commit is contained in:
1 parent
b625d30568
commit
1d9a85bef9
1 file changed
+20
-1
@@ -11,9 +11,14 @@ spec:
|
|||||||
rules:
|
rules:
|
||||||
# No successful webinar check for 5m (~2-3 missed 2-min checks).
|
# No successful webinar check for 5m (~2-3 missed 2-min checks).
|
||||||
# Catches: playwright hangs/timeouts, version skew, site changes, hung job.
|
# Catches: playwright hangs/timeouts, version skew, site changes, hung job.
|
||||||
|
# The last_success > 0 guard is mandatory: checker.py initialises
|
||||||
|
# last_success to 0, so without it `time() - 0` equals the current epoch
|
||||||
|
# and humanizeDuration renders ~20722d on every pod restart. Keep the
|
||||||
|
# duration expression on the left so $value stays the real gap.
|
||||||
- alert: WebinarCheckerNoSuccessfulCheck
|
- alert: WebinarCheckerNoSuccessfulCheck
|
||||||
expr: |
|
expr: |
|
||||||
(time() - webinar_check_last_success_timestamp_seconds > 300)
|
((time() - webinar_check_last_success_timestamp_seconds) > 300)
|
||||||
|
and (webinar_check_last_success_timestamp_seconds > 0)
|
||||||
and (webinar_check_last_run_timestamp_seconds > 0)
|
and (webinar_check_last_run_timestamp_seconds > 0)
|
||||||
for: 2m
|
for: 2m
|
||||||
labels:
|
labels:
|
||||||
@@ -22,6 +27,20 @@ spec:
|
|||||||
summary: "Webinar checker has no successful check for 5m"
|
summary: "Webinar checker has no successful check for 5m"
|
||||||
description: "edu-master/webinar-checker: last successful webinar check was {{ $value | humanizeDuration }} ago. Checks are failing or hanging (see consecutive failures alert). Notifications about new webinars are NOT being sent."
|
description: "edu-master/webinar-checker: last successful webinar check was {{ $value | humanizeDuration }} ago. Checks are failing or hanging (see consecutive failures alert). Notifications about new webinars are NOT being sent."
|
||||||
|
|
||||||
|
# Checks are running but none has ever succeeded since pod start.
|
||||||
|
# Split out from the rule above so a zeroed gauge never feeds
|
||||||
|
# humanizeDuration.
|
||||||
|
- alert: WebinarCheckerNeverSucceeded
|
||||||
|
expr: |
|
||||||
|
(webinar_check_last_success_timestamp_seconds == 0)
|
||||||
|
and (webinar_check_last_run_timestamp_seconds > 0)
|
||||||
|
for: 10m
|
||||||
|
labels:
|
||||||
|
severity: critical
|
||||||
|
annotations:
|
||||||
|
summary: "Webinar checker has never completed a successful check"
|
||||||
|
description: 'edu-master/webinar-checker: checks have been running for 10m but not one has ever succeeded since the pod started, so every check is failing. Check pod logs (Loki: {namespace="edu-master", container="webinar-checker"}).'
|
||||||
|
|
||||||
# Fast path: 3 consecutive failures (~6+ min at 2-min interval).
|
# Fast path: 3 consecutive failures (~6+ min at 2-min interval).
|
||||||
- alert: WebinarCheckerConsecutiveFailures
|
- alert: WebinarCheckerConsecutiveFailures
|
||||||
expr: |
|
expr: |
|
||||||
|
|||||||
Reference in new issue
Block a user