Files
homelab/incident-archive/Overlay FS failure.md
2026-01-21 17:33:18 +01:00

55 lines
1.8 KiB
Markdown

# Resolved: Overlay FS failure (and so containers)
15-12-2025 03:02 AM EET: Degraded control panels' performances, following by full cascade docker failure
15-12-2025 04:36 AM EET: Identified: Services are terminated due to server software (Overlay FS) + hardware issues (HDD).
15-12-2025 08:45 PM EET: Restored NextCloud service with few tweaks to lower I/O
---
16-12-2025 08:34 PM EET: Ordered new HDD, ETA 22nd of December - 2nd of January
---
17-12-2025 02:47 AM EET: To avoid additional I/O into kuma's database, disabled uptime monitoring for non-critical services, such as:
- Game servers
- Gitea (no public projects being hosted yet)
- Landings
- Cloud services
- PenPot
- Auth provider
- Secondary management tools
- Chernuha's non-important infrastructure
These can be identified by seeing ">2m ago" under monitor's heartbeats.
---
09-01-2026 02:32 PM EET: NextCloud's frontend files are corrupted due to the unknown issue. All user data is integrity-verified. To prevent user data corruption, NextCloud service will be restored after new hardware will be available.
---
14-01-2026 08:42 PM EET: After planned updating and restarting server, critical firmware software were corrupted because of physical degradation of the disk. Server inaccessible in any way
---
15-01-2026 11:30 AM EET: A new NAS-Grade HDD (Seagate IronWolf Pro) was ordered. ETA 16-01-2026 EET Before 12:00 PM
---
16-01-2026 12:47 AM EET: New server system is installed, data backed up. Experiencing docker memory leak.
---
17-01-2026 01:27 PM EET: All services except NextCloud and Satisfactory server are online.
17-01-2026 03:48 PM EET: Nextcloud is online. Satisfactory will be provided on-demand. Monitoring status
---
##### Status: All services are online
**Solution: moving all infrastructure onto new NAS-Grade HDD with fresh OS install**