Skip to content

fix: raise healthcheck timeout so monitor can finish - #156

Merged
dgibbs64 merged 2 commits into
mainfrom
fix/healthcheck-timeout
Oct 4, 2026
Merged

dgibbs64 merged 2 commits into
mainfrom
fix/healthcheck-timeout

Conversation

@dgibbs64

@dgibbs64 dgibbs64 commented Oct 3, 2026 •

Copy link
Copy Markdown
Member

Summary

Raises the image HEALTHCHECK timeout from 1m to 5m, in Dockerfile.j2 and all generated dockerfiles/*. This is the same single-line change in each file.

Why

The healthcheck runs /app/entrypoint-healthcheck.sh, which runs LinuxGSM's monitor. When the server doesn't answer, monitor keeps querying for at least 60 seconds by design ("Query will wait up to 60 seconds to confirm server is down", in command_monitor.sh) before it restarts the server. With --timeout=1m, Docker kills the check before it finishes, so:

  • a server that is actually down never gets restarted, because monitor is killed first
  • the container flips to unhealthy on every check (--retries=1)

On the LinuxGSM test fleet, colserver and inssserver show Health check exceeded timeout (1m0s) on every check, about 360 times a day, while their monitor is still partway through querying.

5 minutes covers the 60s query window, per-query timeouts across several IPs, and a stop and start. A healthy server still returns in seconds, so normal checks are unaffected. Interval, start period and retries are unchanged.

Notes

  • Dockerfile.j2 isn't a repo-sync file, so this won't be overwritten.
  • I edited the generated Dockerfiles directly rather than running generate-dockerfiles.sh, which regenerates from LinuxGSM master's server list. The healthcheck line was identical in all 139 files.

@dgibbs64
dgibbs64 force-pushed the fix/healthcheck-timeout branch from 051a188 to 53c319e Compare October 3, 2026 23:52
GitHub Copilot and others added 2 commits October 4, 2026 23:07
The healthcheck runs the LinuxGSM monitor command. Monitor deliberately
keeps querying for 60 seconds before deciding a server is down, then
restarts it. With a 1 minute timeout, Docker killed every check on a
down server before it could restart, so the container sat unhealthy
and never recovered.

Raise the timeout to 5 minutes in Dockerfile.j2 and every generated
Dockerfile, to cover the query window plus a stop and start.
@dgibbs64
dgibbs64 force-pushed the fix/healthcheck-timeout branch from 334caf0 to 5508da6 Compare October 4, 2026 22:08
@dgibbs64
dgibbs64 merged commit 79dc996 into main Oct 4, 2026
2 of 3 checks passed
@dgibbs64
dgibbs64 deleted the fix/healthcheck-timeout branch October 4, 2026 22:09
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant