Repository navigation
fix: raise healthcheck timeout so monitor can finish - #156
Merged
Merged
Conversation
dgibbs64
force-pushed
the
fix/healthcheck-timeout
branch
from
October 3, 2026 23:52
051a188 to
53c319e
Compare
The healthcheck runs the LinuxGSM monitor command. Monitor deliberately keeps querying for 60 seconds before deciding a server is down, then restarts it. With a 1 minute timeout, Docker killed every check on a down server before it could restart, so the container sat unhealthy and never recovered. Raise the timeout to 5 minutes in Dockerfile.j2 and every generated Dockerfile, to cover the query window plus a stop and start.
dgibbs64
force-pushed
the
fix/healthcheck-timeout
branch
from
October 4, 2026 22:08
334caf0 to
5508da6
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Raises the image
HEALTHCHECKtimeout from 1m to 5m, inDockerfile.j2and all generateddockerfiles/*. This is the same single-line change in each file.Why
The healthcheck runs
/app/entrypoint-healthcheck.sh, which runs LinuxGSM'smonitor. When the server doesn't answer,monitorkeeps querying for at least 60 seconds by design ("Query will wait up to 60 seconds to confirm server is down", incommand_monitor.sh) before it restarts the server. With--timeout=1m, Docker kills the check before it finishes, so:--retries=1)On the LinuxGSM test fleet,
colserverandinssservershowHealth check exceeded timeout (1m0s)on every check, about 360 times a day, while their monitor is still partway through querying.5 minutes covers the 60s query window, per-query timeouts across several IPs, and a stop and start. A healthy server still returns in seconds, so normal checks are unaffected. Interval, start period and retries are unchanged.
Notes
Dockerfile.j2isn't a repo-sync file, so this won't be overwritten.generate-dockerfiles.sh, which regenerates from LinuxGSMmaster's server list. The healthcheck line was identical in all 139 files.