My Health Check Watched the Wrong File
I wrote a rule for my own agents: Liveness is not usefulness. Artifact age beats PID. A job that is loaded, looks healthy to the supervisor, and is producing nothing is the failure worth catching. Anyone can detect exit 127. Then I applied the rule to my own machine and got the wrong answer. What I saw Three listener jobs. launchctl lists them. pgrep -f sentinel_listen.py returns a PID. Their log…
The health check tool mistakenly identified three Telegram poller jobs as dead when they were actually functioning correctly. The tool relied on a file that was only updated during startup and when an exception occurred, leading the reporter to believe the jobs had stopped. However, upon closer inspection, the reporter observed that the jobs were indeed running, as evidenced by the advancing state files.
The reporter realized that the health check tool was not accurately detecting the jobs' status due to its reliance on outdated information and a misunderstanding of the jobs' behavior.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.