Uptime Monitoring

We monitor the site with HetrixTools. The dashboard is at hetrixtools.com/dashboard/uptime-monitors (ask a team member for access). It polls the site from several locations around the world and sends an alert when the site stops responding, so we usually hear about an outage before a student does.

What the monitor checks

The monitor requests a dedicated health page:

https://www.hackleyclubz.org/uptime.php

That page is small on purpose, but it proves the whole chain works: DNS, the web server, PHP, the site configuration, and the connection to the database (which lives on a separate DreamHost machine). On a healthy site it returns HTTP 200 and a body like:

success : 2026-09-14 14:03:21 EDT

If anything in that chain is broken it returns HTTP 503 and the body starts with failure : followed by the reason. The monitor treats any non-200 response, or a timeout, as down.

You can open the same URL in a browser at any time to see the current state for yourself.

When an alert arrives

Work down this list; stop as soon as the site comes back.

  1. Confirm it. Open https://www.hackleyclubz.org/uptime.php yourself. A single missed poll from one location is sometimes a network blip, and HetrixTools re-checks from other locations before declaring an outage.
  2. Read the body. If the page loads but says failure :, the reason after the second colon tells you which piece is broken. A database error usually means the database host is down or the server’s IP has lost access to it. The web server and PHP are fine in that case.
  3. Check the services. If the main site is up but chat or images are not, log in as an administrator and go to Admin → Maintenance → Restart Services. That page shows which background services are running and can restart each one without touching the rest of the server.
  4. Reboot. If the page does not load at all and SSH does not respond either, follow How to Reboot.
  5. Rebuild. If the machine does not come back from a reboot, it is time for the recovery runbook in the source repository (docs/how_to_recover_from_backup_from_machine_failure.txt).

Nightly backups are taken automatically and listed under Admin → Maintenance → Backups. A fresh dated file there every morning is the other half of “is everything okay.”