What a hosting control panel should do when a site is down
A site can be down for visitors while the server is healthy. Six things a control panel owes you when that happens, and how KLYRN does each.
Two different questions
"Is the site up?" is two questions, and most tools answer only the first.
Is this server's copy of the site working? nginx is running, the PHP pool answers, the database is there, the disk is not full.
Are visitors being served? The domain resolves to something, that something reaches this server, and a page comes back.
A panel that checks itself on 127.0.0.1 can say yes to the first while the answer to the second is no. Every light is green and the customer is on the phone. Here is what a panel should do instead.
1. Check in the order a request travels
A request passes through the web server, the application runtime, the database and the disk, in that order. A diagnosis that follows the same path finds the first broken link without guessing.
klyrn site diagnose example.com
KLYRN checks whether the site is suspended, the nginx service and this site's vhost, the PHP-FPM pool and its socket or the application service, the database, free space and inodes, the certificate and the names it covers, what the site returns over HTTP, and recent KLYRN jobs for the site.
2. Run every check, even after one fails
A site can be broken in two ways at once. A tool that stops at the first failure makes you fix the site twice. KLYRN runs every probe and sorts the findings with the critical ones first.
It also avoids saying the same thing twice. When a stopped PHP pool already explains a 502, the HTTP check confirms the symptom and does not report a second problem.
3. Name a cause and show the evidence
"502 Bad Gateway" is a symptom. "The pool is not running", "the pool file is missing", "the socket is stale" and "the app is down" are four different causes with four different fixes, and KLYRN's diagnosis tells them apart.
Each finding carries the evidence it was based on. When nginx will not accept its configuration, the finding includes what nginx -t said. You should be able to disagree with a diagnosis, and you cannot disagree with a red dot.
There are three severities: critical, attention and healthy. More would invite arguing about the boundary instead of fixing the site.
4. Connect the breakage to the last change
Most outages follow a change. The panel made many of those changes itself, so it is the one tool that can put "this site broke" next to "this job ran on it an hour ago". That is why recent jobs for the site are part of the diagnosis.
5. Offer a small repair, then check again
When the cause has a known fix, offer it by name:
klyrn site diagnose example.com --repair php.pool.restart
KLYRN has six repairs and no seventh: reload nginx, restart the PHP pool, restart the application, start MariaDB, repair the site's file permissions, and request the certificate again. A repair has no field that could carry a command or a path, so it cannot be turned into something else.
Every repair runs the diagnosis again afterwards and reports the worst remaining finding. A repair that does not say whether it worked is a button, not a fix.
6. Ask what a visitor actually gets
This is the second question from the top, and since 1.0.2 KLYRN's diagnosis asks it. When the domain does not resolve to this server, it fetches the site by its public name and checks whether a request for the domain arrives here at all.
That separates three situations a local check cannot tell apart:
- A proxy answers with an error and the request never reaches this server. The DNS record behind the proxy is wrong. The site is reported as down, with the record to change.
- A proxy reaches this server and HTTPS between them fails. The site needs its certificate.
- The domain is served by another machine. Expected before a migration cut-over, and a finding at any other time.
In each case the report says plainly that the rest of it describes the copy on this server.
And when the panel itself is the problem
A diagnosis that needs the panel to be running is no use when the panel is what broke. klyrn doctor covers the whole server with 22 checks, failures first. It reads the filesystem, systemd and the network directly, so it still answers when the panel is down, and its exit code tells monitoring whether anything is failing.