Architecture

How Long Until You Know Your Site Is Down: What to Monitor

Most site owners hear it from a customer. Why homepage-only checks miss the expensive failures, which pages to watch, and where alerts should go.

E
Eric Founder, Roamer Tech · · 6 min read

Want to start now? One-click deploy your own service

No servers to set up, isolated container, automatic SSL, cancel renewal anytime.

See plans & pricing

One question: if your website went down right now, how long would it take you to find out?

For most people the answer is “when somebody tells me”. And the person who tells you is usually the one who was about to place an order and could not.

The 30-second version

QuestionAnswer
Is monitoring the homepage enoughNo — it misses the expensive failures
The one page to monitorThe page that makes money (checkout, booking, forms)
Is a status code enoughNo — check for a keyword on the page too
Where alerts should goSomewhere you already look
Check intervalToo frequent produces false alarms, which make you mute the alerts

What homepage-only monitoring misses

The most common setup is to monitor the homepage and check whether it returns 200. That catches a server going down entirely, but not the most expensive kinds of failure.

A white screen often returns 200

Under some configurations, a PHP fatal error outputs a blank page while the HTTP status code stays 200. Monitoring says everything is fine; visitors see a blank screen.

The homepage may be propped up by cache

The database is down, but the homepage has a full-page cache and still renders normally. Your monitoring is perfectly happy, while every page that needs live computation — login, cart, checkout — is broken.

A single broken page does not affect the homepage

A plugin update breaks the checkout page while the homepage is completely fine. That kind of failure can run for days without anyone noticing.

What you should monitor

TargetWhy
The homepage (with keyword matching)Not just the status code — check that the text that should be there is there, which catches white screens
The page that makes moneyCheckout, booking, sign-up forms. This is the one to set up first
The login pageNeeds code and a database, so it catches failures that caching hides
SSL certificate expiryAn expired certificate throws a warning on every page, which scares people off worse than being down
Domain expiryForgetting to renew is extremely costly, and recovery is slow

Of these, keyword matching is the one most worth setting up. Rather than asking “did it respond”, ask “does the response contain what it should” — for example, that the word “Checkout” appears on the checkout page. This one trick catches the vast majority of cases where the status code lies.

Alerts have to reach somewhere you actually look

Monitoring configured perfectly but alerting to a mailbox you open once a week is the same as no monitoring at all.

Alerts should go to LINE or Telegram — somewhere you already look. Most monitoring tools support these channels; pick the one you will genuinely see in real time.

Also turn on recovery notifications. If you only ever get “it is down” and never “it is back”, you are permanently unsure what the current state is.

The trade-off in check intervals

Checking every minute is not automatically better. Checks that are too frequent turn brief network hiccups into a pile of false alarms, and once you have had enough false alarms you start ignoring the notifications — which is the real risk.

A more practical approach is to grade by page importance: tighter on the checkout page, looser on ordinary pages. Pair that with a “only alert after N consecutive failures” setting to filter out one-off blips.

Choosing a tool

Monitoring services come in three categories, and for most sites the free tier is enough:

TypeGood forWatch out for
Free SaaS monitoringMonitoring one to three sitesCheck interval and number of monitors are usually capped
Paid SaaS monitoringMonitoring a lot of thingsMostly priced per monitor
Self-hosted open source toolsYou already have a server, or need to monitor an internal networkIt cannot live on the same machine as what it monitors

Self-hosted monitoring has a fundamental contradiction

This deserves its own section: if the monitoring system sits on the same machine as the thing it monitors, you learn nothing when that machine goes down.

And that is precisely the most common kind of failure — the host’s network drops, hardware fails, resources run out and the whole machine stops responding. In every one of those scenarios, same-machine monitoring is useless.

So for self-hosted monitoring to mean anything, it has to live on a different machine, ideally in a different data center. At which point you are running an extra machine just for monitoring, and that machine needs somebody to notice whether it is still alive.

“Who monitors the monitoring system” is the question you cannot get around here, and it is why most small teams simply use an external SaaS monitor.

When you do not need this at all

Honestly: if your site is a personal blog and half a day of downtime costs you nothing, monitoring is just one more thing to manage.

The signal that it is worth setting up is clear — the site has something that is lost when it goes down: orders, bookings, sign-ups, inquiries. If it does, the value of monitoring is not that it fixes the site; it is that it turns “the customer noticed” into “you noticed first”. That gap is usually several hours, plus one apology.

FAQ

Q: Doesn’t managed hosting already monitor this?

A host’s monitoring watches whether the server and services are alive, which is a different question from “are your pages rendering correctly”. When a plugin breaks the checkout page, the server is perfectly healthy.

Q: Do I have to monitor a lot of pages?

No. Start with one — the page that makes money. That page costs the most when it breaks, and it is usually also the most complex and the most likely to break.

Q: Why is a status code not enough?

Because status codes lie. A white screen can return 200, and so can a hacked page. Matching the text that should be on the page is what actually confirms the site is usable.

Q: What do I do about false alarms?

Loosen the check interval and require two consecutive failures before alerting. False alarms are more dangerous than missed ones — they train you to ignore notifications.

Q: What if my site goes down often?

Then monitoring only tells you; it does not fix the problem. Frequent outages usually mean insufficient resources or plugin conflicts; how to tell is covered in when to upgrade your plan and WooCommerce hosting requirements.

Further reading

Want someone to build it for you?

If you would rather not assemble these workflows yourself, or the project is large enough that you want someone planning it with you, Roamer Tech (RoamerHost’s parent company) takes on business process automation and AI agent development work:

Ready to get started?

60 seconds after you subscribe, your service is installed automatically — you just use it, we handle the rest.

See plans & pricing

Billed monthly or yearly by plan · cancel renewal anytime

Hi, I'm Roamer! Tap me anytime with a question and I'll help you out.

Roamer

Roamer - AI assistant

Online
Roamer

Ask me anything, anytime — I'll do my best to help!

Powered by RoamerHost AI