Kuma Maintenance Windows vs Healthchecks Grace: Planned Downtime That Still Pages You

Diana Vos

Diana Vos

September 21, 2026

Kuma Maintenance Windows vs Healthchecks Grace: Planned Downtime That Still Pages You

You scheduled the upgrade. You told yourself the alerts were handled. At 01:14 your phone still lights up because “handled” meant three different things in three different tools. Uptime Kuma’s maintenance windows and Healthchecks.io grace periods both try to silence planned pain. They are not interchangeable, and mixing up which one you set is how planned downtime still pages you.

This is for solo builders who run both a pull-style uptime board and push-style cron heartbeats—or who are about to, after one too many false alarms during a Compose recreate.

Two different clocks for “don’t bother me”

Uptime Kuma maintenance is calendar-shaped. You mark monitors (or groups) as under maintenance for a window. During that window, failed checks should not fire notifications the way a surprise outage would. The mental model is: the service is allowed to be down because a human said so for a time range.

Healthchecks grace is schedule-shaped. A check expects a ping by a deadline. Grace is extra time after that deadline before the check is considered late. The mental model is: jobs sometimes run long, clocks drift, and DST is rude—so wait a bit before yelling.

Maintenance answers “we are breaking this on purpose.” Grace answers “this job is allowed to be late by N minutes.” Using grace as a substitute for maintenance—or maintenance as a substitute for grace—creates the exact pages you were trying to avoid.

Person silencing a phone beside a glowing laptop at night

Why planned downtime still pages people

The usual failure is incomplete coverage. You pause the Immich HTTP monitor in Kuma and forget the keyword monitor on the API. You pause the group that says “media” and leave the reverse-proxy monitor outside the group. You pause Kuma and forget that Healthchecks is watching the backup script that cannot finish while the disk is remounted.

Another failure is timezone math. Maintenance set in the UI for “tonight” on a VPS that thinks it lives in UTC while you think in local time. Grace set in minutes that do not cover a restore that always takes forty minutes when the catalog is large.

A third failure is overlapping systems with different owners in your head. Kuma watches the URL. Healthchecks watches the cron. ntfy still receives both. You muted one channel in Telegram and not the other. Planned downtime is a systems problem, not a single checkbox problem.

Workshop bench cleared during an upgrade under a lamp

When Kuma maintenance windows are the right tool

Use maintenance when the thing that will fail is a pull check: HTTP, TCP, ping, keyword—anything Kuma initiates.

  • Rebooting a NAS, hypervisor, or Docker host.
  • Certificate renewal that briefly breaks a vhost.
  • Migrating a container between machines.
  • Intentional offline windows for power or noise reasons.

Put related monitors in a group before you need the window. The night of the migration is a bad time to discover that “Photos” and “Immich API” are not in the same group. Name groups after blast radius, not after vibes: host-nas01 beats misc.

Write the maintenance note as if future-you will read it during a real incident next month. “Postgres major upgrade on nas01; expect Immich and Nextcloud red” is better than “upgrade.”

End the window on purpose. Open-ended maintenance that you forget to clear trains you to ignore the dashboard. A dashboard you ignore is worse than noisy alerts.

When Healthchecks grace is the right tool

Use grace when the signal is a missable ping from a job—not when a website is allowed to be down.

  • Backup scripts that sometimes run long.
  • Sync jobs that wait on a slow upstream.
  • Weekly tasks that collide with other maintenance.
  • Hosts whose clocks are only mostly honest.

Grace is not “silence forever while I rebuild the server.” If the cron host is offline for three hours, a fifteen-minute grace still pages you—and it should, unless you paused or disabled the check. For host rebuilds, pause the check, use a longer schedule temporary override if your plan supports it, or accept the page as the cost of a true miss.

Keep grace tight enough that a stuck job still hurts. A two-hour grace on a five-minute heartbeat turns Healthchecks into a polite diary. Many people widen grace after one false alarm and never narrow it again. Review grace the way you review retry counts: after the incident, not during the apology text.

The combined runbook that usually works

  1. List every notifier that can reach your phone for the services you will touch.
  2. In Kuma, start maintenance on the group that covers those HTTP/TCP monitors. Confirm with a deliberate fail if you can do it safely.
  3. In Healthchecks, pause checks for jobs that cannot run during the window, or temporarily extend grace only if the job will still run and merely run late.
  4. Mute secondary channels only if you understand what else they carry.
  5. Do the work.
  6. Clear maintenance. Unpause checks. Watch one clean cycle before you sleep.

Step six is where people fail. They clear Kuma, forget Healthchecks, and get paged by a backup that could not run—which is correct behavior that feels like betrayal at 3 a.m.

Maintenance versus retries versus grace

Kuma retries soften blips. They are not maintenance. A three-retry policy will still alert if the service stays down for the whole upgrade. Retries buy you seconds to minutes, not an hour of Compose pulls.

Healthchecks grace softens late pings. It does not mark a check as “expected down.” If you need expected silence, pause.

Status page visibility is a separate choice. Maintenance can show as maintenance on a status page, which is honest for external readers. Pausing a Healthchecks check simply stops the late alert; it does not produce a friendly public banner unless you build that elsewhere.

Edge cases that burn solo operators

Dependent monitors. You maintain the app and leave the dependency monitor live. The app is “supposed” to be down; the database monitor still screams. Either maintain both or accept that dependency noise is part of the night.

Push monitors in Kuma. If you use Kuma push monitors alongside Healthchecks, you now have two heartbeat systems. Pause both or pick one product for cron and delete the duplicate. Dual heartbeats double the ways to forget a pause.

Calendar bots and “quiet hours.” Phone quiet hours are not a substitute for tool-level maintenance. They hide pages you needed and fail when a family member’s shared calendar changes Do Not Disturb behavior.

DST and travel. If you schedule maintenance from a laptop that just changed timezones, verify the server-local interpretation. Screenshot the window times. Boring, effective.

A worked example: Saturday NAS upgrade

You plan to update TrueNAS, reboot, and let Immich and Nextcloud come back when the pools are healthy. The honest inventory might look like this:

  • Kuma: HTTP monitors for Immich, Nextcloud, the NAS UI, and maybe a SMB TCP check.
  • Healthchecks: nightly snapshot script, weekly offsite sync, a SMART report cron.
  • Notifications: ntfy on the phone, email for Healthchecks, Discord for Kuma because you set it up on a whim.

If you only open a Kuma maintenance window for Immich and Nextcloud, the NAS UI monitor and the SMB check still page. If you maintain every Kuma monitor on that host but leave Healthchecks alone, the snapshot ping goes late while the pool is unavailable and Healthchecks correctly calls that late. Grace will not save you unless the job still runs and merely finishes a little after the usual deadline—which it will not, if the job cannot start.

The fix for that night is boring: Kuma maintenance on the whole host group, pause the Healthchecks that cannot run, leave alone any checks for services on other machines, then reverse the pauses after one green cycle. Write those four lines in the same note as the TrueNAS version you are installing. Checklists beat memory after midnight.

What “success” looks like the next morning

Success is not zero notifications forever. Success is zero unexpected notifications during the window, plus at least one expected signal if something you did not pause stays broken. If everything is silent and Immich is still down at brunch, you over-muted. If you got three pages from monitors you intentionally left live because they sit on a different VLAN, that is information, not failure.

Keep a short postmortem habit even for planned work: which alert fired, which tool, which monitor name. After three upgrades you will see the same forgotten monitor twice. That is your cue to fix grouping—not to widen grace globally.

Decision rules

Reach for Kuma maintenance when pull checks will go red on purpose for a bounded time.

Reach for Healthchecks grace when jobs will still run but may finish late.

Reach for pause / disable when the job or host will not run at all during the work.

If you only change one of those and still get paged, you did not fail at “alerting philosophy.” You failed at inventory. Keep a short list of notifiers next to your upgrade checklist. Planned downtime that still pages you is almost always a missing line on that list—not a bug in the product you remembered to configure.

More articles for you