Alert & incident flow
Exactly what HowlOps does when a monitor fails, an external alert arrives, or a heartbeat is missed, from the first signal through notification, escalation, and recovery.
On this page
- The big picture
- Alerts and incidents are two levels
- Step 1: where alerts come from
- Monitor goes down
- External alert arrives
- Heartbeat is missed
- One alert per incident (deduplication)
- Step 2: screenshots (HTTP monitors)
- Step 3: how the alert is delivered
- Configure a routing rule
- A policy governs the alert
- No policy governs the alert
- Step 4: escalation steps and your personal chain
- Step 5: routing labels
- Step 6: recovery and resolve
- Acknowledge, silence, unacknowledge
- What you can configure (summary)
For severities, manual incidents, incident notes, correlation, and status-page updates, see Incidents & severities.
The big picture
The source-specific checks differ, but every accepted signal enters the same routing, paging, response, and recovery pipeline. Select a source to compare the paths.
See what happens after a signal arrives.
Choose a source. The first checks change, but routing, paging, response, and recovery stay connected.
Signal
A probe reports that the configured target failed.
Confirm
Retries, thresholds, and optional regional rules confirm the condition before opening an alert.
Route
Routing rules select a policy; without one, enabled channels receive the alert.
Reach
The policy pages its current target and continues only while the alert is unacknowledged.
Respond
Acknowledge, coordinate, and publish updates. Recovery resolves the alert and sends the all-clear.
The rest of this page owns the detailed behavior. Setup guides link here instead of copying the rules into several places.
Alerts and incidents are two levels
HowlOps separates a raw alert from a confirmed incident:
| Level | What it is | Where you see it |
|---|---|---|
| Alert | A signal that just fired. Most noise lives here. | The Alerts list |
| Incident | An alert that matters enough to track and coordinate. | The Incidents list |
An alert becomes an incident in one of two ways:
- Automatically, when it matches a promotion rule. Out of the box there is one rule: any alert with critical severity is promoted. You can add your own rules that match on severity, source (monitor, heartbeat, integration), or labels.
- Manually, with the Promote to incident button on any alert.
You are paged for the signal itself, not for the promotion. Notifications and escalation start the moment the alert is created, whether or not it ever becomes an incident. Promotion is about triage and tracking, not about whether you get woken up.
Step 1: where alerts come from
Monitor goes down
A single failed check does nothing on its own. A detector evaluates each completed check and opens an alert only when the configured confirmation rule is satisfied:
| Setting | Effect |
|---|---|
| Default | 2 consecutive failed checks |
| Retry on fail | The prober re-checks itself, so 1 confirmed failure is enough |
| Confirmation threshold | A fixed count you set (for example 3) that overrides the above |
For monitors that run from several regions you also get:
| Setting | Effect |
|---|---|
| Deduplicated (default) | One alert per monitor. Failing regions accumulate. Closes when all regions recover. |
| Per region | A separate alert per failing region, plus a global one when everything is down. |
| Region quorum | Only page when at least X% of regions are down. Below that it counts as partial and stays quiet. |
External alert arrives
Alerts can come from Prometheus Alertmanager, Grafana, Datadog, CloudWatch, or a generic webhook. HowlOps normalises the payload to a name, a state, a severity, and a fingerprint (a stable key the source provides). Deduplication is keyed on that fingerprint:
- A new fingerprint opens a new alert.
- A repeat of the same fingerprint updates it and bumps a "fired N times" counter. It does not open a duplicate or re-notify.
External alerts have no monitor of their own, so there is no screenshot, prober log, or recheck for them.
Heartbeat is missed
A heartbeat expects a ping on a schedule. It is marked down when:
- the last ping is older than
interval + grace, or - it was never pinged and is older than twice its interval, or
- it missed its cron-scheduled time plus grace.
The alert is titled with the heartbeat's name and carries its configured priority (default P1). It is idempotent. While one alert is open, missing again does not open another.
One alert per incident (deduplication)
Whatever the source, you are notified once when something breaks and once when it recovers, never on every failed check. Subsequent failures for the same open alert are deduplicated:
14:00 check fails → alert sent (incident opens)
14:10 check fails again → no alert (same incident)
14:20 check fails again → no alert
14:30 check succeeds → recovery alert sent (incident resolves)
You can set a per-monitor repeat alert interval (see Step 4). While an alert stays open and unacknowledged, HowlOps sends another notification after that interval and continues until the alert is acknowledged, silenced, resolved, or deleted.
Step 2: screenshots (HTTP monitors)
If a monitor is HTTP, the workspace has screenshot_monitoring, and you enabled Capture screenshots on failure, HowlOps takes a screenshot of the page after the alert opens. The setting is off by default.
This is intentional. The screenshot lands a few seconds in, so it captures the page during an outage that actually lasts. For a sub-second blip it may be empty or late, which is expected, the screenshot is meant for incidents that persist long enough to investigate.
Step 3: how the alert is delivered
The first decision is whether an escalation policy governs the alert. HowlOps can pick one for you automatically:
How HowlOps chooses between escalation and broadcast
Configure a routing rule
A routing rule is a condition + target pair. When a new alert is created, HowlOps evaluates your rules in ascending priority order (1 → 2 → 3 …). The first matching rule wins and pins its escalation policy; the rest are skipped. If no rule matches, the alert falls back to the monitor's own policy, then the workspace-level default escalation policy.
To create one, go to Alerts → Routing Rules → New Rule, then set a condition, a match value, the escalation policy to apply, and a priority (lower number = evaluated first).
| Condition type | Example value | Matches when… |
|---|---|---|
monitor_tag | payments | The monitor carries the specified tag |
monitor_type | http, heartbeat | The monitor is of the specified type |
Rules are evaluated from the lowest priority number upward. The first match selects its policy and stops evaluation; when nothing matches, the monitor policy and then the workspace default are used.
Combine routing rules with monitor tags for fine-grained triage. For example, route every production-tagged monitor to a stricter escalation policy while staging monitors only post to Slack.
A policy governs the alert
The immediate broadcast is skipped (so you are not notified twice). The escalation engine takes over and pages on-call step by step over time (see Step 4).
No policy governs the alert
The alert is broadcast to every enabled channel plus mobile push. Each channel still passes through a set of filters and can be skipped:
| Filter | What it does |
|---|---|
| Workspace event matrix | Per-workspace mapping from event types to channel types |
| Per-channel events | On a single channel, pick which events it receives. Empty means all. |
| Personal preferences | A channel's owner can opt out of specific events |
| Quiet hours | Non-critical events are suppressed during your quiet window (chat and SMS are dropped; email is held for the daily digest). Critical always pages. |
| Deduplication | The same alert is not repeated to the same channel within a window (default 5 minutes) |
Quiet hours only affect non-critical alerts, a critical alert always pages, day or night (external alerts and missed heartbeats follow the same rule). During your quiet window, non-critical chat and SMS notifications are dropped and non-critical email is held for the daily digest. The per-monitor Respect quiet hours flag (off by default) additionally lets you hold that monitor's entire dispatch during the window. See Noise control for the full per-channel behaviour.
Step 4: escalation steps and your personal chain
When a policy governs the alert, two layers of timing apply, and both are yours to configure.
Escalation steps run in order. Each step has its own delay, measured from when the previous step fired, and its own target. The policy can also repeat from step 1 if nobody acknowledges (see Escalation policies). Your personal chain, alongside it, decides how you are reached: each channel has a priority and a fallback delay, so a single page spreads across channels over time.
Page the on-call engineer
Page the secondary
delay 15 min after step 1Notify the team lead
delay 30 min after step 2Push
ImmediatelySMS
After 10 minVoice
After 15 minExample: a 3-step escalation policy (delays 0 / 15 / 30 min) running alongside a 4-channel personal chain. Both stop the moment you acknowledge.
The chain advances only while the alert is still unacknowledged and not silenced. If you acknowledge at minute 3, the SMS and voice steps never fire.
You can also set a per-monitor renotify interval to re-send an alert every N seconds while it stays unacknowledged.
Step 5: routing labels
Routing rules match on labels. Where those labels come from depends on the source:
- Monitors and heartbeats: HowlOps sets the labels from your configuration (the monitor type, its tags, the severity).
- External integrations: the labels come from the incoming payload. HowlOps reads the labels and tags the source sent and routes on those, so the quality of your routing depends on the source sending good labels.
Step 6: recovery and resolve
When the service recovers, the alert resolves automatically:
- A monitor that returns healthy auto-resolves (multi-region waits for every region).
- An external source that sends a
resolvedevent closes the matching alert by fingerprint. - A heartbeat that pings again recovers and closes its alert.
The all-clear travels the same path the alert did:
- If a policy paged on-call, only the people who were actually paged get the all-clear.
- If it was a broadcast, the all-clear goes to the channels that received the original alert.
Recovery notifications are sent automatically (opt-out per channel). If you only want to hear about problems, turn off the recovery (up) event in Settings, Notifications, or use the "Incidents only, no recovery noise" preset. The alert still disappears from your open list either way.
Acknowledge, silence, unacknowledge
14:32
Monitor DOWN, Production API (status 500)
14:32
Paged on-call via push
14:34
Acknowledged by Sarah, escalation pauses
14:39
Recovered, 7 min, all-clear sent to Sarah
| Action | What happens |
|---|---|
| Acknowledge | Escalation stops advancing. You signal you are on it. |
| Silence (N minutes) | A pause. Escalation and your personal chain both hold, then resume when the window ends. |
| Unacknowledge | Re-arms escalation and your personal chain immediately. |
| Close | Manually closes the alert or incident. |
What you can configure (summary)
| Where | Setting | Effect |
|---|---|---|
| Monitor | Confirmation threshold / retry on fail | How many failures before an alert opens |
| Monitor | Alert dedup mode, region quorum | Multi-region behaviour |
| Monitor | Default severity | Critical, warning, or info, drives priority and quiet hours |
| Monitor | Respect quiet hours | Whether non-critical alerts are held at night |
| Monitor | Capture screenshots on failure | HTTP screenshot on outage |
| Monitor | Renotify interval | Repeat an unacknowledged alert |
| Heartbeat | Interval, grace, cron, priority | When a missed ping becomes an alert |
| Workspace | Promotion rules | Which alerts auto-become incidents |
| Workspace | Alert routing rules | Which policy (or broadcast) handles an alert |
| Workspace | Escalation policies | Steps, delays, targets, repeat |
| Workspace | Quiet hours window | When non-critical alerts are held |
| You | Personal notification chain | Channel order and fallback delays |
| You | Notification preferences | Which events reach you, including recovery |
Was this page helpful?