At-least-once delivery without paging anyone twice
August 2026 · Escalade
Every monitoring system retries. Alertmanager re-sends a firing alert on a fixed interval, and any webhook client that times out will try again without knowing whether the first request landed. So an alerting service is guaranteed to receive duplicates. The only real question is where you absorb them, and the wrong answer is five pages for one outage.
Two different duplicates
These are not the same problem and they do not have the same fix. Conflating them is the common mistake. A dedup key on the incident does nothing for a duplicated send, and per-send idempotency does nothing for a monitor that reports the same outage twice.
Duplicate incidents are the same outage reported more than once, from a retried or repeated webhook. Duplicate deliveries are one incident and one escalation step, where the send is attempted more than once because a worker died mid-flight or the channel call timed out ambiguously.
Absorbing duplicate incidents
The caller supplies a dedupKey. There is a fast pre-check that looks for an
existing open incident, but the pre-check is an optimisation, not the guarantee. The
guarantee is in the schema:
CREATE UNIQUE INDEX ux_incident_open_dedup
ON incident (org_id, dedup_key) WHERE status = 'OPEN';
A partial unique index, scoped to open incidents. Two properties fall out of that, and both are deliberate.
It is race-safe by construction. If two requests both slip past the
pre-check, one INSERT wins and the other raises a constraint violation, which
the service catches and turns into "here is the existing incident," HTTP 200 rather than
201. The database arbitrates. A SELECT-then-INSERT only hopes.
It covers only OPEN. The same key is free again once the
incident closes, so a disk filling up again next Tuesday is a new page rather than a
silently swallowed one. Scoping the index to the open state is what keeps deduplication
from quietly becoming suppression.
The index is on (org_id, dedup_key), so two tenants can use the same key
without colliding. Multi-tenant isolation lives at the query layer, not in the UI.
The Alertmanager receiver reuses this as a correlation key. Each alert's
fingerprint becomes the dedup key, so a firing alert opens an
incident and the matching resolved alert auto-closes it. Alerts that clear on
their own never need a human at all, and Escalade follows the monitor's own view of the
world instead of maintaining a second one.
Absorbing duplicate deliveries
A failed send retries with exponential backoff, 30s then 60s then 120s, capped, up to
max-attempts. Retrying at all means accepting that an ambiguous timeout,
where the send succeeded but the response was lost, will occasionally produce a real
duplicate message.
So what this guarantees is at-least-once, not exactly-once. Exactly-once is not achievable here and I will not claim it. The send is an external side effect, and a crash after the HTTP call but before the commit will re-send on recovery. You cannot make Slack or an email provider deduplicate for you.
What I can do is handle duplication where the system actually owns the boundary, which is
the caller-facing dedup on dedup_key, and accept that a page might rarely
repeat. For an alerting system that is the safe direction to fail. A stronger version
would pass an idempotency token per attempt to the transport, but that only works if the
transport honours one, and the two that matter here do not.
When retries run out
After the attempt cap, the delivery becomes a dead_letter row with a reason.
An explicit, queryable failure state rather than a swallowed exception. Two decisions
matter more than the mechanism.
A dead channel does not halt the ladder. If Slack is down, escalation continues to the next step and the email step still fires. Stopping the ladder on a failed channel would let one outage silently prevent anyone from being paged, which is the exact failure a multi-step policy exists to prevent.
"Nobody responded" is a state, not silence. There are two different "dead" concepts here, and keeping them apart is the whole point:
| Record | Means |
|---|---|
dead_letter row |
One delivery gave up, after exhausting its retries on that channel. |
DEAD_LETTERED incident |
The whole policy ran out of steps and nobody acknowledged. |
The second one is not terminal. Only RESOLVED is. That distinction is
deliberate: automated paging giving up is not the same as the incident being handled, so a
human who arrives late can still acknowledge it. Which is also why a sweeper sets the
status after a grace period rather than the instant the last step fires. Marking it
immediately would make the incident terminal while the on-call is still
reading the page, then reject their acknowledgement seconds later.
Why this is the whole product
An alerting system that silently drops a page is worse than no alerting system, because the team has stopped watching manually and believes it is covered. The partial index, the retry cap, the dead-letter row, the sweeper, the ladder continuing past a dead channel: all of it exists to make failure loud and queryable rather than absent. The CRUD around it is the easy part.
Escalade is on
GitHub. Retry and
dead-letter behaviour lives in EscalationStepProcessor, the sweeper in
DeadLetterSweeper, and the dedup index in V1__init.sql.