The ack-versus-escalate race, and the lock that closes it
August 2026 · Escalade
Someone acknowledges a page at the exact moment the worker fires the next escalation step.
Both transactions are in flight. Both read an incident that is OPEN, and both
are correct about what they read. One of them has to lose, and the system has to decide
which one. Otherwise the person who just acknowledged gets paged again, or their manager
does.
Why this is not a rare case
It looks like a millisecond-wide window, which is the intuition that gets it ignored. Here is why that intuition is wrong. The entire design of an escalation ladder is that a step fires exactly when a human is most likely to be responding. The delay is tuned to "how long before we assume nobody saw it," so acknowledgements do not spread evenly across the window. They cluster at the step boundaries. The system schedules the collision.
The two writers
The ack path is an API call. It sets status = ACKNOWLEDGED, stamps
acked_at, and cancels attempts that have not gone out. The escalation path is
the worker: it marks the current attempt SENT, advances
current_step, and schedules the next notification_attempt.
Both mutate the same incident row, read-then-write, in separate transactions. Left alone, that is a textbook lost update.
How it actually interleaves
| t | Worker transaction | Ack (API) transaction |
|---|---|---|
| 1 |
Claims step 0 with FOR UPDATE (incident version = N); delivers the
page
|
— |
| 2 | Still delivering; holds the attempt row lock |
Reads the incident (version N), tries to UPDATE, blocks on the
worker's row lock
|
| 3 |
Marks SENT, schedules step 1, current_step++ so version
goes N to N+1, commits and releases the lock
|
Unblocks. Its write expected version N but finds N+1, so optimistic locking fails and it rolls back |
| 4 | Done |
Retries: re-reads at version N+1, still OPEN, sets
ACKNOWLEDGED and cancels the pending step 1, commits (HTTP 200)
|
Why it resolves correctly
The mechanism is a version column on incident, mapped with JPA's
@Version. What makes it the right mechanism is that the two writers are not
equals.
The worker advancing a step is routine bookkeeping. If it loses, it re-reads on the next
tick, finds an incident that is no longer OPEN, and cancels itself instead of
paging. A human acknowledging is the signal that stops the page, so the ack is the one
that retries until it wins, up to three attempts, and it wins against an incident that has
moved one step forward but is still open.
So the race resolves to "no unnecessary page" from either direction. Defence in depth on top of that: if the retries are somehow exhausted, the API returns 409, never a 500. A concurrency race must not surface to a caller as a server error.
The version column makes the race detectable. Retrying the ack makes it correct. Without it you get a 500 and keep paging someone who already responded.
What I ruled out
Pessimistic SELECT … FOR UPDATE on the incident. Correct, and
I do use it on the attempt row, where I explicitly want workers claiming disjoint sets. On
the incident it is wrong. The escalation path holds that row across a channel call, so
every acknowledgement for that incident would serialise behind a slow Slack request,
unconditionally, in the overwhelmingly common case where there is no collision at all.
A status check immediately before sending. This narrows the window without closing it. The check and the send are still not atomic, so it is the kind of fix that makes a bug rarer and much harder to reproduce.
Serializable isolation. Also correct, but it charges every transaction in the service for one contended row.
Optimistic locking pays nothing in the common case and detects the collision exactly when it matters. Rare contention, short transactions, and a safe retry path for both losers is the profile where optimistic beats pessimistic.
What this does not fix
A page that is already in flight when the ack lands still goes out. That is inherent to at-least-once delivery against an external service, and I would rather state it than let someone find it. What the version column closes is the unacceptable half: paging someone after they acknowledged. Step 1 gets cancelled. Step 0, already sent, stays sent.
The bug I shipped first
The worker originally processed a whole batch of due attempts in one transaction. That is wrong, for a reason specific to optimistic locking. A single collision anywhere in the batch rolls the whole transaction back, including attempts that had already been delivered. Their rows revert to pending and the next tick re-sends them. One ack colliding with one step could re-page several people.
The fix was one transaction per attempt, so the blast radius of a rollback is now exactly one page. I mention it because the reasoning generalises: under optimistic concurrency, transaction boundaries are not only a performance choice. They decide how much already-completed side-effecting work a rollback can undo.
Proving it
Testcontainers against real Postgres, which is non-negotiable here.
SKIP LOCKED and optimistic locking cannot be exercised on an in-memory
database, so a test running on H2 would prove nothing about the thing being claimed.
The ack-race test injects a channel that blocks on a latch, so the worker is
deterministically inside delivery when the ack fires. Then it releases the latch and
asserts the final state: incident ACKNOWLEDGED, step 0 SENT,
pending step 1 CANCELLED. The latch is what turns a timing race into a
repeatable test rather than a flaky one. CI provides Docker, so it runs on every push
instead of only on my machine.
Escalade is on
GitHub. The version column is
on Incident; the escalation path is EscalationStepProcessor and
the retry loop is IncidentService.withRetryOnCollision. Why the worker claims
jobs the way it does is a separate post.