Why I claim jobs with SKIP LOCKED instead of a message broker
August 2026 · Escalade
Escalade delivers a page at a specific time, from any number of app instances, and must never deliver the same one twice. The obvious answer is a message broker. I used a Postgres table and one line of SQL instead. I think that was right, but only because of properties this particular workload has, and I want to be specific about which ones.
The work
Every escalation step becomes a notification_attempt row with a
due_at. A worker wakes up, finds the attempts that are due, delivers them,
and schedules the next step. So the whole scheduling problem reduces to finding rows where
due_at <= now(). Databases are already very good at that.
I did reach for a broker first, because "background job" pattern-matches to one. What stopped me was writing the requirements down. A few hundred incidents. One datastore. No ordering guarantees needed across incidents. No second consumer wanting the same events. A broker would have been another process to run, deploy, secure and monitor, and it would have bought me none of the three things brokers are actually good at.
The claim query
SELECT * FROM notification_attempt
WHERE status = 'PENDING' AND due_at <= now()
ORDER BY due_at
LIMIT 1
FOR UPDATE SKIP LOCKED
FOR UPDATE locks the row this transaction reads and holds it until commit.
SKIP LOCKED is the important half. Instead of blocking on a row another
transaction already holds, the scan steps over it and takes the next one. Several workers
running this query at the same instant each walk away with a disjoint set. No
coordination, no leader election, no broker.
It helps to be able to say what happens without it. Plain FOR UPDATE does not
skip; it waits. Every worker queues behind whichever one got there first, so the pool
serialises into a single consumer with extra steps, and each waiter then wakes to re-read
a row whose status changed underneath it. Removing that failure mode is what turns a table
into a real competing-consumers queue.
What this buys
Nothing to run, deploy, secure or monitor beyond the database that was already there. For something meant to be self-hosted by whoever adopts it, that counts for more than usual.
The claim and the state change commit together. With a broker they live in different systems, so a crash between acking the message and committing the row leaves the two disagreeing about what happened. That entire class of bug is absent here. Correctness comes from the database rather than from a consensus protocol.
The queue is also queryable. "What is about to page, and when" is a SELECT
against a table, not a request to an admin UI. And scaling out needs no coordination
infrastructure at all: run more instances, and they cooperate because the database says
so.
What it costs
Four things, and I would rather name them than have them found.
Latency is the poll interval. The worker is a @Scheduled
loop rather than a push subscription, so time-to-first-page is bounded below by how often
it ticks. Fine for human-timescale paging. Not fine for a sub-second SLA.
The send happens inside the lock-holding transaction. More on that below. It is the sharpest edge in the design.
Every worker's poll is load on the primary. Free at this size. Not free at every size.
There is no jitter across instances. Workers tick on the same fixed delay, so the poll gets a small thundering herd. Harmless with a handful of instances, and easy to fix. I have not fixed it.
The cost I did not see at first
The HTTP call to Slack or the email provider runs inside the transaction holding the attempt's row lock. I first filed the connect and read timeouts under tuning. They are not tuning. They are correctness. An endpoint that accepts a connection and then hangs would hold that row lock for as long as the socket stays open, stalling that incident's whole escalation ladder behind it. Bounded timeouts, 3s connect and 5s read, turn the worst case into a retryable failure instead of an indefinite stall. That makes them a correct minimum, not a performance knob.
The proper fix, if this needed to scale out, is to stop holding a lock across a network call. Claim the row, commit an in-flight marker, deliver outside the lock, then confirm, with a visibility timeout and a reaper to recover attempts whose worker died mid-flight. It is a known shape. I also know why I have not built it: at this scale the bounded timeout is sufficient, and the reaper is a second failure mode to get right.
When I would switch
Not taste. Three specific signals. Throughput past what a single Postgres can serve. A need for strict per-key ordering. Or a second consumer that wants the same events, which is fan-out, and a queue table is a poor fan-out primitive. Any of those and claiming moves to Kafka or SQS.
I would rather watch claim contention and poll latency in the metrics than guess the threshold up front. The counters are already there. The point of having them is to let the system tell me when the assumption expires.
Escalade is on
GitHub. The claim query lives
in NotificationAttemptRepository; the worker loop is
EscalationEngine and NotificationWorkerScheduler. The race this
design creates with acknowledgements is
its own post.