← Akshit Khandelwal

Why I claim jobs with SKIP LOCKED instead of a message broker

Escalade delivers a page at a specific time, from any number of app instances, and must never deliver the same one twice. The obvious answer is a message broker. I used a Postgres table and one line of SQL instead. I think that was right, but only because of properties this particular workload has, and I want to be specific about which ones.

The work

Every escalation step becomes a notification_attempt row with a due_at. A worker wakes up, finds the attempts that are due, delivers them, and schedules the next step. So the whole scheduling problem reduces to finding rows where due_at <= now(). Databases are already very good at that.

I did reach for a broker first, because "background job" pattern-matches to one. What stopped me was writing the requirements down. A few hundred incidents. One datastore. No ordering guarantees needed across incidents. No second consumer wanting the same events. A broker would have been another process to run, deploy, secure and monitor, and it would have bought me none of the three things brokers are actually good at.

The claim query

SELECT * FROM notification_attempt
WHERE status = 'PENDING' AND due_at <= now()
ORDER BY due_at
LIMIT 1
FOR UPDATE SKIP LOCKED

FOR UPDATE locks the row this transaction reads and holds it until commit. SKIP LOCKED is the important half. Instead of blocking on a row another transaction already holds, the scan steps over it and takes the next one. Several workers running this query at the same instant each walk away with a disjoint set. No coordination, no leader election, no broker.

It helps to be able to say what happens without it. Plain FOR UPDATE does not skip; it waits. Every worker queues behind whichever one got there first, so the pool serialises into a single consumer with extra steps, and each waiter then wakes to re-read a row whose status changed underneath it. Removing that failure mode is what turns a table into a real competing-consumers queue.

What this buys

Nothing to run, deploy, secure or monitor beyond the database that was already there. For something meant to be self-hosted by whoever adopts it, that counts for more than usual.

The claim and the state change commit together. With a broker they live in different systems, so a crash between acking the message and committing the row leaves the two disagreeing about what happened. That entire class of bug is absent here. Correctness comes from the database rather than from a consensus protocol.

The queue is also queryable. "What is about to page, and when" is a SELECT against a table, not a request to an admin UI. And scaling out needs no coordination infrastructure at all: run more instances, and they cooperate because the database says so.

What it costs

Four things, and I would rather name them than have them found.

Latency is the poll interval. The worker is a @Scheduled loop rather than a push subscription, so time-to-first-page is bounded below by how often it ticks. Fine for human-timescale paging. Not fine for a sub-second SLA.

The send happens inside the lock-holding transaction. More on that below. It is the sharpest edge in the design.

Every worker's poll is load on the primary. Free at this size. Not free at every size.

There is no jitter across instances. Workers tick on the same fixed delay, so the poll gets a small thundering herd. Harmless with a handful of instances, and easy to fix. I have not fixed it.

The cost I did not see at first

The HTTP call to Slack or the email provider runs inside the transaction holding the attempt's row lock. I first filed the connect and read timeouts under tuning. They are not tuning. They are correctness. An endpoint that accepts a connection and then hangs would hold that row lock for as long as the socket stays open, stalling that incident's whole escalation ladder behind it. Bounded timeouts, 3s connect and 5s read, turn the worst case into a retryable failure instead of an indefinite stall. That makes them a correct minimum, not a performance knob.

The proper fix, if this needed to scale out, is to stop holding a lock across a network call. Claim the row, commit an in-flight marker, deliver outside the lock, then confirm, with a visibility timeout and a reaper to recover attempts whose worker died mid-flight. It is a known shape. I also know why I have not built it: at this scale the bounded timeout is sufficient, and the reaper is a second failure mode to get right.

When I would switch

Not taste. Three specific signals. Throughput past what a single Postgres can serve. A need for strict per-key ordering. Or a second consumer that wants the same events, which is fan-out, and a queue table is a poor fan-out primitive. Any of those and claiming moves to Kafka or SQS.

I would rather watch claim contention and poll latency in the metrics than guess the threshold up front. The counters are already there. The point of having them is to let the system tell me when the assumption expires.


Escalade is on GitHub. The claim query lives in NotificationAttemptRepository; the worker loop is EscalationEngine and NotificationWorkerScheduler. The race this design creates with acknowledgements is its own post.