Scaling a background worker often starts with a deceptively simple problem.
You have a database table containing work that needs to be processed.
One application instance polls the table, claims some rows, processes them, and moves on.
Then you deploy a second instance.
Then five.
Now every worker is looking at the same rows.
OUTBOX
│
┌───────────┼───────────┐
│ │ │
Pod A Pod B Pod C
│ │ │
└───────────┼───────────┘
│
Same pending rows
At this point, reaching for a distributed coordination mechanism can feel inevitable.
Maybe Redis.
Maybe a distributed lock.
Maybe another service whose job is to decide which worker owns which work.
But there is another question worth asking first:
If the work already lives in your relational database, does the database already have enough information to coordinate the workers?
For many queue-like workloads, PostgreSQL and MySQL provide a remarkably useful answer:
SELECT ... FOR UPDATE SKIP LOCKED
It is only a few words of SQL.
But those few words can fundamentally change how concurrent workers cooperate.
Start with One Worker
Imagine an Outbox table containing events waiting to be published:
+----+----------------------+---------+
| ID | EVENT_TYPE | STATUS |
+----+----------------------+---------+
| 1 | PaymentCreated | PENDING |
| 2 | CustomerUpdated | PENDING |
| 3 | PaymentAuthorized | PENDING |
| 4 | OrderCompleted | PENDING |
+----+----------------------+---------+
With one worker, processing this is straightforward.
Worker
│
▼
SELECT pending events
│
▼
Process events
│
▼
Update state
There is nobody else competing for the rows.
Now deploy three workers.
Worker A ─┐
Worker B ─┼──► SELECT pending events
Worker C ─┘
If all three execute the same ordinary query at roughly the same time, they may all observe the same pending work.
We need a way for a worker to say:
These rows are mine for this transaction. Other workers should find something else to do.
FOR UPDATE Solves Only Half the Problem
A locking read gives us part of the answer:
SELECT id, event_type, payload
FROM outbox_event
WHERE status = 'PENDING'
ORDER BY id
LIMIT 100
FOR UPDATE;
The selected rows are locked for the transaction.
That protects them from being concurrently claimed in the same way by another transaction.
But suppose Worker A locks the first 100 rows.
Worker B arrives immediately afterward and wants work.
With ordinary locking behavior, Worker B can end up waiting for rows Worker A already owns.
Worker A ─────► [E1 E2 E3 E4] 🔒
│
Worker B ────────────────┘
WAIT
There may be thousands of other events available.
Worker B does not necessarily need those particular rows.
It just needs some work.
That distinction is exactly where SKIP LOCKED becomes useful.
Skip Work Another Worker Already Owns
Change the query to:
SELECT id, event_type, payload
FROM outbox_event
WHERE status = 'PENDING'
ORDER BY id
LIMIT 100
FOR UPDATE SKIP LOCKED;
Now the behavior changes.
If Worker A already holds locks on some qualifying rows, Worker B does not wait for those row locks.
It skips them and continues looking for other eligible rows.
OUTBOX
E1 E2 E3 E4 E5 E6
🔒 🔒 🔒 🔒
Worker A ─┘ │ │ │
│ │ │
Worker B ───── SKIP LOCKED ───► E5 E6
Add another worker and the table can naturally distribute available work:
PostgreSQL
┌─────────────┐
│ OUTBOX │
│ │
│ E1 E2 E3 E4 │
│ E5 E6 E7 E8 │
└──────┬──────┘
│
FOR UPDATE
SKIP LOCKED
│
┌────────────┼────────────┐
▼ ▼ ▼
Pod A Pod B Pod C
E1, E2 E3, E4 E5, E6
No worker needs a global lock around the entire queue.
Workers coordinate through the rows representing the work itself.
Why SKIP LOCKED Is Different from an Ordinary Query
There is an important caveat.
SKIP LOCKED intentionally does not give you a normal, globally consistent view of all matching rows.
Some rows may disappear from one worker's result because another transaction currently owns their locks.
That would be undesirable for many normal business queries.
Imagine querying account balances and silently ignoring locked accounts.
Obviously wrong.
But a worker queue has different semantics.
The worker isn't asking:
Show me the complete truth about every pending row.
It is asking:
Give me some currently available work that I can safely claim.
That is why queue-like tables are such a natural fit for this behavior.
This Is Exactly the Problem an Outbox Dispatcher Has
This matters directly to NERV Event.
The Transactional Outbox Pattern gives us durable events.
A business transaction stores both:
Business State
+
Outbox Event
│
▼
COMMIT
After commit, a dispatcher needs to publish those events.
In production, however, we rarely want one application instance to become the permanent global dispatcher.
Spring Boot services run across multiple pods.
Pod A ─┐
Pod B ─┤
Pod C ─┼──► Shared Outbox Table
Pod D ─┤
Pod E ─┘
The challenge becomes:
How can every pod participate in dispatching without every pod processing the same Outbox rows?
Database row claiming gives us an elegant answer.
Each dispatcher transaction claims a bounded batch.
Other dispatchers skip the claimed rows and continue finding work.
The database becomes both:
- the durable store for the Outbox event, and
- the concurrency boundary for claiming it.
That second property deserves more attention than it usually gets.
Why Not Just Put a Distributed Lock in Redis?
We could.
For example:
Worker
│
▼
Acquire Redis Lock
│
▼
Read PostgreSQL
│
▼
Process Work
│
▼
Release Redis Lock
There are workloads where Redis is exactly the right technology.
But adding Redis solely for coordination changes the architecture.
Before:
Worker
│
▼
PostgreSQL
│
├── Work
└── Ownership
After:
┌────► Redis
│ │
Worker ──────┤ Coordination
│
└────► PostgreSQL
│
Work
We now operate another distributed system.
It has its own availability, network communication, failure behavior, timeouts, monitoring, deployment, and consistency questions.
That does not make Redis a bad choice.
It means Redis should solve a problem we actually have.
Infrastructure should earn its place in the architecture.
The Database Has One Particularly Useful Property
Suppose a worker starts a transaction and claims some rows.
BEGIN
SELECT ...
FOR UPDATE SKIP LOCKED;
Then the worker crashes before committing.
The database connection disappears.
The transaction rolls back.
The row locks are released.
Another worker can claim the rows.
Worker A
│
├── BEGIN
├── Lock E1, E2
│
X CRASH
│
▼
ROLLBACK
│
▼
Locks released
│
▼
Worker B can claim E1, E2
Ownership is tied to the transaction.
There is no separate lease record that must independently agree with the transaction's outcome.
But this leads to an equally important warning.
Do Not Hold the Transaction Open While Doing Slow Work
It is easy to misuse this pattern.
Imagine claiming 100 Outbox events and then keeping the transaction open while performing 100 slow network calls.
BEGIN
│
├── Claim 100 rows
│
├── Call Kafka
│
├── Wait
│
├── Call Kafka
│
├── Wait
│
├── ...
│
└── 30 seconds later
COMMIT
Now those row locks and the database connection remain tied to the processing lifetime.
The same principle from long-running workflow design applies here:
Keep database transactions short.
The exact claim-and-dispatch strategy depends on the delivery guarantees of the system, but transaction duration must be deliberate.
SKIP LOCKED is not permission to turn a database transaction into a worker lease that stays open indefinitely.
SKIP LOCKED Does Not Give You Exactly-Once Processing
This is another important distinction.
Suppose an Outbox worker publishes an event successfully.
Then it crashes before recording the successful publication.
Publish to Kafka
│
▼
SUCCESS
│
X
CRASH
│
▼
Database still says PENDING
After recovery, another worker may publish the event again.
The row lock prevented two workers from owning the same row simultaneously.
It did not eliminate the distributed failure window between the database and the broker.
That is why NERV Event still relies on the broader reliability model:
Transactional Outbox
+
At-Least-Once Delivery
+
Idempotent Consumer
+
Inbox
Locking solves concurrent ownership.
It does not magically solve every distributed-systems problem around delivery.
SKIP LOCKED Does Not Automatically Preserve Ordering
Suppose these events belong to the same aggregate:
PaymentCreated
↓
PaymentAuthorized
↓
PaymentCaptured
Now imagine several workers independently claiming rows.
Simply using:
ORDER BY created_at
FOR UPDATE SKIP LOCKED
does not by itself establish a complete business-level ordering guarantee across concurrent transactions.
One worker may have locked an earlier event while another worker sees a later eligible row.
If strict ordering matters, the architecture needs an explicit ordering strategy.
In NERV Event, this is why concepts such as an orderingKey matter.
payment-123
│
├── Event A
├── Event B
└── Event C
payment-456
│
├── Event X
├── Event Y
└── Event Z
Ideally:
payment-123 → A → B → C
payment-456 → X → Y → Z
We preserve ordering within the required boundary while still allowing unrelated keys to progress concurrently.
This is a useful reminder:
Row claiming and event ordering are separate problems.
Indexes Matter More Than the Clever SQL
Once this pattern is running at meaningful scale, the query needs to find available work efficiently.
A query like:
SELECT id
FROM outbox_event
WHERE status = 'PENDING'
AND next_attempt_at <= CURRENT_TIMESTAMP
ORDER BY created_at, id
LIMIT 100
FOR UPDATE SKIP LOCKED;
needs indexes that support the way work is discovered.
Otherwise the database may spend significantly more time scanning rows than actually claiming them.
The correct index depends on the schema and workload, but fields involved in eligibility and ordering deserve deliberate attention.
The lesson is simple:
SKIP LOCKED makes contention behavior better. It does not make a bad query cheap.
Batch Size Is Also an Architectural Decision
Why not claim 10,000 rows at once?
Because larger batches affect more than throughput.
They can increase:
- transaction duration,
- lock lifetime,
- connection hold time,
- recovery work after failure,
- memory consumption, and
- the amount of work monopolized by one worker.
Too small a batch, however, increases database round trips.
There is no universal magic number.
Batch size should be treated as a throughput-versus-fairness-versus-recovery tradeoff and measured under realistic load.
Shopify Recently Faced a Much Larger Version of This Problem
This pattern is not limited to background job tables.
In 2026, Shopify Engineering published an interesting account of redesigning its inventory reservation system.
The previous reservation model used Redis, while the authoritative inventory ledger lived in MySQL.
That meant reservation and inventory state crossed two different systems.
Shopify moved inventory reservations into MySQL so reservation and inventory changes could share an ACID consistency boundary.
A key part of the new design used MySQL 8's SKIP LOCKED behavior so concurrent reservation transactions could skip inventory units already claimed by other transactions instead of waiting on them.
But the interesting part of the story is not simply:
Shopify replaced Redis with MySQL.
That would be an oversimplification.
The interesting lesson is that the team reconsidered whether the extra distributed-system boundary was still necessary for this particular workload.
And making the database approach work required much more than adding SKIP LOCKED.
They also had to reason about:
- primary-key and index design,
- transaction isolation,
- gap locking,
- consistent lock ordering,
- batching,
- database connection consumption, and
- production observability.
One of their most interesting findings was that the eventual throughput ceiling was not simply query execution or database CPU.
Database connection usage elsewhere in the checkout path was part of the real bottleneck.
That is a valuable reminder for any worker architecture:
The database connection is also a finite resource.
Optimizing row locking while ignoring connection hold time can simply move the bottleneck somewhere else.
So Should We Stop Using Redis?
No.
That would be the wrong conclusion.
Redis remains extremely useful for workloads such as:
- high-speed caching,
- ephemeral state,
- TTL-heavy data,
- counters,
- rate limiting,
- specialized data structures, and
- workloads that genuinely benefit from coordination outside the transactional database.
The argument is not:
Never use Redis.
The argument is:
Do not introduce Redis merely because concurrent coordination sounds like something that must require Redis.
First understand what the database can already guarantee.
When Database Coordination Is a Good Fit
FOR UPDATE SKIP LOCKED is particularly attractive when:
- the work already lives in the relational database,
- workers need to claim independent rows,
- the workload is naturally queue-like,
- transactions can remain short,
- the database has enough connection and I/O capacity, and
- keeping state and ownership within one transactional system simplifies correctness.
Outbox and Inbox processing fit these characteristics surprisingly well.
When You Should Look Beyond the Database
There is also a point where the database should not become your answer to every coordination problem.
You should reconsider the architecture when:
- worker throughput begins competing with critical transactional workloads,
- connection consumption becomes unacceptable,
- work distribution requires semantics poorly represented by rows and transactions,
- extremely low-latency ephemeral coordination dominates the workload,
- database contention becomes the limiting factor, or
- a dedicated queue or coordination system better matches the operational requirements.
The goal is not database minimalism.
The goal is deliberate architecture.
The Bigger Lesson
What I like most about SKIP LOCKED is not the SQL syntax.
It is the architectural question it forces us to ask.
Suppose the work is already durable in PostgreSQL.
Suppose transactions already define ownership.
Suppose the database already knows which rows another worker has claimed.
Do we really need another system just to coordinate access to those rows?
Sometimes the answer is absolutely yes.
But sometimes:
Database State
+
Transaction
+
Row Lock
+
SKIP LOCKED
is enough.
That is an important lesson from building infrastructure such as NERV Event.
Production architecture does not become better simply because it contains more specialized infrastructure.
Every additional system introduces another operational and failure boundary.
Sometimes the best distributed system is the one you didn't need to add.
About NERV Event
NERV Event is an open-source event-driven infrastructure library for Spring Boot focused on reliable event publishing and processing.
It provides infrastructure for Transactional Outbox and Inbox patterns, retries, idempotency, ordering, multi-instance processing, Kafka and SQS integration, scheduler resilience, and operational visibility.
Explore the project:
NERV Event is part of NERV — Next-Generation Engineering for Runtime Velocity.
References
- Shopify Engineering — We replaced Redis with MySQL for inventory reservations—and it scaled
- PostgreSQL Documentation — SELECT / Locking Clause
- MySQL Reference Manual — InnoDB Locking Reads

Post a Comment