What Happens to Outbox When the Database Goes Down?

⏱ 8 min read

Every outbox talk speaks about the broker failing.

I've given this talk enough times to know the slides by heart. The broker goes down. The outbox row is already committed, the worker publishes it later, and nothing is lost. It lands well, because it's true.

Then someone asks the follow-up, and I don't have a slide for it. What happens when the database is the thing that's unavailable? Not RabbitMQ or the broker. The database the outbox table lives in.

This is more than a fair question.

What the outbox is actually for 🔗

The outbox pattern exists because of one specific problem. You need to save data and publish an event. Those are two different systems, talking over a network, and they can't share a transaction.

Do the save first and the publish second. A crash in between leaves an order in your database that nobody downstream heard about. Do the publish first, and a failure leaves an event announcing work that never happened. Both are wrong, and no ordering of the two calls fixes it.

The outbox removes the second system from the critical path. The event becomes a row in the same database, in the same transaction as your domain data. Either both commit or neither does. A separate worker reads that table afterwards and publishes. If the broker is down right then, the row waits.

The inbox pattern points the same idea the other way. Record that you received the message, in your own database, before you act on it. Then you can recognise a redelivery.

Notice what both of those depend on. They assume your database is the reliable thing in the picture. But it can't be. It's a system like any other.

Why a missing database is a different kind of failure 🔗

The problem the outbox solves is a partial write. Half the work committed and half didn't, and the two halves sit in systems that can't agree with each other.

An unavailable database doesn't produce a partial write. It produces no write.

The transaction never begins. There's no order row and no outbox row, and nothing sits half-finished anywhere. There's no inconsistency to repair, because nothing got recorded to be inconsistent with. Your service tried to do some work and the work didn't happen. That's an ordinary failure, and your application already knows how to talk about it.

That's the shape of the whole answer. A down broker threatens correctness. A down database threatens availability. They feel the same when you're the one being paged. They need different responses.

The outbox path from producer to consumer. A down broker only delays publishing, because the row is already committed. A down database stops the transaction from starting at all.

On the producer side, the save just fails 🔗

Your handler opens a transaction and the connection throws. The request fails. The caller gets an error.

You retry at the application layer, with backoff and a limit. It's the same retry you'd write for any database that's temporarily unreachable. If it's an HTTP request, the client sees a 503 and decides whether to try again. If it's a message, nothing acknowledges it, so it comes back later. We'll get to that case.

The important part is what didn't happen. Nothing was published, because publishing was never the first step. No downstream service acted on an order that doesn't exist. Your API returned an honest error instead of an acceptance it couldn't back up.

Failing to write is a clean failure. Writing half of it costs you more.

The outbox worker stops and waits 🔗

The worker's whole job is to poll a table. That table lives in the database that just went away. So the worker gets a connection error, logs it, and tries again next cycle. It does this for as long as the outage lasts.

Meanwhile the rows committed before the outage sit there untouched, because that's what committed means. They're not in flight. They're not in a queue that might expire them. They're durable rows in a table, which is the safest place a pending message can be.

When the database comes back, the next poll succeeds and the worker picks up where it stopped. Messages go out late. None go out twice because of the outage, and none go missing.

Check one thing in your own setup. Make sure the worker's failure to reach the database is loud. A component that polls quietly can stay broken for an hour after everything else recovers. Nobody notices, because no user ever sees an error.

On the consumer side, the broker holds the line 🔗

The consumer pulls a message off the queue and tries to write it to the inbox table. That write fails.

So the consumer never acknowledges the message.

That single fact does all the work. Every broker worth using treats an unacknowledged message the same way. After a visibility timeout or a channel close, it goes back on the queue and gets delivered again. The broker never had a reason to believe anyone handled it.

The rule generalises well beyond this scenario. Never acknowledge work you couldn't persist. Acknowledge after the write, never before, and most of your redelivery behaviour becomes correct by construction.

Where it still bites 🔗

The self-healing story is real, and it has an expiry date. This is the part I'd put on a slide.

Queues grow while consumers are stuck. Every message you'd normally process piles up, along with the redeliveries of the ones already being retried. Queue depth tells you how much trouble is waiting when things come back. It's often the first place a database outage becomes visible.

Retries don't run forever. Each redelivery burns an attempt. Once a message exhausts its retry policy it lands in the dead letter queue. An outage longer than your retry budget quietly turns good messages into dead-lettered ones, and a human has to replay them.

Message TTL (time to live) and queue retention keep running during the outage. If your broker discards messages older than an hour, an hour-long outage starts deleting work. That's the failure mode that turns a delay into real data loss. It happens without a single error in your logs.

None of these are outbox problems. They're what happens when a system stays unavailable longer than its timeouts expected. They're also why "it recovers on its own" needs a duration attached to it.

Three zones, and two numbers decide where the boundaries fall. Up to your retry budget, the backlog is just a backlog. Past it, messages dead-letter and a human has to replay them. Past TTL and retention, the broker starts deleting, and there's nothing left to replay.

A timeline of outage duration split into three zones. Queues grow, until the retry budget runs out; then messages land in the dead letter queue, until TTL and retention expire; after that they are deleted.

What this means in practice 🔗

  • Treat a database outage as an availability incident, not a data integrity one. Nothing partial was written, so nothing partial needs undoing.
  • Retry the producer's save at the application layer. Backoff, a limit, and an honest error to the caller.
  • Alert on your outbox worker's failures. Silent polling is how a recovered system stays broken.
  • Acknowledge only after the inbox write succeeds. That one rule covers most redelivery correctness.
  • Know your numbers: retry budget, message TTL, and queue retention. They decide how long an outage stays survivable.

Closing 🔗

This question is a good one because it makes you say out loud which system you're trusting. The outbox pattern trusts your database completely. That's reasonable, right up until someone asks what happens when that trust is unavailable for twenty minutes.

Nothing partial gets written, so nothing partial needs undoing. Everything pauses instead of corrupting. Then a retention policy you configured two years ago decides how long that stays true.

Go and find out how long your broker will hold everything before it starts deleting. Most teams have never measured it, and it's a five-minute answer. Tell me what you find on Twitter/X.

Get the next one by email

Every post here goes out by email too - one a week.
.NET, messaging, and distributed systems, with the trade-offs the docs leave out.