Reliable Delivery
This guide explains how to make sure that events reliably leave the system that produced them and reach every system that depends on them. It covers the dual-write problem, the transactional outbox, delivery guarantees, ordering, and the handling of consumers that fall behind or fail.
Publishing events is only useful if delivery can be relied on. If an event is recorded but never published, other systems silently miss it: a loan is registered, but the reader's account never shows it. If an event is published but its recording fails, other systems react to something that did not happen. Either way, systems drift apart, and nobody notices until something goes wrong.
The Dual-Write Problem#
The root of most delivery problems is the dual write: a system stores a change in its database and then sends a message to a broker. These are two separate operations, and there is no transaction that spans both:
- If the store succeeds and sending fails, the event is lost for everyone else.
- If sending succeeds and the store fails, consumers react to something that never happened.
- If the system crashes between the two, it is unclear which of the two happened.
Retrying does not solve this on its own, because the system may not know whether the first attempt succeeded. The rule that follows is simple:
Recording an event and making it available to others must succeed together – or fail together.
The Transactional Outbox#
A common solution is the transactional outbox. Instead of sending a message directly, the system writes it into an outbox table in the same transaction as the change itself. A separate relay reads the outbox and forwards its entries to the broker, marking them as sent once the broker has confirmed them.
Because the change and the outbox entry are written atomically, neither can exist without the other. If the relay fails, it simply tries again later. The price is a second component to operate, and the possibility of sending a message twice – if the relay crashes after sending but before marking the entry as sent.
The Event Store as the Outbox#
In an event-sourced system, the problem looks different: the event store already is the outbox. Recording an event and making it available are the same operation, because consumers read the events directly from the store, in order, at their own pace.
This removes the dual write entirely, as long as consumers pull from the store. It reappears only when events have to be forwarded to a separate broker – for example, to reach systems that cannot read from the store. In that case, the forwarding component is itself a consumer: it reads from the store, remembers how far it has come, and publishes to the broker, retrying as needed.
Delivery Guarantees#
Between systems, three delivery guarantees are commonly distinguished:
- At most once. Every event is delivered once or not at all. Nothing is ever duplicated, but events can be lost. This is rarely acceptable for business events.
- At least once. Every event is delivered, possibly more than once. Nothing is lost, but consumers must cope with duplicates.
- Exactly once. Every event is delivered exactly once. Across independent systems, this cannot be guaranteed by the transport alone.
In practice, reliable systems combine at-least-once delivery with idempotent consumers: every event is delivered, and processing the same event twice has the same effect as processing it once. The result is often called effectively once. Idempotency can come from the nature of the operation – setting a value is idempotent, incrementing it is not – or from remembering the identifiers of events that have already been processed. How to build such consumers is described in Building Event Handlers.
Ordering#
Many consumers depend on the order of events: a book cannot be returned before it was borrowed. Global ordering across all events is expensive and rarely necessary. What usually matters is the order within one subject – all events of one loan, one reader, one copy.
Keep this order intact along the entire path:
- Deliver the events of one subject in the order they were recorded.
- When consumers process events in parallel, partition by subject, so that events of the same subject are never processed concurrently.
- When events are forwarded to a broker, use the subject as the partition key, so that the broker preserves the order too.
Consumers That Fall Behind#
Consumers work at different speeds. A search index may keep up in milliseconds; a nightly report may process events once a day. Reliable delivery must not depend on every consumer keeping up:
- Backpressure. A slow consumer should slow down its own intake, not force the producer to wait or lose events.
- Retention. Events must remain available long enough for the slowest consumer – and for new consumers that need to start from the beginning.
- Replays. A consumer that has to rebuild its state, or a new consumer that joins later, should be able to process the history again from any point.
An event store that keeps all events provides retention and replays by design. A broker that deletes events after a while does not, which is one reason to treat the event store, not the broker, as the source of truth.
Events That Cannot Be Processed#
Sometimes a consumer cannot process an event: the data is unexpected, a downstream system rejects it, or a bug makes the handler fail. Retrying forever blocks everything behind it; skipping silently loses data. Instead:
- Retry with increasing delays, since many failures are temporary.
- Set failing events aside after a limited number of attempts – in a dead-letter queue or a list of failed events – together with the reason for the failure.
- Make them visible and fix the cause, then process the events again.
When events of a subject must stay in order, setting one aside may mean pausing the entire subject until the problem is solved.
Designing for Reliability#
Reliable delivery is not a feature of a single component, but a property of the whole path from producer to consumer. Record and publish atomically, deliver at least once, make consumers idempotent, preserve the order where it matters, and plan for consumers that are slow, new, or failing. Then the events a system publishes become something others can build on – which is the whole point of publishing them.