Skip to content
01TX
Insights

Architecture

Designing an OMS for Recovery, Not Just Correctness

7 min read

Describe an order management system to someone and you'll almost always describe the happy path: an order comes in, gets validated, gets routed, gets filled, the position updates, everyone moves on. That description is correct, and it's also not the hard part. The hard part is what the system does in the other case — when the process crashes between sending an order and recording its acknowledgment, when a network partition separates the OMS from a venue mid-session, when the system restarts and has to figure out, from scratch, what state the world was actually in.

An OMS that's only been designed for the happy path tends to work fine in testing and demos, because testing and demos rarely interrupt it mid-operation. Production does, eventually, always.

The ambiguity problem

Here's the concrete failure case: an order is sent to a venue, and the process crashes before the acknowledgment is written anywhere durable. On restart, what is that order's status? It might have never reached the venue. It might have reached the venue and been accepted. It might have been filled already. From the OMS's own records, if those records are a simple mutable 'current state' table, there's often no way to tell — the crash happened in the gap between an action and recording that the action happened, and that gap is exactly where ambiguity lives.

This is not a rare edge case that only matters at extreme scale. Any sufficiently long-running OMS will restart — deployments, infrastructure maintenance, actual failures — and every restart re-opens this question for whatever was in flight at that moment.

Treat state as a log, not a snapshot

The structural fix is to stop treating an order's current status as the primary source of truth and start treating the sequence of events that produced it that way. Every transition — submitted, acknowledged, partially filled, filled, canceled, rejected — gets written as an immutable, append-only event before anything else happens. The 'current state' becomes a derived view, computed by replaying those events, not a row that gets overwritten in place.

This reframing is what makes recovery tractable. On restart, the system doesn't have to guess what state it was in — it replays the event log and arrives back at an answer that's provably consistent with everything that was durably recorded. If the crash happened before an event was written, the worst case is that the system doesn't yet know about an action it took — which is a problem you can design around (via reconciliation against the venue) — rather than a problem that corrupts the record going forward.

Idempotency is not optional

Recovery and reconnection logic, done honestly, produces at-least-once delivery almost everywhere: a resent FIX message after a gap fill, a retried API call after a timeout where the first call may or may not have succeeded, a venue that resends a fill confirmation because its own acknowledgment got lost. An OMS that assumes every inbound message is new will eventually double-book a fill or process a cancellation twice.

The fix is a dedupe key on every inbound event — something that uniquely identifies 'this specific fill' or 'this specific acknowledgment' independent of how many times it arrives — checked against the event log before anything is applied. This has to be designed in from the start; retrofitting idempotency onto a system that was built assuming clean, singular delivery of every message is a much larger project than building it in from day one.

Reconciliation belongs in the main design, not bolted on

Even with a well-designed event log and idempotent handling, the OMS's record and the venue's record can still diverge — a message genuinely lost, a bug, a venue-side issue. The only way to catch that reliably is continuous reconciliation against venue state, not a nightly batch job that was added after the fact because someone asked for one.

Treating reconciliation as a core, always-running part of the system — not an operations afterthought — is really the same design principle as the event log, applied at a different boundary: don't trust your own state implicitly; be able to verify it against an independent source, continuously, and know immediately when it's wrong rather than finding out at month-end.

Why this is the actual job

Recovery design and audit-trail design turn out to be the same problem viewed from two directions. Both require that a system's history be an explicit, replayable record rather than something implicit in whatever the current state happens to be. An OMS built this way can answer 'what happened and why' after a crash, a dispute, or a regulatory request, using the same mechanism — because it was never relying on a snapshot that could silently drift from the truth in the first place.

This is also why an OMS is a genuinely different engineering problem from most backend systems: correctness on the happy path is table stakes, and the actual differentiator — the part that determines whether it holds up in production over years of restarts, network blips, and venue-side incidents — is how carefully the failure paths were designed before they were ever needed.

Building something that runs into these problems directly?