Why We Rebuilt Our Entire Fulfillment Stack in Twelve Weeks
When your delivery promise drops below thirty minutes, the old monolith starts breaking in ways nobody predicted.
I spent two weeks last January standing in a Santiago micro-fulfillment center, watching pickers race against a thirty-minute countdown timer we had just rolled out. The system was fast enough on paper — route optimization, inventory lookup, checkout — all under four hundred milliseconds. But the moment three concurrent orders hit the same aisle, the database locks turned our elegant service mesh into a bottleneck nobody had load-tested for. We lost eleven orders that first morning.
The decision to rewrite came at a Tuesday standup where our VP of Engineering slid a single metric across the table: peak-hour order loss had climbed to 3.2%, up from 0.4% six months earlier. That number represented real people waiting at the door of their apartment, groceries stuck somewhere between a shelf and a van. Nobody argued after that.
The Architecture Nobody Wanted
We chose event sourcing — not because it was trendy, but because our core problem was temporal. Every order is a race condition between picking, packing, and dispatch. A traditional CRUD model kept losing the thread: which bag was packed first, which picker had capacity, which substitution the customer had already approved. Event logs gave us a timeline we could actually reason about.
The migration took twelve weeks. We ran the new stack in shadow mode for the first four — every order processed twice, old and new, with automated diffing on timestamps and totals. By week eight, the new system was faster on every metric we cared about. The old one was still serving traffic, but only as a safety net we hoped we would never need.