NIMBUS
Infrastructure

The Infrastructure Ceiling Is a Lie

Three years ago we moved every Nimbus service above the abstraction layer. Here is what broke, what held, and why we are never going back.

Cameron Dray · June 2026 · 8 min read

In January 2023, we ran a single relational datastore on bare metal in a colocation facility outside Nuremberg. It handled auth, billing, and user state for fourteen services. When it went down, twice that winter, we lost everything: logins, payments, session state. The team spent two days rebuilding from logs while customers opened tickets. I remember sitting in an airport lounge at six in the morning thinking: this ceiling we built for ourselves is entirely artificial.

The Abstraction Actually Held

The conventional wisdom in 2022 was that managed compute meant lock-in, cold starts, and debugging hell. We heard it from investors, candidates, and launch threads. What nobody mentioned was that our bare-metal setup required three full-time engineers just to keep replication lag under two seconds. The cold start problem was real, but the alternative was another winter of recovery at dawn.

The moment we stopped treating infrastructure as craft and started treating it as a utility, our velocity tripled.

Today Nimbus runs on managed infrastructure from compute to queue to object storage. Our recovery time dropped from six hours to under ninety seconds. The ceiling was never real; it was just the altitude at which we had grown comfortable breathing thin air.

This is the Oblivion Sky Tower design system, applied by Curio Design — a design-style library for AI agents. Full Oblivion Sky Tower guide → designbycurio.com/learn/oblivion-2013-bubbleship