Outgrowing a provider is a scaling event, not a mistake
Moving FocalHQ from Firestore to PostgreSQL on Neon and putting Cloudflare in front of it - why the first choice was still the right one, and what decides whether the second one costs weeks or months.
FocalHQ ran on Firestore, and that was the correct decision (case study). A document store asks nothing of you up front: no migrations, no schema review, no connection pool to size, no columns to argue about before you know what the business calls things. When the shape of the domain is still moving — when “engagement” means one thing this month and something slightly different next — a database that lets you write whatever you have is not a shortcut, it is the correct tool for a system whose specification is still being discovered.
The decision stopped being correct the day the Suppliers dashboard needed an answer the data model
could not give cheaply: every supplier, with this quarter’s bills totalled next to them, matched
against the engagements they were booked to, split by currency. In PostgreSQL that is one query
with a couple of joins and a GROUP BY. In a document store it is a fan-out read, a set of
denormalized counters someone has to keep correct on every write, and a consistency argument every
time the two disagree. Nothing was failing. The database was answering the question it was shaped
to answer, and the business had started asking a different one.
That gap — between the question the business asks and the question your provider is shaped to answer — is what growth actually does to a stack. It rarely shows up as a load problem, which is what “scalability” is usually taken to mean. It shows up as queries that need three round trips, features that need a background job to stay correct, and a growing folder of code whose only purpose is to compensate for the storage engine. Scaling, at that point, is a procurement decision wearing an engineering costume.
Which leads to the rule that matters more than any individual provider choice: you cannot pick a provider you will never outgrow, so stop optimizing for that. Optimize for the exit cost — and the exit cost is a property of your boundaries, not of any abstraction layer you write.
The instinct, once a team has been burned once, is to write the abstraction on day one: a repository interface over everything, a “database-agnostic” data layer, an adapter per vendor with exactly one implementation. In our experience that fails twice. It costs real money immediately, paid by every feature that now goes through a layer serving no one. And it does not work when the day comes, because a portability layer written against one provider quietly encodes that provider’s semantics — its transaction model, its ID generation, its idea of what a query can do — so the second implementation ends up either impossible or a lie. Speculative portability is the most expensive kind of insurance: high premium, and it does not pay out.
What does pay out is much less glamorous, and mostly consists of things you should be doing anyway.
Two kinds of provider change, and only one of them is expensive
Before anything else, work out which kind of change you are looking at, because the two have nothing in common but the word “provider”.
Some providers sit in front of you. DNS, CDN, WAF, bot protection, edge access control. Putting Cloudflare in front of FocalHQ’s API and web app changed no application code at all — it is a nameserver change and a config surface. Turnstile is a script tag and one server-side verification call. Zero Trust in front of the admin surfaces is a policy, not a refactor. The whole thing is reversible in an afternoon, because nothing in the system is structurally aware that it is there. That is the defining property of this category, and it is why adopting an edge provider is a low-risk way to buy capability you would otherwise have to build: rate limiting, geographic caching, DDoS absorption, and a bot wall are all things a small team should rent rather than operate.
Some providers sit underneath you. Database, auth, object storage, queue. Your data model is shaped by their constraints, your business logic is written against their guarantees, and every read path in the codebase has an opinion about them. This is where migrations turn into quarters.
The practical use of the distinction is at selection time, not migration time. Lock-in in front of you is cheap optionality — buy it freely. Lock-in underneath you is a structural commitment, and it deserves the kind of scrutiny a hiring decision gets.
Choose the standard, not the product
Neon’s decisive feature, for this decision, was not branching or scale-to-zero. It was that Neon is
PostgreSQL: the same wire protocol, the same SQL, the same pg_dump. The proprietary parts are
operational conveniences that live outside the application — nothing in a service’s code knows
which company runs its Postgres. If Neon turns out to be the wrong call in two years, the move is a
dump, a restore and a connection string, over a maintenance window. That is a procurement change.
Compare a provider whose value is its own query API. There, your business logic is written in a dialect only that vendor speaks, and leaving means rewriting every read path in the system. Same category of product; entirely different exit cost.
So the question to ask while choosing, before anything is committed: if this vendor doubled its price tomorrow, what would the migration touch? If the honest answer is “connection configuration and some ops runbooks”, you are buying a service. If it is “every query in the codebase”, you are buying a dependency. Both can be correct — but only one of them should be a surprise.
The four things that made the move cheap
None of them were done for portability. That is the point: the properties that make a provider swappable are the same ones that make a system pleasant to work in, which is why you get them without having to predict the future.
One service owns its data. FocalHQ’s services already kept their own data with no shared tables, for reasons that had nothing to do with migrations. The payoff arrived anyway: the move was not one migration, it was eight small ones, each with its own cutover, its own verification and its own rollback. Invoicing moved while expenses were still on the old store. No weekend, no all-hands freeze — the same service-by-service discipline that retired the original Apps Script back office.
The API contract was frozen. The v1 routes and their DTOs do not change; new capability lands as v2 beside them. So the iOS app and the web frontend never found out that the database underneath them had been replaced. A migration that is invisible at the contract boundary is a migration you can ship on a Tuesday, and a contract that is free to drift is a migration that turns into a coordinated client release.
Data access lives in one place per service. Not an abstraction layer — no interfaces, no adapters, no dependency injection ceremony. Just the ordinary discipline that queries live in a data module and use cases call functions on it, rather than every route handler reaching for the client directly. The metric that predicts your migration cost is not “do we have a repository pattern”. It is how many files import the vendor’s SDK, and you can count that today, in one command, for every provider you depend on. If the number is large, you have found your real migration estimate.
Domain objects are not database rows. A service whose core types are the provider’s response shape has that provider’s data model living inside its business logic, and every function signature becomes a migration site. Where amounts are already integers in minor units and dates are already normalized at one edge, swapping what serializes them is a small, local job.
What does not port, and has to be re-decided
The comfortable version of this story is “we changed the driver”. The honest version is that a storage engine’s semantics leak into design decisions that survive it, and a migration that ports them faithfully carries the workaround instead of the requirement.
Consistency and transactions. Firestore transactions and PostgreSQL transactions guarantee different things at different costs. Retry loops, carefully idempotent writes and compensating-update paths that existed to work around the old model are often simply unnecessary now — and code that defends against a problem the new database does not have is not free, it is a permanent tax on everyone who reads it.
Eventing. This is the one that catches teams out. A document store hands you a change stream for free, and features quietly get built on it. Relational databases do not, so the trigger you were relying on becomes an explicit outbox table and a consumer you own. Inventory that dependency before the migration, not during it.
The read model. Denormalized aggregates and duplicated fields exist to make a document store answer relational questions. Ported one-to-one, they become dead weight that still has to be kept correct. Deleting them is part of the migration, not a follow-up ticket that never gets scheduled.
The cost shape. Per-operation billing and per-connection billing reward opposite habits. Chatty read patterns that were fine before now want batching; serverless compute in front of Postgres now wants a pooler. Neither is hard, but both are decisions, and neither appears in a schema diff.
Buying lock-in on purpose
None of this is an argument for vendor neutrality as a virtue. FocalHQ still authenticates with Firebase Auth, still deploys on Firebase Functions v2, still stores receipts in Google Cloud Storage, and now leans on Cloudflare for a set of protections we have no business building ourselves. Every one of those is lock-in, and every one of them is worth it, because the alternative is operating infrastructure that is not the product.
The test is proportionality, applied twice. Does this lock-in sit in front of the system, where leaving is a config change, or underneath it, where leaving is a rewrite? And is what it buys me proportional to what it would cost to walk away? A bot wall you rent for a config file is a bargain at almost any price. A query language you can never leave had better be doing something extraordinary.
What it adds up to
Growth does not usually break your provider. It changes the questions you ask it, until the answers start requiring code that exists only to compensate — and that code, not a latency graph, is the signal that a boundary has moved.
Nothing about this is avoidable by choosing better the first time, and the teams that suffer most in these migrations are usually the ones that tried: they either picked a heavyweight platform they did not need for two years, or they wrapped everything in an abstraction that made the daily work worse and the eventual move no easier. The teams that move in weeks did four unglamorous things first — kept ownership of data inside service boundaries, froze the contract clients depend on, kept vendor calls in one module per service, and kept domain types independent of storage shapes.
The first provider was right. The second one is also temporary. Build so that finding out costs a sprint, not a quarter.