Firebase Functions v2 in production: what we'd keep and what we'd avoid
A year of running per-concern services on Functions v2 - cold starts, concurrency, and where Cloud Run earns its place.
Firebase Functions v2 is not an incremental update to v1. It is Cloud Run wearing a Firebase deployment workflow: every function you deploy becomes its own Cloud Run service, with Cloud Run’s networking, scaling, and billing underneath. Once you internalize that, most of v2’s behavior — the good and the annoying — stops being surprising. After a year of running production workloads on it, here is where we landed.
What we’d keep
Per-concern services instead of per-endpoint functions. The v1 habit was one function per endpoint: createInvoice, updateInvoice, deleteInvoice, each a separate deployment unit. In v2 that habit gets expensive, because every function is a full Cloud Run service — each with its own cold start behavior, its own min-instance bill if you warm it, and its own slot in an increasingly slow deploy. We consolidated to one service per concern: an invoices function routing internally with a lightweight router, a webhooks function for third-party callbacks, a jobs function for queue consumers. Deploys got faster, warm instances got shared across endpoints that used to cold-start independently, and the mental model got simpler. The trade-off is coarser IAM and scaling settings per concern rather than per endpoint — in practice we never missed the granularity. Where those services draw the line at sharing code with each other is a separate decision, and a stricter one.
Concurrency, deliberately configured. The headline v2 feature: one instance can serve many requests at once (up to 1,000; the default is 80), where v1 was strictly one request per instance. For I/O-bound work — which for us means waiting on Firestore, external accounting APIs, and email providers — this is the single biggest cost and cold-start lever. Ten concurrent requests that would have spawned ten v1 instances now share one warm instance. Fewer instances means fewer cold starts encountered, full stop.
A small min-instance budget on the hot path. Cold starts in v2 are Cloud Run cold starts: container spin-up plus your global scope initialization. For user-facing endpoints we set minInstances: 1 and treat it as what it is — a fixed monthly fee for deleting the worst latency tail. With per-concern consolidation, one warm instance covers a whole concern’s endpoints, which is exactly why the two decisions belong together. Everything async — queue consumers, scheduled jobs, event triggers — runs at minInstances: 0, because nothing user-facing is waiting on it.
Configuration in code. v2 puts memory, CPU, concurrency, timeouts, and min instances in the function definition itself, next to the logic they govern. Combined with Secret Manager integration for credentials, the entire runtime posture of a service is reviewable in a pull request. We would not go back to configuration living in a console.
What we’d avoid
Treating concurrency as free. Concurrency shares one Node event loop. A CPU-heavy request — PDF rendering, XML transformation, a large JSON parse — blocks every other request on that instance while it runs. Our rule after getting burned: concurrency high on I/O-bound services, concurrency 1 (or a dedicated service) for anything that computes. Mixing the two profiles in one service is how you get latency spikes that no dashboard explains.
Global state written for one request at a time. v1 let you be sloppy: module-level variables held per-request state and nothing bad happened, because nothing else was in flight. Under v2 concurrency, that same code is a data race. The audit is boring but mandatory — anything mutable at module scope is either genuinely shared (connection pools, clients: good) or per-request state that must move inside the handler (everything else). Related: connection pools sized for v1’s one-request world need re-sizing when 80 requests share them.
Function sprawl. Even with per-concern discipline, the count creeps — a trigger here, a scheduled job there. Every one is a Cloud Run service, and full deploys slow down roughly linearly with the count while occasionally tripping build quotas. Deploying selectively (--only functions:invoices) helps day to day, but the real fix is resisting the creep: a new function needs a reason it cannot be a route or a job type inside an existing concern.
Assuming defaults are production settings. Default memory allocates default CPU, and under concurrency all requests on the instance share that CPU. The defaults are demo settings. Load-test each concern, then set memory, CPU, and concurrency explicitly — the numbers are three lines of code away.
Where Cloud Run earns its place
Functions v2 being “Cloud Run underneath” invites the question: why not just use Cloud Run? For most of our workload the Firebase layer pays its way — trigger wiring to Firestore and Pub/Sub without Eventarc plumbing, local emulation, one deploy command, no Dockerfile to maintain. But we move a workload to Cloud Run proper when any of these appear:
- A custom runtime or system dependency. Anything needing a specific binary, a non-Node runtime, or an OS package belongs in a container we control.
- Streaming or long-lived connections. WebSockets and server-sent events are Cloud Run territory; the Functions layer is built around request/response.
- Deployment control. Cloud Run revisions with traffic splitting give you gradual rollouts and one-command rollbacks. Functions deploys are all-or-nothing.
- Long-running or heavy jobs. Batch work that runs for many minutes with dedicated CPU fits Cloud Run (or Cloud Run jobs) better than a function pretending to be a worker.
The decision rule we settled on: Firebase-triggered, I/O-bound, request-shaped work stays on Functions v2 with per-concern services; anything with a custom runtime, a long-lived connection, or a rollout-safety requirement graduates to Cloud Run. Since both are the same infrastructure underneath, graduating a service is a contained move — the router-based per-concern structure means the code barely changes, which is one more argument for that structure being the right default.
v2’s real lesson is that the serverless abstraction got thinner, and that is a feature. You are operating small servers that happen to deploy like functions. Configure them like servers — deliberately — and they are the cheapest, lowest-maintenance compute we run. Treat them like v1 functions with a version bump, and the surprises find you in production.