How we cut cold starts in half

A cold start is the time between a request arriving and a fresh instance being ready to answer it. For services that scale to zero, it is the first thing users feel. Over the last quarter we brought the median cold start of a Node service from 410 ms to 190 ms. This post explains the three changes that did most of the work.

Measure before you change anything

We split every cold start into four phases and logged each one: scheduling, image pull, runtime boot, and application init. The breakdown surprised us. Image pulls were already cached on most hosts, but application init, the time your own code spends importing modules and opening connections, was almost half of the total.

cold-start breakdown (p50, before)
scheduling      38 ms
image pull      52 ms
runtime boot   131 ms
app init       189 ms
total          410 ms

Snapshot the runtime after boot

The runtime boot phase does the same work for every instance of a given build. We now take a memory snapshot right after the runtime finishes booting, store it next to the build, and restore it on the next cold start. Restoring a snapshot takes about 30 ms, compared to 131 ms for a full boot.

Snapshots are taken per build, so a new deploy never restores state from an old one.

Load modules when they are first used

Many services import every route handler at startup, even though one request only needs one of them. The runtime now supports lazy route modules: the router knows which file handles each path, and imports it on the first request to that path.

server.ts
import { createServer } from "@cloud/runtime";

createServer({
  routes: {
    "/api/invoices": () => import("./routes/invoices"),
    "/api/reports": () => import("./routes/reports"),
  },
});

Route the first request to a warm neighbor

The last change was in the router. When a new instance is starting, the first request is sent to an existing warm instance if one has capacity, and the new instance only receives traffic once it reports ready. Users no longer wait for a cold start during a scale-up, only during a scale from zero.

Results

The median cold start is now 190 ms and the p95 is 420 ms, down from 980 ms. All three changes are on by default. Lazy route modules need the runtime 2.4 or later.

Share:

Written by

Daniel Okafor

Staff engineer, Runtime

Daniel works on the runtime and the database layer. Daniel writes about cold starts, migrations, and measuring performance in production.

Newsletter

Engineering notes, once a month

Release highlights, deep dives from the platform team, and guides you can use the same day. No spam.

gats-fo-lex

Unsubscribe anytime. Read our Privacy Policy.

Free tier · No credit card

Push code today. Be live before your coffee cools.

Connect a repository, pick a region, and get a production URL with HTTPS, logs, and autoscaling already switched on.

$ git push origin main

  1. Build34 s
  2. Deploy12 s
  3. Health checks3 s

your-app.example.com

Buy NowTheme Details