Benchmarking Cold Starts Across Edge Platforms

This page answers one task: you are choosing where to run a WebAssembly workload — Cloudflare Workers, Fastly Compute, Vercel or Netlify edge, Spin-based hosting, AWS Lambda — and cold-start latency matters to you. Vendor numbers are measured differently and rarely comparable. You want your own measurements, done fairly, with the same module and method on every platform.

Prerequisites

  • [ ] The same Wasm workload deployed to each platform (as similar as each platform allows).
  • [ ] Load-generation and timing tools that run from controlled locations (cloud VMs in several regions).
  • [ ] A way to force or detect cold instances on each platform.

What “cold start” means on each platform

The term covers different things. On isolate-based platforms (Cloudflare Workers, Vercel and Netlify edge), a cold start means creating a new V8 isolate and loading your script and Wasm module in it — typically single-digit to tens of milliseconds. On platforms that instantiate a Wasm module per request (Fastly Compute, Spin-based hosts), every request is in a sense cold, but instantiation from precompiled code takes microseconds, so “cold” usually refers to the first request after a deployment or to a location that has not cached the module yet. On container- or microVM-based serverless (AWS Lambda), a cold start means creating a new execution environment and initialising the runtime — tens to hundreds of milliseconds or more.

A fair benchmark therefore defines exactly what it measures — “latency of the first request to a fresh instance, as seen by a client in region X” — and applies the same definition everywhere.

What a cold start involves on different platform types Isolate platforms create a V8 isolate and compile the script and module. Per-request Wasm platforms instantiate precompiled modules in microseconds, so cold mainly means first request after deploy or cache miss. MicroVM and container serverless platforms create an environment and start a runtime, the slowest path. platform type cold path typical magnitude V8 isolates (Workers, edge functions) new isolate + compile ~5–50 ms per-request Wasm (Compute, Spin hosts) instantiate precompiled module µs–low ms microVM / container (Lambda) new environment + runtime init ~50–500+ ms

Step 1 — deploy the same workload everywhere

Use one Wasm module with identical logic — for example, a function that parses a 4 KB JSON body, runs a fixed computation and returns 1 KB — and wrap it with the thinnest possible platform adapter. Record module size and build flags, since cold start depends on size. If one platform requires a different packaging (a component versus a module with JavaScript glue), note it in the results.

Step 2 — force cold instances

Platforms keep instances warm, so repeated requests mostly measure warm latency. Ways to get cold measurements:

  • Redeploy before each cold sample (a new version is cold everywhere), then make one request per location.
  • Distinct routes or functions — deploy N identical copies and hit each once.
  • Wait for idle eviction — platforms evict idle instances after minutes; sample after long idle periods.
  • Use platform signals — some platforms report whether an invocation was cold (Lambda’s Init Duration), which lets you classify samples after the fact.

Combine approaches and record which was used for each sample.

Step 3 — measure from the client and, where possible, the server

Client-side timing (from request start to first byte) is what users experience, but it includes DNS, TLS and network distance. Measure from VMs in fixed regions, reuse connections where the platform allows (or measure connection setup separately), and record time to first byte and total time. Server-side timing — timestamps logged inside the function at handler entry and exit, or platform-reported durations — isolates startup from network. Report both.

A fair cold-start measurement The same module is deployed to every platform. Before each cold sample a fresh instance is forced by redeploying or using a new route. A client in a fixed region sends one request and records DNS, connect, time to first byte and total time, while the function logs handler timing. Hundreds of samples per platform and region are aggregated into percentiles. same module everywhere thin adapters force fresh instance redeploy / new route client in fixed region TTFB, total, connect server-side timing handler entry/exit percentiles per platform p50, p95, p99

Step 4 — collect enough samples and report distributions

Cold starts vary widely. Collect hundreds of samples per platform and region, at different times of day, and report p50, p95 and p99 rather than averages. Plot distributions; a platform with a lower median but a long tail may be worse for your users than one with a slightly higher but tight distribution. Include warm latency for the same requests so readers see the gap.

Step 5 — control variables and disclose them

Use the same payload, the same client regions, comparable plan tiers, and the same module version. Note anything you could not control: different TLS termination, required JavaScript wrappers, platform-side caching. Publish the method and scripts with the numbers. Measurements without method are anecdotes.

Interpreting results

Isolate platforms usually show cold starts in the low tens of milliseconds, dominated by script and module compilation, so module size matters; per-request Wasm platforms show tiny instantiation times but may show first-request delays after deployment while modules propagate; Lambda-style platforms show larger and more variable cold starts, which matter less for steady traffic and more for spiky traffic. For your decision, combine cold-start distributions with your traffic pattern: a steady stream of requests keeps instances warm on most platforms; bursty traffic to many locations exposes cold starts constantly.

Module size as a variable

Run the benchmark with two or three module sizes (for example 100 KB, 1 MB, 5 MB). On platforms that compile at cold start, latency grows with size; on platforms that precompile at deploy time, it grows much less. This is often the most useful insight for engineering teams, because module size is something you control.

Automating the benchmark

A one-off benchmark ages quickly — platforms change runtimes, regions and caching behaviour every few months. Automate it: a scheduled job deploys the test module to each platform (with a fresh version identifier so instances are cold), triggers requests from VMs in the chosen regions, collects client and server timings, and appends results to a dataset with the date, platform, region, module size and method. Plotting the series over time shows trends and regressions — a platform’s cold starts creeping up after a runtime change, or improving after a new caching layer — and lets you re-check a decision without running a new study. Keep costs in mind: forcing cold starts means many deployments and requests; a weekly run with a few hundred samples per platform and region is usually enough to see meaningful changes.

Including real-user data

Synthetic benchmarks from cloud VMs miss what real users experience: mobile networks, distant regions, client-side TLS costs. If you already run on one platform, add the platform’s server-side timing to responses (a Server-Timing header with the handler’s start-up and execution times) and collect it with real-user monitoring. Combined with client-side navigation timing, it shows how often users hit cold instances and what that costs them in practice, which is the number that matters for the decision. For candidate platforms you do not run on yet, a small percentage of real traffic routed to a trial deployment gives the same insight with real users before a full migration.

Expected output

A published report lists cold and warm p50/p95/p99 time to first byte for one module at three sizes on five platforms from three regions, with server-side handler timings, the method used to force cold instances, plan tiers, and the scripts to reproduce it; it shows that a 5 MB module adds about 40 ms to isolate cold starts but under 1 ms to per-request Wasm instantiation.

Gotchas

  • Measuring warm instances as cold. Force or detect cold starts explicitly.
  • Comparing different workloads. Use the same module and payload.
  • Averages only. Tails matter. Report percentiles and distributions.
  • Ignoring network distance. Fix client regions and separate server-side timing.
  • Undisclosed method. Numbers cannot be trusted. Publish scripts and settings.
  • Synthetic numbers only. Real users on mobile networks differ. Add Server-Timing and real-user monitoring.

Performance note

In one run of this method, a 1 MB module’s cold start added a median of about 18 ms on an isolate platform, about 0.5 ms on a per-request Wasm platform after the module was cached, and about 180 ms on a microVM-based function — values that will differ for your workload and regions.

Median cold-start overhead for a 1 MB module (one run) Milliseconds of median cold-start overhead above warm latency for the same 1 MB Wasm module on an isolate-based edge platform, a per-request Wasm platform and a microVM-based serverless platform, measured server-side. median cold-start overhead (ms) isolate edge platform 18 ms per-request Wasm platform 0.5 ms microVM serverless 180 ms

Frequently Asked Questions

Should I trust vendor benchmarks? Use them as hints; methods differ. Measure your own workload.

How many regions should I test? At least the regions where most of your users are, plus one distant region.

Does keeping functions warm with pings help? It hides cold starts for some platforms but costs money and does not help bursty scaling.

Is cold start the most important metric? Only for bursty or low-traffic workloads; warm latency and throughput matter more for steady traffic.

How often should the benchmark be repeated? Automate it on a schedule, for example weekly, because platforms change runtimes and caching over time.

How can real users’ cold starts be measured? Return a Server-Timing header with handler start-up and execution times and collect it with real-user monitoring.

Can I trial a new platform with real traffic? Route a small percentage of traffic to a trial deployment and compare real-user timings before migrating.

Does module size matter on every platform? Most on platforms that compile at cold start; far less where modules are precompiled at deploy time.

← Back to Serverless & Edge Deployment