Running Wasm on Fastly Compute
This guide answers one task: write, test and deploy a request handler that runs as a WebAssembly module on Fastly Compute, where the platform creates a brand-new instance for every single request.
Prerequisites
- [ ] The
fastlyCLI, authenticated with an API token. - [ ] Rust with the
wasm32-wasip1target, or Node for the JavaScript SDK. - [ ] A Fastly service, or the CLI’s ability to create one.
- [ ] Viceroy, the local runtime, which the CLI installs for you.
An instance per request, genuinely
Most platforms reuse an instance across requests and ask you to be careful about state. Fastly Compute does not: each request gets a fresh WebAssembly instance with zero-initialised memory, created from the compiled module in tens of microseconds and discarded afterwards.
That removes an entire category of bug. There is no cross-request state to leak, no cache to invalidate, no reset to remember. It also means anything your module does at startup happens on every request, so the cost of initialisation is the cost per request — a constructor that builds a lookup table is paid a million times a day rather than once.
A handler in Rust
The Rust SDK maps the platform onto types that will look familiar to anyone who has written an HTTP service, with the important difference that backends are declared rather than dialled.
use fastly::{Error, Request, Response};
use fastly::http::StatusCode;
#[fastly::main]
fn main(mut req: Request) -> Result<Response, Error> {
if req.get_method() != "POST" {
return Ok(Response::from_status(StatusCode::METHOD_NOT_ALLOWED));
}
let body = req.take_body_bytes();
if body.len() > 1 << 20 {
return Ok(Response::from_status(StatusCode::PAYLOAD_TOO_LARGE));
}
let out = transform(&body)?; // your pure logic
Ok(Response::from_status(StatusCode::OK)
.with_content_type(fastly::mime::APPLICATION_JSON)
.with_body(out))
}
Notice the size check before any work. With a fresh instance per request and a memory limit per instance, rejecting an oversized body early is cheaper than discovering the limit halfway through processing it.
Backends are declared, not dialled
The module cannot open a socket. Outbound requests go to backends named in your service configuration, which is the platform’s capability mechanism.
# fastly.toml
name = "engine"
language = "rust"
manifest_version = 3
[local_server.backends.origin]
url = "https://origin.example.com"
[local_server.backends.config_api]
url = "https://config.internal.example.com"
let mut upstream = Request::get("https://origin.example.com/data.json");
upstream.set_ttl(60);
let resp = upstream.send("origin")?; // the backend name, not the URL, grants access
let payload = resp.into_body_bytes();
An attempt to reach a host with no matching backend fails rather than connecting, which is the property that makes the capability real. Declaring backends also gives you a single place to audit what a service can talk to — worth reviewing whenever a new dependency appears.
Streaming, because the platform is built for it
Fastly Compute is a caching platform first, and its APIs assume bodies flow rather than buffer. Streaming through the module keeps memory bounded and lets the client start receiving output before your handler has finished.
use std::io::{BufRead, BufReader, Write};
#[fastly::main]
fn main(req: Request) -> Result<Response, Error> {
let upstream = Request::get("https://origin.example.com/feed.ndjson").send("origin")?;
let mut out = Response::from_status(200).with_content_type(fastly::mime::APPLICATION_JSON);
let mut writer = out.stream_to_client(); // response starts flowing now
let reader = BufReader::new(upstream.into_body());
for line in reader.lines() {
let transformed = transform_line(&line?)?;
writer.write_all(transformed.as_bytes())?;
}
writer.finish()?;
Ok(Response::from_status(200))
}
Streaming a response you have already started means the status and headers are committed before you know whether the rest will succeed. Decide your status early, or buffer just enough to be confident — an error midway through a streamed 200 response can only be signalled by truncating it.
Configuration without a redeploy
Anything baked into the module requires a rebuild and a deploy to change, which is fine for logic and wrong for values that operations needs to adjust. The platform provides config stores — dictionaries of key-value pairs edited independently of the service version — for exactly this.
use fastly::ConfigStore;
let cfg = ConfigStore::open("runtime_settings");
let max_items: usize = cfg.get("max_items").and_then(|v| v.parse().ok()).unwrap_or(100);
let mode = cfg.get("mode").unwrap_or_else(|| "normal".into());
Two habits keep this from becoming its own problem. Always supply a default, because a missing or malformed value must not take the service down — a config store edited by a human at two in the morning will eventually contain a typo. And treat the values as untrusted input: parse, validate, clamp to a sensible range, and log when a value is rejected so the person who set it finds out.
Secrets belong in a secret store rather than a config store, and neither belongs in the compiled module.
A .wasm file is a distributable artifact; any string compiled into it is readable by anyone who
obtains it, which on an edge platform includes rather more parties than you might assume.
Observability without a machine to log into
There is no process to attach to, so logging is the primary diagnostic and it needs to be set up before the first real traffic. The platform ships log lines to an endpoint you configure, and the useful discipline is the same as anywhere: structured lines, one per request, with a correlation identifier.
use fastly::log::Endpoint;
use std::io::Write;
let mut log = Endpoint::from_name("observability");
writeln!(log, r#"{{"req_id":"{}","path":"{}","status":{},"ms":{:.1},"bytes":{}}}"#,
req_id, path, status, elapsed_ms, out_len)?;
Sample rather than logging everything once traffic is meaningful — full logging at edge volumes is expensive and rarely more informative than a one-in-fifty sample plus every error. Include the service version in each line so that during a rollout you can tell which build produced a given result, and keep the field names stable, because a dashboard built on last month’s names is worse than no dashboard.
Test locally with Viceroy
The CLI runs your compiled module under Viceroy, a local implementation of the platform’s host calls,
including backends mapped to whatever you point them at in fastly.toml.
fastly compute build
fastly compute serve # runs on http://127.0.0.1:7676
curl -s -X POST localhost:7676 --data-binary @fixtures/input.json | head -c 100
Local testing catches capability mistakes early: a request to an undeclared backend fails the same way it would in production, and a module built for the wrong target fails to load at all. What it does not reproduce is real network latency to backends or the platform’s caching behaviour, so keep a smoke test against a deployed service for those.
Expected output
A build reports the artifact size, which is the number to watch against the platform’s limit:
fastly compute build
Built package 'engine' (engine.tar.gz)
Wasm binary: 1.42 MiB
Package size: 512.8 KiB
fastly compute deploy
Uploading package... done
Activating service version 7... done
Deployed to https://engine.edgecompute.app
# a request, with the platform's own timing header
curl -sD - -o /dev/null https://engine.edgecompute.app | grep -i x-served
x-served-by: cache-lhr-egll1980072-LHR
x-compute-time: 2.4ms
Gotchas
- Startup work per request. A lazily built table is built every time. Move it to a
constcomputed at build time, or accept the cost knowingly. - Backend not declared. Outbound requests fail with a clear error, which is the system working — add the backend rather than routing around it.
- Wrong target. Build for
wasm32-wasip1; the platform is a WASI host, unlike some others. - Buffering a large body. Instance memory is bounded. Stream, or reject early with a size check.
- Committing a status before you can honour it. Once a streamed response has started, you cannot change the status code.
- Assuming threads. Execution is single-threaded per request; a threaded build will not run.
Performance note
A representative transform over a 20 kB JSON body ran in about 2.4 ms of compute, with instance creation around 35 µs — negligible next to the work. Moving a 40 kB lookup table from runtime construction to a build-time constant removed 1.1 ms per request, which at scale was by far the largest single improvement available and is a direct consequence of the per-request instance model.
Frequently Asked Questions
Can I use the JavaScript SDK instead of Rust? Yes. It compiles JavaScript into a module with an embedded engine, which is heavier than a Rust build and much faster to write. For request shaping and light transformation it is often the right trade; for compute-heavy work, Rust’s smaller and faster output wins.
How does caching interact with my handler? The platform’s cache sits in front of and around your code, and you control it through the request and response APIs — setting a TTL on a backend request, or deciding whether a response is cacheable. Using it well is usually a bigger win than optimising the module.
Is this the same WASI my local runtime provides?
Close enough that a module built for wasm32-wasip1 runs under both, but the platform provides its own
host calls for HTTP, caching and configuration on top. Code using those will not run under plain
wasmtime without a shim.
Related
- Deploying Wasm to Cloudflare Workers — the same job on a different platform.
- Compiling Rust to wasm32-wasip1 — producing the artifact.
- Cold start characteristics of server-side Wasm — where the microseconds go.
← Back to Serverless & Edge Deployment