Tracing Wasm Requests with OpenTelemetry
This page answers one task: a request passes through a host and one or more WebAssembly guests, and when it is slow you want a trace that shows exactly where the time went — in the host, in instantiation, in the guest’s own logic, or in calls the guest made to other services.
Prerequisites
- [ ] A host that runs Wasm guests (an Axum or similar service embedding wasmtime, a Spin app, or a Node service using Wasm).
- [ ] An OpenTelemetry collector or tracing backend (Jaeger, Tempo, Honeycomb, Datadog) and the OpenTelemetry SDK for the host’s language.
- [ ] Familiarity with spans, trace context and the
traceparentheader.
What a useful trace looks like
A distributed trace is a tree of spans, each a timed operation with attributes, linked by a trace id propagated between services. For a host running Wasm, the interesting spans are: the incoming request; compiling or fetching the component (if not cached); instantiating it; the call into the guest; operations inside the guest; and outbound requests the guest makes through the host. Without Wasm-aware instrumentation, the guest is a black box — one long span labelled “handler” — and you cannot tell whether a slow request was slow because of instantiation, guest computation, or a downstream API the guest called.
The host can create most spans by itself, because it performs instantiation and mediates every import. Guest-internal spans need cooperation: either the guest calls a host-provided tracing import, or it uses an observability interface that the host implements. In both cases, trace context flows into the guest so its spans join the request’s trace.
Step 1 — instrument the host around guest calls
Start with host-side spans, which need no guest changes. In a Rust host using the tracing crate with an OpenTelemetry exporter:
use tracing::{info_span, Instrument};
async fn handle_request(state: &AppState, req: Request) -> Response {
let span = info_span!("guest.invoke", guest = %state.guest_id, guest_version = %state.guest_version);
async {
let instance = info_span!("guest.instantiate")
.in_scope(|| state.pre.instantiate(&mut store))?; // InstancePre: fast instantiation
let resp = info_span!("guest.call", export = "handle")
.in_scope(|| instance.call_handle(&mut store, req))?;
Ok(resp)
}
.instrument(span)
.await
}
Record attributes that help diagnosis: guest id and version, fuel consumed (store.get_fuel()), memory size after the call, and whether the call trapped.
These alone often answer “was the guest slow, or was it something else?”.
Step 2 — propagate trace context into the guest
The incoming request’s traceparent header identifies the trace and parent span. Pass it, and the id of the host’s guest.call span, into the guest — as
a request header the guest can read, or through a host function that returns the current context:
interface tracing {
record span-context { trace-id: string, span-id: string }
current: func() -> span-context;
start-span: func(name: string, attributes: list>) -> u64;
end-span: func(id: u64);
}
The host implements current by returning the context of its own active span, so guest spans become its children.
Step 3 — create guest spans through the host
The guest opens and closes spans around its own operations; the host turns them into real OpenTelemetry spans with correct timing and parentage:
// guest
fn price_order(order: &Order) -> Result<Price, Error> {
let span = tracing_host::start_span("price", &[("items", &order.items.len().to_string())]);
let result = compute_price(order);
tracing_host::end_span(span);
result
}
// host implementation of start-span / end-span
fn start_span(&mut self, name: String, attrs: Vec<(String, String)>) -> u64 {
let mut span = self.tracer.start_with_context(name, &self.current_cx);
for (k, v) in attrs.into_iter().take(16) { span.set_attribute(KeyValue::new(k, v)); }
self.spans.insert(span) // returns a handle id
}
fn end_span(&mut self, id: u64) { if let Some(mut s) = self.spans.remove(id) { s.end(); } }
The host owns timing (so the guest cannot fake durations), caps attributes, and ends any spans the guest forgot to close when the call returns. A thin
wrapper crate can expose this as a tracing subscriber inside the guest so guest code uses ordinary #[instrument] attributes.
Step 4 — trace outbound calls automatically
When the guest makes HTTP requests through a host import (wasi:http/outgoing-handler or a custom fetch), the host creates a client span and injects
traceparent into the outgoing headers, so the downstream service’s spans join the same trace. The guest needs no changes: because the host mediates all
network access, every outbound call is traced by construction — one of the observability benefits of the capability model.
Step 5 — use platform support where it exists
Some platforms provide this out of the box. Spin emits OpenTelemetry traces for requests, component execution and outbound calls when
OTEL_EXPORTER_OTLP_ENDPOINT is set, and work on a standard WASI observability interface is under way so that guests can create spans without a
host-specific import. Check your host’s documentation before building your own; where nothing exists, the custom import above is a small amount of code.
Reading traces to find the bottleneck
A trace is only useful if someone reads it with a question in mind. For a slow request, look first at the widest span below the request span and follow
the widest child at each level — the critical path. If guest.instantiate is wide, compilation is probably not cached or the module has expensive start-up
work; precompile components and use pre-instantiation. If guest.call is wide but its guest spans are narrow, the time is in guest code that has no span
yet — add one or two coarse spans to narrow it down, or profile the guest. If an outbound span dominates, the dependency is slow and the guest is merely
waiting. Comparing the trace of a slow request with one of a typical request for the same route usually makes the difference obvious in seconds. Span
attributes such as fuel consumed and memory size let you search for expensive invocations across many traces rather than inspecting them one by one.
Sampling and cost
Tracing every request with every guest span can be expensive at high throughput. Use head-based sampling — decide at the start of the request whether to
record it, and propagate the decision so all services agree — with a rate that keeps volume manageable, and tail-based sampling in the collector to keep
all slow or failed requests regardless of the head decision. Inside the guest, prefer a handful of coarse spans — parse, compute, render — over spans in
tight loops; each span costs a boundary crossing into the host and an allocation in the exporter. When a sampled-out request reaches the guest, the host’s
start-span can return a no-op handle immediately, so unsampled requests pay almost nothing. Measure the overhead with sampling at 100% in a load test,
then choose a rate that keeps it below a percent or two of request time.
Expected output
A trace for a slow request shows the HTTP request span, a 0.2 ms instantiation span, the guest call, guest-created spans for parsing, pricing and rendering, and a 35 ms outbound inventory call continuing into the downstream service — making it obvious that the dependency, not the guest, was slow.
Gotchas
- Guest as a black box. Only one span for the whole guest. Add host spans and a tracing import.
- Lost context. Guest spans appear as separate traces. Propagate the parent context into the guest.
- Guest-controlled timing. Guests could report fake durations. Let the host time spans.
- Unclosed guest spans. Traps leave spans open. End them in the host after each call.
- Unsampled requests paying full cost. Return no-op span handles when the request is not sampled.
- Spans in hot loops. Each span crosses the boundary. Keep guest spans coarse.
Performance note
Host-side spans around instantiation and the guest call added about 4 µs per request. Each guest span through the host import cost about 2 µs. With five guest spans and 10% head sampling, tracing overhead averaged under 0.2% of request time.
Frequently Asked Questions
Can the guest export spans directly to a collector? Only if it has network access, which defeats the capability model. Let the host export.
Does this work in the browser? Yes — the page’s OpenTelemetry Web SDK can create spans around Wasm calls, and a JavaScript import can let the module create child spans.
How do I trace across several guests in one request? The host propagates the same context into each guest call, so all guest spans share the trace.
What attributes should guest spans carry? Low-cardinality descriptors — operation, item counts, cache hit or miss — never user data.
Can I trace instantiation in the browser?
Yes — wrap WebAssembly.instantiate in a span with the browser SDK; it pairs well with the real-user metrics described elsewhere in this topic.
Related
- Emitting structured logs from server-side Wasm — logs correlated with traces.
- Handling HTTP requests with wasi:http — where trace headers arrive.
- Cold-start characteristics of server-side Wasm — what instantiation spans show.
- Profiling Wasm hot paths with perf — going deeper than spans.
← Back to Observability & Error Reporting