Tracing Wasm Requests with OpenTelemetry

This page answers one task: a request passes through a host and one or more WebAssembly guests, and when it is slow you want a trace that shows exactly where the time went — in the host, in instantiation, in the guest’s own logic, or in calls the guest made to other services.

Prerequisites

  • [ ] A host that runs Wasm guests (an Axum or similar service embedding wasmtime, a Spin app, or a Node service using Wasm).
  • [ ] An OpenTelemetry collector or tracing backend (Jaeger, Tempo, Honeycomb, Datadog) and the OpenTelemetry SDK for the host’s language.
  • [ ] Familiarity with spans, trace context and the traceparent header.

What a useful trace looks like

A distributed trace is a tree of spans, each a timed operation with attributes, linked by a trace id propagated between services. For a host running Wasm, the interesting spans are: the incoming request; compiling or fetching the component (if not cached); instantiating it; the call into the guest; operations inside the guest; and outbound requests the guest makes through the host. Without Wasm-aware instrumentation, the guest is a black box — one long span labelled “handler” — and you cannot tell whether a slow request was slow because of instantiation, guest computation, or a downstream API the guest called.

The host can create most spans by itself, because it performs instantiation and mediates every import. Guest-internal spans need cooperation: either the guest calls a host-provided tracing import, or it uses an observability interface that the host implements. In both cases, trace context flows into the guest so its spans join the request’s trace.

Spans for one request through a Wasm host The request span contains a span for instantiating the guest and a span for the guest call. Inside the guest call, guest-created spans cover its own work, and outbound HTTP made through the host appears as a child span with context propagated to the downstream service. HTTP request span host — 48 ms instantiate guest host — 0.2 ms guest call: handle host-created — 46 ms guest spans: parse, price, render via host import — 9 ms outbound GET /inventory host-mediated — 35 ms

Step 1 — instrument the host around guest calls

Start with host-side spans, which need no guest changes. In a Rust host using the tracing crate with an OpenTelemetry exporter:

use tracing::{info_span, Instrument};

async fn handle_request(state: &AppState, req: Request) -> Response {
    let span = info_span!("guest.invoke", guest = %state.guest_id, guest_version = %state.guest_version);
    async {
        let instance = info_span!("guest.instantiate")
            .in_scope(|| state.pre.instantiate(&mut store))?;          // InstancePre: fast instantiation
        let resp = info_span!("guest.call", export = "handle")
            .in_scope(|| instance.call_handle(&mut store, req))?;
        Ok(resp)
    }
    .instrument(span)
    .await
}

Record attributes that help diagnosis: guest id and version, fuel consumed (store.get_fuel()), memory size after the call, and whether the call trapped. These alone often answer “was the guest slow, or was it something else?”.

Step 2 — propagate trace context into the guest

The incoming request’s traceparent header identifies the trace and parent span. Pass it, and the id of the host’s guest.call span, into the guest — as a request header the guest can read, or through a host function that returns the current context:

interface tracing {
  record span-context { trace-id: string, span-id: string }
  current: func() -> span-context;
  start-span: func(name: string, attributes: list>) -> u64;
  end-span: func(id: u64);
}

The host implements current by returning the context of its own active span, so guest spans become its children.

Step 3 — create guest spans through the host

The guest opens and closes spans around its own operations; the host turns them into real OpenTelemetry spans with correct timing and parentage:

// guest
fn price_order(order: &Order) -> Result<Price, Error> {
    let span = tracing_host::start_span("price", &[("items", &order.items.len().to_string())]);
    let result = compute_price(order);
    tracing_host::end_span(span);
    result
}
// host implementation of start-span / end-span
fn start_span(&mut self, name: String, attrs: Vec<(String, String)>) -> u64 {
    let mut span = self.tracer.start_with_context(name, &self.current_cx);
    for (k, v) in attrs.into_iter().take(16) { span.set_attribute(KeyValue::new(k, v)); }
    self.spans.insert(span)                                   // returns a handle id
}
fn end_span(&mut self, id: u64) { if let Some(mut s) = self.spans.remove(id) { s.end(); } }

The host owns timing (so the guest cannot fake durations), caps attributes, and ends any spans the guest forgot to close when the call returns. A thin wrapper crate can expose this as a tracing subscriber inside the guest so guest code uses ordinary #[instrument] attributes.

A guest span created through the host The host starts the guest call span and invokes the guest. The guest asks the host to start a span named price; the host creates it as a child of the call span and returns a handle. When the guest ends the span, the host records its duration and exports it with the rest of the trace. host tracer guest collector call handle() inside span guest.call start-span("price") → id 7 end-span(7) export spans with trace id

Step 4 — trace outbound calls automatically

When the guest makes HTTP requests through a host import (wasi:http/outgoing-handler or a custom fetch), the host creates a client span and injects traceparent into the outgoing headers, so the downstream service’s spans join the same trace. The guest needs no changes: because the host mediates all network access, every outbound call is traced by construction — one of the observability benefits of the capability model.

Step 5 — use platform support where it exists

Some platforms provide this out of the box. Spin emits OpenTelemetry traces for requests, component execution and outbound calls when OTEL_EXPORTER_OTLP_ENDPOINT is set, and work on a standard WASI observability interface is under way so that guests can create spans without a host-specific import. Check your host’s documentation before building your own; where nothing exists, the custom import above is a small amount of code.

Reading traces to find the bottleneck

A trace is only useful if someone reads it with a question in mind. For a slow request, look first at the widest span below the request span and follow the widest child at each level — the critical path. If guest.instantiate is wide, compilation is probably not cached or the module has expensive start-up work; precompile components and use pre-instantiation. If guest.call is wide but its guest spans are narrow, the time is in guest code that has no span yet — add one or two coarse spans to narrow it down, or profile the guest. If an outbound span dominates, the dependency is slow and the guest is merely waiting. Comparing the trace of a slow request with one of a typical request for the same route usually makes the difference obvious in seconds. Span attributes such as fuel consumed and memory size let you search for expensive invocations across many traces rather than inspecting them one by one.

Sampling and cost

Tracing every request with every guest span can be expensive at high throughput. Use head-based sampling — decide at the start of the request whether to record it, and propagate the decision so all services agree — with a rate that keeps volume manageable, and tail-based sampling in the collector to keep all slow or failed requests regardless of the head decision. Inside the guest, prefer a handful of coarse spans — parse, compute, render — over spans in tight loops; each span costs a boundary crossing into the host and an allocation in the exporter. When a sampled-out request reaches the guest, the host’s start-span can return a no-op handle immediately, so unsampled requests pay almost nothing. Measure the overhead with sampling at 100% in a load test, then choose a rate that keeps it below a percent or two of request time.

Expected output

A trace for a slow request shows the HTTP request span, a 0.2 ms instantiation span, the guest call, guest-created spans for parsing, pricing and rendering, and a 35 ms outbound inventory call continuing into the downstream service — making it obvious that the dependency, not the guest, was slow.

Gotchas

  • Guest as a black box. Only one span for the whole guest. Add host spans and a tracing import.
  • Lost context. Guest spans appear as separate traces. Propagate the parent context into the guest.
  • Guest-controlled timing. Guests could report fake durations. Let the host time spans.
  • Unclosed guest spans. Traps leave spans open. End them in the host after each call.
  • Unsampled requests paying full cost. Return no-op span handles when the request is not sampled.
  • Spans in hot loops. Each span crosses the boundary. Keep guest spans coarse.

Performance note

Host-side spans around instantiation and the guest call added about 4 µs per request. Each guest span through the host import cost about 2 µs. With five guest spans and 10% head sampling, tracing overhead averaged under 0.2% of request time.

Tracing overhead per request Microseconds added per request by host-side spans only, by host spans plus five guest spans for sampled requests, and averaged across all requests with ten percent head sampling. microseconds per request host spans only 4 µs host + 5 guest spans (sampled) 14 µs average with 10% sampling 5 µs

Frequently Asked Questions

Can the guest export spans directly to a collector? Only if it has network access, which defeats the capability model. Let the host export.

Does this work in the browser? Yes — the page’s OpenTelemetry Web SDK can create spans around Wasm calls, and a JavaScript import can let the module create child spans.

How do I trace across several guests in one request? The host propagates the same context into each guest call, so all guest spans share the trace.

What attributes should guest spans carry? Low-cardinality descriptors — operation, item counts, cache hit or miss — never user data.

Can I trace instantiation in the browser? Yes — wrap WebAssembly.instantiate in a span with the browser SDK; it pairs well with the real-user metrics described elsewhere in this topic.

← Back to Observability & Error Reporting