Writing a String Function in WAT

This page answers one task: you want to write a function in WAT that processes text — converts it to uppercase, measures it, logs it — and call it from JavaScript with real strings. WebAssembly has no string type, so you need a representation, a byte loop, and glue code on the JavaScript side.

Prerequisites

  • [ ] wat2wasm or wasm-tools parse.
  • [ ] Node.js or a browser.
  • [ ] WAT basics: locals, block/loop/br_if, memory loads and stores.

How strings live in Wasm

To WebAssembly, a string is bytes in linear memory. Two representations are common: pointer and length (two i32s: where the bytes start and how many there are) and NUL-terminated (a pointer; the string ends at the first zero byte, as in C). Pointer and length is safer and faster — no scanning to find the end, and the string may contain zero bytes — and it is what most toolchains and the Component Model use. JavaScript strings are UTF-16, so crossing the boundary means encoding to UTF-8 bytes with TextEncoder on the way in and decoding with TextDecoder on the way out.

A string's round trip through a WAT function JavaScript encodes a string into UTF-8 bytes directly in Wasm memory with TextEncoder.encodeInto. It calls the exported function with the pointer and byte length. The WAT function loops over bytes with load8_u and store8. JavaScript decodes the result bytes back into a string with TextDecoder. JS string UTF-16 encodeInto memory UTF-8 bytes at ptr call upper(ptr, len) byte loop in WAT TextDecoder. decode view over ptr, len JS string result

Step 1 — store text in a data segment

(module
  (import "env" "log" (func $log (param i32 i32)))
  (memory (export "memory") 1)
  (data (i32.const 0) "hello from wat")

The data segment writes the 14 bytes of “hello from wat” at address 0 during instantiation. Bytes after it are zero, which lets the NUL-terminated strlen below find the end. String literals in WAT can contain escapes such as \n, \00 and \e2\82\ac (the UTF-8 bytes of €).

Step 2 — write an uppercase function

  (func $upper (export "upper") (param $p i32) (param $len i32)
    (local $end i32) (local $c i32)
    (local.set $end (i32.add (local.get $p) (local.get $len)))
    (block $done
      (loop $next
        (br_if $done (i32.ge_u (local.get $p) (local.get $end)))
        (local.set $c (i32.load8_u (local.get $p)))
        (if (i32.lt_u (i32.sub (local.get $c) (i32.const 97)) (i32.const 26))
          (then (i32.store8 (local.get $p) (i32.sub (local.get $c) (i32.const 32)))))
        (local.set $p (i32.add (local.get $p) (i32.const 1)))
        (br $next))))

The loop walks from $p to $end, loads each byte unsigned, and lowers ‘a’–‘z’ by 32. The test (c − 97) <u 26 checks the range 97–122 with one comparison: values below 97 wrap around to large unsigned numbers and fail. The function modifies the string in place; an out-of-place version would take a destination pointer as well.

Step 3 — write strlen and call an import

  (func $strlen (export "strlen") (param $p i32) (result i32)
    (local $q i32)
    (local.set $q (local.get $p))
    (block $done
      (loop $next
        (br_if $done (i32.eqz (i32.load8_u (local.get $q))))
        (local.set $q (i32.add (local.get $q) (i32.const 1)))
        (br $next)))
    (i32.sub (local.get $q) (local.get $p)))

  (func (export "greet")
    (call $upper (i32.const 0) (i32.const 5))
    (call $log (i32.const 0) (call $strlen (i32.const 0)))))

greet uppercases the first word in place, measures the NUL-terminated string, and passes pointer and length to the imported log.

Pointer and length versus NUL-terminated strings Pointer and length strings know their size immediately and can contain zero bytes, and are the standard for Wasm interfaces. NUL-terminated strings need a scan to find the length and cannot contain zero bytes, but match C conventions. pointer + length length known instantly zero bytes allowed two values to pass Wasm interfaces NUL-terminated scan to find the end no embedded zeros one value to pass C interop

Step 4 — call it from JavaScript

let memory;
const { instance } = await WebAssembly.instantiate(bytes, { env: {
  log: (ptr, len) => console.log(new TextDecoder().decode(new Uint8Array(memory.buffer, ptr, len))),
}});
memory = instance.exports.memory;
instance.exports.greet();                                // HELLO from wat

const { written } = new TextEncoder().encodeInto("Grüße, wasm!", new Uint8Array(memory.buffer, 2048));
instance.exports.upper(2048, written);
new TextDecoder().decode(new Uint8Array(memory.buffer, 2048, written));   // "GRüßE, WASM!"

encodeInto writes UTF-8 straight into Wasm memory and returns the number of bytes written — 14 here, although the string has 12 characters, because ü and ß take two bytes each. The log import reads memory lazily, so it works even though memory is assigned after instantiation.

Step 5 — handle UTF-8 correctly

The ASCII-only uppercase leaves ü and ß unchanged, which is correct for an ASCII function but not full Unicode case mapping. Byte-wise operations that only touch ASCII bytes are always safe on UTF-8, because every byte of a multi-byte character is ≥ 128 and can never be mistaken for ASCII. Anything that needs character semantics — Unicode case mapping, counting characters, slicing by character — needs a UTF-8 decoder in the module or should stay in JavaScript.

Allocating space for strings

The examples use fixed addresses (0 and 2048), which is fine for experiments. Real modules export an allocator — even a simple bump allocator that hands out increasing addresses from a global — and JavaScript asks it for space before encoding. Reserve enough: encodeInto stops when the view is full, and a string’s UTF-8 size can be up to three times its JavaScript length.

A bump allocator for strings

A real interface needs somewhere to put incoming strings. The smallest allocator that works is a global pointer that only moves forward, plus a reset function the host calls between batches:

  (global $heap (mut i32) (i32.const 4096))
  (func (export "alloc") (param $size i32) (result i32)
    (local $p i32)
    (local.set $p (global.get $heap))
    (global.set $heap (i32.and (i32.add (i32.add (local.get $p) (local.get $size)) (i32.const 7))
                               (i32.const -8)))                       ;; keep 8-byte alignment
    (local.get $p))
  (func (export "reset") (global.set $heap (i32.const 4096)))

JavaScript then asks for space before encoding: const ptr = alloc(s.length * 3), encodeInto(s, new Uint8Array(memory.buffer, ptr, s.length * 3)), and passes ptr with the written count. Starting the heap at 4096 keeps it clear of the data segment and the fixed output area. The allocator never checks for running out of memory; a production version compares against memory.size × 65536 and calls memory.grow when needed — after which JavaScript must recreate its views.

Comparing two strings

Equality is another loop over bytes, but with pointer-and-length strings the first check is free: different lengths mean different strings. Only equal lengths need the byte loop, which can stop at the first mismatch. Return 1 or 0 as an i32; JavaScript treats it as a boolean. For ordering (as in sorting), compare bytes as unsigned values — UTF-8 byte order matches Unicode code point order, so a byte-wise comparison sorts by code point, which is what most code wants unless it needs locale-aware collation.

Debugging string code

Most bugs in hand-written string functions are off-by-one bounds or a forgotten unsigned load. Two habits catch them quickly: test with strings of length 0, 1 and one byte longer than any internal buffer, and include at least one non-ASCII character so signed loads show up as wrong results. When output looks wrong, dump the raw bytes from JavaScript (new Uint8Array(memory.buffer, ptr, len)) before decoding; seeing the bytes usually makes the mistake obvious.

Returning a new string

Functions that produce a string of a different length — trimming, replacing, formatting a number — cannot work in place. Allocate the output with the bump allocator, write the bytes, and return the pointer and length. With multi-value, a WAT function can declare (result i32 i32) and JavaScript receives an array [ptr, len]; without it, write the length to a known address or an exported global. Either way, decode the result immediately and call reset once the batch is done, so temporary strings do not accumulate.

Expected output

The module assembles; greet() logs “HELLO from wat”; encoding “Grüße, wasm!” writes 14 bytes; upper turns it into “GRüßE, WASM!”; and strlen(0) returns 14 for the data segment.

Gotchas

  • Signed byte loads. i32.load8_s makes bytes ≥ 128 negative. Use load8_u.
  • Counting characters as bytes. UTF-8 lengths differ from JavaScript lengths.
  • Fixed addresses in real code. Strings overwrite each other. Export an allocator.
  • Too-small destination views. encodeInto truncates silently. Check read and written.
  • Cached views after growth. Recreate views after memory grows.
  • An allocator that never grows memory. Large inputs overrun the heap. Check size and call memory.grow.

Performance note

Uppercasing a 1 MB ASCII buffer with the byte loop took about 1.2 ms in Chrome, compared with about 3 ms for toUpperCase on the equivalent JavaScript string plus encoding; the encode and decode steps dominate for short strings.

Uppercasing 1 MB of ASCII text Milliseconds to uppercase one megabyte of ASCII text with JavaScript toUpperCase, with the WAT byte loop excluding encoding, and with the WAT byte loop including TextEncoder and TextDecoder. ms for 1 MB JS toUpperCase 3 ms WAT loop only 1.2 ms WAT loop + encode + decode 4.1 ms

Frequently Asked Questions

Can WAT functions return a string? Return a pointer and length (two results with multi-value, or write them to memory) and decode in JavaScript.

Why not use UTF-16 in memory? You can, but UTF-8 is the norm for Wasm toolchains and smaller for mostly-ASCII text.

Is there a string type in Wasm? Not in the core spec; the stringref and JS string builtins proposals add one.

How do I embed non-ASCII text in a data segment? Write the UTF-8 bytes directly or as \hh escapes.

Does byte-wise comparison sort UTF-8 correctly? It sorts by Unicode code point, which matches most needs except locale-aware collation.

How much space should I allocate for an encoded string? Three bytes per JavaScript string unit is the safe upper bound for UTF-8.

How does JavaScript receive two results from one function? With multi-value, an exported function returning two values gives JavaScript an array of both.

← Back to WebAssembly Text Format (WAT) Basics