Writing a String Function in WAT
This page answers one task: you want to write a function in WAT that processes text — converts it to uppercase, measures it, logs it — and call it from JavaScript with real strings. WebAssembly has no string type, so you need a representation, a byte loop, and glue code on the JavaScript side.
Prerequisites
- [ ]
wat2wasmorwasm-tools parse. - [ ] Node.js or a browser.
- [ ] WAT basics: locals,
block/loop/br_if, memory loads and stores.
How strings live in Wasm
To WebAssembly, a string is bytes in linear memory. Two representations are common: pointer and length (two i32s: where the bytes start and how many there
are) and NUL-terminated (a pointer; the string ends at the first zero byte, as in C). Pointer and length is safer and faster — no scanning to find the end,
and the string may contain zero bytes — and it is what most toolchains and the Component Model use. JavaScript strings are UTF-16, so crossing the boundary means
encoding to UTF-8 bytes with TextEncoder on the way in and decoding with TextDecoder on the way out.
Step 1 — store text in a data segment
(module
(import "env" "log" (func $log (param i32 i32)))
(memory (export "memory") 1)
(data (i32.const 0) "hello from wat")
The data segment writes the 14 bytes of “hello from wat” at address 0 during instantiation. Bytes after it are zero, which lets the NUL-terminated strlen
below find the end. String literals in WAT can contain escapes such as \n, \00 and \e2\82\ac (the UTF-8 bytes of €).
Step 2 — write an uppercase function
(func $upper (export "upper") (param $p i32) (param $len i32)
(local $end i32) (local $c i32)
(local.set $end (i32.add (local.get $p) (local.get $len)))
(block $done
(loop $next
(br_if $done (i32.ge_u (local.get $p) (local.get $end)))
(local.set $c (i32.load8_u (local.get $p)))
(if (i32.lt_u (i32.sub (local.get $c) (i32.const 97)) (i32.const 26))
(then (i32.store8 (local.get $p) (i32.sub (local.get $c) (i32.const 32)))))
(local.set $p (i32.add (local.get $p) (i32.const 1)))
(br $next))))
The loop walks from $p to $end, loads each byte unsigned, and lowers ‘a’–‘z’ by 32. The test (c − 97) <u 26 checks the range 97–122 with one comparison:
values below 97 wrap around to large unsigned numbers and fail. The function modifies the string in place; an out-of-place version would take a destination
pointer as well.
Step 3 — write strlen and call an import
(func $strlen (export "strlen") (param $p i32) (result i32)
(local $q i32)
(local.set $q (local.get $p))
(block $done
(loop $next
(br_if $done (i32.eqz (i32.load8_u (local.get $q))))
(local.set $q (i32.add (local.get $q) (i32.const 1)))
(br $next)))
(i32.sub (local.get $q) (local.get $p)))
(func (export "greet")
(call $upper (i32.const 0) (i32.const 5))
(call $log (i32.const 0) (call $strlen (i32.const 0)))))
greet uppercases the first word in place, measures the NUL-terminated string, and passes pointer and length to the imported log.
Step 4 — call it from JavaScript
let memory;
const { instance } = await WebAssembly.instantiate(bytes, { env: {
log: (ptr, len) => console.log(new TextDecoder().decode(new Uint8Array(memory.buffer, ptr, len))),
}});
memory = instance.exports.memory;
instance.exports.greet(); // HELLO from wat
const { written } = new TextEncoder().encodeInto("Grüße, wasm!", new Uint8Array(memory.buffer, 2048));
instance.exports.upper(2048, written);
new TextDecoder().decode(new Uint8Array(memory.buffer, 2048, written)); // "GRüßE, WASM!"
encodeInto writes UTF-8 straight into Wasm memory and returns the number of bytes written — 14 here, although the string has 12 characters, because ü and ß
take two bytes each. The log import reads memory lazily, so it works even though memory is assigned after instantiation.
Step 5 — handle UTF-8 correctly
The ASCII-only uppercase leaves ü and ß unchanged, which is correct for an ASCII function but not full Unicode case mapping. Byte-wise operations that only touch ASCII bytes are always safe on UTF-8, because every byte of a multi-byte character is ≥ 128 and can never be mistaken for ASCII. Anything that needs character semantics — Unicode case mapping, counting characters, slicing by character — needs a UTF-8 decoder in the module or should stay in JavaScript.
Allocating space for strings
The examples use fixed addresses (0 and 2048), which is fine for experiments. Real modules export an allocator — even a simple bump allocator that hands out
increasing addresses from a global — and JavaScript asks it for space before encoding. Reserve enough: encodeInto stops when the view is full, and a
string’s UTF-8 size can be up to three times its JavaScript length.
A bump allocator for strings
A real interface needs somewhere to put incoming strings. The smallest allocator that works is a global pointer that only moves forward, plus a reset function the host calls between batches:
(global $heap (mut i32) (i32.const 4096))
(func (export "alloc") (param $size i32) (result i32)
(local $p i32)
(local.set $p (global.get $heap))
(global.set $heap (i32.and (i32.add (i32.add (local.get $p) (local.get $size)) (i32.const 7))
(i32.const -8))) ;; keep 8-byte alignment
(local.get $p))
(func (export "reset") (global.set $heap (i32.const 4096)))
JavaScript then asks for space before encoding: const ptr = alloc(s.length * 3), encodeInto(s, new Uint8Array(memory.buffer, ptr, s.length * 3)), and passes
ptr with the written count. Starting the heap at 4096 keeps it clear of the data segment and the fixed output area. The allocator never checks for running
out of memory; a production version compares against memory.size × 65536 and calls memory.grow when needed — after which JavaScript must recreate its views.
Comparing two strings
Equality is another loop over bytes, but with pointer-and-length strings the first check is free: different lengths mean different strings. Only equal lengths
need the byte loop, which can stop at the first mismatch. Return 1 or 0 as an i32; JavaScript treats it as a boolean. For ordering (as in sorting), compare
bytes as unsigned values — UTF-8 byte order matches Unicode code point order, so a byte-wise comparison sorts by code point, which is what most code wants unless it
needs locale-aware collation.
Debugging string code
Most bugs in hand-written string functions are off-by-one bounds or a forgotten unsigned load. Two habits catch them quickly: test with strings of length 0, 1
and one byte longer than any internal buffer, and include at least one non-ASCII character so signed loads show up as wrong results. When output looks wrong,
dump the raw bytes from JavaScript (new Uint8Array(memory.buffer, ptr, len)) before decoding; seeing the bytes usually makes the mistake obvious.
Returning a new string
Functions that produce a string of a different length — trimming, replacing, formatting a number — cannot work in place. Allocate the output with the bump
allocator, write the bytes, and return the pointer and length. With multi-value, a WAT function can declare (result i32 i32) and JavaScript receives an array
[ptr, len]; without it, write the length to a known address or an exported global. Either way, decode the result immediately and call reset once the batch
is done, so temporary strings do not accumulate.
Expected output
The module assembles; greet() logs “HELLO from wat”; encoding “Grüße, wasm!” writes 14 bytes; upper turns it into “GRüßE, WASM!”; and strlen(0) returns
14 for the data segment.
Gotchas
- Signed byte loads.
i32.load8_smakes bytes ≥ 128 negative. Useload8_u. - Counting characters as bytes. UTF-8 lengths differ from JavaScript lengths.
- Fixed addresses in real code. Strings overwrite each other. Export an allocator.
- Too-small destination views.
encodeIntotruncates silently. Checkreadandwritten. - Cached views after growth. Recreate views after memory grows.
- An allocator that never grows memory. Large inputs overrun the heap. Check size and call
memory.grow.
Performance note
Uppercasing a 1 MB ASCII buffer with the byte loop took about 1.2 ms in Chrome, compared with about 3 ms for toUpperCase on the equivalent JavaScript string
plus encoding; the encode and decode steps dominate for short strings.
Frequently Asked Questions
Can WAT functions return a string? Return a pointer and length (two results with multi-value, or write them to memory) and decode in JavaScript.
Why not use UTF-16 in memory? You can, but UTF-8 is the norm for Wasm toolchains and smaller for mostly-ASCII text.
Is there a string type in Wasm? Not in the core spec; the stringref and JS string builtins proposals add one.
How do I embed non-ASCII text in a data segment?
Write the UTF-8 bytes directly or as \hh escapes.
Does byte-wise comparison sort UTF-8 correctly? It sorts by Unicode code point, which matches most needs except locale-aware collation.
How much space should I allocate for an encoded string? Three bytes per JavaScript string unit is the safe upper bound for UTF-8.
How does JavaScript receive two results from one function? With multi-value, an exported function returning two values gives JavaScript an array of both.
Related
- Reading and writing memory in WAT — loads and stores.
- Writing loops and branches in WAT — the loop pattern.
- Importing JavaScript functions into WAT — the log import.
- Exporting memory and globals from WAT — memory access.
← Back to WebAssembly Text Format (WAT) Basics