Using the Emscripten File System API
This guide answers one task: make fopen, fread and the rest work inside a browser for code compiled
with Emscripten — choosing the right file system for the job, getting data in and out, and persisting what
should survive a reload.
Prerequisites
- [ ] Emscripten 3.1.60 or later, activated in your shell.
- [ ] C or C++ code that uses standard file APIs.
- [ ] A clear idea of which files are inputs, which are outputs and which must persist.
- [ ] A local server; the virtual file system’s data file will not load from
file://.
Four file systems, four jobs
Emscripten’s virtual file system is a layer under the standard library, and you mount different backends at different paths depending on what each is for.
MEMFS is the default and lives entirely in linear memory. It is fast, it is the right choice for
scratch files and outputs, and it disappears on reload.
IDBFS persists to IndexedDB, and is how a user’s work survives a refresh. It is not automatic: writes
land in memory and reach IndexedDB only when you call syncfs.
NODEFS maps a real directory and exists only under Node, which makes it useful for testing the same module outside a browser.
WORKERFS mounts File and Blob objects a user selected, without copying their contents into
linear memory — the option that makes a two-gigabyte input possible.
Preloading files at build time
The simplest way to give code the files it expects is to bundle them. Emscripten emits a .data file
alongside the module and populates the file system before main runs.
emcc app.c --preload-file assets@/assets -o app.js
FILE *f = fopen("/assets/config.txt", "r"); // just works
The @ separates the host path from the virtual path, which is worth using explicitly — without it the
files land at a path derived from your build directory, and the code then works on your machine and
nowhere else.
The whole .data file downloads before main runs, so this suits tens of megabytes at most. Beyond that,
fetch and write files yourself so the application starts while data arrives.
Writing files from JavaScript
For inputs the user provides, write them into the file system before calling into the module.
const bytes = new Uint8Array(await file.arrayBuffer());
Module.FS.mkdir('/work');
Module.FS.writeFile('/work/input.bin', bytes);
Module.ccall('process_file', 'number', ['string'], ['/work/input.bin']);
const out = Module.FS.readFile('/work/output.bin'); // Uint8Array
Module.FS.unlink('/work/input.bin'); // free the memory
Module.FS.unlink('/work/output.bin');
Note the unlink calls. MEMFS holds file contents in linear memory, so a file that is no longer needed
is memory you cannot reclaim any other way. A long-running session that writes files without removing
them grows until the tab fails.
Export what you need at build time, or these APIs will not be present:
emcc app.c -sEXPORTED_RUNTIME_METHODS='["FS","ccall","cwrap"]' -o app.js
Persisting with IDBFS
IDBFS makes a directory survive a reload, with one significant caveat: nothing is written to IndexedDB until you ask.
Module.FS.mkdir('/data');
Module.FS.mount(Module.IDBFS, {}, '/data');
// load what was persisted, before using the directory
await new Promise((res, rej) => Module.FS.syncfs(true, (e) => (e ? rej(e) : res())));
// … the module reads and writes /data …
// persist back — this is the call people forget
await new Promise((res, rej) => Module.FS.syncfs(false, (e) => (e ? rej(e) : res())));
The boolean argument is the direction: true populates the file system from IndexedDB, false writes it
back. Getting it backwards on startup wipes the user’s data, which is a sufficiently unpleasant bug to be
worth a comment in the code.
Sync at meaningful points — after a save action, on a timer, and on visibilitychange rather than on
beforeunload, which browsers increasingly do not guarantee will run.
Large inputs without copying
WORKERFS mounts File or Blob objects so the module reads them lazily rather than holding them in
linear memory. For a media or archive tool, this is the difference between handling a two-gigabyte input
and failing to allocate.
// in a worker
FS.mkdir('/in');
FS.mount(WORKERFS, { files: [file] }, '/in'); // file is a File from an input element
Module.ccall('scan_archive', 'number', ['string'], [`/in/${file.name}`]);
Reads are synchronous from the module’s point of view and backed by the browser’s file object underneath, so peak memory is whatever the module buffers rather than the file’s size. It is read-only and worker-only, which is the trade.
Lazy loading a file over the network
Between preloading everything and mounting a local file there is a third option: a virtual file backed by an HTTP range request, so the module reads parts of a remote file without downloading all of it.
Module.FS.createLazyFile('/data', 'atlas.bin', '/assets/atlas.bin', true, false);
// the module can now fopen("/data/atlas.bin") and seek anywhere in it
The file appears to the module as an ordinary read-only file. Behind it, each read issues a range request for the bytes the module asked for, which means a hundred-megabyte asset costs only the regions actually touched. A texture atlas, a font collection, a database index and a sprite sheet are all natural fits.
Two conditions have to hold. The server must honour Range requests and expose Accept-Ranges, and — in
the browser’s main thread — the underlying request is synchronous, which is deprecated and blocks. In a
worker the synchronous request is permitted and does not block anything the user sees, which is another
reason this pattern belongs off the main thread.
Measure before adopting it. If the module’s access pattern touches most of the file anyway, the range requests are slower than one sequential download and the lazy file has made things worse. It wins when access is genuinely sparse, which is exactly when a naive preload would have been wasteful.
Expected output
A working setup shows the module reading and writing paths as if they were real:
[fs] mounted IDBFS at /data
[fs] syncfs(populate) restored 3 files, 1.2 MB
[app] opened /assets/config.txt
[app] wrote /work/output.bin (412 kB)
[fs] syncfs(persist) wrote 4 files
console.log(Module.FS.readdir('/data'));
// [ '.', '..', 'project.db', 'settings.json', 'cache.bin' ]
console.log(Module.FS.stat('/data/project.db').size);
// 1048576
Gotchas
FS is not defined. Not exported. AddFStoEXPORTED_RUNTIME_METHODS.syncfsdirection reversed on startup. Overwrites persisted data with an empty file system.- Files never unlinked. MEMFS contents occupy
linear memoryfor the life of the instance. - Preloading hundreds of megabytes. All of it downloads before
mainruns. - WORKERFS on the main thread. Not available; it is worker-only by design.
- Relying on
beforeunloadto sync. Not reliably fired; sync onvisibilitychangeand periodically.
Performance note
Writing a 100 MB file with FS.writeFile took about 90 ms and added 100 MB to the heap permanently until
unlinked. Mounting the same file through WORKERFS took under a millisecond and added nothing, with reads
costing roughly the same per byte as reading from memory once the browser’s cache was warm. For anything
above a few tens of megabytes the mount is not an optimisation but the only workable approach.
Frequently Asked Questions
Can I use OPFS instead? Emscripten has an OPFS backend, and it is attractive for large persistent data because it avoids IndexedDB’s overhead. It carries the same worker requirement as OPFS generally — see persisting a Wasm database to OPFS for how that access model works.
Is there a quota on the persisted data?
Yes, the origin’s storage quota, shared with everything else the origin stores. Request persistent storage
and check navigator.storage.estimate() before writing large amounts, and handle a failed write rather
than assuming space is available.
How do I test file-handling code outside a browser? Build for Node and mount NODEFS over a fixture directory. The same C code then runs against real files in an ordinary test, which is far faster to iterate on than a browser test.
Can I avoid the file system entirely?
Often, and it is worth considering. If the C code’s only use of files is reading one input and writing one
output, exposing functions that take a pointer and a length is smaller, faster and removes the whole
virtual file system from the build — -sFILESYSTEM=0 then strips several kilobytes of runtime. The file
system earns its place when the code genuinely walks directories, seeks, or uses a library that insists on
paths.
Does the file system work with pthreads? Yes, with the usual caveats about which thread performs the operation. Keep file access on one thread rather than sharing handles across them.
Related
- Migrating legacy C code to WebAssembly — the wider porting problem.
- Building Emscripten projects with CMake — wiring preloaded assets into a real build.
- Working with datasets larger than memory — when even a mount is not enough.
The rule of thumb worth remembering: MEMFS for scratch, IDBFS for what must survive, WORKERFS for what is too big to copy, and no file system at all when the code can take a pointer instead.
← Back to C/C++ to Wasm with Emscripten