Reading EXIF and Image Metadata with Wasm

This page answers one task: an application handles user photos — a gallery, an upload form, a photo organiser — and needs metadata such as capture date, camera model, orientation, dimensions and GPS location, quickly and for many files, without decoding the images. You also want to remove sensitive metadata before uploading.

Prerequisites

  • [ ] Image files from a file input, drag-and-drop or the File System Access API.
  • [ ] A metadata parser compiled to Wasm (for example the Rust kamadak-exif or nom-exif crates, or a C library such as libexif or Exiv2 built with Emscripten).
  • [ ] A worker for batch processing.

Where metadata lives

Photo metadata is stored in the file’s container structure, separately from the compressed pixels. In JPEG, EXIF sits in an APP1 segment near the start of the file, as a TIFF-structured block of tagged fields; XMP (XML-based) may follow in another APP1 segment. HEIC and AVIF use the ISO base media file format (boxes), with EXIF in an item referenced from the meta box. PNG stores text chunks and, in newer files, an eXIf chunk; WebP stores EXIF and XMP in RIFF chunks.

Because metadata is near the start (JPEG, PNG, WebP) or referenced from a small index (HEIC), a parser needs only a small part of the file — often the first 64 KB. Reading that slice instead of the whole file, and parsing it in Wasm, makes metadata extraction take milliseconds per photo even for 20 MB files.

Where metadata lives in common image formats JPEG stores EXIF and XMP in APP1 segments near the start of the file. PNG uses text chunks and an eXIf chunk. WebP uses EXIF and XMP chunks in its RIFF container. HEIC and AVIF reference an EXIF item from the meta box. In all cases the pixels are stored separately and need not be decoded. JPEG APP1 Exif + APP1 XMP at start PNG eXIf + tEXt/iTXt chunks WebP EXIF + XMP RIFF chunks HEIC / AVIF meta box → Exif item pixels not needed for metadata

Step 1 — read only the header

async function readHead(file, bytes = 128 * 1024) {
  return new Uint8Array(await file.slice(0, Math.min(bytes, file.size)).arrayBuffer());
}

For HEIC, the EXIF item’s offset is listed in the meta box near the start; the parser may need a second read at that offset. Design the Wasm API so it can say “I need bytes at offset X, length Y” and the JavaScript side reads that slice — that keeps memory small for large files.

Step 2 — parse in Wasm

use wasm_bindgen::prelude::*;
use serde::Serialize;

#[derive(Serialize)]
pub struct PhotoMeta {
    pub taken_at: Option<String>,   // DateTimeOriginal + OffsetTimeOriginal if present
    pub make: Option<String>,
    pub model: Option<String>,
    pub orientation: Option<u16>,
    pub width: Option<u32>,
    pub height: Option<u32>,
    pub gps: Option<(f64, f64)>,
}

#[wasm_bindgen]
pub fn read_meta(head: &[u8]) -> Result<JsValue, JsError> {
    let exif = exif::Reader::new()
        .read_from_container(&mut std::io::Cursor::new(head))
        .map_err(|e| JsError::new(&format!("no readable EXIF: {e}")))?;
    let meta = extract(&exif);      // map tags: DateTimeOriginal, Make, Model, Orientation, GPS*
    Ok(serde_wasm_bindgen::to_value(&meta)?)
}

Parsing is pure byte manipulation, which compiles to small, fast Wasm, and the same crate runs natively for server-side processing, so client and server agree on results.

Step 3 — interpret the values correctly

A few fields need care. Orientation (values 1–8) tells viewers how to rotate or flip the stored pixels; browsers apply it automatically when displaying images (image-orientation: from-image is the default), but code that draws pixels itself must apply it. Capture time is stored as local time without a zone (DateTimeOriginal), with an optional OffsetTimeOriginal; treat it as a local wall-clock time unless the offset is present. GPS coordinates are stored as degrees, minutes and seconds rationals plus hemisphere references, which must be combined into signed decimal degrees.

EXIF fields that need interpretation Orientation values one to eight describe rotation and mirroring that viewers must apply. DateTimeOriginal is local time without a zone unless OffsetTimeOriginal is present. GPS latitude and longitude are degree-minute-second rationals with N S E W references that must be converted to signed decimals. field stored as interpret as Orientation 1–8 rotation + mirror to apply DateTimeOriginal "2024:06:01 14:03:22" local time; zone only via OffsetTimeOriginal GPSLatitude + Ref 3 rationals + N/S signed decimal degrees PixelXDimension integer may differ from actual pixels

Step 4 — strip location before upload

Photos carry GPS coordinates that reveal where they were taken — often a user’s home. Remove them (or all EXIF) before uploading unless the user explicitly wants location shared. For JPEG, removing metadata does not require re-encoding: rewrite the file without the APP1 segments (or with a filtered EXIF block) and keep the compressed pixel data byte-for-byte, which preserves quality and is fast. A Wasm function can do this in a single pass:

#[wasm_bindgen]
pub fn strip_jpeg_metadata(input: &[u8], keep_orientation: bool) -> Vec<u8> {
    // copy SOI, skip APP1 (Exif/XMP) segments, optionally write a minimal Exif with Orientation, copy the rest unchanged
    jpeg::rewrite_without_app1(input, keep_orientation)
}

Keep the orientation (or apply it to the pixels) — stripping it makes portrait photos appear sideways after upload. For HEIC and other formats, the server may need to convert anyway; strip there if client-side rewriting is not supported.

Step 5 — process batches in a worker

For an import of thousands of photos, post files to a worker, read headers in parallel batches, parse in Wasm, and send results back in groups. The bottleneck is usually file reading, not parsing; reading only headers keeps it fast even for large libraries.

Handling malformed and hostile files

Metadata parsers read untrusted input from files of unknown origin, and parsers in C have a history of memory-safety vulnerabilities. In Wasm, a parser bug cannot escape the module’s memory, but it can still trap, hang or return garbage. Run parsing in a worker, use a memory-safe parser where possible (Rust crates), bound the bytes it may read, and catch errors per file so one corrupt photo does not stop a batch. Treat extracted strings (camera model, comments) as untrusted text when displaying them.

Sorting and grouping by capture time

Photo organisers group pictures into days and events by capture time, and getting this right across devices is harder than it looks. Many cameras and older phones record only local time without an offset, so a photo taken at 23:30 in Tokyo and one taken at 23:30 in New York have the same DateTimeOriginal. Newer phones write OffsetTimeOriginal, which allows exact ordering. A practical rule: sort by the instant when an offset is present; otherwise treat the time as local to wherever it was taken and group by calendar date as written, which matches users’ expectations (“photos from Saturday”). When EXIF dates are missing — screenshots, edited exports, messaging apps that strip metadata — fall back to dates in the file name if they follow a known pattern, and only then to the file’s modification time, which often reflects when it was copied rather than taken. Show users when a date was inferred, so they understand why a picture appears where it does.

Deduplication with metadata

Large imports contain duplicates: the same photo synced from two devices, exported twice, or resized by a messaging app. Metadata gives a cheap first pass — identical capture time, camera model and dimensions strongly suggest the same picture — and a perceptual hash of a small thumbnail confirms it, still without decoding full-size images in most cases. Running both in a worker over thousands of files takes seconds, and presenting likely duplicates for the user to confirm is far more pleasant than cleaning up a library afterwards.

Keeping location by choice

Some users want location — for a map view or travel albums. Make sharing location an explicit choice per upload or per album rather than the default, and keep the stripping code path tested so a refactor does not quietly start uploading coordinates.

Expected output

Selecting 2,000 photos extracts capture time, camera, orientation, dimensions and GPS for all of them in about 4 seconds, reading 128 KB per file; times with offsets are shown in local time correctly; uploads strip GPS while keeping orientation, with pixel data unchanged; and three corrupt files are reported individually without stopping the batch.

Gotchas

  • Reading whole files for metadata. Slow and memory-hungry. Read the header slice.
  • Dropping orientation when stripping. Photos turn sideways. Keep or apply it.
  • Treating capture time as UTC. It is local wall-clock time. Use the offset if present.
  • Re-encoding JPEGs to remove metadata. Quality loss. Rewrite segments instead.
  • Trusting metadata strings. They are user-controlled. Escape when displaying.
  • Location stripping without tests. A refactor can start uploading GPS. Test the stripping path.

Performance note

Reading a 128 KB header and parsing EXIF took about 1.5 ms per photo; decoding the full image to read dimensions, as a naive approach might, took about 120 ms for a 12-megapixel JPEG.

Time per photo to obtain metadata Milliseconds per 12-megapixel JPEG to obtain capture time and dimensions by reading a 128 KB header and parsing EXIF in Wasm, and by decoding the full image. ms per photo header slice + Wasm EXIF parse 1.5 ms full image decode 120 ms

Frequently Asked Questions

Can createImageBitmap give me dimensions? Only by decoding the image, which is far slower than reading headers.

Does the browser strip EXIF on upload? No — files upload as-is. Strip metadata yourself if needed.

Is XMP worth parsing? For ratings, keywords and edit history from photo apps, yes; it is XML in a segment.

Can I write metadata too? Yes, with a parser that supports writing; rewrite segments without re-encoding pixels.

What if a photo has no EXIF capture date? Fall back to a date in the file name if it follows a known pattern, then to the file’s modification time, and mark the date as inferred.

How can duplicates be found quickly in a large import? Compare capture time, camera and dimensions first, then confirm with a perceptual hash of a small thumbnail.

← Back to Media Processing & Codecs in Wasm