Design

LivePhoto.

Print a moment. Scan it. Watch it come alive. A mobile web app that plays a video on top of its own printed photo — tracked to the paper, live through the camera, no special hardware. Every piece below is rebuilt from the app source.

LivePhoto makes a printed photograph play. You pick a video, print one of its frames, and from then on pointing the app's camera at that print overlays the video exactly on the paper — matched to its position, tracked through perspective, picking up the moment the print re-enters the frame. Xiaomi ships this as a phone feature backed by their gallery; this is the same experience rebuilt on the open web, end to end: one page you open, one camera permission, and a photo on your fridge that turns into the moment it came from.

Scanning a photo

Open the scanner and every photo in the album is a candidate at once — each item's compiled target is downloaded and merged into a single tracker, so there is nothing to choose before you point. A frosted capsule narrates the wait like a dynamic island: a dithered pixel grid while targets download, a rolling dot-matrix wave while searching, and on lock it expands into a now-playing chip — thumbnail, name, a four-bar equalizer.

The lock itself is engineered for the first impression: the chosen frame appears pinned to the print instantly as a poster, and the video takes over from its own first frame — the printed image and frame zero are the same picture, so the handoff is invisible. Decoders are warmed at startup (a muted play, parked back at frame 0) so the first lock shows motion immediately instead of a decode stall. Lose the print and re-find it within 5 seconds and playback resumes where it was; longer, and it restarts from the top. The clip is prefetched as a file the moment a lock happens, so the share button can hand it to the iOS share sheet inside the tap's user gesture — awaiting a network fetch there would burn the gesture and iOS would refuse the sheet.

Making one

Creating is a four-step editorial flow, and its calendar is honest: the video starts uploading in the background the moment you pick it, so the final Save is a metadata write that lands in a blink. Two ways to make a target: print a frame straight from the video — the recommended path, because a digital frame is noise-free — or enroll a photo you already have. The second path is a small document scanner: a gradient-restricted Hough transform finds the print's quadrilateral nine times a second, an outline settles over it, and after 1.4 seconds of stability the capture takes itself, rectifies the perspective, and trims Instax-style borders. It asks for three captures — straight on, then tilted ±30°, coached by a Face ID-style gauge with haptic ticks — so recognition survives being looked at from an angle.

Then the phone does the heavy lifting that would normally be a server's job: the MindAR compiler runs in the browser at enrollment, distilling the image into a feature fingerprint. It has to — the compiler needs tfjs and WebGL, which the Workers runtime doesn't have — and it turns out to be the right architecture anyway: scanning devices just download the precompiled result. The compile ends in a verdict from the matcher's own arithmetic: 100+ feature points is Good, 40 is Fair, fewer is told to pick a busier photo. Every step of the half-built session persists to IndexedDB, so a phone call mid-enrollment leaves the session waiting for twelve hours.

How it works

The frontend is deliberately plain: vanilla TypeScript views over a 45-line el() helper, three.js for the overlay, and MindAR for image tracking — no framework, and three real dependencies. The backend is one Cloudflare Worker in front of a private R2 bucket. There are no public object URLs: every byte of media streams through the Worker with an ownership check, and video responses honour Range requests because iOS refuses to play video from a server that doesn't. Auth is Sign in with Apple implemented from scratch on WebCrypto — Apple's client secret is an ES256 JWT you mint and sign yourself, and requesting any scope forces the callback into a cross-site form POST, two quirks that auth SaaS normally hides. Sessions are stateless HMAC-signed cookies; the family whitelist is re-checked on every request, so removing an email locks that person out immediately. The whitelist is the whole user store.

Fig. 1How it fits together — and the one hop that happens on paper
Create flow pick · capture · save
MindAR compiler in-browser · tfjs + WebGL
Worker auth + media gateway
Sign in with Apple ES256 secret · WebCrypto
R2 · private bucket video · frame · target.mind
Scanner one tracker, every target
Patched matcher 15 inliers + pose checks
Pose smoother motion-gated dead-band
Compilation happens on the creator's phone — the Workers runtime can't run it, and the scanner shouldn't. The loop lane along the bottom is the printed photo itself: the only edge in the system that travels by hand.
Fig. 2Enroll → scan, one photo's pass through the system
Creator phone Worker R2 Scanning phone 1 video upload — starts at pick, in the background 2 compile target on-device — seconds 3 save: frame + target.mind + buffer id 4 promote uploads/<id> → photos/<id>/ 5 GET /api/library?scope=scan 6 items + precompiled .mind targets 7 merge targets → one tracker 8 lock — ≥15 inliers, pose validated 9 GET video (Range) 10 206 chunks — poster first, video takes over
The upload starts at pick, so Save is a metadata write. On the scanning side everything is fetched precompiled, merged, and matched locally — the network's last job is streaming the video, poster first.

Teaching the matcher to say no

The hardest bug in the project was a matcher that was too eager to say yes. Stock MindAR accepts a detection with 6 inlier feature matches and never validates the recovered homography — so a thumbnail of the target inside a screenshot, or a patch of dense UI text, would lock with a wildly wrong pose. The fix is a patched copy of the matching core, swapped in by a Vite plugin so both the compiler and the tracking worker use it: the inlier floor rises to 15 and — the part stock forgets — is re-checked on the final refined set; the projected quad must be convex, unmirrored, not razor-thin, and cover at least 5% of the camera frame; and the inliers must be spread out, so a tight cluster of points can't dictate a full-frame pose. A standalone bench page compiles a target and runs the real matcher over screen-recording frames under both configs: stock locks onto them, strict rejects each with a named reason. Its observed floor: with 15 inliers required, a print smaller than roughly a tenth of the frame no longer locks — which is the point.

Holding still

MindAR re-estimates the pose fresh every frame, and the estimate is noisy — so the overlay jitters even when the phone is dead still, and the video looks like a sticker coming loose exactly when it should look printed. A plain low-pass can't fix it: enough smoothing to kill the jitter makes real motion visibly trail. The smoother in front of the renderer separates the two regimes instead. A low-pass-filtered velocity estimate tells jitter from motion — jitter is zero-mean, so its filtered velocity is ~0; real motion sustains one. When still, a dead-band (1.2% of marker size, 0.34° of rotation) freezes the pose outright; as motion ramps up the dead-band fades to zero and the filter cutoff opens from 0.3 Hz to 8 Hz, so the overlay lands where the pose says, on time. Position and scale filter as scalars, rotation as a gated slerp — the overlay can drift, but it can never wobble non-rigidly. Rock-steady when the camera is still, snappy when it moves, and every constant tunable from the URL on real hardware.

Both of these fights — the matcher taught to say no, the overlay taught to hold still — plus the first-frame budget that makes a lock feel instant, get a longer write-up with all the numbers defended: How I improved AR image tracking.

One design language, pure black

The app calls its system an editorial gallery: pure black, hairline rules for structure, oversized Outfit display type over Inter body text, and sentence-case labels that lean on weight and dim ink for hierarchy. The library is photographs on black, edge to edge. What the system spends its budget on is motion and state: a large-title header that collapses by FLIP on the compositor, a modal that blurs the page canvas itself (because iOS clips fixed overlays), an island that spring-morphs between measured sizes, and three loaders — Bayer dither, dot-matrix wave, braille spinner — where another app would own one spinner. The whole thing is 2,319 lines of CSS and zero component libraries. There is a light theme too — the same tokens inverted, camera surfaces excepted — and the floating cards on the shelf above wear it.

Even the ornament is engineered: the empty state's drifting pixel-art field is real footage quantized into a 4-bit format (.pxf — band indices only, ~166 KB for eight seconds), recoloured at render time by named five-step palettes tuned per clip — meadow for the windmill, dawn for the sunrise. The film on the chip above is an actual frame of the shipped asset.

What it is, exactly

A stage-1 prototype whose focus areas are written at the top of the README — detection accuracy, first-frame alignment, latency, multi-angle enrollment — and a family album, whitelisted to a handful of Apple IDs, run for the people in the photos. Search and filtering are parked in a commented-out block; desktop gets a soft gate rather than support; the matcher's thresholds were tuned by pointing phones at prints, not by a benchmark suite. And the honest caveat: the recognition bench that proves the matcher's strictness is also how I know its limit — hold the print far enough away and the scanner simply, correctly, declines.

The album's clips are a family mix — stock footage and little pixel-art loops; the Mona Lisa enrolls from a postcard, which feels right.