How I improved AR image tracking
LivePhoto plays a video on top of its own printed photo — you point your phone at the print, and the moment the print came from plays on the paper, tracked to its position and perspective. The demo version of that worked in an afternoon: MindAR compiles an image into a feature target, a three.js plane follows the tracked pose, done.
Making it good took three separate fights, and all three came after the afternoon. The tracker locked onto screenshots of the photo as happily as the photo. The overlay jittered like a sticker coming loose exactly when it should have looked printed. And the first second after a lock — the entire first impression — was a decode stall. This post is what changed, with the numbers, because every one of these fixes is a number that had to be defended on a real phone.
Part one: the matcher that said yes to everything
Stock mind-ar accepts a detection with six inlier feature matches, and it never validates the homography those six points imply. Six coincidental descriptor matches are easy to find in the wild — a thumbnail of the target inside a screenshot, a patch of dense UI text, sometimes a keyboard — and a homography fitted to a coincidence produces poses that are confidently, spectacularly wrong: the overlay would appear at wild sizes on whichever impostor the tracker had latched onto.
The fix is a patched copy of the matching core, swapped in by a Vite plugin so the compiler, the in-page matcher, and the tracking worker all use the same one. It changes the shape of the decision from “enough matches?” to a chain of gates, each of which can reject with a named reason:
The thresholds ship as data, so every one of them can be overridden from the scan URL for on-device tuning:
1const DEFAULT_TUNING = {2 /** min inlier matches at every stage (stock mind-ar: 6, and unchecked on the final pass) */3 minInliers: 15,4 /** best/second-best hamming distance ratio for a descriptor match (lower = stricter) */5 hammingRatio: 0.7,6 /** projected target quad must cover at least this fraction of the camera frame */7 minScreenAreaFraction: 0.05,8 /** ... and at most this multiple of the frame (rejects wild extrapolations) */9 maxScreenAreaFraction: 8,10 /** longest/shortest projected edge ratio bound (rejects razor-thin quads) */11 maxEdgeRatio: 4,12 /** inlier bounding box must span this fraction of the frame's short side, both axes */13 minSpreadFraction: 0.05,14};Three details matter more than the raw numbers.
The bar applies to the refined set. Matching runs in two passes: raw descriptor matches fit a first homography, then a homography-guided second pass re-matches every query point and typically grows the inlier set ~1.5×. Stock mind-ar checks its (tiny) minimum only on pass one. The patch gates pass one at 60% of the target — so borderline-true matches survive to refinement — and applies the full 15-inlier bar to the final refined set, where it means something:
1// Pass 1 (raw descriptor matches + first RANSAC) systematically undercounts:2// the homography-guided second pass typically grows the inlier set ~1.5x.3// Gate pass 1 at a fraction of the target so borderline-true matches survive4// to refinement; the full minInliers bar applies to the FINAL refined set.5const pass1Min = Math.max(8, Math.round(tuning.minInliers * 0.6));The pose has to be geometrically possible. The recovered homography projects the target’s
corners into the camera frame, and the projected quad must be convex with consistent winding —
a sign flip means a mirrored pose, which no camera pointed at the front of a print can
produce — must cover between 5% and 8× of the frame, and must stay well-proportioned (longest edge at
most 4× the shortest). Each failure returns a name: not-convex, mirrored, too-small,
skewed.
Points have to be spread out. A tight cluster of inliers can’t pin down scale, so its pose
is meaningless no matter how many points are in it. The inlier bounding box must span at least
5% of the frame’s short side in both axes, or the detection is rejected as clustered.
To keep the thresholds honest there’s a standalone bench page — one URL, open in any browser —
that compiles a target and runs the real matcher over screen-recording frames under two
configs: stock (6 inliers, no validation) and strict (the shipping defaults). Stock locks
onto the screenshots; strict rejects each with a reason. Synthetic composites with known ground
truth — the target drawn into UI clutter at controlled sizes — verify the cover gate cuts in
where intended: with 15 inliers required, the observed floor is a print around 8–10% of the
frame area. Below that, the scanner simply, correctly, declines.
Part two: tracking the frame without the wobble
MindAR estimates the pose fresh every tracker frame, and the estimate is noisy. On a dead-still phone the physical scene doesn’t move, but the overlay’s edges do — a one-pixel shimmer that reads as fake faster than any tracking failure, because paper doesn’t shimmer.
Two things defeat the obvious fixes. MindAR’s built-in One-Euro filter runs on the 16 matrix elements independently, which leaves non-rigid wobble — the quad’s corners disagree with each other, so the overlay breathes. And a plain low-pass strong enough to kill the jitter adds so much lag that real motion visibly trails the print.
The observation that unlocks it: jitter and motion are different regimes, and a low-pass-filtered velocity estimate can tell them apart. Jitter is zero-mean — filtered, its velocity is roughly zero. Real motion sustains a velocity. So each channel of the decomposed pose (position and scale as scalars, rotation as an angle) runs through a motion-gated dead-band:
1/** Motion-gated dead-band low-pass for one scalar channel. */2filter(target: number, dt: number, ref: number): number {3 const rawVel = (target - this.prev) / dt;4 this.prev = target;5 this.dHat += alphaOf(this.dCutoff, dt) * (rawVel - this.dHat);67 const g = smoothstep(this.vLow, this.vHigh, Math.abs(this.dHat) / ref);8 const bandAbs = this.band * ref * (1 - g); // dead-band fades out as motion ramps up9 const delta = target - this.disp;10 const soft = Math.abs(delta) <= bandAbs ? this.disp : target - Math.sign(delta) * bandAbs;1112 const hz = this.stillHz + (this.moveHz - this.stillHz) * g;13 this.disp += (soft - this.disp) * alphaOf(hz, dt);14 return this.disp;15}When still, the dead-band freezes sub-threshold movement outright — 1.2% of marker size for
position and scale, 0.34° for rotation — and the filter cutoff sits at a glacial 0.3 Hz. As
the motion gate g ramps up, the dead-band fades to zero (so the display settles onto the
latest pose) and the cutoff opens to 8 Hz (so the overlay keeps up). Rotation gets the same
treatment as a gated slerp, with the target quaternion flipped into the displayed one’s
hemisphere first so the slerp always takes the short way around. The frame delta is clamped to
[1/240 s, 0.1 s] so a backgrounded tab can’t produce a huge catch-up step, and a degenerate
pose — MindAR zeroes the matrix for a frame before it admits the target is lost — hides the
overlay instead of collapsing it.
The result is the whole point of the product: rock-steady when the camera is still, snappy when it moves, and rigid to the corners — the overlay can drift a hair, but always as one piece.
MindAR’s own filter constants moved too, in the responsive direction: filterMinCF
0.001 → 0.01 and filterBeta 1000 → 10000 — its defaults favour smoothness, and with a
dedicated smoother downstream the tracker’s job is raw speed. A warmup tolerance
of 4 frames kills one-frame flicker locks; a miss tolerance of 8 rides through brief
occlusions and holds the lock.
Part three: the first-frame budget
The lock is the first impression, and the budget for it is roughly one second of human patience. Three decisions get it under that.
The poster trick. The frame you print is frame zero of the video — the same pixels. So
on lock, a poster plane with the frame texture appears pinned to the print instantly, in the
same render tick, and the video takes over from its own first frame the moment the decoder
reports playing. The handoff is invisible because the two images are the same image. If the
print is lost and re-found within 5 seconds, playback resumes where it was; longer, and it
restarts from the top — a moment, replayed whole.
Decoders are warmed before they’re needed. iOS in particular will happily make you pay a
few hundred milliseconds of decode spin-up on first play. At scanner startup, each video gets a
muted play(), then is parked back at frame 0 after 200 ms — staggered 300 ms apart so the
warm-ups take their turns — with a dataset.live guard so a real lock mid-warm-up plays on
through its own priming.
One tracker for the whole album. Every item’s compiled .mind target is fetched and
merged — the format is msgpack, so merging is decode, concatenate, re-encode — into a single
tracker instance. One detector pass per frame scans the entire library instead of running one
tracker per item, and #/scan?id=… narrows to a single target when you know what you’re
pointing at, which measurably shortens time-to-lock. The expensive step, compiling images into
targets, happens once at enrollment in the creator’s browser (the compiler needs tfjs and
WebGL, which a server worker doesn’t have) — scanning devices only ever download precompiled
results.
All of this was tuned by instrument. The scan page has a debug HUD (#/scan?hud=1) that surfaces
exactly the numbers this prototype exists to measure — camera warm-up, targets fetched,
camera-to-first-lock, lock-to-video, live FPS — and the same URL grammar overrides every
matcher and smoother constant per scan. Every threshold in this post earned its value by being
pointed at a real print, on a real phone, with the HUD on.
The write-up of the whole app — every component of the interface rebuilt from source, from the pixel-film empty state to the dynamic-island scanner — lives on the design page: LivePhoto, taken apart.
← Back to the journal