← All notes

How I improved AR image tracking

webarcomputer-visionperformancetypescript


LivePhoto plays a video on top of its own printed photo — you point your phone at the print, and the moment the print came from plays on the paper, tracked to its position and perspective. The demo version of that worked in an afternoon: MindAR compiles an image into a feature target, a three.js plane follows the tracked pose, done.

Making it good took three separate fights, and all three came after the afternoon. The tracker locked onto screenshots of the photo as happily as the photo. The overlay jittered like a sticker coming loose exactly when it should have looked printed. And the first second after a lock — the entire first impression — was a decode stall. This post is what changed, with the numbers, because every one of these fixes is a number that had to be defended on a real phone.

Part one: the matcher that said yes to everything

Stock mind-ar accepts a detection with six inlier feature matches, and it never validates the homography those six points imply. Six coincidental descriptor matches are easy to find in the wild — a thumbnail of the target inside a screenshot, a patch of dense UI text, sometimes a keyboard — and a homography fitted to a coincidence produces poses that are confidently, spectacularly wrong: the overlay would appear at wild sizes on whichever impostor the tracker had latched onto.

The fix is a patched copy of the matching core, swapped in by a Vite plugin so the compiler, the in-page matcher, and the tracking worker all use the same one. It changes the shape of the decision from “enough matches?” to a chain of gates, each of which can reject with a named reason:

one detection attempt
Camera frame every tracker tick
Descriptor match hamming ratio < 0.7
Hough + RANSAC first homography
Pose gates convex · area · spread
Guided re-match inliers grow ~1.5×
Final bar ≥ 15 inliers + gates again
Lock pose → smoother
Rejected named reason · next frame
Two passes, gated twice. The first pass earns the right to refine; the refined set has to clear the full bar — stock mind-ar forgets to re-check it.

The thresholds ship as data, so every one of them can be overridden from the scan URL for on-device tuning:

js
1const DEFAULT_TUNING = {2  /** min inlier matches at every stage (stock mind-ar: 6, and unchecked on the final pass) */3  minInliers: 15,4  /** best/second-best hamming distance ratio for a descriptor match (lower = stricter) */5  hammingRatio: 0.7,6  /** projected target quad must cover at least this fraction of the camera frame */7  minScreenAreaFraction: 0.05,8  /** ... and at most this multiple of the frame (rejects wild extrapolations) */9  maxScreenAreaFraction: 8,10  /** longest/shortest projected edge ratio bound (rejects razor-thin quads) */11  maxEdgeRatio: 4,12  /** inlier bounding box must span this fraction of the frame's short side, both axes */13  minSpreadFraction: 0.05,14};

Three details matter more than the raw numbers.

The bar applies to the refined set. Matching runs in two passes: raw descriptor matches fit a first homography, then a homography-guided second pass re-matches every query point and typically grows the inlier set ~1.5×. Stock mind-ar checks its (tiny) minimum only on pass one. The patch gates pass one at 60% of the target — so borderline-true matches survive to refinement — and applies the full 15-inlier bar to the final refined set, where it means something:

js
1// Pass 1 (raw descriptor matches + first RANSAC) systematically undercounts:2// the homography-guided second pass typically grows the inlier set ~1.5x.3// Gate pass 1 at a fraction of the target so borderline-true matches survive4// to refinement; the full minInliers bar applies to the FINAL refined set.5const pass1Min = Math.max(8, Math.round(tuning.minInliers * 0.6));

The pose has to be geometrically possible. The recovered homography projects the target’s corners into the camera frame, and the projected quad must be convex with consistent winding — a sign flip means a mirrored pose, which no camera pointed at the front of a print can produce — must cover between 5% and 8× of the frame, and must stay well-proportioned (longest edge at most 4× the shortest). Each failure returns a name: not-convex, mirrored, too-small, skewed.

Points have to be spread out. A tight cluster of inliers can’t pin down scale, so its pose is meaningless no matter how many points are in it. The inlier bounding box must span at least 5% of the frame’s short side in both axes, or the detection is rejected as clustered.

To keep the thresholds honest there’s a standalone bench page — one URL, open in any browser — that compiles a target and runs the real matcher over screen-recording frames under two configs: stock (6 inliers, no validation) and strict (the shipping defaults). Stock locks onto the screenshots; strict rejects each with a reason. Synthetic composites with known ground truth — the target drawn into UI clutter at controlled sizes — verify the cover gate cuts in where intended: with 15 inliers required, the observed floor is a print around 8–10% of the frame area. Below that, the scanner simply, correctly, declines.

Part two: tracking the frame without the wobble

MindAR estimates the pose fresh every tracker frame, and the estimate is noisy. On a dead-still phone the physical scene doesn’t move, but the overlay’s edges do — a one-pixel shimmer that reads as fake faster than any tracking failure, because paper doesn’t shimmer.

Two things defeat the obvious fixes. MindAR’s built-in One-Euro filter runs on the 16 matrix elements independently, which leaves non-rigid wobble — the quad’s corners disagree with each other, so the overlay breathes. And a plain low-pass strong enough to kill the jitter adds so much lag that real motion visibly trails the print.

The observation that unlocks it: jitter and motion are different regimes, and a low-pass-filtered velocity estimate can tell them apart. Jitter is zero-mean — filtered, its velocity is roughly zero. Real motion sustains a velocity. So each channel of the decomposed pose (position and scale as scalars, rotation as an angle) runs through a motion-gated dead-band:

ts
1/** Motion-gated dead-band low-pass for one scalar channel. */2filter(target: number, dt: number, ref: number): number {3  const rawVel = (target - this.prev) / dt;4  this.prev = target;5  this.dHat += alphaOf(this.dCutoff, dt) * (rawVel - this.dHat);6​7  const g = smoothstep(this.vLow, this.vHigh, Math.abs(this.dHat) / ref);8  const bandAbs = this.band * ref * (1 - g); // dead-band fades out as motion ramps up9  const delta = target - this.disp;10  const soft = Math.abs(delta) <= bandAbs ? this.disp : target - Math.sign(delta) * bandAbs;11​12  const hz = this.stillHz + (this.moveHz - this.stillHz) * g;13  this.disp += (soft - this.disp) * alphaOf(hz, dt);14  return this.disp;15}

When still, the dead-band freezes sub-threshold movement outright — 1.2% of marker size for position and scale, 0.34° for rotation — and the filter cutoff sits at a glacial 0.3 Hz. As the motion gate g ramps up, the dead-band fades to zero (so the display settles onto the latest pose) and the cutoff opens to 8 Hz (so the overlay keeps up). Rotation gets the same treatment as a gated slerp, with the target quaternion flipped into the displayed one’s hemisphere first so the slerp always takes the short way around. The frame delta is clamped to [1/240 s, 0.1 s] so a backgrounded tab can’t produce a huge catch-up step, and a degenerate pose — MindAR zeroes the matrix for a frame before it admits the target is lost — hides the overlay instead of collapsing it.

The result is the whole point of the product: rock-steady when the camera is still, snappy when it moves, and rigid to the corners — the overlay can drift a hair, but always as one piece.

MindAR’s own filter constants moved too, in the responsive direction: filterMinCF 0.001 → 0.01 and filterBeta 1000 → 10000 — its defaults favour smoothness, and with a dedicated smoother downstream the tracker’s job is raw speed. A warmup tolerance of 4 frames kills one-frame flicker locks; a miss tolerance of 8 rides through brief occlusions and holds the lock.

Part three: the first-frame budget

The lock is the first impression, and the budget for it is roughly one second of human patience. Three decisions get it under that.

The poster trick. The frame you print is frame zero of the video — the same pixels. So on lock, a poster plane with the frame texture appears pinned to the print instantly, in the same render tick, and the video takes over from its own first frame the moment the decoder reports playing. The handoff is invisible because the two images are the same image. If the print is lost and re-found within 5 seconds, playback resumes where it was; longer, and it restarts from the top — a moment, replayed whole.

Decoders are warmed before they’re needed. iOS in particular will happily make you pay a few hundred milliseconds of decode spin-up on first play. At scanner startup, each video gets a muted play(), then is parked back at frame 0 after 200 ms — staggered 300 ms apart so the warm-ups take their turns — with a dataset.live guard so a real lock mid-warm-up plays on through its own priming.

One tracker for the whole album. Every item’s compiled .mind target is fetched and merged — the format is msgpack, so merging is decode, concatenate, re-encode — into a single tracker instance. One detector pass per frame scans the entire library instead of running one tracker per item, and #/scan?id=… narrows to a single target when you know what you’re pointing at, which measurably shortens time-to-lock. The expensive step, compiling images into targets, happens once at enrollment in the creator’s browser (the compiler needs tfjs and WebGL, which a server worker doesn’t have) — scanning devices only ever download precompiled results.

All of this was tuned by instrument. The scan page has a debug HUD (#/scan?hud=1) that surfaces exactly the numbers this prototype exists to measure — camera warm-up, targets fetched, camera-to-first-lock, lock-to-video, live FPS — and the same URL grammar overrides every matcher and smoother constant per scan. Every threshold in this post earned its value by being pointed at a real print, on a real phone, with the HUD on.


The write-up of the whole app — every component of the interface rebuilt from source, from the pixel-film empty state to the dynamic-island scanner — lives on the design page: LivePhoto, taken apart.


← Back to the journal