Measure the video, not a PNG per frame

Detection now walks a browser-seekable H.264 proxy in MediaPipe's VIDEO
running mode. The PNG sequence it replaces was 112MB for 7.6 seconds at
1440x1920 and 1.1GB at the 900-frame limit; the proxy is 6MB, and landmarks
detected off decoded H.264 rather than off the PNGs moved at most 0.0033 of
frame width.

Three things had to be true for video mode to work, and each was measured
against the same footage decoded to PNGs:

/blob/<digest> answers byte ranges. Django's FileResponse does no Range
handling, and a media element handed 200 with no Accept-Ranges reports an
empty `seekable`, no-ops every currentTime write, and detects frame one
ninety times without raising.

A seek aims at the MIDDLE of its frame. Aiming at i/fps sits on a frame
boundary and landed one frame early 31 times in 91; (i + 0.5)/fps was exact
on all 91.

Timestamps are strictly increasing footage milliseconds. Video mode is a
tracker: a repeat leaves the graph in an error state every later call
re-throws, so the landmarker is discarded on failure, and passing the frame
index instead of i*1000/fps moved landmarks six times further from the
per-frame answer.

Frames are verified rather than trusted. requestVideoFrameCallback states
which frame it handed over, the walker discards any other and fails loudly
if the one it asked for never arrives — a stale presentation from the tail
of a previous seek is what produced "asked for frame 1 and it presented
frame 2" on a video whose seeks were in fact exact.

The proxy is re-encoded even when the upload is already H.264: HEVC is not
decodable everywhere, and footage identity is the proxy's digest. The JPEG
stills beside it are tracing references, outside the footage digest because
re-rendering them at another size is not different footage.

Verified end to end in a real browser against real footage: 228/228 frames
detected, a drawn roto face, 37 backend and 234 frontend tests green.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
Olive Vaughn 2026-09-28 11:32:01 -04:00
parent 686f897401
commit 83d106bbc5
14 changed files with 748 additions and 142 deletions

View file

@ -145,8 +145,8 @@ boundary.**
| # | Stage | In | Out | Cost |
| --- | --- | --- | --- | --- |
| 1 | **ingest** | video | footage: frames, audio, manifest | minutes, in-app |
| 2 | **detect** | footage | raw landmarks per frame | minutes, **cached** |
| 1 | **ingest** | video | footage: a seekable H.264 proxy, tracing stills, audio, manifest | minutes, in-app |
| 2 | **detect** | the proxy, walked one frame at a time | raw landmarks per frame | minutes, **cached** |
| 3 | **measure** | landmarks | anchor fit, residual, head-local rings, signals, interior pixels | seconds |
| 4 | **condition** | measurements | smoothed transforms and contours | milliseconds |
| 5 | **key** | conditioned signals + policy | channels: sparse keys, quantised holds, kept frames | milliseconds |
@ -575,7 +575,8 @@ PUT /api/analyses/<key> link dense landmarks, mask, crops
POST /api/blocks/missing {keys} -> {missing}
POST /api/blocks {key, descriptor, data, state}
GET /api/blocks/<key>
GET /api/footage/<id> the manifest, with a URL per frame
GET /api/footage/<id> the manifest: the proxy to measure, audio, a URL per tracing still
GET /blob/<digest> immutable bytes, and RANGE-capable so a <video> can seek one
POST /api/sources multipart video upload
POST /api/extractions idempotent decode job
GET /api/extractions/<key> job state and footage id
@ -600,10 +601,28 @@ separately.
**A manifest names frames; it does not locate them.** Until step 9 the client
fetched `manifest.json` and built `frames/0001.png` itself, which made the frame
layout a shared secret between a shell script and a ClojureScript namespace. The
manifest now carries a URL per frame, so the frames can live in the blob store —
or be produced by the app's video upload and server-side ffmpeg extraction. The
upload path needs no `manifest.json` file: the server builds the footage response
from the extracted frame records. The producer changes; the shape does not.
manifest now carries every URL, so the bytes can live in the blob store — or be
produced by the app's video upload and server-side ffmpeg extraction. The upload
path needs no `manifest.json` file: the server builds the footage response from
its own records. The producer changes; the shape does not.
**Tier 3 keeps a video, not a frame per file.** Stage 1 used to decode a PNG per
source frame: 112MB for 7.6 seconds at 1440x1920, and 1.1GB at the 900-frame
limit, for pixels whose only consumer was a canvas MediaPipe read once. It now
writes one browser-safe H.264 proxy — 6MB for the same take — and the page seeks
THAT, frame by frame, in MediaPipe's video running mode. The JPEG stills beside it
are reference images for tracing; nothing measures them, so they are deliberately
outside the footage digest and re-rendering them at another size does not
invalidate an analysis.
Two things make that trustworthy rather than merely smaller. `/blob/<digest>`
answers byte ranges, because a media element handed 200 with no `Accept-Ranges`
reports an empty `seekable` and silently refuses to move — Django's `FileResponse`
does no Range handling, so this is code we own. And every frame is CHECKED:
`requestVideoFrameCallback` states the `mediaTime` of the frame it hands over, the
walker compares it to the frame it asked for, and a mismatch ends the run. Content
addressing over landmarks whose frame alignment was assumed would be addressing a
guess.
**The document stores what a block IS, not what it holds.** A block's element type
is in its own descriptor, which is the only place it is written down: an