Decode uploaded footage in order with WebCodecs

This commit is contained in:
Olive Vaughn 2026-09-28 14:18:09 -04:00
parent 131b39bff0
commit 65ad67c129
12 changed files with 454 additions and 364 deletions

View file

@ -17,7 +17,7 @@ modern conveniences belong in the workflow, not the output. See
## ClojureScript port
The active port plays the synthetic take, accepts video uploads, transcodes them
to a browser-seekable proxy plus audio and tracing stills, analyzes real footage
to an H.264 proxy and decodable stream plus audio and tracing stills, analyzes real footage
for mouth, eyes, brows and pixel-derived teeth, and saves the project with
reusable analysis data. The step 8 data model
represents persistent feature IDs, eye pairs and feature-level observation gaps;
@ -64,7 +64,7 @@ Three tiers, cut by mutability and size — the full argument is in
| --- | --- | --- |
| 1 **authored** | the scene: nodes, channels, features, time maps | the database, as independently addressed leaves. Kilobytes |
| 2 **derived** | detected landmarks, raw mouth crops, and dense channel blocks | `var/blobs`, addressed by analysis and block inputs, including the detector version |
| 3 **source** | the uploaded video, the H.264 proxy measured from it, its tracing stills, and audio | the same blob store, by the hash of their bytes |
| 3 **source** | the uploaded video, H.264 proxy and elementary stream, tracing stills, and audio | the same blob store, by the hash of their bytes |
Only tier 1 is the document. Tier 2 is a pure function of tiers 1 and 3, so a
saved project names its blocks rather than carrying them, and a knob change gives
@ -90,28 +90,25 @@ wasm, which is fetched from a CDN on first use.
For real footage:
Upload it in the app. `./extract.sh` still writes the old PNG-sequence bundle and
`ingest_bundle` still registers it, but footage ingested that way has no proxy and
the loader will say so — the measured pixels come out of the video now.
`ingest_bundle` still registers it, but footage ingested that way has no decodable
stream and the loader will say so — the measured pixels come out of the video now.
MediaPipe's wasm and `face_landmarker.task` are both local; nothing in detection
touches the network.
Detection reads the VIDEO, not a frame per file. The page seeks the proxy to the
MIDDLE of each frame — `(i + 0.5) / fps` — and waits for
`requestVideoFrameCallback` to hand the frame over, then checks the `mediaTime` it
reports against the frame it asked for. Both halves are load-bearing and both were
measured against the same footage decoded to PNGs: aiming at `i / fps` sits on a
frame boundary and landed one frame early 31 times in 91, and aiming at the middle
was exact on all 91. A run that gets a frame it did not ask for stops and says so,
because a one-frame slip between the landmarks and the audio is not something
anyone finds by looking at the result.
Detection reads the H.264 elementary stream with WebCodecs, one coded frame at a
time. The proxy has no B-frames, so decode order is frame order. Each decoded
frame reaches MediaPipe in VIDEO running mode at its footage timestamp. The
decoder and detector advance together, with a pause between frames so progress
can paint. Saved analyses reuse their stored crop pixels and measure them with
the same pauses.
This is what replaced the PNG sequence, which was 112MB for 7.6 seconds and would
be 1.1GB at the 900-frame limit. The proxy is 6MB, and the landmarks barely
notice: detected off decoded H.264 rather than off the PNGs, they moved at most
0.0033 of frame width.
`manifest.json` records the source rate. The extractor keeps every source frame;
The server's footage manifest records the proxy's frame rate and frame count;
the page reads that rate because a guessed fps desynchronises audio from picture.
Choosing a lower picture rate happens after analysis.