Measure the video, not a PNG per frame
Detection now walks a browser-seekable H.264 proxy in MediaPipe's VIDEO running mode. The PNG sequence it replaces was 112MB for 7.6 seconds at 1440x1920 and 1.1GB at the 900-frame limit; the proxy is 6MB, and landmarks detected off decoded H.264 rather than off the PNGs moved at most 0.0033 of frame width. Three things had to be true for video mode to work, and each was measured against the same footage decoded to PNGs: /blob/<digest> answers byte ranges. Django's FileResponse does no Range handling, and a media element handed 200 with no Accept-Ranges reports an empty `seekable`, no-ops every currentTime write, and detects frame one ninety times without raising. A seek aims at the MIDDLE of its frame. Aiming at i/fps sits on a frame boundary and landed one frame early 31 times in 91; (i + 0.5)/fps was exact on all 91. Timestamps are strictly increasing footage milliseconds. Video mode is a tracker: a repeat leaves the graph in an error state every later call re-throws, so the landmarker is discarded on failure, and passing the frame index instead of i*1000/fps moved landmarks six times further from the per-frame answer. Frames are verified rather than trusted. requestVideoFrameCallback states which frame it handed over, the walker discards any other and fails loudly if the one it asked for never arrives — a stale presentation from the tail of a previous seek is what produced "asked for frame 1 and it presented frame 2" on a video whose seeks were in fact exact. The proxy is re-encoded even when the upload is already H.264: HEVC is not decodable everywhere, and footage identity is the proxy's digest. The JPEG stills beside it are tracing references, outside the footage digest because re-rendering them at another size is not different footage. Verified end to end in a real browser against real footage: 228/228 frames detected, a drawn roto face, 37 backend and 234 frontend tests green. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
parent
686f897401
commit
83d106bbc5
14 changed files with 748 additions and 142 deletions
36
README.md
36
README.md
|
|
@ -16,9 +16,10 @@ modern conveniences belong in the workflow, not the output. See
|
|||
|
||||
## ClojureScript port
|
||||
|
||||
The active port plays the synthetic take, accepts video uploads, extracts their
|
||||
frames and audio, analyzes real footage for mouth, eyes, brows and pixel-derived
|
||||
teeth, and saves the project with reusable analysis data. The step 8 data model
|
||||
The active port plays the synthetic take, accepts video uploads, transcodes them
|
||||
to a browser-seekable proxy plus audio and tracing stills, analyzes real footage
|
||||
for mouth, eyes, brows and pixel-derived teeth, and saves the project with
|
||||
reusable analysis data. The step 8 data model
|
||||
represents persistent feature IDs, eye pairs and feature-level observation gaps;
|
||||
its controls are still pending. See the [port plan](docs/port-plan.md).
|
||||
|
||||
|
|
@ -63,7 +64,7 @@ Three tiers, cut by mutability and size — the full argument is in
|
|||
| --- | --- | --- |
|
||||
| 1 **authored** | the scene: nodes, channels, features, time maps | the database, as independently addressed leaves. Kilobytes |
|
||||
| 2 **derived** | detected landmarks, raw mouth crops, and dense channel blocks | `var/blobs`, addressed by analysis and block inputs, including the detector version |
|
||||
| 3 **source** | uploaded video, extracted frames, and audio | the same blob store, by the hash of their bytes |
|
||||
| 3 **source** | the uploaded video, the H.264 proxy measured from it, its tracing stills, and audio | the same blob store, by the hash of their bytes |
|
||||
|
||||
Only tier 1 is the document. Tier 2 is a pure function of tiers 1 and 3, so a
|
||||
saved project names its blocks rather than carrying them, and a knob change gives
|
||||
|
|
@ -88,16 +89,27 @@ wasm, which is fetched from a CDN on first use.
|
|||
|
||||
For real footage:
|
||||
|
||||
```sh
|
||||
./extract.sh /path/to/clip.mov # -> frames/*.png, audio.wav, manifest.json
|
||||
```
|
||||
Upload it in the app. `./extract.sh` still writes the old PNG-sequence bundle and
|
||||
`ingest_bundle` still registers it, but footage ingested that way has no proxy and
|
||||
the loader will say so — the measured pixels come out of the video now.
|
||||
|
||||
then **Load frames**. MediaPipe's wasm is fetched from jsdelivr on first use;
|
||||
`face_landmarker.task` is local.
|
||||
MediaPipe's wasm and `face_landmarker.task` are both local; nothing in detection
|
||||
touches the network.
|
||||
|
||||
Frames are pre-extracted rather than decoded in the page because browser video
|
||||
seeking is approximate and `requestVideoFrameCallback` only delivers frames at
|
||||
playback speed — neither gives a deterministic per-frame pass.
|
||||
Detection reads the VIDEO, not a frame per file. The page seeks the proxy to the
|
||||
MIDDLE of each frame — `(i + 0.5) / fps` — and waits for
|
||||
`requestVideoFrameCallback` to hand the frame over, then checks the `mediaTime` it
|
||||
reports against the frame it asked for. Both halves are load-bearing and both were
|
||||
measured against the same footage decoded to PNGs: aiming at `i / fps` sits on a
|
||||
frame boundary and landed one frame early 31 times in 91, and aiming at the middle
|
||||
was exact on all 91. A run that gets a frame it did not ask for stops and says so,
|
||||
because a one-frame slip between the landmarks and the audio is not something
|
||||
anyone finds by looking at the result.
|
||||
|
||||
This is what replaced the PNG sequence, which was 112MB for 7.6 seconds and would
|
||||
be 1.1GB at the 900-frame limit. The proxy is 6MB, and the landmarks barely
|
||||
notice: detected off decoded H.264 rather than off the PNGs, they moved at most
|
||||
0.0033 of frame width.
|
||||
|
||||
`manifest.json` records the source rate. The extractor keeps every source frame;
|
||||
the page reads that rate because a guessed fps desynchronises audio from picture.
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue