Decode uploaded footage in order with WebCodecs
This commit is contained in:
parent
131b39bff0
commit
65ad67c129
12 changed files with 454 additions and 364 deletions
25
README.md
25
README.md
|
|
@ -17,7 +17,7 @@ modern conveniences belong in the workflow, not the output. See
|
|||
## ClojureScript port
|
||||
|
||||
The active port plays the synthetic take, accepts video uploads, transcodes them
|
||||
to a browser-seekable proxy plus audio and tracing stills, analyzes real footage
|
||||
to an H.264 proxy and decodable stream plus audio and tracing stills, analyzes real footage
|
||||
for mouth, eyes, brows and pixel-derived teeth, and saves the project with
|
||||
reusable analysis data. The step 8 data model
|
||||
represents persistent feature IDs, eye pairs and feature-level observation gaps;
|
||||
|
|
@ -64,7 +64,7 @@ Three tiers, cut by mutability and size — the full argument is in
|
|||
| --- | --- | --- |
|
||||
| 1 **authored** | the scene: nodes, channels, features, time maps | the database, as independently addressed leaves. Kilobytes |
|
||||
| 2 **derived** | detected landmarks, raw mouth crops, and dense channel blocks | `var/blobs`, addressed by analysis and block inputs, including the detector version |
|
||||
| 3 **source** | the uploaded video, the H.264 proxy measured from it, its tracing stills, and audio | the same blob store, by the hash of their bytes |
|
||||
| 3 **source** | the uploaded video, H.264 proxy and elementary stream, tracing stills, and audio | the same blob store, by the hash of their bytes |
|
||||
|
||||
Only tier 1 is the document. Tier 2 is a pure function of tiers 1 and 3, so a
|
||||
saved project names its blocks rather than carrying them, and a knob change gives
|
||||
|
|
@ -90,28 +90,25 @@ wasm, which is fetched from a CDN on first use.
|
|||
For real footage:
|
||||
|
||||
Upload it in the app. `./extract.sh` still writes the old PNG-sequence bundle and
|
||||
`ingest_bundle` still registers it, but footage ingested that way has no proxy and
|
||||
the loader will say so — the measured pixels come out of the video now.
|
||||
`ingest_bundle` still registers it, but footage ingested that way has no decodable
|
||||
stream and the loader will say so — the measured pixels come out of the video now.
|
||||
|
||||
MediaPipe's wasm and `face_landmarker.task` are both local; nothing in detection
|
||||
touches the network.
|
||||
|
||||
Detection reads the VIDEO, not a frame per file. The page seeks the proxy to the
|
||||
MIDDLE of each frame — `(i + 0.5) / fps` — and waits for
|
||||
`requestVideoFrameCallback` to hand the frame over, then checks the `mediaTime` it
|
||||
reports against the frame it asked for. Both halves are load-bearing and both were
|
||||
measured against the same footage decoded to PNGs: aiming at `i / fps` sits on a
|
||||
frame boundary and landed one frame early 31 times in 91, and aiming at the middle
|
||||
was exact on all 91. A run that gets a frame it did not ask for stops and says so,
|
||||
because a one-frame slip between the landmarks and the audio is not something
|
||||
anyone finds by looking at the result.
|
||||
Detection reads the H.264 elementary stream with WebCodecs, one coded frame at a
|
||||
time. The proxy has no B-frames, so decode order is frame order. Each decoded
|
||||
frame reaches MediaPipe in VIDEO running mode at its footage timestamp. The
|
||||
decoder and detector advance together, with a pause between frames so progress
|
||||
can paint. Saved analyses reuse their stored crop pixels and measure them with
|
||||
the same pauses.
|
||||
|
||||
This is what replaced the PNG sequence, which was 112MB for 7.6 seconds and would
|
||||
be 1.1GB at the 900-frame limit. The proxy is 6MB, and the landmarks barely
|
||||
notice: detected off decoded H.264 rather than off the PNGs, they moved at most
|
||||
0.0033 of frame width.
|
||||
|
||||
`manifest.json` records the source rate. The extractor keeps every source frame;
|
||||
The server's footage manifest records the proxy's frame rate and frame count;
|
||||
the page reads that rate because a guessed fps desynchronises audio from picture.
|
||||
Choosing a lower picture rate happens after analysis.
|
||||
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue