Decode uploaded footage in order with WebCodecs

This commit is contained in:
Olive Vaughn 2026-09-28 14:18:09 -04:00
parent 131b39bff0
commit 65ad67c129
12 changed files with 454 additions and 364 deletions

View file

@ -145,8 +145,8 @@ boundary.**
| # | Stage | In | Out | Cost |
| --- | --- | --- | --- | --- |
| 1 | **ingest** | video | footage: a seekable H.264 proxy, tracing stills, audio, manifest | minutes, in-app |
| 2 | **detect** | the proxy, walked one frame at a time | raw landmarks per frame | minutes, **cached** |
| 1 | **ingest** | video | footage: an H.264 proxy and raw stream, tracing stills, audio, manifest | minutes, in-app |
| 2 | **detect** | the raw stream, decoded one frame at a time | raw landmarks per frame | minutes, **cached** |
| 3 | **measure** | landmarks | anchor fit, residual, head-local rings, signals, interior pixels | seconds |
| 4 | **condition** | measurements | smoothed transforms and contours | milliseconds |
| 5 | **key** | conditioned signals + policy | channels: sparse keys, quantised holds, kept frames | milliseconds |
@ -575,8 +575,8 @@ PUT /api/analyses/<key> link dense landmarks, mask, crops
POST /api/blocks/missing {keys} -> {missing}
POST /api/blocks {key, descriptor, data, state}
GET /api/blocks/<key>
GET /api/footage/<id> the manifest: the proxy to measure, audio, a URL per tracing still
GET /blob/<digest> immutable bytes, and RANGE-capable so a <video> can seek one
GET /api/footage/<id> the manifest: video and stream URLs, audio, a URL per tracing still
GET /blob/<digest> immutable bytes, with byte ranges for video playback
POST /api/sources multipart video upload
POST /api/extractions idempotent decode job
GET /api/extractions/<key> job state and footage id
@ -609,20 +609,17 @@ its own records. The producer changes; the shape does not.
**Tier 3 keeps a video, not a frame per file.** Stage 1 used to decode a PNG per
source frame: 112MB for 7.6 seconds at 1440x1920, and 1.1GB at the 900-frame
limit, for pixels whose only consumer was a canvas MediaPipe read once. It now
writes one browser-safe H.264 proxy — 6MB for the same take — and the page seeks
THAT, frame by frame, in MediaPipe's video running mode. The JPEG stills beside it
writes a browser-safe H.264 proxy — 6MB for the same take — and copies its coded
frames into an Annex-B stream. WebCodecs decodes that stream in order, and the
page gives each frame to MediaPipe in video running mode. The JPEG stills beside it
are reference images for tracing; nothing measures them, so they are deliberately
outside the footage digest and re-rendering them at another size does not
invalidate an analysis.
Two things make that trustworthy rather than merely smaller. `/blob/<digest>`
answers byte ranges, because a media element handed 200 with no `Accept-Ranges`
reports an empty `seekable` and silently refuses to move — Django's `FileResponse`
does no Range handling, so this is code we own. And every frame is CHECKED:
`requestVideoFrameCallback` states the `mediaTime` of the frame it hands over, the
walker compares it to the frame it asked for, and a mismatch ends the run. Content
addressing over landmarks whose frame alignment was assumed would be addressing a
guess.
The proxy is encoded without B-frames, so decode order matches presentation
order. The client checks that the stream has exactly the manifest's frame count
before detection. `/blob/<digest>` also answers byte ranges for ordinary video
playback; Django's `FileResponse` does no Range handling, so this is code we own.
**The document stores what a block IS, not what it holds.** A block's element type
is in its own descriptor, which is the only place it is written down: an