Commit graph

6 commits

Author SHA1 Message Date
Olive Vaughn
ae03b61dca sound stuff 2026-09-30 03:31:19 -04:00
Olive Vaughn
65ad67c129 Decode uploaded footage in order with WebCodecs 2026-09-28 14:18:09 -04:00
Olive Vaughn
131b39bff0 Calibrate the video's frame origin instead of assuming it is zero
"the video never presented frame 1; it offered 2" was this function waiting
for a frame it had not asked for a second time. The walk discarded a wrong
frame and re-armed the callback WITHOUT seeking again, and nothing further is
ever presented to a paused element that has not been asked to move — so one
wrong answer starved until the timeout and reported it as the browser refusing.

Underneath that, the wrong answer was not wrong. A container can carry an edit
list, and `currentTime` then counts from the start of the edited presentation
while a frame's `mediaTime` counts from the start of the media. The two differ
by a constant, so the frame at currentTime 0 can honestly report a mediaTime
two frames in. Seeking cannot correct for it in the positive direction: source
frame 0 would have to be found before the start of the video, every attempt
clamps at zero, and the walk offers frame 2 forever.

So the constant is measured once and subtracted. `calibrate!` takes whatever
the browser calls the first frame it shows and makes that the origin, which is
the honest definition anyway — it is what a viewer sees at time zero, and the
audio clock this take plays against starts in the same place.

Calibration also leaves the element on frame 0, so the walk starts holding it.
`frame!` returns immediately for a frame already on screen rather than seeking
to where it already is, which presents nothing and would hang.

Residual error still re-seeks, corrected by exactly the measured miss, and
still fails loudly with every frame that was offered and where it was asked
from, so a next failure is diagnosable in one shot rather than four.

The proxy also drops B-frames now. That removes the edit list at the source
rather than only coping with it, and makes decode order presentation order
should this ever be fed to WebCodecs. 14% larger, and extraction refuses a
proxy whose timeline is shifted so it cannot come back silently. Honest note:
this was my first diagnosis and it did NOT reproduce the failure — both
proxies walk correctly here on hardware decode — so it is hardening, not the
fix.

Verified by forcing the fault: a harness that offsets the reported timeline by
-2, -1, 0, +1, +2 and +5 frames failed on every positive offset before and
recovers on all six now, first attempt. Real app in headed Chrome with Metal
hardware decode: 280/280. 41 backend and 234 frontend tests green.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-28 12:50:11 -04:00
Olive Vaughn
44976cbb4b Stop refusing ordinary phone footage as variable-frame-rate
Uploading a clip shot straight from the iPhone camera app failed with
"variable-frame-rate video needs timestamp-aware playback". The file was not
variable: its container reports avg_frame_rate 8670/299 and nb_frames 289 over
a stream whose decoded timestamps are 280 frames exactly 1/30s apart. The
guard compared two pieces of container metadata and rejected CFR video on the
strength of a summary the container had got wrong about its own contents.

The guard was also obsolete. It dates from when the page measured the source's
own frames, where a wandering frame duration really does break
`frame = floor(t * fps)`. Nothing measures the source now — ffmpeg resamples it
onto a constant rate and the proxy is re-probed after it is written — so
variable input is a thing this converts rather than a thing it refuses.

So: probe picks a rate instead of validating one. It takes the nominal rate,
which is the rate every timestamp in the stream can be expressed at and so the
one that keeps every distinct source frame, and carries it as an exact fraction
because 30000/1001 is not a float and a rounded -r is how a long take drifts.
The disagreement is still recorded as `vfr`, just not fatal.

The frame-count cross-check went with it. It compared the proxy against the
source's nb_frames, which is the number this whole bug proves can lie, and a
resample to a constant rate legitimately changes the count. It now checks the
proxy's DURATION against the source's, because what must not drift is how long
the picture lasts against how long the audio lasts.

Verified on the reported file: 280 frames at 30fps, picture 9.3333s against
audio 9.3167s — half a frame — and 280/280 detected in the real app. 41 backend
tests green, including a genuinely variable fixture end to end.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-28 12:37:21 -04:00
Olive Vaughn
83d106bbc5 Measure the video, not a PNG per frame
Detection now walks a browser-seekable H.264 proxy in MediaPipe's VIDEO
running mode. The PNG sequence it replaces was 112MB for 7.6 seconds at
1440x1920 and 1.1GB at the 900-frame limit; the proxy is 6MB, and landmarks
detected off decoded H.264 rather than off the PNGs moved at most 0.0033 of
frame width.

Three things had to be true for video mode to work, and each was measured
against the same footage decoded to PNGs:

/blob/<digest> answers byte ranges. Django's FileResponse does no Range
handling, and a media element handed 200 with no Accept-Ranges reports an
empty `seekable`, no-ops every currentTime write, and detects frame one
ninety times without raising.

A seek aims at the MIDDLE of its frame. Aiming at i/fps sits on a frame
boundary and landed one frame early 31 times in 91; (i + 0.5)/fps was exact
on all 91.

Timestamps are strictly increasing footage milliseconds. Video mode is a
tracker: a repeat leaves the graph in an error state every later call
re-throws, so the landmarker is discarded on failure, and passing the frame
index instead of i*1000/fps moved landmarks six times further from the
per-frame answer.

Frames are verified rather than trusted. requestVideoFrameCallback states
which frame it handed over, the walker discards any other and fails loudly
if the one it asked for never arrives — a stale presentation from the tail
of a previous seek is what produced "asked for frame 1 and it presented
frame 2" on a video whose seeks were in fact exact.

The proxy is re-encoded even when the upload is already H.264: HEVC is not
decodable everywhere, and footage identity is the proxy's digest. The JPEG
stills beside it are tracing references, outside the footage digest because
re-rendering them at another size is not different footage.

Verified end to end in a real browser against real footage: 228/228 frames
detected, a drawn roto face, 37 backend and 234 frontend tests green.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-28 11:32:01 -04:00
Olive Vaughn
686f897401 Add video upload, extraction progress, and reusable analysis sources 2026-09-28 10:45:51 -04:00