Calibrate the video's frame origin instead of assuming it is zero

"the video never presented frame 1; it offered 2" was this function waiting
for a frame it had not asked for a second time. The walk discarded a wrong
frame and re-armed the callback WITHOUT seeking again, and nothing further is
ever presented to a paused element that has not been asked to move — so one
wrong answer starved until the timeout and reported it as the browser refusing.

Underneath that, the wrong answer was not wrong. A container can carry an edit
list, and `currentTime` then counts from the start of the edited presentation
while a frame's `mediaTime` counts from the start of the media. The two differ
by a constant, so the frame at currentTime 0 can honestly report a mediaTime
two frames in. Seeking cannot correct for it in the positive direction: source
frame 0 would have to be found before the start of the video, every attempt
clamps at zero, and the walk offers frame 2 forever.

So the constant is measured once and subtracted. `calibrate!` takes whatever
the browser calls the first frame it shows and makes that the origin, which is
the honest definition anyway — it is what a viewer sees at time zero, and the
audio clock this take plays against starts in the same place.

Calibration also leaves the element on frame 0, so the walk starts holding it.
`frame!` returns immediately for a frame already on screen rather than seeking
to where it already is, which presents nothing and would hang.

Residual error still re-seeks, corrected by exactly the measured miss, and
still fails loudly with every frame that was offered and where it was asked
from, so a next failure is diagnosable in one shot rather than four.

The proxy also drops B-frames now. That removes the edit list at the source
rather than only coping with it, and makes decode order presentation order
should this ever be fed to WebCodecs. 14% larger, and extraction refuses a
proxy whose timeline is shifted so it cannot come back silently. Honest note:
this was my first diagnosis and it did NOT reproduce the failure — both
proxies walk correctly here on hardware decode — so it is hardening, not the
fix.

Verified by forcing the fault: a harness that offsets the reported timeline by
-2, -1, 0, +1, +2 and +5 frames failed on every positive offset before and
recovers on all six now, first attempt. Real app in headed Chrome with Metal
hardware decode: 280/280. 41 backend and 234 frontend tests green.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
Olive Vaughn 2026-09-28 12:50:11 -04:00
parent 44976cbb4b
commit 131b39bff0
3 changed files with 204 additions and 75 deletions

View file

@ -106,6 +106,18 @@ def _encode_proxy(job, source_path, proxy_path, facts, root):
# step that makes the thing the page measures not be.
"-fps_mode", "cfr", "-r", facts.get("rate") or str(facts["fps"]),
"-c:v", "libx264", "-preset", "veryfast", "-crf", PROXY_CRF,
# NO B-FRAMES, AND THIS IS THE LOAD-BEARING FLAG. With them x264 has a
# two-frame reordering delay, ffmpeg compensates by writing an edit list
# (`elst` media_time 1024 at timebase 1/15360 — exactly two frames), and
# the browser then lives on two timelines at once: `currentTime` obeys the
# edit list and the `mediaTime` reported by requestVideoFrameCallback does
# not. Seek to frame 0 and the browser correctly hands back a frame whose
# mediaTime says 2. Software decoding hides it; hardware decoding does
# not, which is the worst possible way for it to be wrong. Without
# B-frames DTS equals PTS, no edit list is written, and the two timelines
# are the same one. It also makes decode order presentation order, should
# this ever be fed to a WebCodecs VideoDecoder.
"-bf", "0",
# yuv420p and an even frame size are what makes this playable everywhere
# rather than only in the browser that happened to be tested.
"-pix_fmt", "yuv420p", "-vf", "scale=trunc(iw/2)*2:trunc(ih/2)*2",
@ -183,6 +195,27 @@ def probe(path):
"vfr": nominal != average}
def _refuse_a_shifted_timeline(path):
"""The proxy must put frame `i` at `i / fps` on BOTH of the browser's clocks.
Asserted rather than assumed, because the failure is silent and the symptom is
unrecognisable. An encoder delay makes ffmpeg write an edit list, `currentTime`
then obeys it while `requestVideoFrameCallback`'s `mediaTime` does not, and the
page's frame walk is uniformly off by the delay — on hardware decoding only. It
cost two wrong diagnoses to find, so it does not get to come back silently if
somebody changes an encoder flag.
"""
data = json.loads(_command(["ffprobe", "-v", "error", "-select_streams", "v:0",
"-show_streams", "-of", "json", str(path)]))
stream = data["streams"][0]
if int(stream.get("has_b_frames") or 0):
raise ValueError(
"the proxy was encoded with B-frames, whose reordering delay makes the "
"browser's seek clock and its frame-timestamp clock disagree")
if float(stream.get("start_time") or 0) != 0:
raise ValueError(f"the proxy starts at {stream['start_time']}s rather than 0")
def count_frames(path):
"""How many frames a file really holds, counted rather than reported.
@ -255,6 +288,7 @@ def run(key):
# Everything downstream describes the PROXY, not the upload.
proxy_facts = probe(proxy_path)
_refuse_a_shifted_timeline(proxy_path)
frames = count_frames(proxy_path)
if not 1 <= frames <= 900:
raise ValueError(f"the proxy holds {frames} frames; the limit is 1–900")