arthur/clips/extraction.py

502 lines
26 KiB
Python
Raw Normal View History

Measure the video, not a PNG per frame Detection now walks a browser-seekable H.264 proxy in MediaPipe's VIDEO running mode. The PNG sequence it replaces was 112MB for 7.6 seconds at 1440x1920 and 1.1GB at the 900-frame limit; the proxy is 6MB, and landmarks detected off decoded H.264 rather than off the PNGs moved at most 0.0033 of frame width. Three things had to be true for video mode to work, and each was measured against the same footage decoded to PNGs: /blob/<digest> answers byte ranges. Django's FileResponse does no Range handling, and a media element handed 200 with no Accept-Ranges reports an empty `seekable`, no-ops every currentTime write, and detects frame one ninety times without raising. A seek aims at the MIDDLE of its frame. Aiming at i/fps sits on a frame boundary and landed one frame early 31 times in 91; (i + 0.5)/fps was exact on all 91. Timestamps are strictly increasing footage milliseconds. Video mode is a tracker: a repeat leaves the graph in an error state every later call re-throws, so the landmarker is discarded on failure, and passing the frame index instead of i*1000/fps moved landmarks six times further from the per-frame answer. Frames are verified rather than trusted. requestVideoFrameCallback states which frame it handed over, the walker discards any other and fails loudly if the one it asked for never arrives — a stale presentation from the tail of a previous seek is what produced "asked for frame 1 and it presented frame 2" on a video whose seeks were in fact exact. The proxy is re-encoded even when the upload is already H.264: HEVC is not decodable everywhere, and footage identity is the proxy's digest. The JPEG stills beside it are tracing references, outside the footage digest because re-rendering them at another size is not different footage. Verified end to end in a real browser against real footage: 228/228 frames detected, a drawn roto face, 37 backend and 234 frontend tests green. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-28 11:32:01 -04:00
"""Upload a video once, then turn it into the two things the app actually reads.
WHAT CHANGED AND WHY. This used to decode one PNG per source frame and store every
one of them. A 7.6-second 1440x1920 take is 112MB that way, and the 900-frame limit
is 1.1GB — for pixels whose only consumer was a canvas that MediaPipe then read
once. The page now detects from the video itself (see `frontend/src/arthur/flow/
ingest.cljs`), so this produces:
THE PROXY. One browser-safe H.264/yuv420p MP4, CFR, `+faststart`. The same take
is 6MB. This is the analysis source, and it is re-encoded RATHER THAN KEPT AS
UPLOADED even when the upload is already H.264, for two reasons that are both
about not guessing: an iPhone's HEVC is not decodable in every browser, and the
footage's identity is the digest of this file — one produced by one ffmpeg
invocation, not one that depends on which branch the source happened to take.
THE TRACING STILLS. One JPEG per frame, long edge capped, for the tracing editor
to draw over. Reference images; nothing measures them. They are not in the
footage digest — see `models.Footage`.
The proxy is probed after it is written rather than before. `width`, `height` and
`frames` are properties of the file the browser will decode, and taking them from
the source instead is how a scaler or a dropped frame becomes a silent one-frame
offset between the landmarks and the audio.
"""
import hashlib
import json
import subprocess
import tempfile
import threading
import time
from fractions import Fraction
from pathlib import Path
from django.db import close_old_connections, transaction
from . import blobs
from .models import Blob, Extraction, Footage, FootageFrame
_active = set()
_lock = threading.Lock()
TIMEOUT = 3600
Measure the video, not a PNG per frame Detection now walks a browser-seekable H.264 proxy in MediaPipe's VIDEO running mode. The PNG sequence it replaces was 112MB for 7.6 seconds at 1440x1920 and 1.1GB at the 900-frame limit; the proxy is 6MB, and landmarks detected off decoded H.264 rather than off the PNGs moved at most 0.0033 of frame width. Three things had to be true for video mode to work, and each was measured against the same footage decoded to PNGs: /blob/<digest> answers byte ranges. Django's FileResponse does no Range handling, and a media element handed 200 with no Accept-Ranges reports an empty `seekable`, no-ops every currentTime write, and detects frame one ninety times without raising. A seek aims at the MIDDLE of its frame. Aiming at i/fps sits on a frame boundary and landed one frame early 31 times in 91; (i + 0.5)/fps was exact on all 91. Timestamps are strictly increasing footage milliseconds. Video mode is a tracker: a repeat leaves the graph in an error state every later call re-throws, so the landmarker is discarded on failure, and passing the frame index instead of i*1000/fps moved landmarks six times further from the per-frame answer. Frames are verified rather than trusted. requestVideoFrameCallback states which frame it handed over, the walker discards any other and fails loudly if the one it asked for never arrives — a stale presentation from the tail of a previous seek is what produced "asked for frame 1 and it presented frame 2" on a video whose seeks were in fact exact. The proxy is re-encoded even when the upload is already H.264: HEVC is not decodable everywhere, and footage identity is the proxy's digest. The JPEG stills beside it are tracing references, outside the footage digest because re-rendering them at another size is not different footage. Verified end to end in a real browser against real footage: 228/228 frames detected, a drawn roto face, 37 backend and 234 frontend tests green. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-28 11:32:01 -04:00
# Visually lossless enough that landmarks do not move: measured against the same
# frames as PNGs, IMAGE-mode landmarks shifted at most 0.0033 of frame width.
PROXY_CRF = "18"
# The long edge of a tracing still. The proxy keeps full resolution because the
# detector reads it; a still only has to be good enough to draw a cel over.
TRACING_EDGE = 1280
TRACING_QUALITY = "4"
def _command(args):
result = subprocess.run(args, capture_output=True, text=True, timeout=TIMEOUT)
if result.returncode:
raise ValueError((result.stderr or result.stdout or "media tool failed")[-1200:])
return result.stdout
Measure the video, not a PNG per frame Detection now walks a browser-seekable H.264 proxy in MediaPipe's VIDEO running mode. The PNG sequence it replaces was 112MB for 7.6 seconds at 1440x1920 and 1.1GB at the 900-frame limit; the proxy is 6MB, and landmarks detected off decoded H.264 rather than off the PNGs moved at most 0.0033 of frame width. Three things had to be true for video mode to work, and each was measured against the same footage decoded to PNGs: /blob/<digest> answers byte ranges. Django's FileResponse does no Range handling, and a media element handed 200 with no Accept-Ranges reports an empty `seekable`, no-ops every currentTime write, and detects frame one ninety times without raising. A seek aims at the MIDDLE of its frame. Aiming at i/fps sits on a frame boundary and landed one frame early 31 times in 91; (i + 0.5)/fps was exact on all 91. Timestamps are strictly increasing footage milliseconds. Video mode is a tracker: a repeat leaves the graph in an error state every later call re-throws, so the landmarker is discarded on failure, and passing the frame index instead of i*1000/fps moved landmarks six times further from the per-frame answer. Frames are verified rather than trusted. requestVideoFrameCallback states which frame it handed over, the walker discards any other and fails loudly if the one it asked for never arrives — a stale presentation from the tail of a previous seek is what produced "asked for frame 1 and it presented frame 2" on a video whose seeks were in fact exact. The proxy is re-encoded even when the upload is already H.264: HEVC is not decodable everywhere, and footage identity is the proxy's digest. The JPEG stills beside it are tracing references, outside the footage digest because re-rendering them at another size is not different footage. Verified end to end in a real browser against real footage: 228/228 frames detected, a drawn roto face, 37 backend and 234 frontend tests green. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-28 11:32:01 -04:00
def _run_with_progress(job, args, root, name, total, span):
"""Run one ffmpeg and publish its live frame count as `span` of the job.
ffmpeg's `-progress` file is the only honest source for this: parsing its
stderr means parsing a format that is explicitly not an interface, and a
spinner that is not attached to frames is a spinner that lies on a long take.
"""
progress_path = root / f"{name}.progress"
log_path = root / f"{name}.log"
first, last = span
args = ["ffmpeg", "-hide_banner", "-loglevel", "error", "-y",
Measure the video, not a PNG per frame Detection now walks a browser-seekable H.264 proxy in MediaPipe's VIDEO running mode. The PNG sequence it replaces was 112MB for 7.6 seconds at 1440x1920 and 1.1GB at the 900-frame limit; the proxy is 6MB, and landmarks detected off decoded H.264 rather than off the PNGs moved at most 0.0033 of frame width. Three things had to be true for video mode to work, and each was measured against the same footage decoded to PNGs: /blob/<digest> answers byte ranges. Django's FileResponse does no Range handling, and a media element handed 200 with no Accept-Ranges reports an empty `seekable`, no-ops every currentTime write, and detects frame one ninety times without raising. A seek aims at the MIDDLE of its frame. Aiming at i/fps sits on a frame boundary and landed one frame early 31 times in 91; (i + 0.5)/fps was exact on all 91. Timestamps are strictly increasing footage milliseconds. Video mode is a tracker: a repeat leaves the graph in an error state every later call re-throws, so the landmarker is discarded on failure, and passing the frame index instead of i*1000/fps moved landmarks six times further from the per-frame answer. Frames are verified rather than trusted. requestVideoFrameCallback states which frame it handed over, the walker discards any other and fails loudly if the one it asked for never arrives — a stale presentation from the tail of a previous seek is what produced "asked for frame 1 and it presented frame 2" on a video whose seeks were in fact exact. The proxy is re-encoded even when the upload is already H.264: HEVC is not decodable everywhere, and footage identity is the proxy's digest. The JPEG stills beside it are tracing references, outside the footage digest because re-rendering them at another size is not different footage. Verified end to end in a real browser against real footage: 228/228 frames detected, a drawn roto face, 37 backend and 234 frontend tests green. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-28 11:32:01 -04:00
"-stats_period", "0.25", "-progress", str(progress_path)] + args
with open(log_path, "wb") as log:
proc = subprocess.Popen(args, stdout=log, stderr=subprocess.STDOUT)
deadline = time.monotonic() + TIMEOUT
try:
while proc.poll() is None:
if time.monotonic() >= deadline:
Measure the video, not a PNG per frame Detection now walks a browser-seekable H.264 proxy in MediaPipe's VIDEO running mode. The PNG sequence it replaces was 112MB for 7.6 seconds at 1440x1920 and 1.1GB at the 900-frame limit; the proxy is 6MB, and landmarks detected off decoded H.264 rather than off the PNGs moved at most 0.0033 of frame width. Three things had to be true for video mode to work, and each was measured against the same footage decoded to PNGs: /blob/<digest> answers byte ranges. Django's FileResponse does no Range handling, and a media element handed 200 with no Accept-Ranges reports an empty `seekable`, no-ops every currentTime write, and detects frame one ninety times without raising. A seek aims at the MIDDLE of its frame. Aiming at i/fps sits on a frame boundary and landed one frame early 31 times in 91; (i + 0.5)/fps was exact on all 91. Timestamps are strictly increasing footage milliseconds. Video mode is a tracker: a repeat leaves the graph in an error state every later call re-throws, so the landmarker is discarded on failure, and passing the frame index instead of i*1000/fps moved landmarks six times further from the per-frame answer. Frames are verified rather than trusted. requestVideoFrameCallback states which frame it handed over, the walker discards any other and fails loudly if the one it asked for never arrives — a stale presentation from the tail of a previous seek is what produced "asked for frame 1 and it presented frame 2" on a video whose seeks were in fact exact. The proxy is re-encoded even when the upload is already H.264: HEVC is not decodable everywhere, and footage identity is the proxy's digest. The JPEG stills beside it are tracing references, outside the footage digest because re-rendering them at another size is not different footage. Verified end to end in a real browser against real footage: 228/228 frames detected, a drawn roto face, 37 backend and 234 frontend tests green. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-28 11:32:01 -04:00
raise TimeoutError(f"{name} timed out")
if progress_path.exists():
lines = progress_path.read_text(errors="replace").splitlines()
count = next((int(line[6:].strip()) for line in reversed(lines)
if line.startswith("frame=") and
line[6:].strip().isdigit()), 0)
if count and total:
Measure the video, not a PNG per frame Detection now walks a browser-seekable H.264 proxy in MediaPipe's VIDEO running mode. The PNG sequence it replaces was 112MB for 7.6 seconds at 1440x1920 and 1.1GB at the 900-frame limit; the proxy is 6MB, and landmarks detected off decoded H.264 rather than off the PNGs moved at most 0.0033 of frame width. Three things had to be true for video mode to work, and each was measured against the same footage decoded to PNGs: /blob/<digest> answers byte ranges. Django's FileResponse does no Range handling, and a media element handed 200 with no Accept-Ranges reports an empty `seekable`, no-ops every currentTime write, and detects frame one ninety times without raising. A seek aims at the MIDDLE of its frame. Aiming at i/fps sits on a frame boundary and landed one frame early 31 times in 91; (i + 0.5)/fps was exact on all 91. Timestamps are strictly increasing footage milliseconds. Video mode is a tracker: a repeat leaves the graph in an error state every later call re-throws, so the landmarker is discarded on failure, and passing the frame index instead of i*1000/fps moved landmarks six times further from the per-frame answer. Frames are verified rather than trusted. requestVideoFrameCallback states which frame it handed over, the walker discards any other and fails loudly if the one it asked for never arrives — a stale presentation from the tail of a previous seek is what produced "asked for frame 1 and it presented frame 2" on a video whose seeks were in fact exact. The proxy is re-encoded even when the upload is already H.264: HEVC is not decodable everywhere, and footage identity is the proxy's digest. The JPEG stills beside it are tracing references, outside the footage digest because re-rendering them at another size is not different footage. Verified end to end in a real browser against real footage: 228/228 frames detected, a drawn roto face, 37 backend and 234 frontend tests green. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-28 11:32:01 -04:00
reached = first + int((last - first) * min(1.0, count / total))
if reached > job.progress:
job.progress = reached
job.save(update_fields=["progress", "updated"])
time.sleep(0.2)
finally:
if proc.poll() is None:
proc.kill()
proc.wait()
if proc.returncode:
Measure the video, not a PNG per frame Detection now walks a browser-seekable H.264 proxy in MediaPipe's VIDEO running mode. The PNG sequence it replaces was 112MB for 7.6 seconds at 1440x1920 and 1.1GB at the 900-frame limit; the proxy is 6MB, and landmarks detected off decoded H.264 rather than off the PNGs moved at most 0.0033 of frame width. Three things had to be true for video mode to work, and each was measured against the same footage decoded to PNGs: /blob/<digest> answers byte ranges. Django's FileResponse does no Range handling, and a media element handed 200 with no Accept-Ranges reports an empty `seekable`, no-ops every currentTime write, and detects frame one ninety times without raising. A seek aims at the MIDDLE of its frame. Aiming at i/fps sits on a frame boundary and landed one frame early 31 times in 91; (i + 0.5)/fps was exact on all 91. Timestamps are strictly increasing footage milliseconds. Video mode is a tracker: a repeat leaves the graph in an error state every later call re-throws, so the landmarker is discarded on failure, and passing the frame index instead of i*1000/fps moved landmarks six times further from the per-frame answer. Frames are verified rather than trusted. requestVideoFrameCallback states which frame it handed over, the walker discards any other and fails loudly if the one it asked for never arrives — a stale presentation from the tail of a previous seek is what produced "asked for frame 1 and it presented frame 2" on a video whose seeks were in fact exact. The proxy is re-encoded even when the upload is already H.264: HEVC is not decodable everywhere, and footage identity is the proxy's digest. The JPEG stills beside it are tracing references, outside the footage digest because re-rendering them at another size is not different footage. Verified end to end in a real browser against real footage: 228/228 frames detected, a drawn roto face, 37 backend and 234 frontend tests green. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-28 11:32:01 -04:00
raise ValueError(log_path.read_text(errors="replace")[-1200:] or f"{name} failed")
def _encode_proxy(job, source_path, proxy_path, facts, root):
"""The uploaded video -> one H.264 file every browser can decode and seek."""
total = facts.get("reported_frames") or round(facts["duration"] * facts["fps"])
_run_with_progress(
job,
["-i", str(source_path), "-an",
Stop refusing ordinary phone footage as variable-frame-rate Uploading a clip shot straight from the iPhone camera app failed with "variable-frame-rate video needs timestamp-aware playback". The file was not variable: its container reports avg_frame_rate 8670/299 and nb_frames 289 over a stream whose decoded timestamps are 280 frames exactly 1/30s apart. The guard compared two pieces of container metadata and rejected CFR video on the strength of a summary the container had got wrong about its own contents. The guard was also obsolete. It dates from when the page measured the source's own frames, where a wandering frame duration really does break `frame = floor(t * fps)`. Nothing measures the source now — ffmpeg resamples it onto a constant rate and the proxy is re-probed after it is written — so variable input is a thing this converts rather than a thing it refuses. So: probe picks a rate instead of validating one. It takes the nominal rate, which is the rate every timestamp in the stream can be expressed at and so the one that keeps every distinct source frame, and carries it as an exact fraction because 30000/1001 is not a float and a rounded -r is how a long take drifts. The disagreement is still recorded as `vfr`, just not fatal. The frame-count cross-check went with it. It compared the proxy against the source's nb_frames, which is the number this whole bug proves can lie, and a resample to a constant rate legitimately changes the count. It now checks the proxy's DURATION against the source's, because what must not drift is how long the picture lasts against how long the audio lasts. Verified on the reported file: 280 frames at 30fps, picture 9.3333s against audio 9.3167s — half a frame — and 280/280 detected in the real app. 41 backend tests green, including a genuinely variable fixture end to end. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-28 12:37:21 -04:00
# Constant frame rate at the rate `probe` chose. This RESAMPLES rather
# than asserts: the upload is allowed to be variable, and this is the
# step that makes the thing the page measures not be.
"-fps_mode", "cfr", "-r", facts.get("rate") or str(facts["fps"]),
Measure the video, not a PNG per frame Detection now walks a browser-seekable H.264 proxy in MediaPipe's VIDEO running mode. The PNG sequence it replaces was 112MB for 7.6 seconds at 1440x1920 and 1.1GB at the 900-frame limit; the proxy is 6MB, and landmarks detected off decoded H.264 rather than off the PNGs moved at most 0.0033 of frame width. Three things had to be true for video mode to work, and each was measured against the same footage decoded to PNGs: /blob/<digest> answers byte ranges. Django's FileResponse does no Range handling, and a media element handed 200 with no Accept-Ranges reports an empty `seekable`, no-ops every currentTime write, and detects frame one ninety times without raising. A seek aims at the MIDDLE of its frame. Aiming at i/fps sits on a frame boundary and landed one frame early 31 times in 91; (i + 0.5)/fps was exact on all 91. Timestamps are strictly increasing footage milliseconds. Video mode is a tracker: a repeat leaves the graph in an error state every later call re-throws, so the landmarker is discarded on failure, and passing the frame index instead of i*1000/fps moved landmarks six times further from the per-frame answer. Frames are verified rather than trusted. requestVideoFrameCallback states which frame it handed over, the walker discards any other and fails loudly if the one it asked for never arrives — a stale presentation from the tail of a previous seek is what produced "asked for frame 1 and it presented frame 2" on a video whose seeks were in fact exact. The proxy is re-encoded even when the upload is already H.264: HEVC is not decodable everywhere, and footage identity is the proxy's digest. The JPEG stills beside it are tracing references, outside the footage digest because re-rendering them at another size is not different footage. Verified end to end in a real browser against real footage: 228/228 frames detected, a drawn roto face, 37 backend and 234 frontend tests green. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-28 11:32:01 -04:00
"-c:v", "libx264", "-preset", "veryfast", "-crf", PROXY_CRF,
# NO B-FRAMES, AND THIS IS THE LOAD-BEARING FLAG. It is what makes
# decode order presentation order, so the page can treat access unit k
# of the elementary stream as frame k without demuxing a container or
# consulting a timestamp. With them x264 has a
Calibrate the video's frame origin instead of assuming it is zero "the video never presented frame 1; it offered 2" was this function waiting for a frame it had not asked for a second time. The walk discarded a wrong frame and re-armed the callback WITHOUT seeking again, and nothing further is ever presented to a paused element that has not been asked to move — so one wrong answer starved until the timeout and reported it as the browser refusing. Underneath that, the wrong answer was not wrong. A container can carry an edit list, and `currentTime` then counts from the start of the edited presentation while a frame's `mediaTime` counts from the start of the media. The two differ by a constant, so the frame at currentTime 0 can honestly report a mediaTime two frames in. Seeking cannot correct for it in the positive direction: source frame 0 would have to be found before the start of the video, every attempt clamps at zero, and the walk offers frame 2 forever. So the constant is measured once and subtracted. `calibrate!` takes whatever the browser calls the first frame it shows and makes that the origin, which is the honest definition anyway — it is what a viewer sees at time zero, and the audio clock this take plays against starts in the same place. Calibration also leaves the element on frame 0, so the walk starts holding it. `frame!` returns immediately for a frame already on screen rather than seeking to where it already is, which presents nothing and would hang. Residual error still re-seeks, corrected by exactly the measured miss, and still fails loudly with every frame that was offered and where it was asked from, so a next failure is diagnosable in one shot rather than four. The proxy also drops B-frames now. That removes the edit list at the source rather than only coping with it, and makes decode order presentation order should this ever be fed to WebCodecs. 14% larger, and extraction refuses a proxy whose timeline is shifted so it cannot come back silently. Honest note: this was my first diagnosis and it did NOT reproduce the failure — both proxies walk correctly here on hardware decode — so it is hardening, not the fix. Verified by forcing the fault: a harness that offsets the reported timeline by -2, -1, 0, +1, +2 and +5 frames failed on every positive offset before and recovers on all six now, first attempt. Real app in headed Chrome with Metal hardware decode: 280/280. 41 backend and 234 frontend tests green. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-28 12:50:11 -04:00
# two-frame reordering delay, ffmpeg compensates by writing an edit list
# (`elst` media_time 1024 at timebase 1/15360 — exactly two frames), and
# the browser then lives on two timelines at once: `currentTime` obeys the
# edit list and the `mediaTime` reported by requestVideoFrameCallback does
# not. Seek to frame 0 and the browser correctly hands back a frame whose
# mediaTime says 2. Software decoding hides it; hardware decoding does
# not, which is the worst possible way for it to be wrong. Without
# B-frames DTS equals PTS, no edit list is written, and the two timelines
# are the same one. It also makes decode order presentation order, should
# this ever be fed to a WebCodecs VideoDecoder.
"-bf", "0",
Measure the video, not a PNG per frame Detection now walks a browser-seekable H.264 proxy in MediaPipe's VIDEO running mode. The PNG sequence it replaces was 112MB for 7.6 seconds at 1440x1920 and 1.1GB at the 900-frame limit; the proxy is 6MB, and landmarks detected off decoded H.264 rather than off the PNGs moved at most 0.0033 of frame width. Three things had to be true for video mode to work, and each was measured against the same footage decoded to PNGs: /blob/<digest> answers byte ranges. Django's FileResponse does no Range handling, and a media element handed 200 with no Accept-Ranges reports an empty `seekable`, no-ops every currentTime write, and detects frame one ninety times without raising. A seek aims at the MIDDLE of its frame. Aiming at i/fps sits on a frame boundary and landed one frame early 31 times in 91; (i + 0.5)/fps was exact on all 91. Timestamps are strictly increasing footage milliseconds. Video mode is a tracker: a repeat leaves the graph in an error state every later call re-throws, so the landmarker is discarded on failure, and passing the frame index instead of i*1000/fps moved landmarks six times further from the per-frame answer. Frames are verified rather than trusted. requestVideoFrameCallback states which frame it handed over, the walker discards any other and fails loudly if the one it asked for never arrives — a stale presentation from the tail of a previous seek is what produced "asked for frame 1 and it presented frame 2" on a video whose seeks were in fact exact. The proxy is re-encoded even when the upload is already H.264: HEVC is not decodable everywhere, and footage identity is the proxy's digest. The JPEG stills beside it are tracing references, outside the footage digest because re-rendering them at another size is not different footage. Verified end to end in a real browser against real footage: 228/228 frames detected, a drawn roto face, 37 backend and 234 frontend tests green. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-28 11:32:01 -04:00
# yuv420p and an even frame size are what makes this playable everywhere
# rather than only in the browser that happened to be tested.
"-pix_fmt", "yuv420p", "-vf", "scale=trunc(iw/2)*2:trunc(ih/2)*2",
"-movflags", "+faststart", str(proxy_path)],
root, "proxy", total, (0, 55))
def _elementary_stream(proxy_path, out_path):
"""The proxy's video, unwrapped into a raw Annex-B H.264 stream.
A STREAM COPY, not a second encode: the same coded frames as the MP4, with
the container's length-prefixed NAL units rewritten as start-code-delimited
ones. It costs a file read and nothing else.
This exists because the page decodes with WebCodecs, and `VideoDecoder` takes
demuxed chunks rather than a container. Handing it Annex-B means the client
needs no demuxer: NAL start codes are findable in a loop, and because the
proxy is encoded with no B-frames, decode order is presentation order — so
access unit k IS frame k, with no container timing to consult and no clock to
reconcile. That is the whole reason this file is worth the bytes it costs.
"""
_command(["ffmpeg", "-hide_banner", "-loglevel", "error", "-y",
"-i", str(proxy_path), "-an", "-c:v", "copy",
"-bsf:v", "h264_mp4toannexb", "-f", "h264", str(out_path)])
Measure the video, not a PNG per frame Detection now walks a browser-seekable H.264 proxy in MediaPipe's VIDEO running mode. The PNG sequence it replaces was 112MB for 7.6 seconds at 1440x1920 and 1.1GB at the 900-frame limit; the proxy is 6MB, and landmarks detected off decoded H.264 rather than off the PNGs moved at most 0.0033 of frame width. Three things had to be true for video mode to work, and each was measured against the same footage decoded to PNGs: /blob/<digest> answers byte ranges. Django's FileResponse does no Range handling, and a media element handed 200 with no Accept-Ranges reports an empty `seekable`, no-ops every currentTime write, and detects frame one ninety times without raising. A seek aims at the MIDDLE of its frame. Aiming at i/fps sits on a frame boundary and landed one frame early 31 times in 91; (i + 0.5)/fps was exact on all 91. Timestamps are strictly increasing footage milliseconds. Video mode is a tracker: a repeat leaves the graph in an error state every later call re-throws, so the landmarker is discarded on failure, and passing the frame index instead of i*1000/fps moved landmarks six times further from the per-frame answer. Frames are verified rather than trusted. requestVideoFrameCallback states which frame it handed over, the walker discards any other and fails loudly if the one it asked for never arrives — a stale presentation from the tail of a previous seek is what produced "asked for frame 1 and it presented frame 2" on a video whose seeks were in fact exact. The proxy is re-encoded even when the upload is already H.264: HEVC is not decodable everywhere, and footage identity is the proxy's digest. The JPEG stills beside it are tracing references, outside the footage digest because re-rendering them at another size is not different footage. Verified end to end in a real browser against real footage: 228/228 frames detected, a drawn roto face, 37 backend and 234 frontend tests green. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-28 11:32:01 -04:00
def _extract_stills(job, proxy_path, frames_dir, frames, root):
"""The proxy -> one tracing JPEG per frame, long edge capped."""
_run_with_progress(
job,
["-i", str(proxy_path), "-fps_mode", "passthrough",
"-vf", f"scale='if(gt(iw,ih),min({TRACING_EDGE},iw),-2)':"
f"'if(gt(iw,ih),-2,min({TRACING_EDGE},ih))'",
"-q:v", TRACING_QUALITY, str(frames_dir / "%04d.jpg")],
root, "stills", frames, (55, 85))
Stop refusing ordinary phone footage as variable-frame-rate Uploading a clip shot straight from the iPhone camera app failed with "variable-frame-rate video needs timestamp-aware playback". The file was not variable: its container reports avg_frame_rate 8670/299 and nb_frames 289 over a stream whose decoded timestamps are 280 frames exactly 1/30s apart. The guard compared two pieces of container metadata and rejected CFR video on the strength of a summary the container had got wrong about its own contents. The guard was also obsolete. It dates from when the page measured the source's own frames, where a wandering frame duration really does break `frame = floor(t * fps)`. Nothing measures the source now — ffmpeg resamples it onto a constant rate and the proxy is re-probed after it is written — so variable input is a thing this converts rather than a thing it refuses. So: probe picks a rate instead of validating one. It takes the nominal rate, which is the rate every timestamp in the stream can be expressed at and so the one that keeps every distinct source frame, and carries it as an exact fraction because 30000/1001 is not a float and a rounded -r is how a long take drifts. The disagreement is still recorded as `vfr`, just not fatal. The frame-count cross-check went with it. It compared the proxy against the source's nb_frames, which is the number this whole bug proves can lie, and a resample to a constant rate legitimately changes the count. It now checks the proxy's DURATION against the source's, because what must not drift is how long the picture lasts against how long the audio lasts. Verified on the reported file: 280 frames at 30fps, picture 9.3333s against audio 9.3167s — half a frame — and 280/280 detected in the real app. 41 backend tests green, including a genuinely variable fixture end to end. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-28 12:37:21 -04:00
MAX_RATE = 120 # a capture rate; past this the container is describing something else
Measure the frame rate rather than believing the container `probe` took `r_frame_rate` whenever it was at or under the cap, on the grounds that it is the rate that keeps every distinct source frame. It is not a claim about frames at all: ordinary iPhone footage declares 120 over a stream whose timestamps are 1/30s apart, and resampling it up turned an 11-second clip into 1293 proxy frames instead of 323 — four times the encode, four times the tracing stills (91MB against 23MB), four times the blobs and the rows, for 970 frames that are copies of their neighbours. So `_measured_rate` reads the timestamps and `_choose_rate` keeps whichever declared rate they bear out. Two details carry it: the times are sorted before differencing, because an HEVC stream arrives in decode order and differencing that measures the reordering delay instead of the rate; and the statistic is the MEDIAN interval, which is what keeps the property the nominal rate was being taken for — a take held on one frame still reports the rate of the parts that move, so no distinct frame is dropped. A genuine 120fps capture still extracts at 120, and there is a test on that specifically. `Source.probe` also stopped being the place a reading goes to be preserved. The facts are a pure function of bytes that are the row's own identity, so a re-upload re-reads them: otherwise every already-uploaded source would have gone on resampling to four times the frames with no way to correct it short of deleting the row. Already-extracted footage is untouched — `extraction_key` still says scheme 3, so those jobs stay done and reachable. Bumping it re-extracts everything at the corrected rate. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-10-04 23:22:50 -04:00
# How many packet timestamps `_measured_rate` reads, and the fewest intervals it
# will draw a conclusion from. 300 is a flat cost on a long take and still a
# wide enough sample for a median; below 8 intervals there is not enough of a
# stream to outvote one odd timestamp, so the metadata is left to speak.
RATE_SAMPLE = 300
RATE_MINIMUM = 8
# How far a declared rate may sit from the measured one and still be taken as
# what the stream is: 2% covers 30 against 30000/1001 and nothing like 120
# against 30.
RATE_TOLERANCE = 0.02
Stop refusing ordinary phone footage as variable-frame-rate Uploading a clip shot straight from the iPhone camera app failed with "variable-frame-rate video needs timestamp-aware playback". The file was not variable: its container reports avg_frame_rate 8670/299 and nb_frames 289 over a stream whose decoded timestamps are 280 frames exactly 1/30s apart. The guard compared two pieces of container metadata and rejected CFR video on the strength of a summary the container had got wrong about its own contents. The guard was also obsolete. It dates from when the page measured the source's own frames, where a wandering frame duration really does break `frame = floor(t * fps)`. Nothing measures the source now — ffmpeg resamples it onto a constant rate and the proxy is re-probed after it is written — so variable input is a thing this converts rather than a thing it refuses. So: probe picks a rate instead of validating one. It takes the nominal rate, which is the rate every timestamp in the stream can be expressed at and so the one that keeps every distinct source frame, and carries it as an exact fraction because 30000/1001 is not a float and a rounded -r is how a long take drifts. The disagreement is still recorded as `vfr`, just not fatal. The frame-count cross-check went with it. It compared the proxy against the source's nb_frames, which is the number this whole bug proves can lie, and a resample to a constant rate legitimately changes the count. It now checks the proxy's DURATION against the source's, because what must not drift is how long the picture lasts against how long the audio lasts. Verified on the reported file: 280 frames at 30fps, picture 9.3333s against audio 9.3167s — half a frame — and 280/280 detected in the real app. 41 backend tests green, including a genuinely variable fixture end to end. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-28 12:37:21 -04:00
def probe_image(path):
"""An uploaded still's pixel size as (width, height), refusing anything that
is not one picture."""
data = json.loads(_command(["ffprobe", "-v", "error", "-show_streams",
"-of", "json", str(path)]))
video = [s for s in data.get("streams", []) if s.get("codec_type") == "video"]
if len(video) != 1 or not (video[0].get("width") and video[0].get("height")):
raise ValueError("the uploaded file is not an image")
return int(video[0]["width"]), int(video[0]["height"])
2026-09-30 03:31:19 -04:00
def probe_audio(path):
"""The length of an uploaded sound in seconds, refusing a file with no audio."""
data = json.loads(_command(["ffprobe", "-v", "error", "-show_streams",
"-show_format", "-of", "json", str(path)]))
if not any(s.get("codec_type") == "audio" for s in data.get("streams", [])):
raise ValueError("the uploaded file has no audio stream")
duration = float(data.get("format", {}).get("duration") or 0)
if duration <= 0:
raise ValueError("the sound's length is unknown")
return duration
Measure the frame rate rather than believing the container `probe` took `r_frame_rate` whenever it was at or under the cap, on the grounds that it is the rate that keeps every distinct source frame. It is not a claim about frames at all: ordinary iPhone footage declares 120 over a stream whose timestamps are 1/30s apart, and resampling it up turned an 11-second clip into 1293 proxy frames instead of 323 — four times the encode, four times the tracing stills (91MB against 23MB), four times the blobs and the rows, for 970 frames that are copies of their neighbours. So `_measured_rate` reads the timestamps and `_choose_rate` keeps whichever declared rate they bear out. Two details carry it: the times are sorted before differencing, because an HEVC stream arrives in decode order and differencing that measures the reordering delay instead of the rate; and the statistic is the MEDIAN interval, which is what keeps the property the nominal rate was being taken for — a take held on one frame still reports the rate of the parts that move, so no distinct frame is dropped. A genuine 120fps capture still extracts at 120, and there is a test on that specifically. `Source.probe` also stopped being the place a reading goes to be preserved. The facts are a pure function of bytes that are the row's own identity, so a re-upload re-reads them: otherwise every already-uploaded source would have gone on resampling to four times the frames with no way to correct it short of deleting the row. Already-extracted footage is untouched — `extraction_key` still says scheme 3, so those jobs stay done and reachable. Bumping it re-extracts everything at the corrected rate. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-10-04 23:22:50 -04:00
def _measured_rate(path):
"""The rate the stream's own packet timestamps imply, or None.
THE CONTAINER'S SUMMARY OF ITSELF IS NOT EVIDENCE, and this is the function
that goes and looks. An iPhone's `r_frame_rate` is 120 on footage whose
timestamps are 1/30s apart, which is the difference between 323 frames and
1293 — four times the encode, four times the tracing stills, four times the
blobs, for 970 frames that are copies of their neighbours.
It reads TIMESTAMPS, not frames: `-show_entries packet=pts_time` demuxes
without decoding, so this costs a file read and no pixels. The times are
SORTED before differencing because a stream with B-frames arrives in decode
order — an HEVC clip's first packets come out 0, 0.133, 0.067, 0.033 — and
differencing that order measures the reordering rather than the rate.
THE MEDIAN INTERVAL, which is what makes this safe on genuinely variable
input. It answers "how far apart are two frames normally", so a take held on
one frame for a second still reports the rate of the parts that move, and
choosing it keeps every distinct frame — the property `probe` used to reach
for by taking the nominal rate. Only the last few intervals of the sample are
unreliable (a frame whose turn comes after the window is missing from it), and
a median does not care.
Returning None is the honest answer for a clip too short to sample, and this
also swallows a probe that fails outright: the rate the metadata declares is
the documented fallback, so an optimisation must not be able to refuse an
upload that would otherwise have been accepted.
"""
try:
text = _command(["ffprobe", "-v", "error", "-select_streams", "v:0",
"-show_entries", "packet=pts_time", "-of", "json",
"-read_intervals", f"%+#{RATE_SAMPLE}", str(path)])
packets = json.loads(text).get("packets") or []
times = sorted(float(packet["pts_time"]) for packet in packets
if (packet.get("pts_time") or "N/A") != "N/A")
except (ValueError, OSError):
return None
intervals = sorted(b - a for a, b in zip(times, times[1:]) if b > a)
if len(intervals) < RATE_MINIMUM:
return None
median = intervals[len(intervals) // 2]
return 1.0 / median if median > 0 else None
def _choose_rate(nominal, average, measured):
"""The rate to resample onto, as an exact Fraction.
A DECLARED RATE IS PREFERRED WHEN IT AGREES WITH THE TIMESTAMPS, because it is
the exact rational the stream was authored at — 30000/1001 is not a float, and
`limit_denominator` on a measured 29.97 is a guess at a number the container
already states. So the measured rate is used to CHOOSE between what the
container declares, and only stands in itself when neither declaration
describes the stream.
"""
candidates = [rate for rate in (nominal, average) if 0 < rate <= MAX_RATE]
if measured:
agreeing = [rate for rate in candidates
if abs(float(rate) - measured) <= RATE_TOLERANCE * measured]
if agreeing:
return min(agreeing, key=lambda rate: abs(float(rate) - measured))
from_timestamps = Fraction(measured).limit_denominator(1001)
if 0 < from_timestamps <= MAX_RATE:
return from_timestamps
# Nothing to go on but the metadata, and nominal first keeps the rate that
# drops no distinct frame. An unusable pair falls through to the refusal
# below, which names the rate the file claimed rather than one of these.
return nominal if 0 < nominal <= MAX_RATE else average
def probe(path):
Stop refusing ordinary phone footage as variable-frame-rate Uploading a clip shot straight from the iPhone camera app failed with "variable-frame-rate video needs timestamp-aware playback". The file was not variable: its container reports avg_frame_rate 8670/299 and nb_frames 289 over a stream whose decoded timestamps are 280 frames exactly 1/30s apart. The guard compared two pieces of container metadata and rejected CFR video on the strength of a summary the container had got wrong about its own contents. The guard was also obsolete. It dates from when the page measured the source's own frames, where a wandering frame duration really does break `frame = floor(t * fps)`. Nothing measures the source now — ffmpeg resamples it onto a constant rate and the proxy is re-probed after it is written — so variable input is a thing this converts rather than a thing it refuses. So: probe picks a rate instead of validating one. It takes the nominal rate, which is the rate every timestamp in the stream can be expressed at and so the one that keeps every distinct source frame, and carries it as an exact fraction because 30000/1001 is not a float and a rounded -r is how a long take drifts. The disagreement is still recorded as `vfr`, just not fatal. The frame-count cross-check went with it. It compared the proxy against the source's nb_frames, which is the number this whole bug proves can lie, and a resample to a constant rate legitimately changes the count. It now checks the proxy's DURATION against the source's, because what must not drift is how long the picture lasts against how long the audio lasts. Verified on the reported file: 280 frames at 30fps, picture 9.3333s against audio 9.3167s — half a frame — and 280/280 detected in the real app. 41 backend tests green, including a genuinely variable fixture end to end. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-28 12:37:21 -04:00
"""What the upload is, as far as choosing a proxy rate goes.
IT NO LONGER REFUSES VARIABLE-FRAME-RATE INPUT, and the reason is the proxy.
That refusal was written when the page measured the source's own frames, where
a wandering frame duration really does break `frame = floor(t * fps)`. Nothing
measures the source now: ffmpeg resamples it onto a constant rate, and the
proxy — constant by construction, and re-probed after it is written — is the
only timeline anything downstream sees.
Keeping the check would have been worse than useless, because the thing it
tested is not reliable. Ordinary iPhone footage, shot straight from the camera
app, reports `avg_frame_rate` 8670/299 and `nb_frames` 289 on a stream whose
decoded timestamps are 280 frames exactly 1/30s apart. The container's summary
of itself disagreed with the container's own contents, so the guard rejected
CFR video for being variable.
Measure the frame rate rather than believing the container `probe` took `r_frame_rate` whenever it was at or under the cap, on the grounds that it is the rate that keeps every distinct source frame. It is not a claim about frames at all: ordinary iPhone footage declares 120 over a stream whose timestamps are 1/30s apart, and resampling it up turned an 11-second clip into 1293 proxy frames instead of 323 — four times the encode, four times the tracing stills (91MB against 23MB), four times the blobs and the rows, for 970 frames that are copies of their neighbours. So `_measured_rate` reads the timestamps and `_choose_rate` keeps whichever declared rate they bear out. Two details carry it: the times are sorted before differencing, because an HEVC stream arrives in decode order and differencing that measures the reordering delay instead of the rate; and the statistic is the MEDIAN interval, which is what keeps the property the nominal rate was being taken for — a take held on one frame still reports the rate of the parts that move, so no distinct frame is dropped. A genuine 120fps capture still extracts at 120, and there is a test on that specifically. `Source.probe` also stopped being the place a reading goes to be preserved. The facts are a pure function of bytes that are the row's own identity, so a re-upload re-reads them: otherwise every already-uploaded source would have gone on resampling to four times the frames with no way to correct it short of deleting the row. Already-extracted footage is untouched — `extraction_key` still says scheme 3, so those jobs stay done and reachable. Bumping it re-extracts everything at the corrected rate. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-10-04 23:22:50 -04:00
THE RATE IS MEASURED AND THE DECLARATIONS ARE VOTED ON, which is the same
distrust applied to the one number that still comes from here. This used to
take `r_frame_rate` outright — the rate every timestamp in the stream can be
expressed at, and so the rate that keeps every distinct source frame. The
trouble is that it is not a claim about frames at all: the file above declares
120 and holds 30, and resampling it up cost four times the encode, four times
the tracing stills and four times the blobs for 970 duplicated frames. So
`_measured_rate` reads the timestamps, `_choose_rate` keeps whichever declared
rate they bear out, and the nominal rate is believed when it is true rather
than because it is nominal.
Duration is preserved either way — ffmpeg's CFR conversion is driven by
timestamps, so the audio stays in sync at any rate — and the median interval
keeps the no-distinct-frame-dropped property that taking the nominal rate was
reaching for. See `_measured_rate`.
Stop refusing ordinary phone footage as variable-frame-rate Uploading a clip shot straight from the iPhone camera app failed with "variable-frame-rate video needs timestamp-aware playback". The file was not variable: its container reports avg_frame_rate 8670/299 and nb_frames 289 over a stream whose decoded timestamps are 280 frames exactly 1/30s apart. The guard compared two pieces of container metadata and rejected CFR video on the strength of a summary the container had got wrong about its own contents. The guard was also obsolete. It dates from when the page measured the source's own frames, where a wandering frame duration really does break `frame = floor(t * fps)`. Nothing measures the source now — ffmpeg resamples it onto a constant rate and the proxy is re-probed after it is written — so variable input is a thing this converts rather than a thing it refuses. So: probe picks a rate instead of validating one. It takes the nominal rate, which is the rate every timestamp in the stream can be expressed at and so the one that keeps every distinct source frame, and carries it as an exact fraction because 30000/1001 is not a float and a rounded -r is how a long take drifts. The disagreement is still recorded as `vfr`, just not fatal. The frame-count cross-check went with it. It compared the proxy against the source's nb_frames, which is the number this whole bug proves can lie, and a resample to a constant rate legitimately changes the count. It now checks the proxy's DURATION against the source's, because what must not drift is how long the picture lasts against how long the audio lasts. Verified on the reported file: 280 frames at 30fps, picture 9.3333s against audio 9.3167s — half a frame — and 280/280 detected in the real app. 41 backend tests green, including a genuinely variable fixture end to end. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-28 12:37:21 -04:00
"""
data = json.loads(_command(["ffprobe", "-v", "error", "-show_streams",
"-show_format", "-of", "json", str(path)]))
video = next((s for s in data.get("streams", []) if s.get("codec_type") == "video"), None)
if not video:
raise ValueError("the uploaded file has no video stream")
nominal = Fraction(video.get("r_frame_rate") or "0")
average = Fraction(video.get("avg_frame_rate") or "0")
Stop refusing ordinary phone footage as variable-frame-rate Uploading a clip shot straight from the iPhone camera app failed with "variable-frame-rate video needs timestamp-aware playback". The file was not variable: its container reports avg_frame_rate 8670/299 and nb_frames 289 over a stream whose decoded timestamps are 280 frames exactly 1/30s apart. The guard compared two pieces of container metadata and rejected CFR video on the strength of a summary the container had got wrong about its own contents. The guard was also obsolete. It dates from when the page measured the source's own frames, where a wandering frame duration really does break `frame = floor(t * fps)`. Nothing measures the source now — ffmpeg resamples it onto a constant rate and the proxy is re-probed after it is written — so variable input is a thing this converts rather than a thing it refuses. So: probe picks a rate instead of validating one. It takes the nominal rate, which is the rate every timestamp in the stream can be expressed at and so the one that keeps every distinct source frame, and carries it as an exact fraction because 30000/1001 is not a float and a rounded -r is how a long take drifts. The disagreement is still recorded as `vfr`, just not fatal. The frame-count cross-check went with it. It compared the proxy against the source's nb_frames, which is the number this whole bug proves can lie, and a resample to a constant rate legitimately changes the count. It now checks the proxy's DURATION against the source's, because what must not drift is how long the picture lasts against how long the audio lasts. Verified on the reported file: 280 frames at 30fps, picture 9.3333s against audio 9.3167s — half a frame — and 280/280 detected in the real app. 41 backend tests green, including a genuinely variable fixture end to end. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-28 12:37:21 -04:00
if nominal <= 0 and average <= 0:
raise ValueError("the video's frame rate is unknown")
Measure the frame rate rather than believing the container `probe` took `r_frame_rate` whenever it was at or under the cap, on the grounds that it is the rate that keeps every distinct source frame. It is not a claim about frames at all: ordinary iPhone footage declares 120 over a stream whose timestamps are 1/30s apart, and resampling it up turned an 11-second clip into 1293 proxy frames instead of 323 — four times the encode, four times the tracing stills (91MB against 23MB), four times the blobs and the rows, for 970 frames that are copies of their neighbours. So `_measured_rate` reads the timestamps and `_choose_rate` keeps whichever declared rate they bear out. Two details carry it: the times are sorted before differencing, because an HEVC stream arrives in decode order and differencing that measures the reordering delay instead of the rate; and the statistic is the MEDIAN interval, which is what keeps the property the nominal rate was being taken for — a take held on one frame still reports the rate of the parts that move, so no distinct frame is dropped. A genuine 120fps capture still extracts at 120, and there is a test on that specifically. `Source.probe` also stopped being the place a reading goes to be preserved. The facts are a pure function of bytes that are the row's own identity, so a re-upload re-reads them: otherwise every already-uploaded source would have gone on resampling to four times the frames with no way to correct it short of deleting the row. Already-extracted footage is untouched — `extraction_key` still says scheme 3, so those jobs stay done and reachable. Bumping it re-extracts everything at the corrected rate. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-10-04 23:22:50 -04:00
measured = _measured_rate(path)
rate = _choose_rate(nominal, average, measured)
Stop refusing ordinary phone footage as variable-frame-rate Uploading a clip shot straight from the iPhone camera app failed with "variable-frame-rate video needs timestamp-aware playback". The file was not variable: its container reports avg_frame_rate 8670/299 and nb_frames 289 over a stream whose decoded timestamps are 280 frames exactly 1/30s apart. The guard compared two pieces of container metadata and rejected CFR video on the strength of a summary the container had got wrong about its own contents. The guard was also obsolete. It dates from when the page measured the source's own frames, where a wandering frame duration really does break `frame = floor(t * fps)`. Nothing measures the source now — ffmpeg resamples it onto a constant rate and the proxy is re-probed after it is written — so variable input is a thing this converts rather than a thing it refuses. So: probe picks a rate instead of validating one. It takes the nominal rate, which is the rate every timestamp in the stream can be expressed at and so the one that keeps every distinct source frame, and carries it as an exact fraction because 30000/1001 is not a float and a rounded -r is how a long take drifts. The disagreement is still recorded as `vfr`, just not fatal. The frame-count cross-check went with it. It compared the proxy against the source's nb_frames, which is the number this whole bug proves can lie, and a resample to a constant rate legitimately changes the count. It now checks the proxy's DURATION against the source's, because what must not drift is how long the picture lasts against how long the audio lasts. Verified on the reported file: 280 frames at 30fps, picture 9.3333s against audio 9.3167s — half a frame — and 280/280 detected in the real app. 41 backend tests green, including a genuinely variable fixture end to end. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-28 12:37:21 -04:00
if not 0 < rate <= MAX_RATE:
raise ValueError(f"the video reports a frame rate of {float(rate):g}, which is "
"not a rate footage can be measured at")
duration = float(data.get("format", {}).get("duration") or 0)
2026-10-03 00:39:07 -04:00
# if duration > 0 and duration * float(rate) > 901:
# raise ValueError("video is longer than the 900-frame footage limit")
Stop refusing ordinary phone footage as variable-frame-rate Uploading a clip shot straight from the iPhone camera app failed with "variable-frame-rate video needs timestamp-aware playback". The file was not variable: its container reports avg_frame_rate 8670/299 and nb_frames 289 over a stream whose decoded timestamps are 280 frames exactly 1/30s apart. The guard compared two pieces of container metadata and rejected CFR video on the strength of a summary the container had got wrong about its own contents. The guard was also obsolete. It dates from when the page measured the source's own frames, where a wandering frame duration really does break `frame = floor(t * fps)`. Nothing measures the source now — ffmpeg resamples it onto a constant rate and the proxy is re-probed after it is written — so variable input is a thing this converts rather than a thing it refuses. So: probe picks a rate instead of validating one. It takes the nominal rate, which is the rate every timestamp in the stream can be expressed at and so the one that keeps every distinct source frame, and carries it as an exact fraction because 30000/1001 is not a float and a rounded -r is how a long take drifts. The disagreement is still recorded as `vfr`, just not fatal. The frame-count cross-check went with it. It compared the proxy against the source's nb_frames, which is the number this whole bug proves can lie, and a resample to a constant rate legitimately changes the count. It now checks the proxy's DURATION against the source's, because what must not drift is how long the picture lasts against how long the audio lasts. Verified on the reported file: 280 frames at 30fps, picture 9.3333s against audio 9.3167s — half a frame — and 280/280 detected in the real app. 41 backend tests green, including a genuinely variable fixture end to end. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-28 12:37:21 -04:00
frames = video.get("nb_frames")
return {"fps": float(rate),
# The exact rate, for ffmpeg. 30000/1001 is not a float, and handing
# `-r` a rounded one is how a long take drifts out of sync.
"rate": f"{rate.numerator}/{rate.denominator}",
"nominal_fps": float(nominal), "average_fps": float(average),
Measure the frame rate rather than believing the container `probe` took `r_frame_rate` whenever it was at or under the cap, on the grounds that it is the rate that keeps every distinct source frame. It is not a claim about frames at all: ordinary iPhone footage declares 120 over a stream whose timestamps are 1/30s apart, and resampling it up turned an 11-second clip into 1293 proxy frames instead of 323 — four times the encode, four times the tracing stills (91MB against 23MB), four times the blobs and the rows, for 970 frames that are copies of their neighbours. So `_measured_rate` reads the timestamps and `_choose_rate` keeps whichever declared rate they bear out. Two details carry it: the times are sorted before differencing, because an HEVC stream arrives in decode order and differencing that measures the reordering delay instead of the rate; and the statistic is the MEDIAN interval, which is what keeps the property the nominal rate was being taken for — a take held on one frame still reports the rate of the parts that move, so no distinct frame is dropped. A genuine 120fps capture still extracts at 120, and there is a test on that specifically. `Source.probe` also stopped being the place a reading goes to be preserved. The facts are a pure function of bytes that are the row's own identity, so a re-upload re-reads them: otherwise every already-uploaded source would have gone on resampling to four times the frames with no way to correct it short of deleting the row. Already-extracted footage is untouched — `extraction_key` still says scheme 3, so those jobs stay done and reachable. Bumping it re-extracts everything at the corrected rate. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-10-04 23:22:50 -04:00
# What the timestamps said, and null when there were too few to ask.
# Recorded because it is the input to a decision this file used not to
# make, and the one number that explains a chosen rate matching
# neither declaration.
"measured_fps": measured,
"width": int(video["width"]), "height": int(video["height"]),
"duration": duration,
Stop refusing ordinary phone footage as variable-frame-rate Uploading a clip shot straight from the iPhone camera app failed with "variable-frame-rate video needs timestamp-aware playback". The file was not variable: its container reports avg_frame_rate 8670/299 and nb_frames 289 over a stream whose decoded timestamps are 280 frames exactly 1/30s apart. The guard compared two pieces of container metadata and rejected CFR video on the strength of a summary the container had got wrong about its own contents. The guard was also obsolete. It dates from when the page measured the source's own frames, where a wandering frame duration really does break `frame = floor(t * fps)`. Nothing measures the source now — ffmpeg resamples it onto a constant rate and the proxy is re-probed after it is written — so variable input is a thing this converts rather than a thing it refuses. So: probe picks a rate instead of validating one. It takes the nominal rate, which is the rate every timestamp in the stream can be expressed at and so the one that keeps every distinct source frame, and carries it as an exact fraction because 30000/1001 is not a float and a rounded -r is how a long take drifts. The disagreement is still recorded as `vfr`, just not fatal. The frame-count cross-check went with it. It compared the proxy against the source's nb_frames, which is the number this whole bug proves can lie, and a resample to a constant rate legitimately changes the count. It now checks the proxy's DURATION against the source's, because what must not drift is how long the picture lasts against how long the audio lasts. Verified on the reported file: 280 frames at 30fps, picture 9.3333s against audio 9.3167s — half a frame — and 280/280 detected in the real app. 41 backend tests green, including a genuinely variable fixture end to end. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-28 12:37:21 -04:00
# KEPT, AND NO LONGER TRUSTED AS A COUNT. See the docstring: this is
# the container's claim about itself, it is wrong on ordinary phone
# footage, and `run` checks the proxy's DURATION instead.
"reported_frames": int(frames) if frames and frames.isdigit() else None,
"has_audio": any(s.get("codec_type") == "audio" for s in data.get("streams", [])),
Stop refusing ordinary phone footage as variable-frame-rate Uploading a clip shot straight from the iPhone camera app failed with "variable-frame-rate video needs timestamp-aware playback". The file was not variable: its container reports avg_frame_rate 8670/299 and nb_frames 289 over a stream whose decoded timestamps are 280 frames exactly 1/30s apart. The guard compared two pieces of container metadata and rejected CFR video on the strength of a summary the container had got wrong about its own contents. The guard was also obsolete. It dates from when the page measured the source's own frames, where a wandering frame duration really does break `frame = floor(t * fps)`. Nothing measures the source now — ffmpeg resamples it onto a constant rate and the proxy is re-probed after it is written — so variable input is a thing this converts rather than a thing it refuses. So: probe picks a rate instead of validating one. It takes the nominal rate, which is the rate every timestamp in the stream can be expressed at and so the one that keeps every distinct source frame, and carries it as an exact fraction because 30000/1001 is not a float and a rounded -r is how a long take drifts. The disagreement is still recorded as `vfr`, just not fatal. The frame-count cross-check went with it. It compared the proxy against the source's nb_frames, which is the number this whole bug proves can lie, and a resample to a constant rate legitimately changes the count. It now checks the proxy's DURATION against the source's, because what must not drift is how long the picture lasts against how long the audio lasts. Verified on the reported file: 280 frames at 30fps, picture 9.3333s against audio 9.3167s — half a frame — and 280/280 detected in the real app. 41 backend tests green, including a genuinely variable fixture end to end. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-28 12:37:21 -04:00
"vfr": nominal != average}
Calibrate the video's frame origin instead of assuming it is zero "the video never presented frame 1; it offered 2" was this function waiting for a frame it had not asked for a second time. The walk discarded a wrong frame and re-armed the callback WITHOUT seeking again, and nothing further is ever presented to a paused element that has not been asked to move — so one wrong answer starved until the timeout and reported it as the browser refusing. Underneath that, the wrong answer was not wrong. A container can carry an edit list, and `currentTime` then counts from the start of the edited presentation while a frame's `mediaTime` counts from the start of the media. The two differ by a constant, so the frame at currentTime 0 can honestly report a mediaTime two frames in. Seeking cannot correct for it in the positive direction: source frame 0 would have to be found before the start of the video, every attempt clamps at zero, and the walk offers frame 2 forever. So the constant is measured once and subtracted. `calibrate!` takes whatever the browser calls the first frame it shows and makes that the origin, which is the honest definition anyway — it is what a viewer sees at time zero, and the audio clock this take plays against starts in the same place. Calibration also leaves the element on frame 0, so the walk starts holding it. `frame!` returns immediately for a frame already on screen rather than seeking to where it already is, which presents nothing and would hang. Residual error still re-seeks, corrected by exactly the measured miss, and still fails loudly with every frame that was offered and where it was asked from, so a next failure is diagnosable in one shot rather than four. The proxy also drops B-frames now. That removes the edit list at the source rather than only coping with it, and makes decode order presentation order should this ever be fed to WebCodecs. 14% larger, and extraction refuses a proxy whose timeline is shifted so it cannot come back silently. Honest note: this was my first diagnosis and it did NOT reproduce the failure — both proxies walk correctly here on hardware decode — so it is hardening, not the fix. Verified by forcing the fault: a harness that offsets the reported timeline by -2, -1, 0, +1, +2 and +5 frames failed on every positive offset before and recovers on all six now, first attempt. Real app in headed Chrome with Metal hardware decode: 280/280. 41 backend and 234 frontend tests green. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-28 12:50:11 -04:00
def _refuse_a_shifted_timeline(path):
"""The proxy must put frame `i` at `i / fps` on BOTH of the browser's clocks.
Asserted rather than assumed, because the failure is silent and the symptom is
unrecognisable. An encoder delay makes ffmpeg write an edit list, `currentTime`
then obeys it while `requestVideoFrameCallback`'s `mediaTime` does not, and the
page's frame walk is uniformly off by the delay — on hardware decoding only. It
cost two wrong diagnoses to find, so it does not get to come back silently if
somebody changes an encoder flag.
"""
data = json.loads(_command(["ffprobe", "-v", "error", "-select_streams", "v:0",
"-show_streams", "-of", "json", str(path)]))
stream = data["streams"][0]
if int(stream.get("has_b_frames") or 0):
raise ValueError(
"the proxy was encoded with B-frames, whose reordering delay makes the "
"browser's seek clock and its frame-timestamp clock disagree")
if float(stream.get("start_time") or 0) != 0:
raise ValueError(f"the proxy starts at {stream['start_time']}s rather than 0")
Measure the video, not a PNG per frame Detection now walks a browser-seekable H.264 proxy in MediaPipe's VIDEO running mode. The PNG sequence it replaces was 112MB for 7.6 seconds at 1440x1920 and 1.1GB at the 900-frame limit; the proxy is 6MB, and landmarks detected off decoded H.264 rather than off the PNGs moved at most 0.0033 of frame width. Three things had to be true for video mode to work, and each was measured against the same footage decoded to PNGs: /blob/<digest> answers byte ranges. Django's FileResponse does no Range handling, and a media element handed 200 with no Accept-Ranges reports an empty `seekable`, no-ops every currentTime write, and detects frame one ninety times without raising. A seek aims at the MIDDLE of its frame. Aiming at i/fps sits on a frame boundary and landed one frame early 31 times in 91; (i + 0.5)/fps was exact on all 91. Timestamps are strictly increasing footage milliseconds. Video mode is a tracker: a repeat leaves the graph in an error state every later call re-throws, so the landmarker is discarded on failure, and passing the frame index instead of i*1000/fps moved landmarks six times further from the per-frame answer. Frames are verified rather than trusted. requestVideoFrameCallback states which frame it handed over, the walker discards any other and fails loudly if the one it asked for never arrives — a stale presentation from the tail of a previous seek is what produced "asked for frame 1 and it presented frame 2" on a video whose seeks were in fact exact. The proxy is re-encoded even when the upload is already H.264: HEVC is not decodable everywhere, and footage identity is the proxy's digest. The JPEG stills beside it are tracing references, outside the footage digest because re-rendering them at another size is not different footage. Verified end to end in a real browser against real footage: 228/228 frames detected, a drawn roto face, 37 backend and 234 frontend tests green. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-28 11:32:01 -04:00
def count_frames(path):
"""How many frames a file really holds, counted rather than reported.
`nb_frames` is a container's claim. This is the decoder's answer, and it is
what the page will get when it walks the proxy — so a disagreement between the
two has to be settled before the count reaches a manifest, not after it has
become a one-frame audio offset nobody can find.
"""
text = _command(["ffprobe", "-v", "error", "-select_streams", "v:0",
"-count_frames", "-show_entries", "stream=nb_read_frames",
"-of", "default=nokey=1:noprint_wrappers=1", str(path)])
counted = text.strip()
if not counted.isdigit():
raise ValueError("could not count the proxy's frames")
return int(counted)
def extraction_key(source, settings):
# Scheme 3: the extraction now also produces the elementary stream the page
# decodes, so a job run under scheme 2 did not make everything this one does.
text = json.dumps({"scheme": 3, "source": source.blob_id, "settings": settings},
sort_keys=True, separators=(",", ":"))
return "sha256:" + hashlib.sha256(text.encode()).hexdigest()
def _register(job, proxy_path, stream_path, stills, audio_path, facts):
Measure the video, not a PNG per frame Detection now walks a browser-seekable H.264 proxy in MediaPipe's VIDEO running mode. The PNG sequence it replaces was 112MB for 7.6 seconds at 1440x1920 and 1.1GB at the 900-frame limit; the proxy is 6MB, and landmarks detected off decoded H.264 rather than off the PNGs moved at most 0.0033 of frame width. Three things had to be true for video mode to work, and each was measured against the same footage decoded to PNGs: /blob/<digest> answers byte ranges. Django's FileResponse does no Range handling, and a media element handed 200 with no Accept-Ranges reports an empty `seekable`, no-ops every currentTime write, and detects frame one ninety times without raising. A seek aims at the MIDDLE of its frame. Aiming at i/fps sits on a frame boundary and landed one frame early 31 times in 91; (i + 0.5)/fps was exact on all 91. Timestamps are strictly increasing footage milliseconds. Video mode is a tracker: a repeat leaves the graph in an error state every later call re-throws, so the landmarker is discarded on failure, and passing the frame index instead of i*1000/fps moved landmarks six times further from the per-frame answer. Frames are verified rather than trusted. requestVideoFrameCallback states which frame it handed over, the walker discards any other and fails loudly if the one it asked for never arrives — a stale presentation from the tail of a previous seek is what produced "asked for frame 1 and it presented frame 2" on a video whose seeks were in fact exact. The proxy is re-encoded even when the upload is already H.264: HEVC is not decodable everywhere, and footage identity is the proxy's digest. The JPEG stills beside it are tracing references, outside the footage digest because re-rendering them at another size is not different footage. Verified end to end in a real browser against real footage: 228/228 frames detected, a drawn roto face, 37 backend and 234 frontend tests green. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-28 11:32:01 -04:00
proxy_digest, proxy_size = blobs.adopt(proxy_path)
stream_digest, stream_size = blobs.adopt(stream_path)
audio_digest, audio_size = blobs.adopt(audio_path)
Measure the video, not a PNG per frame Detection now walks a browser-seekable H.264 proxy in MediaPipe's VIDEO running mode. The PNG sequence it replaces was 112MB for 7.6 seconds at 1440x1920 and 1.1GB at the 900-frame limit; the proxy is 6MB, and landmarks detected off decoded H.264 rather than off the PNGs moved at most 0.0033 of frame width. Three things had to be true for video mode to work, and each was measured against the same footage decoded to PNGs: /blob/<digest> answers byte ranges. Django's FileResponse does no Range handling, and a media element handed 200 with no Accept-Ranges reports an empty `seekable`, no-ops every currentTime write, and detects frame one ninety times without raising. A seek aims at the MIDDLE of its frame. Aiming at i/fps sits on a frame boundary and landed one frame early 31 times in 91; (i + 0.5)/fps was exact on all 91. Timestamps are strictly increasing footage milliseconds. Video mode is a tracker: a repeat leaves the graph in an error state every later call re-throws, so the landmarker is discarded on failure, and passing the frame index instead of i*1000/fps moved landmarks six times further from the per-frame answer. Frames are verified rather than trusted. requestVideoFrameCallback states which frame it handed over, the walker discards any other and fails loudly if the one it asked for never arrives — a stale presentation from the tail of a previous seek is what produced "asked for frame 1 and it presented frame 2" on a video whose seeks were in fact exact. The proxy is re-encoded even when the upload is already H.264: HEVC is not decodable everywhere, and footage identity is the proxy's digest. The JPEG stills beside it are tracing references, outside the footage digest because re-rendering them at another size is not different footage. Verified end to end in a real browser against real footage: 228/228 frames detected, a drawn roto face, 37 backend and 234 frontend tests green. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-28 11:32:01 -04:00
still_blobs = [(index, *blobs.adopt(path)) for index, path in enumerate(stills)]
width, height, fps, frames = facts["width"], facts["height"], facts["fps"], facts["frames"]
# The footage's own identity: the bytes the page will measure, the audio it
# will clock against, and the rate that ties them together. Scheme 2 — scheme
# 1 hashed a PNG per frame, and those footages name pixels this no longer has.
h = hashlib.sha256()
Measure the video, not a PNG per frame Detection now walks a browser-seekable H.264 proxy in MediaPipe's VIDEO running mode. The PNG sequence it replaces was 112MB for 7.6 seconds at 1440x1920 and 1.1GB at the 900-frame limit; the proxy is 6MB, and landmarks detected off decoded H.264 rather than off the PNGs moved at most 0.0033 of frame width. Three things had to be true for video mode to work, and each was measured against the same footage decoded to PNGs: /blob/<digest> answers byte ranges. Django's FileResponse does no Range handling, and a media element handed 200 with no Accept-Ranges reports an empty `seekable`, no-ops every currentTime write, and detects frame one ninety times without raising. A seek aims at the MIDDLE of its frame. Aiming at i/fps sits on a frame boundary and landed one frame early 31 times in 91; (i + 0.5)/fps was exact on all 91. Timestamps are strictly increasing footage milliseconds. Video mode is a tracker: a repeat leaves the graph in an error state every later call re-throws, so the landmarker is discarded on failure, and passing the frame index instead of i*1000/fps moved landmarks six times further from the per-frame answer. Frames are verified rather than trusted. requestVideoFrameCallback states which frame it handed over, the walker discards any other and fails loudly if the one it asked for never arrives — a stale presentation from the tail of a previous seek is what produced "asked for frame 1 and it presented frame 2" on a video whose seeks were in fact exact. The proxy is re-encoded even when the upload is already H.264: HEVC is not decodable everywhere, and footage identity is the proxy's digest. The JPEG stills beside it are tracing references, outside the footage digest because re-rendering them at another size is not different footage. Verified end to end in a real browser against real footage: 228/228 frames detected, a drawn roto face, 37 backend and 234 frontend tests green. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-28 11:32:01 -04:00
h.update(f"arthur-footage-2/{fps}/{frames}/{width}x{height}\n".encode())
h.update(proxy_digest.encode())
h.update(audio_digest.encode())
Measure the video, not a PNG per frame Detection now walks a browser-seekable H.264 proxy in MediaPipe's VIDEO running mode. The PNG sequence it replaces was 112MB for 7.6 seconds at 1440x1920 and 1.1GB at the 900-frame limit; the proxy is 6MB, and landmarks detected off decoded H.264 rather than off the PNGs moved at most 0.0033 of frame width. Three things had to be true for video mode to work, and each was measured against the same footage decoded to PNGs: /blob/<digest> answers byte ranges. Django's FileResponse does no Range handling, and a media element handed 200 with no Accept-Ranges reports an empty `seekable`, no-ops every currentTime write, and detects frame one ninety times without raising. A seek aims at the MIDDLE of its frame. Aiming at i/fps sits on a frame boundary and landed one frame early 31 times in 91; (i + 0.5)/fps was exact on all 91. Timestamps are strictly increasing footage milliseconds. Video mode is a tracker: a repeat leaves the graph in an error state every later call re-throws, so the landmarker is discarded on failure, and passing the frame index instead of i*1000/fps moved landmarks six times further from the per-frame answer. Frames are verified rather than trusted. requestVideoFrameCallback states which frame it handed over, the walker discards any other and fails loudly if the one it asked for never arrives — a stale presentation from the tail of a previous seek is what produced "asked for frame 1 and it presented frame 2" on a video whose seeks were in fact exact. The proxy is re-encoded even when the upload is already H.264: HEVC is not decodable everywhere, and footage identity is the proxy's digest. The JPEG stills beside it are tracing references, outside the footage digest because re-rendering them at another size is not different footage. Verified end to end in a real browser against real footage: 228/228 frames detected, a drawn roto face, 37 backend and 234 frontend tests green. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-28 11:32:01 -04:00
with transaction.atomic():
Measure the video, not a PNG per frame Detection now walks a browser-seekable H.264 proxy in MediaPipe's VIDEO running mode. The PNG sequence it replaces was 112MB for 7.6 seconds at 1440x1920 and 1.1GB at the 900-frame limit; the proxy is 6MB, and landmarks detected off decoded H.264 rather than off the PNGs moved at most 0.0033 of frame width. Three things had to be true for video mode to work, and each was measured against the same footage decoded to PNGs: /blob/<digest> answers byte ranges. Django's FileResponse does no Range handling, and a media element handed 200 with no Accept-Ranges reports an empty `seekable`, no-ops every currentTime write, and detects frame one ninety times without raising. A seek aims at the MIDDLE of its frame. Aiming at i/fps sits on a frame boundary and landed one frame early 31 times in 91; (i + 0.5)/fps was exact on all 91. Timestamps are strictly increasing footage milliseconds. Video mode is a tracker: a repeat leaves the graph in an error state every later call re-throws, so the landmarker is discarded on failure, and passing the frame index instead of i*1000/fps moved landmarks six times further from the per-frame answer. Frames are verified rather than trusted. requestVideoFrameCallback states which frame it handed over, the walker discards any other and fails loudly if the one it asked for never arrives — a stale presentation from the tail of a previous seek is what produced "asked for frame 1 and it presented frame 2" on a video whose seeks were in fact exact. The proxy is re-encoded even when the upload is already H.264: HEVC is not decodable everywhere, and footage identity is the proxy's digest. The JPEG stills beside it are tracing references, outside the footage digest because re-rendering them at another size is not different footage. Verified end to end in a real browser against real footage: 228/228 frames detected, a drawn roto face, 37 backend and 234 frontend tests green. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-28 11:32:01 -04:00
proxy_blob, _ = Blob.objects.get_or_create(
digest=proxy_digest, defaults={"size": proxy_size, "media_type": "video/mp4"})
stream_blob, _ = Blob.objects.get_or_create(
digest=stream_digest, defaults={"size": stream_size, "media_type": "video/h264"})
audio_blob, _ = Blob.objects.get_or_create(
digest=audio_digest, defaults={"size": audio_size, "media_type": "audio/wav"})
footage, created = Footage.objects.get_or_create(
digest=h.hexdigest(),
defaults={"label": job.source.filename[:200], "source": job.source.filename[:200],
Measure the video, not a PNG per frame Detection now walks a browser-seekable H.264 proxy in MediaPipe's VIDEO running mode. The PNG sequence it replaces was 112MB for 7.6 seconds at 1440x1920 and 1.1GB at the 900-frame limit; the proxy is 6MB, and landmarks detected off decoded H.264 rather than off the PNGs moved at most 0.0033 of frame width. Three things had to be true for video mode to work, and each was measured against the same footage decoded to PNGs: /blob/<digest> answers byte ranges. Django's FileResponse does no Range handling, and a media element handed 200 with no Accept-Ranges reports an empty `seekable`, no-ops every currentTime write, and detects frame one ninety times without raising. A seek aims at the MIDDLE of its frame. Aiming at i/fps sits on a frame boundary and landed one frame early 31 times in 91; (i + 0.5)/fps was exact on all 91. Timestamps are strictly increasing footage milliseconds. Video mode is a tracker: a repeat leaves the graph in an error state every later call re-throws, so the landmarker is discarded on failure, and passing the frame index instead of i*1000/fps moved landmarks six times further from the per-frame answer. Frames are verified rather than trusted. requestVideoFrameCallback states which frame it handed over, the walker discards any other and fails loudly if the one it asked for never arrives — a stale presentation from the tail of a previous seek is what produced "asked for frame 1 and it presented frame 2" on a video whose seeks were in fact exact. The proxy is re-encoded even when the upload is already H.264: HEVC is not decodable everywhere, and footage identity is the proxy's digest. The JPEG stills beside it are tracing references, outside the footage digest because re-rendering them at another size is not different footage. Verified end to end in a real browser against real footage: 228/228 frames detected, a drawn roto face, 37 backend and 234 frontend tests green. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-28 11:32:01 -04:00
"fps": fps, "frames": frames, "width": width, "height": height,
"audio": audio_blob, "video": proxy_blob, "stream": stream_blob})
if not created and not footage.stream_id:
# The same footage by identity, extracted before the elementary
# stream existed. Its digest is over the proxy and the audio, which
# have not changed — so this is the same footage gaining a file it
# was always entitled to, not a different one.
footage.stream = stream_blob
if not footage.video_id:
footage.video = proxy_blob
footage.save(update_fields=["stream", "video"])
if created:
rows = []
Measure the video, not a PNG per frame Detection now walks a browser-seekable H.264 proxy in MediaPipe's VIDEO running mode. The PNG sequence it replaces was 112MB for 7.6 seconds at 1440x1920 and 1.1GB at the 900-frame limit; the proxy is 6MB, and landmarks detected off decoded H.264 rather than off the PNGs moved at most 0.0033 of frame width. Three things had to be true for video mode to work, and each was measured against the same footage decoded to PNGs: /blob/<digest> answers byte ranges. Django's FileResponse does no Range handling, and a media element handed 200 with no Accept-Ranges reports an empty `seekable`, no-ops every currentTime write, and detects frame one ninety times without raising. A seek aims at the MIDDLE of its frame. Aiming at i/fps sits on a frame boundary and landed one frame early 31 times in 91; (i + 0.5)/fps was exact on all 91. Timestamps are strictly increasing footage milliseconds. Video mode is a tracker: a repeat leaves the graph in an error state every later call re-throws, so the landmarker is discarded on failure, and passing the frame index instead of i*1000/fps moved landmarks six times further from the per-frame answer. Frames are verified rather than trusted. requestVideoFrameCallback states which frame it handed over, the walker discards any other and fails loudly if the one it asked for never arrives — a stale presentation from the tail of a previous seek is what produced "asked for frame 1 and it presented frame 2" on a video whose seeks were in fact exact. The proxy is re-encoded even when the upload is already H.264: HEVC is not decodable everywhere, and footage identity is the proxy's digest. The JPEG stills beside it are tracing references, outside the footage digest because re-rendering them at another size is not different footage. Verified end to end in a real browser against real footage: 228/228 frames detected, a drawn roto face, 37 backend and 234 frontend tests green. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-28 11:32:01 -04:00
for index, digest, size in still_blobs:
blob, _ = Blob.objects.get_or_create(
Measure the video, not a PNG per frame Detection now walks a browser-seekable H.264 proxy in MediaPipe's VIDEO running mode. The PNG sequence it replaces was 112MB for 7.6 seconds at 1440x1920 and 1.1GB at the 900-frame limit; the proxy is 6MB, and landmarks detected off decoded H.264 rather than off the PNGs moved at most 0.0033 of frame width. Three things had to be true for video mode to work, and each was measured against the same footage decoded to PNGs: /blob/<digest> answers byte ranges. Django's FileResponse does no Range handling, and a media element handed 200 with no Accept-Ranges reports an empty `seekable`, no-ops every currentTime write, and detects frame one ninety times without raising. A seek aims at the MIDDLE of its frame. Aiming at i/fps sits on a frame boundary and landed one frame early 31 times in 91; (i + 0.5)/fps was exact on all 91. Timestamps are strictly increasing footage milliseconds. Video mode is a tracker: a repeat leaves the graph in an error state every later call re-throws, so the landmarker is discarded on failure, and passing the frame index instead of i*1000/fps moved landmarks six times further from the per-frame answer. Frames are verified rather than trusted. requestVideoFrameCallback states which frame it handed over, the walker discards any other and fails loudly if the one it asked for never arrives — a stale presentation from the tail of a previous seek is what produced "asked for frame 1 and it presented frame 2" on a video whose seeks were in fact exact. The proxy is re-encoded even when the upload is already H.264: HEVC is not decodable everywhere, and footage identity is the proxy's digest. The JPEG stills beside it are tracing references, outside the footage digest because re-rendering them at another size is not different footage. Verified end to end in a real browser against real footage: 228/228 frames detected, a drawn roto face, 37 backend and 234 frontend tests green. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-28 11:32:01 -04:00
digest=digest, defaults={"size": size, "media_type": "image/jpeg"})
rows.append(FootageFrame(footage=footage, index=index, blob=blob))
FootageFrame.objects.bulk_create(rows)
return footage
def run(key):
close_old_connections()
try:
job = Extraction.objects.select_related("source", "source__blob").get(key=key)
job.state, job.progress, job.error = "running", 0, ""
job.save(update_fields=["state", "progress", "error", "updated"])
facts = job.source.probe
source_path = blobs.path_for(job.source.blob_id)
with tempfile.TemporaryDirectory(prefix="arthur-extract-") as directory:
root = Path(directory)
Measure the video, not a PNG per frame Detection now walks a browser-seekable H.264 proxy in MediaPipe's VIDEO running mode. The PNG sequence it replaces was 112MB for 7.6 seconds at 1440x1920 and 1.1GB at the 900-frame limit; the proxy is 6MB, and landmarks detected off decoded H.264 rather than off the PNGs moved at most 0.0033 of frame width. Three things had to be true for video mode to work, and each was measured against the same footage decoded to PNGs: /blob/<digest> answers byte ranges. Django's FileResponse does no Range handling, and a media element handed 200 with no Accept-Ranges reports an empty `seekable`, no-ops every currentTime write, and detects frame one ninety times without raising. A seek aims at the MIDDLE of its frame. Aiming at i/fps sits on a frame boundary and landed one frame early 31 times in 91; (i + 0.5)/fps was exact on all 91. Timestamps are strictly increasing footage milliseconds. Video mode is a tracker: a repeat leaves the graph in an error state every later call re-throws, so the landmarker is discarded on failure, and passing the frame index instead of i*1000/fps moved landmarks six times further from the per-frame answer. Frames are verified rather than trusted. requestVideoFrameCallback states which frame it handed over, the walker discards any other and fails loudly if the one it asked for never arrives — a stale presentation from the tail of a previous seek is what produced "asked for frame 1 and it presented frame 2" on a video whose seeks were in fact exact. The proxy is re-encoded even when the upload is already H.264: HEVC is not decodable everywhere, and footage identity is the proxy's digest. The JPEG stills beside it are tracing references, outside the footage digest because re-rendering them at another size is not different footage. Verified end to end in a real browser against real footage: 228/228 frames detected, a drawn roto face, 37 backend and 234 frontend tests green. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-28 11:32:01 -04:00
proxy_path = root / "proxy.mp4"
_encode_proxy(job, source_path, proxy_path, facts, root)
# Everything downstream describes the PROXY, not the upload.
proxy_facts = probe(proxy_path)
Calibrate the video's frame origin instead of assuming it is zero "the video never presented frame 1; it offered 2" was this function waiting for a frame it had not asked for a second time. The walk discarded a wrong frame and re-armed the callback WITHOUT seeking again, and nothing further is ever presented to a paused element that has not been asked to move — so one wrong answer starved until the timeout and reported it as the browser refusing. Underneath that, the wrong answer was not wrong. A container can carry an edit list, and `currentTime` then counts from the start of the edited presentation while a frame's `mediaTime` counts from the start of the media. The two differ by a constant, so the frame at currentTime 0 can honestly report a mediaTime two frames in. Seeking cannot correct for it in the positive direction: source frame 0 would have to be found before the start of the video, every attempt clamps at zero, and the walk offers frame 2 forever. So the constant is measured once and subtracted. `calibrate!` takes whatever the browser calls the first frame it shows and makes that the origin, which is the honest definition anyway — it is what a viewer sees at time zero, and the audio clock this take plays against starts in the same place. Calibration also leaves the element on frame 0, so the walk starts holding it. `frame!` returns immediately for a frame already on screen rather than seeking to where it already is, which presents nothing and would hang. Residual error still re-seeks, corrected by exactly the measured miss, and still fails loudly with every frame that was offered and where it was asked from, so a next failure is diagnosable in one shot rather than four. The proxy also drops B-frames now. That removes the edit list at the source rather than only coping with it, and makes decode order presentation order should this ever be fed to WebCodecs. 14% larger, and extraction refuses a proxy whose timeline is shifted so it cannot come back silently. Honest note: this was my first diagnosis and it did NOT reproduce the failure — both proxies walk correctly here on hardware decode — so it is hardening, not the fix. Verified by forcing the fault: a harness that offsets the reported timeline by -2, -1, 0, +1, +2 and +5 frames failed on every positive offset before and recovers on all six now, first attempt. Real app in headed Chrome with Metal hardware decode: 280/280. 41 backend and 234 frontend tests green. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-28 12:50:11 -04:00
_refuse_a_shifted_timeline(proxy_path)
Measure the video, not a PNG per frame Detection now walks a browser-seekable H.264 proxy in MediaPipe's VIDEO running mode. The PNG sequence it replaces was 112MB for 7.6 seconds at 1440x1920 and 1.1GB at the 900-frame limit; the proxy is 6MB, and landmarks detected off decoded H.264 rather than off the PNGs moved at most 0.0033 of frame width. Three things had to be true for video mode to work, and each was measured against the same footage decoded to PNGs: /blob/<digest> answers byte ranges. Django's FileResponse does no Range handling, and a media element handed 200 with no Accept-Ranges reports an empty `seekable`, no-ops every currentTime write, and detects frame one ninety times without raising. A seek aims at the MIDDLE of its frame. Aiming at i/fps sits on a frame boundary and landed one frame early 31 times in 91; (i + 0.5)/fps was exact on all 91. Timestamps are strictly increasing footage milliseconds. Video mode is a tracker: a repeat leaves the graph in an error state every later call re-throws, so the landmarker is discarded on failure, and passing the frame index instead of i*1000/fps moved landmarks six times further from the per-frame answer. Frames are verified rather than trusted. requestVideoFrameCallback states which frame it handed over, the walker discards any other and fails loudly if the one it asked for never arrives — a stale presentation from the tail of a previous seek is what produced "asked for frame 1 and it presented frame 2" on a video whose seeks were in fact exact. The proxy is re-encoded even when the upload is already H.264: HEVC is not decodable everywhere, and footage identity is the proxy's digest. The JPEG stills beside it are tracing references, outside the footage digest because re-rendering them at another size is not different footage. Verified end to end in a real browser against real footage: 228/228 frames detected, a drawn roto face, 37 backend and 234 frontend tests green. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-28 11:32:01 -04:00
frames = count_frames(proxy_path)
2026-10-03 00:39:07 -04:00
#if not 1 <= frames <= 900:
# raise ValueError(f"the proxy holds {frames} frames; the limit is 1–900")
Stop refusing ordinary phone footage as variable-frame-rate Uploading a clip shot straight from the iPhone camera app failed with "variable-frame-rate video needs timestamp-aware playback". The file was not variable: its container reports avg_frame_rate 8670/299 and nb_frames 289 over a stream whose decoded timestamps are 280 frames exactly 1/30s apart. The guard compared two pieces of container metadata and rejected CFR video on the strength of a summary the container had got wrong about its own contents. The guard was also obsolete. It dates from when the page measured the source's own frames, where a wandering frame duration really does break `frame = floor(t * fps)`. Nothing measures the source now — ffmpeg resamples it onto a constant rate and the proxy is re-probed after it is written — so variable input is a thing this converts rather than a thing it refuses. So: probe picks a rate instead of validating one. It takes the nominal rate, which is the rate every timestamp in the stream can be expressed at and so the one that keeps every distinct source frame, and carries it as an exact fraction because 30000/1001 is not a float and a rounded -r is how a long take drifts. The disagreement is still recorded as `vfr`, just not fatal. The frame-count cross-check went with it. It compared the proxy against the source's nb_frames, which is the number this whole bug proves can lie, and a resample to a constant rate legitimately changes the count. It now checks the proxy's DURATION against the source's, because what must not drift is how long the picture lasts against how long the audio lasts. Verified on the reported file: 280 frames at 30fps, picture 9.3333s against audio 9.3167s — half a frame — and 280/280 detected in the real app. 41 backend tests green, including a genuinely variable fixture end to end. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-28 12:37:21 -04:00
# CHECKED AS A DURATION, not as a frame count. The page's clock is
# `frame = floor(audio.currentTime * fps)`, so what must not drift is
# how long the picture lasts against how long the audio lasts — and
# the source's own frame count is a number this has already caught
# lying. A resample to a constant rate legitimately changes the count
# and must not change the duration.
drift = abs(frames / proxy_facts["fps"] - facts["duration"])
if facts["duration"] > 0 and drift > 0.5:
Measure the video, not a PNG per frame Detection now walks a browser-seekable H.264 proxy in MediaPipe's VIDEO running mode. The PNG sequence it replaces was 112MB for 7.6 seconds at 1440x1920 and 1.1GB at the 900-frame limit; the proxy is 6MB, and landmarks detected off decoded H.264 rather than off the PNGs moved at most 0.0033 of frame width. Three things had to be true for video mode to work, and each was measured against the same footage decoded to PNGs: /blob/<digest> answers byte ranges. Django's FileResponse does no Range handling, and a media element handed 200 with no Accept-Ranges reports an empty `seekable`, no-ops every currentTime write, and detects frame one ninety times without raising. A seek aims at the MIDDLE of its frame. Aiming at i/fps sits on a frame boundary and landed one frame early 31 times in 91; (i + 0.5)/fps was exact on all 91. Timestamps are strictly increasing footage milliseconds. Video mode is a tracker: a repeat leaves the graph in an error state every later call re-throws, so the landmarker is discarded on failure, and passing the frame index instead of i*1000/fps moved landmarks six times further from the per-frame answer. Frames are verified rather than trusted. requestVideoFrameCallback states which frame it handed over, the walker discards any other and fails loudly if the one it asked for never arrives — a stale presentation from the tail of a previous seek is what produced "asked for frame 1 and it presented frame 2" on a video whose seeks were in fact exact. The proxy is re-encoded even when the upload is already H.264: HEVC is not decodable everywhere, and footage identity is the proxy's digest. The JPEG stills beside it are tracing references, outside the footage digest because re-rendering them at another size is not different footage. Verified end to end in a real browser against real footage: 228/228 frames detected, a drawn roto face, 37 backend and 234 frontend tests green. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-28 11:32:01 -04:00
raise ValueError(
Stop refusing ordinary phone footage as variable-frame-rate Uploading a clip shot straight from the iPhone camera app failed with "variable-frame-rate video needs timestamp-aware playback". The file was not variable: its container reports avg_frame_rate 8670/299 and nb_frames 289 over a stream whose decoded timestamps are 280 frames exactly 1/30s apart. The guard compared two pieces of container metadata and rejected CFR video on the strength of a summary the container had got wrong about its own contents. The guard was also obsolete. It dates from when the page measured the source's own frames, where a wandering frame duration really does break `frame = floor(t * fps)`. Nothing measures the source now — ffmpeg resamples it onto a constant rate and the proxy is re-probed after it is written — so variable input is a thing this converts rather than a thing it refuses. So: probe picks a rate instead of validating one. It takes the nominal rate, which is the rate every timestamp in the stream can be expressed at and so the one that keeps every distinct source frame, and carries it as an exact fraction because 30000/1001 is not a float and a rounded -r is how a long take drifts. The disagreement is still recorded as `vfr`, just not fatal. The frame-count cross-check went with it. It compared the proxy against the source's nb_frames, which is the number this whole bug proves can lie, and a resample to a constant rate legitimately changes the count. It now checks the proxy's DURATION against the source's, because what must not drift is how long the picture lasts against how long the audio lasts. Verified on the reported file: 280 frames at 30fps, picture 9.3333s against audio 9.3167s — half a frame — and 280/280 detected in the real app. 41 backend tests green, including a genuinely variable fixture end to end. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-28 12:37:21 -04:00
f"the proxy runs {frames / proxy_facts['fps']:.2f}s and the upload "
f"runs {facts['duration']:.2f}s; refusing footage whose picture and "
"audio would drift")
Measure the video, not a PNG per frame Detection now walks a browser-seekable H.264 proxy in MediaPipe's VIDEO running mode. The PNG sequence it replaces was 112MB for 7.6 seconds at 1440x1920 and 1.1GB at the 900-frame limit; the proxy is 6MB, and landmarks detected off decoded H.264 rather than off the PNGs moved at most 0.0033 of frame width. Three things had to be true for video mode to work, and each was measured against the same footage decoded to PNGs: /blob/<digest> answers byte ranges. Django's FileResponse does no Range handling, and a media element handed 200 with no Accept-Ranges reports an empty `seekable`, no-ops every currentTime write, and detects frame one ninety times without raising. A seek aims at the MIDDLE of its frame. Aiming at i/fps sits on a frame boundary and landed one frame early 31 times in 91; (i + 0.5)/fps was exact on all 91. Timestamps are strictly increasing footage milliseconds. Video mode is a tracker: a repeat leaves the graph in an error state every later call re-throws, so the landmarker is discarded on failure, and passing the frame index instead of i*1000/fps moved landmarks six times further from the per-frame answer. Frames are verified rather than trusted. requestVideoFrameCallback states which frame it handed over, the walker discards any other and fails loudly if the one it asked for never arrives — a stale presentation from the tail of a previous seek is what produced "asked for frame 1 and it presented frame 2" on a video whose seeks were in fact exact. The proxy is re-encoded even when the upload is already H.264: HEVC is not decodable everywhere, and footage identity is the proxy's digest. The JPEG stills beside it are tracing references, outside the footage digest because re-rendering them at another size is not different footage. Verified end to end in a real browser against real footage: 228/228 frames detected, a drawn roto face, 37 backend and 234 frontend tests green. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-28 11:32:01 -04:00
proxy_facts["frames"] = frames
stream_path = root / "proxy.h264"
_elementary_stream(proxy_path, stream_path)
Measure the video, not a PNG per frame Detection now walks a browser-seekable H.264 proxy in MediaPipe's VIDEO running mode. The PNG sequence it replaces was 112MB for 7.6 seconds at 1440x1920 and 1.1GB at the 900-frame limit; the proxy is 6MB, and landmarks detected off decoded H.264 rather than off the PNGs moved at most 0.0033 of frame width. Three things had to be true for video mode to work, and each was measured against the same footage decoded to PNGs: /blob/<digest> answers byte ranges. Django's FileResponse does no Range handling, and a media element handed 200 with no Accept-Ranges reports an empty `seekable`, no-ops every currentTime write, and detects frame one ninety times without raising. A seek aims at the MIDDLE of its frame. Aiming at i/fps sits on a frame boundary and landed one frame early 31 times in 91; (i + 0.5)/fps was exact on all 91. Timestamps are strictly increasing footage milliseconds. Video mode is a tracker: a repeat leaves the graph in an error state every later call re-throws, so the landmarker is discarded on failure, and passing the frame index instead of i*1000/fps moved landmarks six times further from the per-frame answer. Frames are verified rather than trusted. requestVideoFrameCallback states which frame it handed over, the walker discards any other and fails loudly if the one it asked for never arrives — a stale presentation from the tail of a previous seek is what produced "asked for frame 1 and it presented frame 2" on a video whose seeks were in fact exact. The proxy is re-encoded even when the upload is already H.264: HEVC is not decodable everywhere, and footage identity is the proxy's digest. The JPEG stills beside it are tracing references, outside the footage digest because re-rendering them at another size is not different footage. Verified end to end in a real browser against real footage: 228/228 frames detected, a drawn roto face, 37 backend and 234 frontend tests green. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-28 11:32:01 -04:00
frames_dir = root / "stills"
frames_dir.mkdir()
_extract_stills(job, proxy_path, frames_dir, frames, root)
stills = sorted(frames_dir.glob("*.jpg"))
if len(stills) != frames:
raise ValueError(f"wrote {len(stills)} tracing stills for {frames} frames")
job.progress = 85
job.save(update_fields=["progress", "updated"])
audio_path = root / "audio.wav"
if facts["has_audio"]:
_command(["ffmpeg", "-hide_banner", "-loglevel", "error", "-y",
"-i", str(source_path), "-vn", "-ac", "1", "-ar", "44100",
str(audio_path)])
else:
_command(["ffmpeg", "-hide_banner", "-loglevel", "error", "-y",
"-f", "lavfi", "-i", "anullsrc=r=44100:cl=mono",
Measure the video, not a PNG per frame Detection now walks a browser-seekable H.264 proxy in MediaPipe's VIDEO running mode. The PNG sequence it replaces was 112MB for 7.6 seconds at 1440x1920 and 1.1GB at the 900-frame limit; the proxy is 6MB, and landmarks detected off decoded H.264 rather than off the PNGs moved at most 0.0033 of frame width. Three things had to be true for video mode to work, and each was measured against the same footage decoded to PNGs: /blob/<digest> answers byte ranges. Django's FileResponse does no Range handling, and a media element handed 200 with no Accept-Ranges reports an empty `seekable`, no-ops every currentTime write, and detects frame one ninety times without raising. A seek aims at the MIDDLE of its frame. Aiming at i/fps sits on a frame boundary and landed one frame early 31 times in 91; (i + 0.5)/fps was exact on all 91. Timestamps are strictly increasing footage milliseconds. Video mode is a tracker: a repeat leaves the graph in an error state every later call re-throws, so the landmarker is discarded on failure, and passing the frame index instead of i*1000/fps moved landmarks six times further from the per-frame answer. Frames are verified rather than trusted. requestVideoFrameCallback states which frame it handed over, the walker discards any other and fails loudly if the one it asked for never arrives — a stale presentation from the tail of a previous seek is what produced "asked for frame 1 and it presented frame 2" on a video whose seeks were in fact exact. The proxy is re-encoded even when the upload is already H.264: HEVC is not decodable everywhere, and footage identity is the proxy's digest. The JPEG stills beside it are tracing references, outside the footage digest because re-rendering them at another size is not different footage. Verified end to end in a real browser against real footage: 228/228 frames detected, a drawn roto face, 37 backend and 234 frontend tests green. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-28 11:32:01 -04:00
"-t", str(frames / proxy_facts["fps"]), "-c:a", "pcm_s16le",
str(audio_path)])
footage = _register(job, proxy_path, stream_path, stills, audio_path, proxy_facts)
job.footage, job.state, job.progress = footage, "done", 100
job.save(update_fields=["footage", "state", "progress", "updated"])
except Exception as exc:
Extraction.objects.filter(key=key).update(state="failed", error=str(exc)[:2000])
finally:
with _lock:
_active.discard(key)
close_old_connections()
def enqueue(key):
with _lock:
if key in _active:
return
_active.add(key)
threading.Thread(target=run, args=(key,), daemon=True,
name=f"arthur-extract-{key[7:15]}").start()