Measure the video, not a PNG per frame
Detection now walks a browser-seekable H.264 proxy in MediaPipe's VIDEO
running mode. The PNG sequence it replaces was 112MB for 7.6 seconds at
1440x1920 and 1.1GB at the 900-frame limit; the proxy is 6MB, and landmarks
detected off decoded H.264 rather than off the PNGs moved at most 0.0033 of
frame width.
Three things had to be true for video mode to work, and each was measured
against the same footage decoded to PNGs:
/blob/<digest> answers byte ranges. Django's FileResponse does no Range
handling, and a media element handed 200 with no Accept-Ranges reports an
empty `seekable`, no-ops every currentTime write, and detects frame one
ninety times without raising.
A seek aims at the MIDDLE of its frame. Aiming at i/fps sits on a frame
boundary and landed one frame early 31 times in 91; (i + 0.5)/fps was exact
on all 91.
Timestamps are strictly increasing footage milliseconds. Video mode is a
tracker: a repeat leaves the graph in an error state every later call
re-throws, so the landmarker is discarded on failure, and passing the frame
index instead of i*1000/fps moved landmarks six times further from the
per-frame answer.
Frames are verified rather than trusted. requestVideoFrameCallback states
which frame it handed over, the walker discards any other and fails loudly
if the one it asked for never arrives — a stale presentation from the tail
of a previous seek is what produced "asked for frame 1 and it presented
frame 2" on a video whose seeks were in fact exact.
The proxy is re-encoded even when the upload is already H.264: HEVC is not
decodable everywhere, and footage identity is the proxy's digest. The JPEG
stills beside it are tracing references, outside the footage digest because
re-rendering them at another size is not different footage.
Verified end to end in a real browser against real footage: 228/228 frames
detected, a drawn roto face, 37 backend and 234 frontend tests green.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-28 11:32:01 -04:00
|
|
|
|
"""Upload a video once, then turn it into the two things the app actually reads.
|
|
|
|
|
|
|
|
|
|
|
|
WHAT CHANGED AND WHY. This used to decode one PNG per source frame and store every
|
|
|
|
|
|
one of them. A 7.6-second 1440x1920 take is 112MB that way, and the 900-frame limit
|
|
|
|
|
|
is 1.1GB — for pixels whose only consumer was a canvas that MediaPipe then read
|
|
|
|
|
|
once. The page now detects from the video itself (see `frontend/src/arthur/flow/
|
|
|
|
|
|
ingest.cljs`), so this produces:
|
|
|
|
|
|
|
|
|
|
|
|
THE PROXY. One browser-safe H.264/yuv420p MP4, CFR, `+faststart`. The same take
|
|
|
|
|
|
is 6MB. This is the analysis source, and it is re-encoded RATHER THAN KEPT AS
|
|
|
|
|
|
UPLOADED even when the upload is already H.264, for two reasons that are both
|
|
|
|
|
|
about not guessing: an iPhone's HEVC is not decodable in every browser, and the
|
|
|
|
|
|
footage's identity is the digest of this file — one produced by one ffmpeg
|
|
|
|
|
|
invocation, not one that depends on which branch the source happened to take.
|
|
|
|
|
|
|
|
|
|
|
|
THE TRACING STILLS. One JPEG per frame, long edge capped, for the tracing editor
|
|
|
|
|
|
to draw over. Reference images; nothing measures them. They are not in the
|
|
|
|
|
|
footage digest — see `models.Footage`.
|
|
|
|
|
|
|
|
|
|
|
|
The proxy is probed after it is written rather than before. `width`, `height` and
|
|
|
|
|
|
`frames` are properties of the file the browser will decode, and taking them from
|
|
|
|
|
|
the source instead is how a scaler or a dropped frame becomes a silent one-frame
|
|
|
|
|
|
offset between the landmarks and the audio.
|
|
|
|
|
|
"""
|
2026-09-28 09:38:49 -04:00
|
|
|
|
|
|
|
|
|
|
import hashlib
|
|
|
|
|
|
import json
|
|
|
|
|
|
import subprocess
|
|
|
|
|
|
import tempfile
|
|
|
|
|
|
import threading
|
|
|
|
|
|
import time
|
|
|
|
|
|
from fractions import Fraction
|
|
|
|
|
|
from pathlib import Path
|
|
|
|
|
|
|
|
|
|
|
|
from django.db import close_old_connections, transaction
|
|
|
|
|
|
|
|
|
|
|
|
from . import blobs
|
|
|
|
|
|
from .models import Blob, Extraction, Footage, FootageFrame
|
|
|
|
|
|
|
|
|
|
|
|
_active = set()
|
|
|
|
|
|
_lock = threading.Lock()
|
|
|
|
|
|
TIMEOUT = 3600
|
|
|
|
|
|
|
Measure the video, not a PNG per frame
Detection now walks a browser-seekable H.264 proxy in MediaPipe's VIDEO
running mode. The PNG sequence it replaces was 112MB for 7.6 seconds at
1440x1920 and 1.1GB at the 900-frame limit; the proxy is 6MB, and landmarks
detected off decoded H.264 rather than off the PNGs moved at most 0.0033 of
frame width.
Three things had to be true for video mode to work, and each was measured
against the same footage decoded to PNGs:
/blob/<digest> answers byte ranges. Django's FileResponse does no Range
handling, and a media element handed 200 with no Accept-Ranges reports an
empty `seekable`, no-ops every currentTime write, and detects frame one
ninety times without raising.
A seek aims at the MIDDLE of its frame. Aiming at i/fps sits on a frame
boundary and landed one frame early 31 times in 91; (i + 0.5)/fps was exact
on all 91.
Timestamps are strictly increasing footage milliseconds. Video mode is a
tracker: a repeat leaves the graph in an error state every later call
re-throws, so the landmarker is discarded on failure, and passing the frame
index instead of i*1000/fps moved landmarks six times further from the
per-frame answer.
Frames are verified rather than trusted. requestVideoFrameCallback states
which frame it handed over, the walker discards any other and fails loudly
if the one it asked for never arrives — a stale presentation from the tail
of a previous seek is what produced "asked for frame 1 and it presented
frame 2" on a video whose seeks were in fact exact.
The proxy is re-encoded even when the upload is already H.264: HEVC is not
decodable everywhere, and footage identity is the proxy's digest. The JPEG
stills beside it are tracing references, outside the footage digest because
re-rendering them at another size is not different footage.
Verified end to end in a real browser against real footage: 228/228 frames
detected, a drawn roto face, 37 backend and 234 frontend tests green.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-28 11:32:01 -04:00
|
|
|
|
# Visually lossless enough that landmarks do not move: measured against the same
|
|
|
|
|
|
# frames as PNGs, IMAGE-mode landmarks shifted at most 0.0033 of frame width.
|
|
|
|
|
|
PROXY_CRF = "18"
|
|
|
|
|
|
# The long edge of a tracing still. The proxy keeps full resolution because the
|
|
|
|
|
|
# detector reads it; a still only has to be good enough to draw a cel over.
|
|
|
|
|
|
TRACING_EDGE = 1280
|
|
|
|
|
|
TRACING_QUALITY = "4"
|
|
|
|
|
|
|
2026-09-28 09:38:49 -04:00
|
|
|
|
|
|
|
|
|
|
def _command(args):
|
|
|
|
|
|
result = subprocess.run(args, capture_output=True, text=True, timeout=TIMEOUT)
|
|
|
|
|
|
if result.returncode:
|
|
|
|
|
|
raise ValueError((result.stderr or result.stdout or "media tool failed")[-1200:])
|
|
|
|
|
|
return result.stdout
|
|
|
|
|
|
|
|
|
|
|
|
|
Measure the video, not a PNG per frame
Detection now walks a browser-seekable H.264 proxy in MediaPipe's VIDEO
running mode. The PNG sequence it replaces was 112MB for 7.6 seconds at
1440x1920 and 1.1GB at the 900-frame limit; the proxy is 6MB, and landmarks
detected off decoded H.264 rather than off the PNGs moved at most 0.0033 of
frame width.
Three things had to be true for video mode to work, and each was measured
against the same footage decoded to PNGs:
/blob/<digest> answers byte ranges. Django's FileResponse does no Range
handling, and a media element handed 200 with no Accept-Ranges reports an
empty `seekable`, no-ops every currentTime write, and detects frame one
ninety times without raising.
A seek aims at the MIDDLE of its frame. Aiming at i/fps sits on a frame
boundary and landed one frame early 31 times in 91; (i + 0.5)/fps was exact
on all 91.
Timestamps are strictly increasing footage milliseconds. Video mode is a
tracker: a repeat leaves the graph in an error state every later call
re-throws, so the landmarker is discarded on failure, and passing the frame
index instead of i*1000/fps moved landmarks six times further from the
per-frame answer.
Frames are verified rather than trusted. requestVideoFrameCallback states
which frame it handed over, the walker discards any other and fails loudly
if the one it asked for never arrives — a stale presentation from the tail
of a previous seek is what produced "asked for frame 1 and it presented
frame 2" on a video whose seeks were in fact exact.
The proxy is re-encoded even when the upload is already H.264: HEVC is not
decodable everywhere, and footage identity is the proxy's digest. The JPEG
stills beside it are tracing references, outside the footage digest because
re-rendering them at another size is not different footage.
Verified end to end in a real browser against real footage: 228/228 frames
detected, a drawn roto face, 37 backend and 234 frontend tests green.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-28 11:32:01 -04:00
|
|
|
|
def _run_with_progress(job, args, root, name, total, span):
|
|
|
|
|
|
"""Run one ffmpeg and publish its live frame count as `span` of the job.
|
|
|
|
|
|
|
|
|
|
|
|
ffmpeg's `-progress` file is the only honest source for this: parsing its
|
|
|
|
|
|
stderr means parsing a format that is explicitly not an interface, and a
|
|
|
|
|
|
spinner that is not attached to frames is a spinner that lies on a long take.
|
|
|
|
|
|
"""
|
|
|
|
|
|
progress_path = root / f"{name}.progress"
|
|
|
|
|
|
log_path = root / f"{name}.log"
|
|
|
|
|
|
first, last = span
|
2026-09-28 09:38:49 -04:00
|
|
|
|
args = ["ffmpeg", "-hide_banner", "-loglevel", "error", "-y",
|
Measure the video, not a PNG per frame
Detection now walks a browser-seekable H.264 proxy in MediaPipe's VIDEO
running mode. The PNG sequence it replaces was 112MB for 7.6 seconds at
1440x1920 and 1.1GB at the 900-frame limit; the proxy is 6MB, and landmarks
detected off decoded H.264 rather than off the PNGs moved at most 0.0033 of
frame width.
Three things had to be true for video mode to work, and each was measured
against the same footage decoded to PNGs:
/blob/<digest> answers byte ranges. Django's FileResponse does no Range
handling, and a media element handed 200 with no Accept-Ranges reports an
empty `seekable`, no-ops every currentTime write, and detects frame one
ninety times without raising.
A seek aims at the MIDDLE of its frame. Aiming at i/fps sits on a frame
boundary and landed one frame early 31 times in 91; (i + 0.5)/fps was exact
on all 91.
Timestamps are strictly increasing footage milliseconds. Video mode is a
tracker: a repeat leaves the graph in an error state every later call
re-throws, so the landmarker is discarded on failure, and passing the frame
index instead of i*1000/fps moved landmarks six times further from the
per-frame answer.
Frames are verified rather than trusted. requestVideoFrameCallback states
which frame it handed over, the walker discards any other and fails loudly
if the one it asked for never arrives — a stale presentation from the tail
of a previous seek is what produced "asked for frame 1 and it presented
frame 2" on a video whose seeks were in fact exact.
The proxy is re-encoded even when the upload is already H.264: HEVC is not
decodable everywhere, and footage identity is the proxy's digest. The JPEG
stills beside it are tracing references, outside the footage digest because
re-rendering them at another size is not different footage.
Verified end to end in a real browser against real footage: 228/228 frames
detected, a drawn roto face, 37 backend and 234 frontend tests green.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-28 11:32:01 -04:00
|
|
|
|
"-stats_period", "0.25", "-progress", str(progress_path)] + args
|
2026-09-28 09:38:49 -04:00
|
|
|
|
with open(log_path, "wb") as log:
|
|
|
|
|
|
proc = subprocess.Popen(args, stdout=log, stderr=subprocess.STDOUT)
|
|
|
|
|
|
deadline = time.monotonic() + TIMEOUT
|
|
|
|
|
|
try:
|
|
|
|
|
|
while proc.poll() is None:
|
|
|
|
|
|
if time.monotonic() >= deadline:
|
Measure the video, not a PNG per frame
Detection now walks a browser-seekable H.264 proxy in MediaPipe's VIDEO
running mode. The PNG sequence it replaces was 112MB for 7.6 seconds at
1440x1920 and 1.1GB at the 900-frame limit; the proxy is 6MB, and landmarks
detected off decoded H.264 rather than off the PNGs moved at most 0.0033 of
frame width.
Three things had to be true for video mode to work, and each was measured
against the same footage decoded to PNGs:
/blob/<digest> answers byte ranges. Django's FileResponse does no Range
handling, and a media element handed 200 with no Accept-Ranges reports an
empty `seekable`, no-ops every currentTime write, and detects frame one
ninety times without raising.
A seek aims at the MIDDLE of its frame. Aiming at i/fps sits on a frame
boundary and landed one frame early 31 times in 91; (i + 0.5)/fps was exact
on all 91.
Timestamps are strictly increasing footage milliseconds. Video mode is a
tracker: a repeat leaves the graph in an error state every later call
re-throws, so the landmarker is discarded on failure, and passing the frame
index instead of i*1000/fps moved landmarks six times further from the
per-frame answer.
Frames are verified rather than trusted. requestVideoFrameCallback states
which frame it handed over, the walker discards any other and fails loudly
if the one it asked for never arrives — a stale presentation from the tail
of a previous seek is what produced "asked for frame 1 and it presented
frame 2" on a video whose seeks were in fact exact.
The proxy is re-encoded even when the upload is already H.264: HEVC is not
decodable everywhere, and footage identity is the proxy's digest. The JPEG
stills beside it are tracing references, outside the footage digest because
re-rendering them at another size is not different footage.
Verified end to end in a real browser against real footage: 228/228 frames
detected, a drawn roto face, 37 backend and 234 frontend tests green.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-28 11:32:01 -04:00
|
|
|
|
raise TimeoutError(f"{name} timed out")
|
2026-09-28 09:38:49 -04:00
|
|
|
|
if progress_path.exists():
|
|
|
|
|
|
lines = progress_path.read_text(errors="replace").splitlines()
|
|
|
|
|
|
count = next((int(line[6:].strip()) for line in reversed(lines)
|
|
|
|
|
|
if line.startswith("frame=") and
|
|
|
|
|
|
line[6:].strip().isdigit()), 0)
|
|
|
|
|
|
if count and total:
|
Measure the video, not a PNG per frame
Detection now walks a browser-seekable H.264 proxy in MediaPipe's VIDEO
running mode. The PNG sequence it replaces was 112MB for 7.6 seconds at
1440x1920 and 1.1GB at the 900-frame limit; the proxy is 6MB, and landmarks
detected off decoded H.264 rather than off the PNGs moved at most 0.0033 of
frame width.
Three things had to be true for video mode to work, and each was measured
against the same footage decoded to PNGs:
/blob/<digest> answers byte ranges. Django's FileResponse does no Range
handling, and a media element handed 200 with no Accept-Ranges reports an
empty `seekable`, no-ops every currentTime write, and detects frame one
ninety times without raising.
A seek aims at the MIDDLE of its frame. Aiming at i/fps sits on a frame
boundary and landed one frame early 31 times in 91; (i + 0.5)/fps was exact
on all 91.
Timestamps are strictly increasing footage milliseconds. Video mode is a
tracker: a repeat leaves the graph in an error state every later call
re-throws, so the landmarker is discarded on failure, and passing the frame
index instead of i*1000/fps moved landmarks six times further from the
per-frame answer.
Frames are verified rather than trusted. requestVideoFrameCallback states
which frame it handed over, the walker discards any other and fails loudly
if the one it asked for never arrives — a stale presentation from the tail
of a previous seek is what produced "asked for frame 1 and it presented
frame 2" on a video whose seeks were in fact exact.
The proxy is re-encoded even when the upload is already H.264: HEVC is not
decodable everywhere, and footage identity is the proxy's digest. The JPEG
stills beside it are tracing references, outside the footage digest because
re-rendering them at another size is not different footage.
Verified end to end in a real browser against real footage: 228/228 frames
detected, a drawn roto face, 37 backend and 234 frontend tests green.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-28 11:32:01 -04:00
|
|
|
|
reached = first + int((last - first) * min(1.0, count / total))
|
|
|
|
|
|
if reached > job.progress:
|
|
|
|
|
|
job.progress = reached
|
2026-09-28 09:38:49 -04:00
|
|
|
|
job.save(update_fields=["progress", "updated"])
|
|
|
|
|
|
time.sleep(0.2)
|
|
|
|
|
|
finally:
|
|
|
|
|
|
if proc.poll() is None:
|
|
|
|
|
|
proc.kill()
|
|
|
|
|
|
proc.wait()
|
|
|
|
|
|
if proc.returncode:
|
Measure the video, not a PNG per frame
Detection now walks a browser-seekable H.264 proxy in MediaPipe's VIDEO
running mode. The PNG sequence it replaces was 112MB for 7.6 seconds at
1440x1920 and 1.1GB at the 900-frame limit; the proxy is 6MB, and landmarks
detected off decoded H.264 rather than off the PNGs moved at most 0.0033 of
frame width.
Three things had to be true for video mode to work, and each was measured
against the same footage decoded to PNGs:
/blob/<digest> answers byte ranges. Django's FileResponse does no Range
handling, and a media element handed 200 with no Accept-Ranges reports an
empty `seekable`, no-ops every currentTime write, and detects frame one
ninety times without raising.
A seek aims at the MIDDLE of its frame. Aiming at i/fps sits on a frame
boundary and landed one frame early 31 times in 91; (i + 0.5)/fps was exact
on all 91.
Timestamps are strictly increasing footage milliseconds. Video mode is a
tracker: a repeat leaves the graph in an error state every later call
re-throws, so the landmarker is discarded on failure, and passing the frame
index instead of i*1000/fps moved landmarks six times further from the
per-frame answer.
Frames are verified rather than trusted. requestVideoFrameCallback states
which frame it handed over, the walker discards any other and fails loudly
if the one it asked for never arrives — a stale presentation from the tail
of a previous seek is what produced "asked for frame 1 and it presented
frame 2" on a video whose seeks were in fact exact.
The proxy is re-encoded even when the upload is already H.264: HEVC is not
decodable everywhere, and footage identity is the proxy's digest. The JPEG
stills beside it are tracing references, outside the footage digest because
re-rendering them at another size is not different footage.
Verified end to end in a real browser against real footage: 228/228 frames
detected, a drawn roto face, 37 backend and 234 frontend tests green.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-28 11:32:01 -04:00
|
|
|
|
raise ValueError(log_path.read_text(errors="replace")[-1200:] or f"{name} failed")
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def _encode_proxy(job, source_path, proxy_path, facts, root):
|
|
|
|
|
|
"""The uploaded video -> one H.264 file every browser can decode and seek."""
|
|
|
|
|
|
total = facts.get("reported_frames") or round(facts["duration"] * facts["fps"])
|
|
|
|
|
|
_run_with_progress(
|
|
|
|
|
|
job,
|
|
|
|
|
|
["-i", str(source_path), "-an",
|
2026-09-28 12:37:21 -04:00
|
|
|
|
# Constant frame rate at the rate `probe` chose. This RESAMPLES rather
|
|
|
|
|
|
# than asserts: the upload is allowed to be variable, and this is the
|
|
|
|
|
|
# step that makes the thing the page measures not be.
|
|
|
|
|
|
"-fps_mode", "cfr", "-r", facts.get("rate") or str(facts["fps"]),
|
Measure the video, not a PNG per frame
Detection now walks a browser-seekable H.264 proxy in MediaPipe's VIDEO
running mode. The PNG sequence it replaces was 112MB for 7.6 seconds at
1440x1920 and 1.1GB at the 900-frame limit; the proxy is 6MB, and landmarks
detected off decoded H.264 rather than off the PNGs moved at most 0.0033 of
frame width.
Three things had to be true for video mode to work, and each was measured
against the same footage decoded to PNGs:
/blob/<digest> answers byte ranges. Django's FileResponse does no Range
handling, and a media element handed 200 with no Accept-Ranges reports an
empty `seekable`, no-ops every currentTime write, and detects frame one
ninety times without raising.
A seek aims at the MIDDLE of its frame. Aiming at i/fps sits on a frame
boundary and landed one frame early 31 times in 91; (i + 0.5)/fps was exact
on all 91.
Timestamps are strictly increasing footage milliseconds. Video mode is a
tracker: a repeat leaves the graph in an error state every later call
re-throws, so the landmarker is discarded on failure, and passing the frame
index instead of i*1000/fps moved landmarks six times further from the
per-frame answer.
Frames are verified rather than trusted. requestVideoFrameCallback states
which frame it handed over, the walker discards any other and fails loudly
if the one it asked for never arrives — a stale presentation from the tail
of a previous seek is what produced "asked for frame 1 and it presented
frame 2" on a video whose seeks were in fact exact.
The proxy is re-encoded even when the upload is already H.264: HEVC is not
decodable everywhere, and footage identity is the proxy's digest. The JPEG
stills beside it are tracing references, outside the footage digest because
re-rendering them at another size is not different footage.
Verified end to end in a real browser against real footage: 228/228 frames
detected, a drawn roto face, 37 backend and 234 frontend tests green.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-28 11:32:01 -04:00
|
|
|
|
"-c:v", "libx264", "-preset", "veryfast", "-crf", PROXY_CRF,
|
|
|
|
|
|
# yuv420p and an even frame size are what makes this playable everywhere
|
|
|
|
|
|
# rather than only in the browser that happened to be tested.
|
|
|
|
|
|
"-pix_fmt", "yuv420p", "-vf", "scale=trunc(iw/2)*2:trunc(ih/2)*2",
|
|
|
|
|
|
"-movflags", "+faststart", str(proxy_path)],
|
|
|
|
|
|
root, "proxy", total, (0, 55))
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def _extract_stills(job, proxy_path, frames_dir, frames, root):
|
|
|
|
|
|
"""The proxy -> one tracing JPEG per frame, long edge capped."""
|
|
|
|
|
|
_run_with_progress(
|
|
|
|
|
|
job,
|
|
|
|
|
|
["-i", str(proxy_path), "-fps_mode", "passthrough",
|
|
|
|
|
|
"-vf", f"scale='if(gt(iw,ih),min({TRACING_EDGE},iw),-2)':"
|
|
|
|
|
|
f"'if(gt(iw,ih),-2,min({TRACING_EDGE},ih))'",
|
|
|
|
|
|
"-q:v", TRACING_QUALITY, str(frames_dir / "%04d.jpg")],
|
|
|
|
|
|
root, "stills", frames, (55, 85))
|
2026-09-28 09:38:49 -04:00
|
|
|
|
|
|
|
|
|
|
|
2026-09-28 12:37:21 -04:00
|
|
|
|
MAX_RATE = 120 # a capture rate; past this the container is describing something else
|
|
|
|
|
|
|
|
|
|
|
|
|
2026-09-28 09:38:49 -04:00
|
|
|
|
def probe(path):
|
2026-09-28 12:37:21 -04:00
|
|
|
|
"""What the upload is, as far as choosing a proxy rate goes.
|
|
|
|
|
|
|
|
|
|
|
|
IT NO LONGER REFUSES VARIABLE-FRAME-RATE INPUT, and the reason is the proxy.
|
|
|
|
|
|
That refusal was written when the page measured the source's own frames, where
|
|
|
|
|
|
a wandering frame duration really does break `frame = floor(t * fps)`. Nothing
|
|
|
|
|
|
measures the source now: ffmpeg resamples it onto a constant rate, and the
|
|
|
|
|
|
proxy — constant by construction, and re-probed after it is written — is the
|
|
|
|
|
|
only timeline anything downstream sees.
|
|
|
|
|
|
|
|
|
|
|
|
Keeping the check would have been worse than useless, because the thing it
|
|
|
|
|
|
tested is not reliable. Ordinary iPhone footage, shot straight from the camera
|
|
|
|
|
|
app, reports `avg_frame_rate` 8670/299 and `nb_frames` 289 on a stream whose
|
|
|
|
|
|
decoded timestamps are 280 frames exactly 1/30s apart. The container's summary
|
|
|
|
|
|
of itself disagreed with the container's own contents, so the guard rejected
|
|
|
|
|
|
CFR video for being variable.
|
|
|
|
|
|
|
|
|
|
|
|
THE RATE IS THE NOMINAL ONE. `r_frame_rate` is the rate every timestamp in the
|
|
|
|
|
|
stream can be expressed at, which is the rate that keeps every distinct source
|
|
|
|
|
|
frame; resampling to the average would drop some. Duration is preserved either
|
|
|
|
|
|
way — ffmpeg's CFR conversion is driven by timestamps, so the audio stays in
|
|
|
|
|
|
sync at any rate — so this trades a possible duplicated frame against a
|
|
|
|
|
|
certainly lost one.
|
|
|
|
|
|
"""
|
2026-09-28 09:38:49 -04:00
|
|
|
|
data = json.loads(_command(["ffprobe", "-v", "error", "-show_streams",
|
|
|
|
|
|
"-show_format", "-of", "json", str(path)]))
|
|
|
|
|
|
video = next((s for s in data.get("streams", []) if s.get("codec_type") == "video"), None)
|
|
|
|
|
|
if not video:
|
|
|
|
|
|
raise ValueError("the uploaded file has no video stream")
|
|
|
|
|
|
nominal = Fraction(video.get("r_frame_rate") or "0")
|
|
|
|
|
|
average = Fraction(video.get("avg_frame_rate") or "0")
|
2026-09-28 12:37:21 -04:00
|
|
|
|
if nominal <= 0 and average <= 0:
|
2026-09-28 09:38:49 -04:00
|
|
|
|
raise ValueError("the video's frame rate is unknown")
|
2026-09-28 12:37:21 -04:00
|
|
|
|
rate = nominal if 0 < nominal <= MAX_RATE else average
|
|
|
|
|
|
if not 0 < rate <= MAX_RATE:
|
|
|
|
|
|
raise ValueError(f"the video reports a frame rate of {float(rate):g}, which is "
|
|
|
|
|
|
"not a rate footage can be measured at")
|
2026-09-28 09:38:49 -04:00
|
|
|
|
duration = float(data.get("format", {}).get("duration") or 0)
|
2026-09-28 12:37:21 -04:00
|
|
|
|
if duration > 0 and duration * float(rate) > 901:
|
2026-09-28 09:38:49 -04:00
|
|
|
|
raise ValueError("video is longer than the 900-frame footage limit")
|
2026-09-28 12:37:21 -04:00
|
|
|
|
frames = video.get("nb_frames")
|
|
|
|
|
|
return {"fps": float(rate),
|
|
|
|
|
|
# The exact rate, for ffmpeg. 30000/1001 is not a float, and handing
|
|
|
|
|
|
# `-r` a rounded one is how a long take drifts out of sync.
|
|
|
|
|
|
"rate": f"{rate.numerator}/{rate.denominator}",
|
|
|
|
|
|
"nominal_fps": float(nominal), "average_fps": float(average),
|
2026-09-28 09:38:49 -04:00
|
|
|
|
"width": int(video["width"]), "height": int(video["height"]),
|
|
|
|
|
|
"duration": duration,
|
2026-09-28 12:37:21 -04:00
|
|
|
|
# KEPT, AND NO LONGER TRUSTED AS A COUNT. See the docstring: this is
|
|
|
|
|
|
# the container's claim about itself, it is wrong on ordinary phone
|
|
|
|
|
|
# footage, and `run` checks the proxy's DURATION instead.
|
2026-09-28 09:38:49 -04:00
|
|
|
|
"reported_frames": int(frames) if frames and frames.isdigit() else None,
|
|
|
|
|
|
"has_audio": any(s.get("codec_type") == "audio" for s in data.get("streams", [])),
|
2026-09-28 12:37:21 -04:00
|
|
|
|
"vfr": nominal != average}
|
2026-09-28 09:38:49 -04:00
|
|
|
|
|
|
|
|
|
|
|
Measure the video, not a PNG per frame
Detection now walks a browser-seekable H.264 proxy in MediaPipe's VIDEO
running mode. The PNG sequence it replaces was 112MB for 7.6 seconds at
1440x1920 and 1.1GB at the 900-frame limit; the proxy is 6MB, and landmarks
detected off decoded H.264 rather than off the PNGs moved at most 0.0033 of
frame width.
Three things had to be true for video mode to work, and each was measured
against the same footage decoded to PNGs:
/blob/<digest> answers byte ranges. Django's FileResponse does no Range
handling, and a media element handed 200 with no Accept-Ranges reports an
empty `seekable`, no-ops every currentTime write, and detects frame one
ninety times without raising.
A seek aims at the MIDDLE of its frame. Aiming at i/fps sits on a frame
boundary and landed one frame early 31 times in 91; (i + 0.5)/fps was exact
on all 91.
Timestamps are strictly increasing footage milliseconds. Video mode is a
tracker: a repeat leaves the graph in an error state every later call
re-throws, so the landmarker is discarded on failure, and passing the frame
index instead of i*1000/fps moved landmarks six times further from the
per-frame answer.
Frames are verified rather than trusted. requestVideoFrameCallback states
which frame it handed over, the walker discards any other and fails loudly
if the one it asked for never arrives — a stale presentation from the tail
of a previous seek is what produced "asked for frame 1 and it presented
frame 2" on a video whose seeks were in fact exact.
The proxy is re-encoded even when the upload is already H.264: HEVC is not
decodable everywhere, and footage identity is the proxy's digest. The JPEG
stills beside it are tracing references, outside the footage digest because
re-rendering them at another size is not different footage.
Verified end to end in a real browser against real footage: 228/228 frames
detected, a drawn roto face, 37 backend and 234 frontend tests green.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-28 11:32:01 -04:00
|
|
|
|
def count_frames(path):
|
|
|
|
|
|
"""How many frames a file really holds, counted rather than reported.
|
|
|
|
|
|
|
|
|
|
|
|
`nb_frames` is a container's claim. This is the decoder's answer, and it is
|
|
|
|
|
|
what the page will get when it walks the proxy — so a disagreement between the
|
|
|
|
|
|
two has to be settled before the count reaches a manifest, not after it has
|
|
|
|
|
|
become a one-frame audio offset nobody can find.
|
|
|
|
|
|
"""
|
|
|
|
|
|
text = _command(["ffprobe", "-v", "error", "-select_streams", "v:0",
|
|
|
|
|
|
"-count_frames", "-show_entries", "stream=nb_read_frames",
|
|
|
|
|
|
"-of", "default=nokey=1:noprint_wrappers=1", str(path)])
|
|
|
|
|
|
counted = text.strip()
|
|
|
|
|
|
if not counted.isdigit():
|
|
|
|
|
|
raise ValueError("could not count the proxy's frames")
|
|
|
|
|
|
return int(counted)
|
|
|
|
|
|
|
|
|
|
|
|
|
2026-09-28 09:38:49 -04:00
|
|
|
|
def extraction_key(source, settings):
|
Measure the video, not a PNG per frame
Detection now walks a browser-seekable H.264 proxy in MediaPipe's VIDEO
running mode. The PNG sequence it replaces was 112MB for 7.6 seconds at
1440x1920 and 1.1GB at the 900-frame limit; the proxy is 6MB, and landmarks
detected off decoded H.264 rather than off the PNGs moved at most 0.0033 of
frame width.
Three things had to be true for video mode to work, and each was measured
against the same footage decoded to PNGs:
/blob/<digest> answers byte ranges. Django's FileResponse does no Range
handling, and a media element handed 200 with no Accept-Ranges reports an
empty `seekable`, no-ops every currentTime write, and detects frame one
ninety times without raising.
A seek aims at the MIDDLE of its frame. Aiming at i/fps sits on a frame
boundary and landed one frame early 31 times in 91; (i + 0.5)/fps was exact
on all 91.
Timestamps are strictly increasing footage milliseconds. Video mode is a
tracker: a repeat leaves the graph in an error state every later call
re-throws, so the landmarker is discarded on failure, and passing the frame
index instead of i*1000/fps moved landmarks six times further from the
per-frame answer.
Frames are verified rather than trusted. requestVideoFrameCallback states
which frame it handed over, the walker discards any other and fails loudly
if the one it asked for never arrives — a stale presentation from the tail
of a previous seek is what produced "asked for frame 1 and it presented
frame 2" on a video whose seeks were in fact exact.
The proxy is re-encoded even when the upload is already H.264: HEVC is not
decodable everywhere, and footage identity is the proxy's digest. The JPEG
stills beside it are tracing references, outside the footage digest because
re-rendering them at another size is not different footage.
Verified end to end in a real browser against real footage: 228/228 frames
detected, a drawn roto face, 37 backend and 234 frontend tests green.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-28 11:32:01 -04:00
|
|
|
|
text = json.dumps({"scheme": 2, "source": source.blob_id, "settings": settings},
|
2026-09-28 09:38:49 -04:00
|
|
|
|
sort_keys=True, separators=(",", ":"))
|
|
|
|
|
|
return "sha256:" + hashlib.sha256(text.encode()).hexdigest()
|
|
|
|
|
|
|
|
|
|
|
|
|
Measure the video, not a PNG per frame
Detection now walks a browser-seekable H.264 proxy in MediaPipe's VIDEO
running mode. The PNG sequence it replaces was 112MB for 7.6 seconds at
1440x1920 and 1.1GB at the 900-frame limit; the proxy is 6MB, and landmarks
detected off decoded H.264 rather than off the PNGs moved at most 0.0033 of
frame width.
Three things had to be true for video mode to work, and each was measured
against the same footage decoded to PNGs:
/blob/<digest> answers byte ranges. Django's FileResponse does no Range
handling, and a media element handed 200 with no Accept-Ranges reports an
empty `seekable`, no-ops every currentTime write, and detects frame one
ninety times without raising.
A seek aims at the MIDDLE of its frame. Aiming at i/fps sits on a frame
boundary and landed one frame early 31 times in 91; (i + 0.5)/fps was exact
on all 91.
Timestamps are strictly increasing footage milliseconds. Video mode is a
tracker: a repeat leaves the graph in an error state every later call
re-throws, so the landmarker is discarded on failure, and passing the frame
index instead of i*1000/fps moved landmarks six times further from the
per-frame answer.
Frames are verified rather than trusted. requestVideoFrameCallback states
which frame it handed over, the walker discards any other and fails loudly
if the one it asked for never arrives — a stale presentation from the tail
of a previous seek is what produced "asked for frame 1 and it presented
frame 2" on a video whose seeks were in fact exact.
The proxy is re-encoded even when the upload is already H.264: HEVC is not
decodable everywhere, and footage identity is the proxy's digest. The JPEG
stills beside it are tracing references, outside the footage digest because
re-rendering them at another size is not different footage.
Verified end to end in a real browser against real footage: 228/228 frames
detected, a drawn roto face, 37 backend and 234 frontend tests green.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-28 11:32:01 -04:00
|
|
|
|
def _register(job, proxy_path, stills, audio_path, facts):
|
|
|
|
|
|
proxy_digest, proxy_size = blobs.adopt(proxy_path)
|
2026-09-28 09:38:49 -04:00
|
|
|
|
audio_digest, audio_size = blobs.adopt(audio_path)
|
Measure the video, not a PNG per frame
Detection now walks a browser-seekable H.264 proxy in MediaPipe's VIDEO
running mode. The PNG sequence it replaces was 112MB for 7.6 seconds at
1440x1920 and 1.1GB at the 900-frame limit; the proxy is 6MB, and landmarks
detected off decoded H.264 rather than off the PNGs moved at most 0.0033 of
frame width.
Three things had to be true for video mode to work, and each was measured
against the same footage decoded to PNGs:
/blob/<digest> answers byte ranges. Django's FileResponse does no Range
handling, and a media element handed 200 with no Accept-Ranges reports an
empty `seekable`, no-ops every currentTime write, and detects frame one
ninety times without raising.
A seek aims at the MIDDLE of its frame. Aiming at i/fps sits on a frame
boundary and landed one frame early 31 times in 91; (i + 0.5)/fps was exact
on all 91.
Timestamps are strictly increasing footage milliseconds. Video mode is a
tracker: a repeat leaves the graph in an error state every later call
re-throws, so the landmarker is discarded on failure, and passing the frame
index instead of i*1000/fps moved landmarks six times further from the
per-frame answer.
Frames are verified rather than trusted. requestVideoFrameCallback states
which frame it handed over, the walker discards any other and fails loudly
if the one it asked for never arrives — a stale presentation from the tail
of a previous seek is what produced "asked for frame 1 and it presented
frame 2" on a video whose seeks were in fact exact.
The proxy is re-encoded even when the upload is already H.264: HEVC is not
decodable everywhere, and footage identity is the proxy's digest. The JPEG
stills beside it are tracing references, outside the footage digest because
re-rendering them at another size is not different footage.
Verified end to end in a real browser against real footage: 228/228 frames
detected, a drawn roto face, 37 backend and 234 frontend tests green.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-28 11:32:01 -04:00
|
|
|
|
still_blobs = [(index, *blobs.adopt(path)) for index, path in enumerate(stills)]
|
|
|
|
|
|
width, height, fps, frames = facts["width"], facts["height"], facts["fps"], facts["frames"]
|
|
|
|
|
|
|
|
|
|
|
|
# The footage's own identity: the bytes the page will measure, the audio it
|
|
|
|
|
|
# will clock against, and the rate that ties them together. Scheme 2 — scheme
|
|
|
|
|
|
# 1 hashed a PNG per frame, and those footages name pixels this no longer has.
|
2026-09-28 09:38:49 -04:00
|
|
|
|
h = hashlib.sha256()
|
Measure the video, not a PNG per frame
Detection now walks a browser-seekable H.264 proxy in MediaPipe's VIDEO
running mode. The PNG sequence it replaces was 112MB for 7.6 seconds at
1440x1920 and 1.1GB at the 900-frame limit; the proxy is 6MB, and landmarks
detected off decoded H.264 rather than off the PNGs moved at most 0.0033 of
frame width.
Three things had to be true for video mode to work, and each was measured
against the same footage decoded to PNGs:
/blob/<digest> answers byte ranges. Django's FileResponse does no Range
handling, and a media element handed 200 with no Accept-Ranges reports an
empty `seekable`, no-ops every currentTime write, and detects frame one
ninety times without raising.
A seek aims at the MIDDLE of its frame. Aiming at i/fps sits on a frame
boundary and landed one frame early 31 times in 91; (i + 0.5)/fps was exact
on all 91.
Timestamps are strictly increasing footage milliseconds. Video mode is a
tracker: a repeat leaves the graph in an error state every later call
re-throws, so the landmarker is discarded on failure, and passing the frame
index instead of i*1000/fps moved landmarks six times further from the
per-frame answer.
Frames are verified rather than trusted. requestVideoFrameCallback states
which frame it handed over, the walker discards any other and fails loudly
if the one it asked for never arrives — a stale presentation from the tail
of a previous seek is what produced "asked for frame 1 and it presented
frame 2" on a video whose seeks were in fact exact.
The proxy is re-encoded even when the upload is already H.264: HEVC is not
decodable everywhere, and footage identity is the proxy's digest. The JPEG
stills beside it are tracing references, outside the footage digest because
re-rendering them at another size is not different footage.
Verified end to end in a real browser against real footage: 228/228 frames
detected, a drawn roto face, 37 backend and 234 frontend tests green.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-28 11:32:01 -04:00
|
|
|
|
h.update(f"arthur-footage-2/{fps}/{frames}/{width}x{height}\n".encode())
|
|
|
|
|
|
h.update(proxy_digest.encode())
|
2026-09-28 09:38:49 -04:00
|
|
|
|
h.update(audio_digest.encode())
|
Measure the video, not a PNG per frame
Detection now walks a browser-seekable H.264 proxy in MediaPipe's VIDEO
running mode. The PNG sequence it replaces was 112MB for 7.6 seconds at
1440x1920 and 1.1GB at the 900-frame limit; the proxy is 6MB, and landmarks
detected off decoded H.264 rather than off the PNGs moved at most 0.0033 of
frame width.
Three things had to be true for video mode to work, and each was measured
against the same footage decoded to PNGs:
/blob/<digest> answers byte ranges. Django's FileResponse does no Range
handling, and a media element handed 200 with no Accept-Ranges reports an
empty `seekable`, no-ops every currentTime write, and detects frame one
ninety times without raising.
A seek aims at the MIDDLE of its frame. Aiming at i/fps sits on a frame
boundary and landed one frame early 31 times in 91; (i + 0.5)/fps was exact
on all 91.
Timestamps are strictly increasing footage milliseconds. Video mode is a
tracker: a repeat leaves the graph in an error state every later call
re-throws, so the landmarker is discarded on failure, and passing the frame
index instead of i*1000/fps moved landmarks six times further from the
per-frame answer.
Frames are verified rather than trusted. requestVideoFrameCallback states
which frame it handed over, the walker discards any other and fails loudly
if the one it asked for never arrives — a stale presentation from the tail
of a previous seek is what produced "asked for frame 1 and it presented
frame 2" on a video whose seeks were in fact exact.
The proxy is re-encoded even when the upload is already H.264: HEVC is not
decodable everywhere, and footage identity is the proxy's digest. The JPEG
stills beside it are tracing references, outside the footage digest because
re-rendering them at another size is not different footage.
Verified end to end in a real browser against real footage: 228/228 frames
detected, a drawn roto face, 37 backend and 234 frontend tests green.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-28 11:32:01 -04:00
|
|
|
|
|
2026-09-28 09:38:49 -04:00
|
|
|
|
with transaction.atomic():
|
Measure the video, not a PNG per frame
Detection now walks a browser-seekable H.264 proxy in MediaPipe's VIDEO
running mode. The PNG sequence it replaces was 112MB for 7.6 seconds at
1440x1920 and 1.1GB at the 900-frame limit; the proxy is 6MB, and landmarks
detected off decoded H.264 rather than off the PNGs moved at most 0.0033 of
frame width.
Three things had to be true for video mode to work, and each was measured
against the same footage decoded to PNGs:
/blob/<digest> answers byte ranges. Django's FileResponse does no Range
handling, and a media element handed 200 with no Accept-Ranges reports an
empty `seekable`, no-ops every currentTime write, and detects frame one
ninety times without raising.
A seek aims at the MIDDLE of its frame. Aiming at i/fps sits on a frame
boundary and landed one frame early 31 times in 91; (i + 0.5)/fps was exact
on all 91.
Timestamps are strictly increasing footage milliseconds. Video mode is a
tracker: a repeat leaves the graph in an error state every later call
re-throws, so the landmarker is discarded on failure, and passing the frame
index instead of i*1000/fps moved landmarks six times further from the
per-frame answer.
Frames are verified rather than trusted. requestVideoFrameCallback states
which frame it handed over, the walker discards any other and fails loudly
if the one it asked for never arrives — a stale presentation from the tail
of a previous seek is what produced "asked for frame 1 and it presented
frame 2" on a video whose seeks were in fact exact.
The proxy is re-encoded even when the upload is already H.264: HEVC is not
decodable everywhere, and footage identity is the proxy's digest. The JPEG
stills beside it are tracing references, outside the footage digest because
re-rendering them at another size is not different footage.
Verified end to end in a real browser against real footage: 228/228 frames
detected, a drawn roto face, 37 backend and 234 frontend tests green.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-28 11:32:01 -04:00
|
|
|
|
proxy_blob, _ = Blob.objects.get_or_create(
|
|
|
|
|
|
digest=proxy_digest, defaults={"size": proxy_size, "media_type": "video/mp4"})
|
2026-09-28 09:38:49 -04:00
|
|
|
|
audio_blob, _ = Blob.objects.get_or_create(
|
|
|
|
|
|
digest=audio_digest, defaults={"size": audio_size, "media_type": "audio/wav"})
|
|
|
|
|
|
footage, created = Footage.objects.get_or_create(
|
|
|
|
|
|
digest=h.hexdigest(),
|
|
|
|
|
|
defaults={"label": job.source.filename[:200], "source": job.source.filename[:200],
|
Measure the video, not a PNG per frame
Detection now walks a browser-seekable H.264 proxy in MediaPipe's VIDEO
running mode. The PNG sequence it replaces was 112MB for 7.6 seconds at
1440x1920 and 1.1GB at the 900-frame limit; the proxy is 6MB, and landmarks
detected off decoded H.264 rather than off the PNGs moved at most 0.0033 of
frame width.
Three things had to be true for video mode to work, and each was measured
against the same footage decoded to PNGs:
/blob/<digest> answers byte ranges. Django's FileResponse does no Range
handling, and a media element handed 200 with no Accept-Ranges reports an
empty `seekable`, no-ops every currentTime write, and detects frame one
ninety times without raising.
A seek aims at the MIDDLE of its frame. Aiming at i/fps sits on a frame
boundary and landed one frame early 31 times in 91; (i + 0.5)/fps was exact
on all 91.
Timestamps are strictly increasing footage milliseconds. Video mode is a
tracker: a repeat leaves the graph in an error state every later call
re-throws, so the landmarker is discarded on failure, and passing the frame
index instead of i*1000/fps moved landmarks six times further from the
per-frame answer.
Frames are verified rather than trusted. requestVideoFrameCallback states
which frame it handed over, the walker discards any other and fails loudly
if the one it asked for never arrives — a stale presentation from the tail
of a previous seek is what produced "asked for frame 1 and it presented
frame 2" on a video whose seeks were in fact exact.
The proxy is re-encoded even when the upload is already H.264: HEVC is not
decodable everywhere, and footage identity is the proxy's digest. The JPEG
stills beside it are tracing references, outside the footage digest because
re-rendering them at another size is not different footage.
Verified end to end in a real browser against real footage: 228/228 frames
detected, a drawn roto face, 37 backend and 234 frontend tests green.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-28 11:32:01 -04:00
|
|
|
|
"fps": fps, "frames": frames, "width": width, "height": height,
|
|
|
|
|
|
"audio": audio_blob, "video": proxy_blob})
|
2026-09-28 09:38:49 -04:00
|
|
|
|
if created:
|
|
|
|
|
|
rows = []
|
Measure the video, not a PNG per frame
Detection now walks a browser-seekable H.264 proxy in MediaPipe's VIDEO
running mode. The PNG sequence it replaces was 112MB for 7.6 seconds at
1440x1920 and 1.1GB at the 900-frame limit; the proxy is 6MB, and landmarks
detected off decoded H.264 rather than off the PNGs moved at most 0.0033 of
frame width.
Three things had to be true for video mode to work, and each was measured
against the same footage decoded to PNGs:
/blob/<digest> answers byte ranges. Django's FileResponse does no Range
handling, and a media element handed 200 with no Accept-Ranges reports an
empty `seekable`, no-ops every currentTime write, and detects frame one
ninety times without raising.
A seek aims at the MIDDLE of its frame. Aiming at i/fps sits on a frame
boundary and landed one frame early 31 times in 91; (i + 0.5)/fps was exact
on all 91.
Timestamps are strictly increasing footage milliseconds. Video mode is a
tracker: a repeat leaves the graph in an error state every later call
re-throws, so the landmarker is discarded on failure, and passing the frame
index instead of i*1000/fps moved landmarks six times further from the
per-frame answer.
Frames are verified rather than trusted. requestVideoFrameCallback states
which frame it handed over, the walker discards any other and fails loudly
if the one it asked for never arrives — a stale presentation from the tail
of a previous seek is what produced "asked for frame 1 and it presented
frame 2" on a video whose seeks were in fact exact.
The proxy is re-encoded even when the upload is already H.264: HEVC is not
decodable everywhere, and footage identity is the proxy's digest. The JPEG
stills beside it are tracing references, outside the footage digest because
re-rendering them at another size is not different footage.
Verified end to end in a real browser against real footage: 228/228 frames
detected, a drawn roto face, 37 backend and 234 frontend tests green.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-28 11:32:01 -04:00
|
|
|
|
for index, digest, size in still_blobs:
|
2026-09-28 09:38:49 -04:00
|
|
|
|
blob, _ = Blob.objects.get_or_create(
|
Measure the video, not a PNG per frame
Detection now walks a browser-seekable H.264 proxy in MediaPipe's VIDEO
running mode. The PNG sequence it replaces was 112MB for 7.6 seconds at
1440x1920 and 1.1GB at the 900-frame limit; the proxy is 6MB, and landmarks
detected off decoded H.264 rather than off the PNGs moved at most 0.0033 of
frame width.
Three things had to be true for video mode to work, and each was measured
against the same footage decoded to PNGs:
/blob/<digest> answers byte ranges. Django's FileResponse does no Range
handling, and a media element handed 200 with no Accept-Ranges reports an
empty `seekable`, no-ops every currentTime write, and detects frame one
ninety times without raising.
A seek aims at the MIDDLE of its frame. Aiming at i/fps sits on a frame
boundary and landed one frame early 31 times in 91; (i + 0.5)/fps was exact
on all 91.
Timestamps are strictly increasing footage milliseconds. Video mode is a
tracker: a repeat leaves the graph in an error state every later call
re-throws, so the landmarker is discarded on failure, and passing the frame
index instead of i*1000/fps moved landmarks six times further from the
per-frame answer.
Frames are verified rather than trusted. requestVideoFrameCallback states
which frame it handed over, the walker discards any other and fails loudly
if the one it asked for never arrives — a stale presentation from the tail
of a previous seek is what produced "asked for frame 1 and it presented
frame 2" on a video whose seeks were in fact exact.
The proxy is re-encoded even when the upload is already H.264: HEVC is not
decodable everywhere, and footage identity is the proxy's digest. The JPEG
stills beside it are tracing references, outside the footage digest because
re-rendering them at another size is not different footage.
Verified end to end in a real browser against real footage: 228/228 frames
detected, a drawn roto face, 37 backend and 234 frontend tests green.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-28 11:32:01 -04:00
|
|
|
|
digest=digest, defaults={"size": size, "media_type": "image/jpeg"})
|
2026-09-28 09:38:49 -04:00
|
|
|
|
rows.append(FootageFrame(footage=footage, index=index, blob=blob))
|
|
|
|
|
|
FootageFrame.objects.bulk_create(rows)
|
|
|
|
|
|
return footage
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def run(key):
|
|
|
|
|
|
close_old_connections()
|
|
|
|
|
|
try:
|
|
|
|
|
|
job = Extraction.objects.select_related("source", "source__blob").get(key=key)
|
|
|
|
|
|
job.state, job.progress, job.error = "running", 0, ""
|
|
|
|
|
|
job.save(update_fields=["state", "progress", "error", "updated"])
|
|
|
|
|
|
facts = job.source.probe
|
|
|
|
|
|
source_path = blobs.path_for(job.source.blob_id)
|
|
|
|
|
|
with tempfile.TemporaryDirectory(prefix="arthur-extract-") as directory:
|
|
|
|
|
|
root = Path(directory)
|
Measure the video, not a PNG per frame
Detection now walks a browser-seekable H.264 proxy in MediaPipe's VIDEO
running mode. The PNG sequence it replaces was 112MB for 7.6 seconds at
1440x1920 and 1.1GB at the 900-frame limit; the proxy is 6MB, and landmarks
detected off decoded H.264 rather than off the PNGs moved at most 0.0033 of
frame width.
Three things had to be true for video mode to work, and each was measured
against the same footage decoded to PNGs:
/blob/<digest> answers byte ranges. Django's FileResponse does no Range
handling, and a media element handed 200 with no Accept-Ranges reports an
empty `seekable`, no-ops every currentTime write, and detects frame one
ninety times without raising.
A seek aims at the MIDDLE of its frame. Aiming at i/fps sits on a frame
boundary and landed one frame early 31 times in 91; (i + 0.5)/fps was exact
on all 91.
Timestamps are strictly increasing footage milliseconds. Video mode is a
tracker: a repeat leaves the graph in an error state every later call
re-throws, so the landmarker is discarded on failure, and passing the frame
index instead of i*1000/fps moved landmarks six times further from the
per-frame answer.
Frames are verified rather than trusted. requestVideoFrameCallback states
which frame it handed over, the walker discards any other and fails loudly
if the one it asked for never arrives — a stale presentation from the tail
of a previous seek is what produced "asked for frame 1 and it presented
frame 2" on a video whose seeks were in fact exact.
The proxy is re-encoded even when the upload is already H.264: HEVC is not
decodable everywhere, and footage identity is the proxy's digest. The JPEG
stills beside it are tracing references, outside the footage digest because
re-rendering them at another size is not different footage.
Verified end to end in a real browser against real footage: 228/228 frames
detected, a drawn roto face, 37 backend and 234 frontend tests green.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-28 11:32:01 -04:00
|
|
|
|
proxy_path = root / "proxy.mp4"
|
|
|
|
|
|
_encode_proxy(job, source_path, proxy_path, facts, root)
|
|
|
|
|
|
|
|
|
|
|
|
# Everything downstream describes the PROXY, not the upload.
|
|
|
|
|
|
proxy_facts = probe(proxy_path)
|
|
|
|
|
|
frames = count_frames(proxy_path)
|
|
|
|
|
|
if not 1 <= frames <= 900:
|
|
|
|
|
|
raise ValueError(f"the proxy holds {frames} frames; the limit is 1–900")
|
2026-09-28 12:37:21 -04:00
|
|
|
|
# CHECKED AS A DURATION, not as a frame count. The page's clock is
|
|
|
|
|
|
# `frame = floor(audio.currentTime * fps)`, so what must not drift is
|
|
|
|
|
|
# how long the picture lasts against how long the audio lasts — and
|
|
|
|
|
|
# the source's own frame count is a number this has already caught
|
|
|
|
|
|
# lying. A resample to a constant rate legitimately changes the count
|
|
|
|
|
|
# and must not change the duration.
|
|
|
|
|
|
drift = abs(frames / proxy_facts["fps"] - facts["duration"])
|
|
|
|
|
|
if facts["duration"] > 0 and drift > 0.5:
|
Measure the video, not a PNG per frame
Detection now walks a browser-seekable H.264 proxy in MediaPipe's VIDEO
running mode. The PNG sequence it replaces was 112MB for 7.6 seconds at
1440x1920 and 1.1GB at the 900-frame limit; the proxy is 6MB, and landmarks
detected off decoded H.264 rather than off the PNGs moved at most 0.0033 of
frame width.
Three things had to be true for video mode to work, and each was measured
against the same footage decoded to PNGs:
/blob/<digest> answers byte ranges. Django's FileResponse does no Range
handling, and a media element handed 200 with no Accept-Ranges reports an
empty `seekable`, no-ops every currentTime write, and detects frame one
ninety times without raising.
A seek aims at the MIDDLE of its frame. Aiming at i/fps sits on a frame
boundary and landed one frame early 31 times in 91; (i + 0.5)/fps was exact
on all 91.
Timestamps are strictly increasing footage milliseconds. Video mode is a
tracker: a repeat leaves the graph in an error state every later call
re-throws, so the landmarker is discarded on failure, and passing the frame
index instead of i*1000/fps moved landmarks six times further from the
per-frame answer.
Frames are verified rather than trusted. requestVideoFrameCallback states
which frame it handed over, the walker discards any other and fails loudly
if the one it asked for never arrives — a stale presentation from the tail
of a previous seek is what produced "asked for frame 1 and it presented
frame 2" on a video whose seeks were in fact exact.
The proxy is re-encoded even when the upload is already H.264: HEVC is not
decodable everywhere, and footage identity is the proxy's digest. The JPEG
stills beside it are tracing references, outside the footage digest because
re-rendering them at another size is not different footage.
Verified end to end in a real browser against real footage: 228/228 frames
detected, a drawn roto face, 37 backend and 234 frontend tests green.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-28 11:32:01 -04:00
|
|
|
|
raise ValueError(
|
2026-09-28 12:37:21 -04:00
|
|
|
|
f"the proxy runs {frames / proxy_facts['fps']:.2f}s and the upload "
|
|
|
|
|
|
f"runs {facts['duration']:.2f}s; refusing footage whose picture and "
|
|
|
|
|
|
"audio would drift")
|
Measure the video, not a PNG per frame
Detection now walks a browser-seekable H.264 proxy in MediaPipe's VIDEO
running mode. The PNG sequence it replaces was 112MB for 7.6 seconds at
1440x1920 and 1.1GB at the 900-frame limit; the proxy is 6MB, and landmarks
detected off decoded H.264 rather than off the PNGs moved at most 0.0033 of
frame width.
Three things had to be true for video mode to work, and each was measured
against the same footage decoded to PNGs:
/blob/<digest> answers byte ranges. Django's FileResponse does no Range
handling, and a media element handed 200 with no Accept-Ranges reports an
empty `seekable`, no-ops every currentTime write, and detects frame one
ninety times without raising.
A seek aims at the MIDDLE of its frame. Aiming at i/fps sits on a frame
boundary and landed one frame early 31 times in 91; (i + 0.5)/fps was exact
on all 91.
Timestamps are strictly increasing footage milliseconds. Video mode is a
tracker: a repeat leaves the graph in an error state every later call
re-throws, so the landmarker is discarded on failure, and passing the frame
index instead of i*1000/fps moved landmarks six times further from the
per-frame answer.
Frames are verified rather than trusted. requestVideoFrameCallback states
which frame it handed over, the walker discards any other and fails loudly
if the one it asked for never arrives — a stale presentation from the tail
of a previous seek is what produced "asked for frame 1 and it presented
frame 2" on a video whose seeks were in fact exact.
The proxy is re-encoded even when the upload is already H.264: HEVC is not
decodable everywhere, and footage identity is the proxy's digest. The JPEG
stills beside it are tracing references, outside the footage digest because
re-rendering them at another size is not different footage.
Verified end to end in a real browser against real footage: 228/228 frames
detected, a drawn roto face, 37 backend and 234 frontend tests green.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-28 11:32:01 -04:00
|
|
|
|
proxy_facts["frames"] = frames
|
|
|
|
|
|
|
|
|
|
|
|
frames_dir = root / "stills"
|
|
|
|
|
|
frames_dir.mkdir()
|
|
|
|
|
|
_extract_stills(job, proxy_path, frames_dir, frames, root)
|
|
|
|
|
|
stills = sorted(frames_dir.glob("*.jpg"))
|
|
|
|
|
|
if len(stills) != frames:
|
|
|
|
|
|
raise ValueError(f"wrote {len(stills)} tracing stills for {frames} frames")
|
|
|
|
|
|
|
|
|
|
|
|
job.progress = 85
|
2026-09-28 09:38:49 -04:00
|
|
|
|
job.save(update_fields=["progress", "updated"])
|
|
|
|
|
|
audio_path = root / "audio.wav"
|
|
|
|
|
|
if facts["has_audio"]:
|
|
|
|
|
|
_command(["ffmpeg", "-hide_banner", "-loglevel", "error", "-y",
|
|
|
|
|
|
"-i", str(source_path), "-vn", "-ac", "1", "-ar", "44100",
|
|
|
|
|
|
str(audio_path)])
|
|
|
|
|
|
else:
|
|
|
|
|
|
_command(["ffmpeg", "-hide_banner", "-loglevel", "error", "-y",
|
|
|
|
|
|
"-f", "lavfi", "-i", "anullsrc=r=44100:cl=mono",
|
Measure the video, not a PNG per frame
Detection now walks a browser-seekable H.264 proxy in MediaPipe's VIDEO
running mode. The PNG sequence it replaces was 112MB for 7.6 seconds at
1440x1920 and 1.1GB at the 900-frame limit; the proxy is 6MB, and landmarks
detected off decoded H.264 rather than off the PNGs moved at most 0.0033 of
frame width.
Three things had to be true for video mode to work, and each was measured
against the same footage decoded to PNGs:
/blob/<digest> answers byte ranges. Django's FileResponse does no Range
handling, and a media element handed 200 with no Accept-Ranges reports an
empty `seekable`, no-ops every currentTime write, and detects frame one
ninety times without raising.
A seek aims at the MIDDLE of its frame. Aiming at i/fps sits on a frame
boundary and landed one frame early 31 times in 91; (i + 0.5)/fps was exact
on all 91.
Timestamps are strictly increasing footage milliseconds. Video mode is a
tracker: a repeat leaves the graph in an error state every later call
re-throws, so the landmarker is discarded on failure, and passing the frame
index instead of i*1000/fps moved landmarks six times further from the
per-frame answer.
Frames are verified rather than trusted. requestVideoFrameCallback states
which frame it handed over, the walker discards any other and fails loudly
if the one it asked for never arrives — a stale presentation from the tail
of a previous seek is what produced "asked for frame 1 and it presented
frame 2" on a video whose seeks were in fact exact.
The proxy is re-encoded even when the upload is already H.264: HEVC is not
decodable everywhere, and footage identity is the proxy's digest. The JPEG
stills beside it are tracing references, outside the footage digest because
re-rendering them at another size is not different footage.
Verified end to end in a real browser against real footage: 228/228 frames
detected, a drawn roto face, 37 backend and 234 frontend tests green.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-28 11:32:01 -04:00
|
|
|
|
"-t", str(frames / proxy_facts["fps"]), "-c:a", "pcm_s16le",
|
2026-09-28 09:38:49 -04:00
|
|
|
|
str(audio_path)])
|
Measure the video, not a PNG per frame
Detection now walks a browser-seekable H.264 proxy in MediaPipe's VIDEO
running mode. The PNG sequence it replaces was 112MB for 7.6 seconds at
1440x1920 and 1.1GB at the 900-frame limit; the proxy is 6MB, and landmarks
detected off decoded H.264 rather than off the PNGs moved at most 0.0033 of
frame width.
Three things had to be true for video mode to work, and each was measured
against the same footage decoded to PNGs:
/blob/<digest> answers byte ranges. Django's FileResponse does no Range
handling, and a media element handed 200 with no Accept-Ranges reports an
empty `seekable`, no-ops every currentTime write, and detects frame one
ninety times without raising.
A seek aims at the MIDDLE of its frame. Aiming at i/fps sits on a frame
boundary and landed one frame early 31 times in 91; (i + 0.5)/fps was exact
on all 91.
Timestamps are strictly increasing footage milliseconds. Video mode is a
tracker: a repeat leaves the graph in an error state every later call
re-throws, so the landmarker is discarded on failure, and passing the frame
index instead of i*1000/fps moved landmarks six times further from the
per-frame answer.
Frames are verified rather than trusted. requestVideoFrameCallback states
which frame it handed over, the walker discards any other and fails loudly
if the one it asked for never arrives — a stale presentation from the tail
of a previous seek is what produced "asked for frame 1 and it presented
frame 2" on a video whose seeks were in fact exact.
The proxy is re-encoded even when the upload is already H.264: HEVC is not
decodable everywhere, and footage identity is the proxy's digest. The JPEG
stills beside it are tracing references, outside the footage digest because
re-rendering them at another size is not different footage.
Verified end to end in a real browser against real footage: 228/228 frames
detected, a drawn roto face, 37 backend and 234 frontend tests green.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-28 11:32:01 -04:00
|
|
|
|
footage = _register(job, proxy_path, stills, audio_path, proxy_facts)
|
2026-09-28 09:38:49 -04:00
|
|
|
|
job.footage, job.state, job.progress = footage, "done", 100
|
|
|
|
|
|
job.save(update_fields=["footage", "state", "progress", "updated"])
|
|
|
|
|
|
except Exception as exc:
|
|
|
|
|
|
Extraction.objects.filter(key=key).update(state="failed", error=str(exc)[:2000])
|
|
|
|
|
|
finally:
|
|
|
|
|
|
with _lock:
|
|
|
|
|
|
_active.discard(key)
|
|
|
|
|
|
close_old_connections()
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def enqueue(key):
|
|
|
|
|
|
with _lock:
|
|
|
|
|
|
if key in _active:
|
|
|
|
|
|
return
|
|
|
|
|
|
_active.add(key)
|
|
|
|
|
|
threading.Thread(target=run, args=(key,), daemon=True,
|
|
|
|
|
|
name=f"arthur-extract-{key[7:15]}").start()
|