Measure the video, not a PNG per frame
Detection now walks a browser-seekable H.264 proxy in MediaPipe's VIDEO running mode. The PNG sequence it replaces was 112MB for 7.6 seconds at 1440x1920 and 1.1GB at the 900-frame limit; the proxy is 6MB, and landmarks detected off decoded H.264 rather than off the PNGs moved at most 0.0033 of frame width. Three things had to be true for video mode to work, and each was measured against the same footage decoded to PNGs: /blob/<digest> answers byte ranges. Django's FileResponse does no Range handling, and a media element handed 200 with no Accept-Ranges reports an empty `seekable`, no-ops every currentTime write, and detects frame one ninety times without raising. A seek aims at the MIDDLE of its frame. Aiming at i/fps sits on a frame boundary and landed one frame early 31 times in 91; (i + 0.5)/fps was exact on all 91. Timestamps are strictly increasing footage milliseconds. Video mode is a tracker: a repeat leaves the graph in an error state every later call re-throws, so the landmarker is discarded on failure, and passing the frame index instead of i*1000/fps moved landmarks six times further from the per-frame answer. Frames are verified rather than trusted. requestVideoFrameCallback states which frame it handed over, the walker discards any other and fails loudly if the one it asked for never arrives — a stale presentation from the tail of a previous seek is what produced "asked for frame 1 and it presented frame 2" on a video whose seeks were in fact exact. The proxy is re-encoded even when the upload is already H.264: HEVC is not decodable everywhere, and footage identity is the proxy's digest. The JPEG stills beside it are tracing references, outside the footage digest because re-rendering them at another size is not different footage. Verified end to end in a real browser against real footage: 228/228 frames detected, a drawn roto face, 37 backend and 234 frontend tests green. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
parent
686f897401
commit
83d106bbc5
14 changed files with 748 additions and 142 deletions
|
|
@ -1,4 +1,27 @@
|
|||
"""Upload a video once, then decode it into the existing footage model."""
|
||||
"""Upload a video once, then turn it into the two things the app actually reads.
|
||||
|
||||
WHAT CHANGED AND WHY. This used to decode one PNG per source frame and store every
|
||||
one of them. A 7.6-second 1440x1920 take is 112MB that way, and the 900-frame limit
|
||||
is 1.1GB — for pixels whose only consumer was a canvas that MediaPipe then read
|
||||
once. The page now detects from the video itself (see `frontend/src/arthur/flow/
|
||||
ingest.cljs`), so this produces:
|
||||
|
||||
THE PROXY. One browser-safe H.264/yuv420p MP4, CFR, `+faststart`. The same take
|
||||
is 6MB. This is the analysis source, and it is re-encoded RATHER THAN KEPT AS
|
||||
UPLOADED even when the upload is already H.264, for two reasons that are both
|
||||
about not guessing: an iPhone's HEVC is not decodable in every browser, and the
|
||||
footage's identity is the digest of this file — one produced by one ffmpeg
|
||||
invocation, not one that depends on which branch the source happened to take.
|
||||
|
||||
THE TRACING STILLS. One JPEG per frame, long edge capped, for the tracing editor
|
||||
to draw over. Reference images; nothing measures them. They are not in the
|
||||
footage digest — see `models.Footage`.
|
||||
|
||||
The proxy is probed after it is written rather than before. `width`, `height` and
|
||||
`frames` are properties of the file the browser will decode, and taking them from
|
||||
the source instead is how a scaler or a dropped frame becomes a silent one-frame
|
||||
offset between the landmarks and the audio.
|
||||
"""
|
||||
|
||||
import hashlib
|
||||
import json
|
||||
|
|
@ -18,6 +41,14 @@ _active = set()
|
|||
_lock = threading.Lock()
|
||||
TIMEOUT = 3600
|
||||
|
||||
# Visually lossless enough that landmarks do not move: measured against the same
|
||||
# frames as PNGs, IMAGE-mode landmarks shifted at most 0.0033 of frame width.
|
||||
PROXY_CRF = "18"
|
||||
# The long edge of a tracing still. The proxy keeps full resolution because the
|
||||
# detector reads it; a still only has to be good enough to draw a cel over.
|
||||
TRACING_EDGE = 1280
|
||||
TRACING_QUALITY = "4"
|
||||
|
||||
|
||||
def _command(args):
|
||||
result = subprocess.run(args, capture_output=True, text=True, timeout=TIMEOUT)
|
||||
|
|
@ -26,31 +57,34 @@ def _command(args):
|
|||
return result.stdout
|
||||
|
||||
|
||||
def _decode_frames(job, source_path, frames_dir, facts, root):
|
||||
"""Decode one frame per source frame and publish ffmpeg's live frame count."""
|
||||
progress_path = root / "frames.progress"
|
||||
log_path = root / "frames.log"
|
||||
total = facts.get("reported_frames") or round(facts["duration"] * facts["fps"])
|
||||
def _run_with_progress(job, args, root, name, total, span):
|
||||
"""Run one ffmpeg and publish its live frame count as `span` of the job.
|
||||
|
||||
ffmpeg's `-progress` file is the only honest source for this: parsing its
|
||||
stderr means parsing a format that is explicitly not an interface, and a
|
||||
spinner that is not attached to frames is a spinner that lies on a long take.
|
||||
"""
|
||||
progress_path = root / f"{name}.progress"
|
||||
log_path = root / f"{name}.log"
|
||||
first, last = span
|
||||
args = ["ffmpeg", "-hide_banner", "-loglevel", "error", "-y",
|
||||
"-stats_period", "0.25", "-progress", str(progress_path),
|
||||
"-i", str(source_path), "-fps_mode", "passthrough",
|
||||
str(frames_dir / "%04d.png")]
|
||||
"-stats_period", "0.25", "-progress", str(progress_path)] + args
|
||||
with open(log_path, "wb") as log:
|
||||
proc = subprocess.Popen(args, stdout=log, stderr=subprocess.STDOUT)
|
||||
deadline = time.monotonic() + TIMEOUT
|
||||
try:
|
||||
while proc.poll() is None:
|
||||
if time.monotonic() >= deadline:
|
||||
raise TimeoutError("video frame extraction timed out")
|
||||
raise TimeoutError(f"{name} timed out")
|
||||
if progress_path.exists():
|
||||
lines = progress_path.read_text(errors="replace").splitlines()
|
||||
count = next((int(line[6:].strip()) for line in reversed(lines)
|
||||
if line.startswith("frame=") and
|
||||
line[6:].strip().isdigit()), 0)
|
||||
if count and total:
|
||||
progress = min(59, int(60 * count / total))
|
||||
if progress > job.progress:
|
||||
job.progress = progress
|
||||
reached = first + int((last - first) * min(1.0, count / total))
|
||||
if reached > job.progress:
|
||||
job.progress = reached
|
||||
job.save(update_fields=["progress", "updated"])
|
||||
time.sleep(0.2)
|
||||
finally:
|
||||
|
|
@ -58,8 +92,35 @@ def _decode_frames(job, source_path, frames_dir, facts, root):
|
|||
proc.kill()
|
||||
proc.wait()
|
||||
if proc.returncode:
|
||||
raise ValueError(log_path.read_text(errors="replace")[-1200:] or
|
||||
"video frame extraction failed")
|
||||
raise ValueError(log_path.read_text(errors="replace")[-1200:] or f"{name} failed")
|
||||
|
||||
|
||||
def _encode_proxy(job, source_path, proxy_path, facts, root):
|
||||
"""The uploaded video -> one H.264 file every browser can decode and seek."""
|
||||
total = facts.get("reported_frames") or round(facts["duration"] * facts["fps"])
|
||||
_run_with_progress(
|
||||
job,
|
||||
["-i", str(source_path), "-an",
|
||||
# Constant frame rate at the source's own rate. `probe` has already
|
||||
# refused VFR, so this asserts that rather than resampling.
|
||||
"-fps_mode", "cfr", "-r", str(facts["fps"]),
|
||||
"-c:v", "libx264", "-preset", "veryfast", "-crf", PROXY_CRF,
|
||||
# yuv420p and an even frame size are what makes this playable everywhere
|
||||
# rather than only in the browser that happened to be tested.
|
||||
"-pix_fmt", "yuv420p", "-vf", "scale=trunc(iw/2)*2:trunc(ih/2)*2",
|
||||
"-movflags", "+faststart", str(proxy_path)],
|
||||
root, "proxy", total, (0, 55))
|
||||
|
||||
|
||||
def _extract_stills(job, proxy_path, frames_dir, frames, root):
|
||||
"""The proxy -> one tracing JPEG per frame, long edge capped."""
|
||||
_run_with_progress(
|
||||
job,
|
||||
["-i", str(proxy_path), "-fps_mode", "passthrough",
|
||||
"-vf", f"scale='if(gt(iw,ih),min({TRACING_EDGE},iw),-2)':"
|
||||
f"'if(gt(iw,ih),-2,min({TRACING_EDGE},ih))'",
|
||||
"-q:v", TRACING_QUALITY, str(frames_dir / "%04d.jpg")],
|
||||
root, "stills", frames, (55, 85))
|
||||
|
||||
|
||||
def probe(path):
|
||||
|
|
@ -88,40 +149,58 @@ def probe(path):
|
|||
"vfr": False}
|
||||
|
||||
|
||||
def count_frames(path):
|
||||
"""How many frames a file really holds, counted rather than reported.
|
||||
|
||||
`nb_frames` is a container's claim. This is the decoder's answer, and it is
|
||||
what the page will get when it walks the proxy — so a disagreement between the
|
||||
two has to be settled before the count reaches a manifest, not after it has
|
||||
become a one-frame audio offset nobody can find.
|
||||
"""
|
||||
text = _command(["ffprobe", "-v", "error", "-select_streams", "v:0",
|
||||
"-count_frames", "-show_entries", "stream=nb_read_frames",
|
||||
"-of", "default=nokey=1:noprint_wrappers=1", str(path)])
|
||||
counted = text.strip()
|
||||
if not counted.isdigit():
|
||||
raise ValueError("could not count the proxy's frames")
|
||||
return int(counted)
|
||||
|
||||
|
||||
def extraction_key(source, settings):
|
||||
text = json.dumps({"scheme": 1, "source": source.blob_id, "settings": settings},
|
||||
text = json.dumps({"scheme": 2, "source": source.blob_id, "settings": settings},
|
||||
sort_keys=True, separators=(",", ":"))
|
||||
return "sha256:" + hashlib.sha256(text.encode()).hexdigest()
|
||||
|
||||
|
||||
def _register(job, frames, audio_path, facts):
|
||||
width, height = blobs.png_size(frames[0])
|
||||
frame_blobs = []
|
||||
for index, path in enumerate(frames):
|
||||
if blobs.png_size(path) != (width, height):
|
||||
raise ValueError(f"decoded frame {index + 1} has different dimensions")
|
||||
digest, size = blobs.adopt(path)
|
||||
frame_blobs.append((index, digest, size))
|
||||
def _register(job, proxy_path, stills, audio_path, facts):
|
||||
proxy_digest, proxy_size = blobs.adopt(proxy_path)
|
||||
audio_digest, audio_size = blobs.adopt(audio_path)
|
||||
still_blobs = [(index, *blobs.adopt(path)) for index, path in enumerate(stills)]
|
||||
width, height, fps, frames = facts["width"], facts["height"], facts["fps"], facts["frames"]
|
||||
|
||||
# The footage's own identity: the bytes the page will measure, the audio it
|
||||
# will clock against, and the rate that ties them together. Scheme 2 — scheme
|
||||
# 1 hashed a PNG per frame, and those footages name pixels this no longer has.
|
||||
h = hashlib.sha256()
|
||||
h.update(f"arthur-footage-1/{facts['fps']}/{len(frames)}/{width}x{height}\n".encode())
|
||||
for _, digest, _ in frame_blobs:
|
||||
h.update(digest.encode())
|
||||
h.update(f"arthur-footage-2/{fps}/{frames}/{width}x{height}\n".encode())
|
||||
h.update(proxy_digest.encode())
|
||||
h.update(audio_digest.encode())
|
||||
|
||||
with transaction.atomic():
|
||||
proxy_blob, _ = Blob.objects.get_or_create(
|
||||
digest=proxy_digest, defaults={"size": proxy_size, "media_type": "video/mp4"})
|
||||
audio_blob, _ = Blob.objects.get_or_create(
|
||||
digest=audio_digest, defaults={"size": audio_size, "media_type": "audio/wav"})
|
||||
footage, created = Footage.objects.get_or_create(
|
||||
digest=h.hexdigest(),
|
||||
defaults={"label": job.source.filename[:200], "source": job.source.filename[:200],
|
||||
"fps": facts["fps"],
|
||||
"frames": len(frames), "width": width,
|
||||
"height": height, "audio": audio_blob})
|
||||
"fps": fps, "frames": frames, "width": width, "height": height,
|
||||
"audio": audio_blob, "video": proxy_blob})
|
||||
if created:
|
||||
rows = []
|
||||
for index, digest, size in frame_blobs:
|
||||
for index, digest, size in still_blobs:
|
||||
blob, _ = Blob.objects.get_or_create(
|
||||
digest=digest, defaults={"size": size, "media_type": "image/png"})
|
||||
digest=digest, defaults={"size": size, "media_type": "image/jpeg"})
|
||||
rows.append(FootageFrame(footage=footage, index=index, blob=blob))
|
||||
FootageFrame.objects.bulk_create(rows)
|
||||
return footage
|
||||
|
|
@ -137,14 +216,29 @@ def run(key):
|
|||
source_path = blobs.path_for(job.source.blob_id)
|
||||
with tempfile.TemporaryDirectory(prefix="arthur-extract-") as directory:
|
||||
root = Path(directory)
|
||||
frames_dir = root / "frames"
|
||||
frames_dir.mkdir()
|
||||
_decode_frames(job, source_path, frames_dir, facts, root)
|
||||
frames = sorted(frames_dir.glob("*.png"))
|
||||
proxy_path = root / "proxy.mp4"
|
||||
_encode_proxy(job, source_path, proxy_path, facts, root)
|
||||
|
||||
# Everything downstream describes the PROXY, not the upload.
|
||||
proxy_facts = probe(proxy_path)
|
||||
frames = count_frames(proxy_path)
|
||||
if not 1 <= frames <= 900:
|
||||
raise ValueError(f"the proxy holds {frames} frames; the limit is 1–900")
|
||||
expected = facts.get("reported_frames")
|
||||
if not frames or len(frames) > 900 or (expected and len(frames) != expected):
|
||||
raise ValueError(f"decoded {len(frames)} frames; expected {expected or '1–900'}")
|
||||
job.progress = 60
|
||||
if expected and frames != expected:
|
||||
raise ValueError(
|
||||
f"the proxy holds {frames} frames and the upload reports {expected}; "
|
||||
"refusing footage whose picture and audio would drift")
|
||||
proxy_facts["frames"] = frames
|
||||
|
||||
frames_dir = root / "stills"
|
||||
frames_dir.mkdir()
|
||||
_extract_stills(job, proxy_path, frames_dir, frames, root)
|
||||
stills = sorted(frames_dir.glob("*.jpg"))
|
||||
if len(stills) != frames:
|
||||
raise ValueError(f"wrote {len(stills)} tracing stills for {frames} frames")
|
||||
|
||||
job.progress = 85
|
||||
job.save(update_fields=["progress", "updated"])
|
||||
audio_path = root / "audio.wav"
|
||||
if facts["has_audio"]:
|
||||
|
|
@ -154,9 +248,9 @@ def run(key):
|
|||
else:
|
||||
_command(["ffmpeg", "-hide_banner", "-loglevel", "error", "-y",
|
||||
"-f", "lavfi", "-i", "anullsrc=r=44100:cl=mono",
|
||||
"-t", str(len(frames) / facts["fps"]), "-c:a", "pcm_s16le",
|
||||
"-t", str(frames / proxy_facts["fps"]), "-c:a", "pcm_s16le",
|
||||
str(audio_path)])
|
||||
footage = _register(job, frames, audio_path, facts)
|
||||
footage = _register(job, proxy_path, stills, audio_path, proxy_facts)
|
||||
job.footage, job.state, job.progress = footage, "done", 100
|
||||
job.save(update_fields=["footage", "state", "progress", "updated"])
|
||||
except Exception as exc:
|
||||
|
|
|
|||
|
|
@ -0,0 +1,24 @@
|
|||
# Generated by Django 5.2.17 on 2026-09-28 15:10
|
||||
|
||||
import django.db.models.deletion
|
||||
from django.db import migrations, models
|
||||
|
||||
|
||||
class Migration(migrations.Migration):
|
||||
|
||||
dependencies = [
|
||||
('clips', '0003_source_extraction'),
|
||||
]
|
||||
|
||||
operations = [
|
||||
migrations.AddField(
|
||||
model_name='footage',
|
||||
name='video',
|
||||
field=models.ForeignKey(blank=True, help_text='the browser-safe proxy the page detects from; null on pre-proxy footage', null=True, on_delete=django.db.models.deletion.PROTECT, related_name='video_for', to='clips.blob'),
|
||||
),
|
||||
migrations.AlterField(
|
||||
model_name='footageframe',
|
||||
name='index',
|
||||
field=models.PositiveIntegerField(help_text="0-based; source frame index + 1 is the JPEG's name"),
|
||||
),
|
||||
]
|
||||
|
|
@ -70,9 +70,21 @@ class Extraction(models.Model):
|
|||
class Footage(models.Model):
|
||||
"""Tier 3: the frames and audio of one extraction, immutable.
|
||||
|
||||
`digest` is over the ordered frame digests plus the audio's, so two
|
||||
extractions of the same clip at the same rate are one footage and the same
|
||||
analysis can be reused across both.
|
||||
`digest` is over the PROXY VIDEO's digest plus the audio's and the rate, so
|
||||
two extractions of the same clip at the same settings are one footage and the
|
||||
same analysis can be reused across both.
|
||||
|
||||
THE PROXY IS THE ANALYSIS SOURCE AND THE FRAMES ARE NOT. `video` is one
|
||||
browser-safe H.264 file, and it is what the page seeks through to detect
|
||||
landmarks. `frame_set` is a JPEG per frame at tracing size: reference stills
|
||||
for the tracing editor, never the thing measured. The two are not
|
||||
interchangeable, and which one carries the pixels an analysis was computed
|
||||
from is the difference between a 6MB take and a 1.1GB one.
|
||||
|
||||
So the frame JPEGs are deliberately NOT in `digest`. They are a rendering of
|
||||
this footage for a human to trace over; re-rendering them at another size does
|
||||
not make it different footage, and putting them in the identity would throw
|
||||
away every analysis when the tracing size changed.
|
||||
|
||||
`feature_absence` is the manifest annotation step 8 introduced: known
|
||||
occlusion intervals, one-based and inclusive, expanded into presence tracks by
|
||||
|
|
@ -88,6 +100,10 @@ class Footage(models.Model):
|
|||
width = models.PositiveIntegerField()
|
||||
height = models.PositiveIntegerField()
|
||||
audio = models.ForeignKey(Blob, on_delete=models.PROTECT, related_name="audio_for")
|
||||
video = models.ForeignKey(
|
||||
Blob, null=True, blank=True, on_delete=models.PROTECT, related_name="video_for",
|
||||
help_text="the browser-safe proxy the page detects from; null on pre-proxy footage",
|
||||
)
|
||||
feature_absence = models.JSONField(default=dict, blank=True)
|
||||
created = models.DateTimeField(auto_now_add=True)
|
||||
|
||||
|
|
@ -99,12 +115,14 @@ class Footage(models.Model):
|
|||
|
||||
|
||||
class FootageFrame(models.Model):
|
||||
"""One source frame. A row rather than an entry in a JSON list, because a
|
||||
"""One tracing still. A row rather than an entry in a JSON list, because a
|
||||
frame is a thing the server serves, and because a blob's references have to be
|
||||
countable before anything can be collected."""
|
||||
countable before anything can be collected.
|
||||
|
||||
A REFERENCE IMAGE, NOT A MEASUREMENT INPUT. See `Footage.video`."""
|
||||
|
||||
footage = models.ForeignKey(Footage, on_delete=models.CASCADE, related_name="frame_set")
|
||||
index = models.PositiveIntegerField(help_text="0-based; source frame index + 1 is the PNG's name")
|
||||
index = models.PositiveIntegerField(help_text="0-based; source frame index + 1 is the JPEG's name")
|
||||
blob = models.ForeignKey(Blob, on_delete=models.PROTECT, related_name="frame_for")
|
||||
|
||||
class Meta:
|
||||
|
|
|
|||
|
|
@ -105,6 +105,46 @@ class BlobStoreTests(TestCase):
|
|||
self.assertEqual(404, self.client.get("/blob/" + "0" * 64).status_code)
|
||||
self.assertEqual(404, self.client.get("/blob/nonsense").status_code)
|
||||
|
||||
def test_a_blob_serves_byte_ranges(self):
|
||||
# NOT AN OPTIMISATION. A <video> that is handed 200 with no Accept-Ranges
|
||||
# reports an empty `seekable`, every currentTime write is a no-op, and the
|
||||
# detector then measures frame one over and over without anything raising.
|
||||
# Django's FileResponse does no Range handling, so this is the whole of
|
||||
# what makes the analysis source seekable.
|
||||
digest, _ = blobs.write(b"0123456789")
|
||||
Blob.objects.create(digest=digest, size=10, media_type="video/mp4")
|
||||
|
||||
whole = self.client.get(f"/blob/{digest}")
|
||||
self.assertEqual(200, whole.status_code)
|
||||
self.assertEqual("bytes", whole["Accept-Ranges"])
|
||||
|
||||
part = self.client.get(f"/blob/{digest}", headers={"range": "bytes=2-5"})
|
||||
self.assertEqual(206, part.status_code)
|
||||
self.assertEqual("bytes 2-5/10", part["Content-Range"])
|
||||
self.assertEqual("4", part["Content-Length"])
|
||||
self.assertEqual(b"2345", b"".join(part.streaming_content))
|
||||
|
||||
# An open end, which is what a media element actually sends first.
|
||||
tail = self.client.get(f"/blob/{digest}", headers={"range": "bytes=7-"})
|
||||
self.assertEqual(206, tail.status_code)
|
||||
self.assertEqual("bytes 7-9/10", tail["Content-Range"])
|
||||
self.assertEqual(b"789", b"".join(tail.streaming_content))
|
||||
|
||||
# A suffix range asks a different question: the LAST n bytes.
|
||||
suffix = self.client.get(f"/blob/{digest}", headers={"range": "bytes=-3"})
|
||||
self.assertEqual(206, suffix.status_code)
|
||||
self.assertEqual("bytes 7-9/10", suffix["Content-Range"])
|
||||
|
||||
# Past the end is a 416 with the real length, so the client can recover.
|
||||
over = self.client.get(f"/blob/{digest}", headers={"range": "bytes=50-60"})
|
||||
self.assertEqual(416, over.status_code)
|
||||
self.assertEqual("bytes */10", over["Content-Range"])
|
||||
|
||||
# Unparsable is not an error: RFC 9110 says ignore it and send it all.
|
||||
junk = self.client.get(f"/blob/{digest}", headers={"range": "furlongs=1-2"})
|
||||
self.assertEqual(200, junk.status_code)
|
||||
self.assertEqual(b"0123456789", b"".join(junk.streaming_content))
|
||||
|
||||
|
||||
@override_settings(BLOB_ROOT=BLOB_DIR)
|
||||
class Tier2Tests(TestCase):
|
||||
|
|
@ -509,11 +549,12 @@ class PageTests(TestCase):
|
|||
@skipUnless(shutil.which("ffmpeg") and shutil.which("ffprobe"), "ffmpeg is required")
|
||||
@override_settings(BLOB_ROOT=BLOB_DIR)
|
||||
class UploadTests(TestCase):
|
||||
def test_frame_decode_reports_live_progress(self):
|
||||
def test_an_ffmpeg_stage_reports_live_progress_within_its_own_span(self):
|
||||
# The job's percentage is shared between the encode and the stills, so a
|
||||
# stage reports its own fraction of its own span rather than of the job.
|
||||
# Half of the frames through a stage that owns 0-55 is 27.
|
||||
with tempfile.TemporaryDirectory() as directory:
|
||||
root = Path(directory)
|
||||
frames = root / "frames"
|
||||
frames.mkdir()
|
||||
job = Mock(progress=0)
|
||||
|
||||
class FakeProcess:
|
||||
|
|
@ -523,7 +564,7 @@ class UploadTests(TestCase):
|
|||
def poll(self):
|
||||
self.calls += 1
|
||||
if self.calls == 1:
|
||||
(root / "frames.progress").write_text("frame=2\nprogress=continue\n")
|
||||
(root / "proxy.progress").write_text("frame=2\nprogress=continue\n")
|
||||
return None
|
||||
return 0
|
||||
|
||||
|
|
@ -532,10 +573,9 @@ class UploadTests(TestCase):
|
|||
|
||||
with patch("clips.extraction.subprocess.Popen", return_value=FakeProcess()), \
|
||||
patch("clips.extraction.time.sleep"):
|
||||
extraction._decode_frames(job, root / "source.mp4", frames,
|
||||
{"reported_frames": 4, "duration": 1, "fps": 4},
|
||||
root)
|
||||
self.assertEqual(30, job.progress)
|
||||
extraction._run_with_progress(job, ["-i", "in.mp4", "out.mp4"],
|
||||
root, "proxy", 4, (0, 55))
|
||||
self.assertEqual(27, job.progress)
|
||||
job.save.assert_called_once_with(update_fields=["progress", "updated"])
|
||||
|
||||
def test_uploaded_video_extracts_to_reopenable_footage(self):
|
||||
|
|
@ -568,6 +608,61 @@ class UploadTests(TestCase):
|
|||
self.assertEqual("done", job["state"], job)
|
||||
footage = self.client.get(f"/api/footage/{job['footage']}").json()
|
||||
self.assertEqual((4, 64, 48), (footage["frames"], footage["width"], footage["height"]))
|
||||
|
||||
# THE PROXY IS THE ANALYSIS SOURCE. The page seeks this URL frame by
|
||||
# frame, so it has to exist, be a video, and answer a Range request —
|
||||
# without the last of those a media element cannot seek it at all.
|
||||
self.assertTrue(footage["video"].startswith("/blob/"), footage)
|
||||
proxy = self.client.get(footage["video"])
|
||||
self.assertEqual(200, proxy.status_code)
|
||||
self.assertEqual("video/mp4", proxy["Content-Type"])
|
||||
self.assertEqual("bytes", proxy["Accept-Ranges"])
|
||||
self.assertEqual(206, self.client.get(footage["video"],
|
||||
headers={"range": "bytes=0-31"}).status_code)
|
||||
|
||||
# And the stills beside it are JPEGs for tracing, one per frame.
|
||||
self.assertEqual(4, len(footage["urls"]))
|
||||
self.assertEqual(200, self.client.get(footage["urls"][0]).status_code)
|
||||
still = self.client.get(footage["urls"][0])
|
||||
self.assertEqual(200, still.status_code)
|
||||
self.assertEqual("image/jpeg", still["Content-Type"])
|
||||
self.assertEqual(200, self.client.get(footage["audio"]).status_code)
|
||||
|
||||
def test_the_proxy_is_re_encoded_rather_than_the_upload_re_served(self):
|
||||
# The footage's identity is the proxy's digest, and the proxy is produced
|
||||
# by one ffmpeg invocation whatever the upload was. If the upload were
|
||||
# passed through when it happened to be playable, identity would depend on
|
||||
# which branch ran — and an HEVC upload would reach a browser that cannot
|
||||
# decode it.
|
||||
with tempfile.TemporaryDirectory() as directory:
|
||||
path = Path(directory) / "already-h264.mp4"
|
||||
subprocess.run([
|
||||
"ffmpeg", "-hide_banner", "-loglevel", "error", "-y",
|
||||
"-f", "lavfi", "-i", "testsrc=s=64x48:r=4:d=1",
|
||||
"-c:v", "libx264", "-pix_fmt", "yuv420p", str(path),
|
||||
], check=True, capture_output=True)
|
||||
payload = path.read_bytes()
|
||||
|
||||
uploaded = self.client.post("/api/sources", {
|
||||
"file": SimpleUploadedFile("already-h264.mp4", payload, content_type="video/mp4")})
|
||||
with patch("clips.extraction.enqueue", side_effect=extraction.run):
|
||||
queued = self.client.post("/api/extractions", json.dumps({
|
||||
"source": uploaded.json()["id"], "settings": {},
|
||||
}), content_type="application/json")
|
||||
job = self.client.get(f"/api/extractions/{queued.json()['key']}").json()
|
||||
self.assertEqual("done", job["state"], job)
|
||||
|
||||
footage = Footage.objects.get(id=job["footage"])
|
||||
self.assertIsNotNone(footage.video)
|
||||
self.assertNotEqual(Source.objects.get(id=uploaded.json()["id"]).blob_id,
|
||||
footage.video_id)
|
||||
|
||||
def test_footage_without_a_proxy_says_so_rather_than_serving_nothing(self):
|
||||
# Footage ingested before the proxy existed. The manifest reports a null
|
||||
# video so the loader can name the fix; it does not omit the field and let
|
||||
# the client discover it somewhere inside MediaPipe.
|
||||
audio, size = blobs.write(b"RIFF....WAVEfmt ")
|
||||
blob = Blob.objects.create(digest=audio, size=size, media_type="audio/wav")
|
||||
footage = Footage.objects.create(
|
||||
digest="e" * 64, fps=12, frames=3, width=8, height=6, audio=blob)
|
||||
manifest = self.client.get(f"/api/footage/{footage.id}").json()
|
||||
self.assertIsNone(manifest["video"])
|
||||
|
|
|
|||
|
|
@ -25,6 +25,7 @@ tool got worse", with no event to attach it to.
|
|||
"""
|
||||
import hashlib
|
||||
import json
|
||||
import re
|
||||
from functools import lru_cache
|
||||
from pathlib import Path
|
||||
from uuid import UUID
|
||||
|
|
@ -235,6 +236,10 @@ def _footage_json(footage: Footage, urls=True):
|
|||
"height": footage.height,
|
||||
"footage": f"sha256:{footage.digest}",
|
||||
"audio": f"/blob/{footage.audio.digest}",
|
||||
# The analysis source. `null` on footage ingested before the proxy
|
||||
# existed, which the loader reports as "re-extract this" rather than
|
||||
# failing somewhere inside MediaPipe.
|
||||
"video": f"/blob/{footage.video.digest}" if footage.video_id else None,
|
||||
"feature-absence": footage.feature_absence or {},
|
||||
}
|
||||
if urls:
|
||||
|
|
@ -252,26 +257,100 @@ def footage_list(request):
|
|||
@require_http_methods(["GET"])
|
||||
def footage_detail(request, footage_id):
|
||||
try:
|
||||
footage = Footage.objects.select_related("audio").get(id=footage_id)
|
||||
footage = Footage.objects.select_related("audio", "video").get(id=footage_id)
|
||||
except Footage.DoesNotExist:
|
||||
return JsonResponse({"error": "no such footage"}, status=404)
|
||||
return JsonResponse(_footage_json(footage))
|
||||
|
||||
|
||||
_RANGE = re.compile(r"^bytes=(\d*)-(\d*)$")
|
||||
|
||||
|
||||
class _Slice:
|
||||
"""A file, readable only up to `remaining` bytes from where it was seeked."""
|
||||
|
||||
def __init__(self, handle, remaining):
|
||||
self.handle, self.remaining = handle, remaining
|
||||
|
||||
def read(self, size=-1):
|
||||
if self.remaining <= 0:
|
||||
return b""
|
||||
if size < 0 or size > self.remaining:
|
||||
size = self.remaining
|
||||
data = self.handle.read(size)
|
||||
self.remaining -= len(data)
|
||||
return data
|
||||
|
||||
def close(self):
|
||||
self.handle.close()
|
||||
|
||||
|
||||
def _byte_range(header, size):
|
||||
"""One `Range` header -> (start, end) inclusive, or None for the whole blob.
|
||||
|
||||
A syntactically broken header is NOT an error: RFC 9110 says an unparsable
|
||||
Range is ignored and the whole representation is sent, which is what a client
|
||||
that meant nothing by it wants. `False` is the third answer — a range that
|
||||
parses and cannot be satisfied — because that one is a 416.
|
||||
"""
|
||||
if not header:
|
||||
return None
|
||||
match = _RANGE.match(header.strip())
|
||||
if not match or match.group(1) == "" and match.group(2) == "":
|
||||
return None
|
||||
first, last = match.group(1), match.group(2)
|
||||
if first == "":
|
||||
# `bytes=-500`: the LAST 500 bytes, which is a different question.
|
||||
length = int(last)
|
||||
if length == 0:
|
||||
return False
|
||||
return (max(0, size - length), size - 1)
|
||||
start = int(first)
|
||||
end = int(last) if last else size - 1
|
||||
end = min(end, size - 1)
|
||||
if start >= size or start > end:
|
||||
return False
|
||||
return (start, end)
|
||||
|
||||
|
||||
@require_http_methods(["GET"])
|
||||
def blob(request, digest):
|
||||
"""Raw bytes, immutable.
|
||||
"""Raw bytes, immutable, and serveable a slice at a time.
|
||||
|
||||
`immutable` is not optimism here, it is the definition: the name IS the hash of
|
||||
the content, so a cached copy cannot be stale. That is what makes serving 600
|
||||
frames out of this cheap enough to do on every load.
|
||||
the content, so a cached copy cannot be stale. That is what makes serving a
|
||||
take's frames out of this cheap enough to do on every load.
|
||||
|
||||
RANGE IS NOT AN OPTIMISATION HERE, IT IS THE FEATURE. Since the analysis source
|
||||
became a video file, a `<video>` element seeks this URL, and a media element
|
||||
that is handed 200OK with no `Accept-Ranges` cannot seek: it reports an empty
|
||||
`seekable` range, every `currentTime` write is a no-op, and detection then runs
|
||||
ninety times over frame one without anything raising. Django's `FileResponse`
|
||||
does not do this for us — there is no Range handling anywhere in it — so the
|
||||
absence of these thirty lines presents as "MediaPipe's video mode is broken".
|
||||
"""
|
||||
try:
|
||||
row = Blob.objects.get(digest=digest)
|
||||
path = blobs.path_for(digest)
|
||||
except (Blob.DoesNotExist, ValueError):
|
||||
return JsonResponse({"error": "no such blob"}, status=404)
|
||||
response = FileResponse(open(path, "rb"), content_type=row.media_type)
|
||||
|
||||
size = path.stat().st_size
|
||||
span = _byte_range(request.headers.get("Range"), size)
|
||||
if span is False:
|
||||
response = HttpResponse(status=416)
|
||||
response["Content-Range"] = f"bytes */{size}"
|
||||
elif span is None:
|
||||
response = FileResponse(open(path, "rb"), content_type=row.media_type)
|
||||
else:
|
||||
start, end = span
|
||||
handle = open(path, "rb")
|
||||
handle.seek(start)
|
||||
response = FileResponse(_Slice(handle, end - start + 1),
|
||||
status=206, content_type=row.media_type)
|
||||
response["Content-Range"] = f"bytes {start}-{end}/{size}"
|
||||
response["Content-Length"] = str(end - start + 1)
|
||||
response["Accept-Ranges"] = "bytes"
|
||||
response["Cache-Control"] = "public, max-age=31536000, immutable"
|
||||
response["ETag"] = f'"{digest}"'
|
||||
return response
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue