Measure the video, not a PNG per frame

Detection now walks a browser-seekable H.264 proxy in MediaPipe's VIDEO
running mode. The PNG sequence it replaces was 112MB for 7.6 seconds at
1440x1920 and 1.1GB at the 900-frame limit; the proxy is 6MB, and landmarks
detected off decoded H.264 rather than off the PNGs moved at most 0.0033 of
frame width.

Three things had to be true for video mode to work, and each was measured
against the same footage decoded to PNGs:

/blob/<digest> answers byte ranges. Django's FileResponse does no Range
handling, and a media element handed 200 with no Accept-Ranges reports an
empty `seekable`, no-ops every currentTime write, and detects frame one
ninety times without raising.

A seek aims at the MIDDLE of its frame. Aiming at i/fps sits on a frame
boundary and landed one frame early 31 times in 91; (i + 0.5)/fps was exact
on all 91.

Timestamps are strictly increasing footage milliseconds. Video mode is a
tracker: a repeat leaves the graph in an error state every later call
re-throws, so the landmarker is discarded on failure, and passing the frame
index instead of i*1000/fps moved landmarks six times further from the
per-frame answer.

Frames are verified rather than trusted. requestVideoFrameCallback states
which frame it handed over, the walker discards any other and fails loudly
if the one it asked for never arrives — a stale presentation from the tail
of a previous seek is what produced "asked for frame 1 and it presented
frame 2" on a video whose seeks were in fact exact.

The proxy is re-encoded even when the upload is already H.264: HEVC is not
decodable everywhere, and footage identity is the proxy's digest. The JPEG
stills beside it are tracing references, outside the footage digest because
re-rendering them at another size is not different footage.

Verified end to end in a real browser against real footage: 228/228 frames
detected, a drawn roto face, 37 backend and 234 frontend tests green.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
Olive Vaughn 2026-09-28 11:32:01 -04:00
parent 686f897401
commit 83d106bbc5
14 changed files with 748 additions and 142 deletions

View file

@ -25,6 +25,7 @@ tool got worse", with no event to attach it to.
"""
import hashlib
import json
import re
from functools import lru_cache
from pathlib import Path
from uuid import UUID
@ -235,6 +236,10 @@ def _footage_json(footage: Footage, urls=True):
"height": footage.height,
"footage": f"sha256:{footage.digest}",
"audio": f"/blob/{footage.audio.digest}",
# The analysis source. `null` on footage ingested before the proxy
# existed, which the loader reports as "re-extract this" rather than
# failing somewhere inside MediaPipe.
"video": f"/blob/{footage.video.digest}" if footage.video_id else None,
"feature-absence": footage.feature_absence or {},
}
if urls:
@ -252,26 +257,100 @@ def footage_list(request):
@require_http_methods(["GET"])
def footage_detail(request, footage_id):
try:
footage = Footage.objects.select_related("audio").get(id=footage_id)
footage = Footage.objects.select_related("audio", "video").get(id=footage_id)
except Footage.DoesNotExist:
return JsonResponse({"error": "no such footage"}, status=404)
return JsonResponse(_footage_json(footage))
_RANGE = re.compile(r"^bytes=(\d*)-(\d*)$")
class _Slice:
"""A file, readable only up to `remaining` bytes from where it was seeked."""
def __init__(self, handle, remaining):
self.handle, self.remaining = handle, remaining
def read(self, size=-1):
if self.remaining <= 0:
return b""
if size < 0 or size > self.remaining:
size = self.remaining
data = self.handle.read(size)
self.remaining -= len(data)
return data
def close(self):
self.handle.close()
def _byte_range(header, size):
"""One `Range` header -> (start, end) inclusive, or None for the whole blob.
A syntactically broken header is NOT an error: RFC 9110 says an unparsable
Range is ignored and the whole representation is sent, which is what a client
that meant nothing by it wants. `False` is the third answer — a range that
parses and cannot be satisfied — because that one is a 416.
"""
if not header:
return None
match = _RANGE.match(header.strip())
if not match or match.group(1) == "" and match.group(2) == "":
return None
first, last = match.group(1), match.group(2)
if first == "":
# `bytes=-500`: the LAST 500 bytes, which is a different question.
length = int(last)
if length == 0:
return False
return (max(0, size - length), size - 1)
start = int(first)
end = int(last) if last else size - 1
end = min(end, size - 1)
if start >= size or start > end:
return False
return (start, end)
@require_http_methods(["GET"])
def blob(request, digest):
"""Raw bytes, immutable.
"""Raw bytes, immutable, and serveable a slice at a time.
`immutable` is not optimism here, it is the definition: the name IS the hash of
the content, so a cached copy cannot be stale. That is what makes serving 600
frames out of this cheap enough to do on every load.
the content, so a cached copy cannot be stale. That is what makes serving a
take's frames out of this cheap enough to do on every load.
RANGE IS NOT AN OPTIMISATION HERE, IT IS THE FEATURE. Since the analysis source
became a video file, a `<video>` element seeks this URL, and a media element
that is handed 200OK with no `Accept-Ranges` cannot seek: it reports an empty
`seekable` range, every `currentTime` write is a no-op, and detection then runs
ninety times over frame one without anything raising. Django's `FileResponse`
does not do this for us — there is no Range handling anywhere in it — so the
absence of these thirty lines presents as "MediaPipe's video mode is broken".
"""
try:
row = Blob.objects.get(digest=digest)
path = blobs.path_for(digest)
except (Blob.DoesNotExist, ValueError):
return JsonResponse({"error": "no such blob"}, status=404)
response = FileResponse(open(path, "rb"), content_type=row.media_type)
size = path.stat().st_size
span = _byte_range(request.headers.get("Range"), size)
if span is False:
response = HttpResponse(status=416)
response["Content-Range"] = f"bytes */{size}"
elif span is None:
response = FileResponse(open(path, "rb"), content_type=row.media_type)
else:
start, end = span
handle = open(path, "rb")
handle.seek(start)
response = FileResponse(_Slice(handle, end - start + 1),
status=206, content_type=row.media_type)
response["Content-Range"] = f"bytes {start}-{end}/{size}"
response["Content-Length"] = str(end - start + 1)
response["Accept-Ranges"] = "bytes"
response["Cache-Control"] = "public, max-age=31536000, immutable"
response["ETag"] = f'"{digest}"'
return response