Measure the video, not a PNG per frame

Detection now walks a browser-seekable H.264 proxy in MediaPipe's VIDEO
running mode. The PNG sequence it replaces was 112MB for 7.6 seconds at
1440x1920 and 1.1GB at the 900-frame limit; the proxy is 6MB, and landmarks
detected off decoded H.264 rather than off the PNGs moved at most 0.0033 of
frame width.

Three things had to be true for video mode to work, and each was measured
against the same footage decoded to PNGs:

/blob/<digest> answers byte ranges. Django's FileResponse does no Range
handling, and a media element handed 200 with no Accept-Ranges reports an
empty `seekable`, no-ops every currentTime write, and detects frame one
ninety times without raising.

A seek aims at the MIDDLE of its frame. Aiming at i/fps sits on a frame
boundary and landed one frame early 31 times in 91; (i + 0.5)/fps was exact
on all 91.

Timestamps are strictly increasing footage milliseconds. Video mode is a
tracker: a repeat leaves the graph in an error state every later call
re-throws, so the landmarker is discarded on failure, and passing the frame
index instead of i*1000/fps moved landmarks six times further from the
per-frame answer.

Frames are verified rather than trusted. requestVideoFrameCallback states
which frame it handed over, the walker discards any other and fails loudly
if the one it asked for never arrives — a stale presentation from the tail
of a previous seek is what produced "asked for frame 1 and it presented
frame 2" on a video whose seeks were in fact exact.

The proxy is re-encoded even when the upload is already H.264: HEVC is not
decodable everywhere, and footage identity is the proxy's digest. The JPEG
stills beside it are tracing references, outside the footage digest because
re-rendering them at another size is not different footage.

Verified end to end in a real browser against real footage: 228/228 frames
detected, a drawn roto face, 37 backend and 234 frontend tests green.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
Olive Vaughn 2026-09-28 11:32:01 -04:00
parent 686f897401
commit 83d106bbc5
14 changed files with 748 additions and 142 deletions

View file

@ -70,9 +70,21 @@ class Extraction(models.Model):
class Footage(models.Model):
"""Tier 3: the frames and audio of one extraction, immutable.
`digest` is over the ordered frame digests plus the audio's, so two
extractions of the same clip at the same rate are one footage and the same
analysis can be reused across both.
`digest` is over the PROXY VIDEO's digest plus the audio's and the rate, so
two extractions of the same clip at the same settings are one footage and the
same analysis can be reused across both.
THE PROXY IS THE ANALYSIS SOURCE AND THE FRAMES ARE NOT. `video` is one
browser-safe H.264 file, and it is what the page seeks through to detect
landmarks. `frame_set` is a JPEG per frame at tracing size: reference stills
for the tracing editor, never the thing measured. The two are not
interchangeable, and which one carries the pixels an analysis was computed
from is the difference between a 6MB take and a 1.1GB one.
So the frame JPEGs are deliberately NOT in `digest`. They are a rendering of
this footage for a human to trace over; re-rendering them at another size does
not make it different footage, and putting them in the identity would throw
away every analysis when the tracing size changed.
`feature_absence` is the manifest annotation step 8 introduced: known
occlusion intervals, one-based and inclusive, expanded into presence tracks by
@ -88,6 +100,10 @@ class Footage(models.Model):
width = models.PositiveIntegerField()
height = models.PositiveIntegerField()
audio = models.ForeignKey(Blob, on_delete=models.PROTECT, related_name="audio_for")
video = models.ForeignKey(
Blob, null=True, blank=True, on_delete=models.PROTECT, related_name="video_for",
help_text="the browser-safe proxy the page detects from; null on pre-proxy footage",
)
feature_absence = models.JSONField(default=dict, blank=True)
created = models.DateTimeField(auto_now_add=True)
@ -99,12 +115,14 @@ class Footage(models.Model):
class FootageFrame(models.Model):
"""One source frame. A row rather than an entry in a JSON list, because a
"""One tracing still. A row rather than an entry in a JSON list, because a
frame is a thing the server serves, and because a blob's references have to be
countable before anything can be collected."""
countable before anything can be collected.
A REFERENCE IMAGE, NOT A MEASUREMENT INPUT. See `Footage.video`."""
footage = models.ForeignKey(Footage, on_delete=models.CASCADE, related_name="frame_set")
index = models.PositiveIntegerField(help_text="0-based; source frame index + 1 is the PNG's name")
index = models.PositiveIntegerField(help_text="0-based; source frame index + 1 is the JPEG's name")
blob = models.ForeignKey(Blob, on_delete=models.PROTECT, related_name="frame_for")
class Meta: