Measure the video, not a PNG per frame
Detection now walks a browser-seekable H.264 proxy in MediaPipe's VIDEO running mode. The PNG sequence it replaces was 112MB for 7.6 seconds at 1440x1920 and 1.1GB at the 900-frame limit; the proxy is 6MB, and landmarks detected off decoded H.264 rather than off the PNGs moved at most 0.0033 of frame width. Three things had to be true for video mode to work, and each was measured against the same footage decoded to PNGs: /blob/<digest> answers byte ranges. Django's FileResponse does no Range handling, and a media element handed 200 with no Accept-Ranges reports an empty `seekable`, no-ops every currentTime write, and detects frame one ninety times without raising. A seek aims at the MIDDLE of its frame. Aiming at i/fps sits on a frame boundary and landed one frame early 31 times in 91; (i + 0.5)/fps was exact on all 91. Timestamps are strictly increasing footage milliseconds. Video mode is a tracker: a repeat leaves the graph in an error state every later call re-throws, so the landmarker is discarded on failure, and passing the frame index instead of i*1000/fps moved landmarks six times further from the per-frame answer. Frames are verified rather than trusted. requestVideoFrameCallback states which frame it handed over, the walker discards any other and fails loudly if the one it asked for never arrives — a stale presentation from the tail of a previous seek is what produced "asked for frame 1 and it presented frame 2" on a video whose seeks were in fact exact. The proxy is re-encoded even when the upload is already H.264: HEVC is not decodable everywhere, and footage identity is the proxy's digest. The JPEG stills beside it are tracing references, outside the footage digest because re-rendering them at another size is not different footage. Verified end to end in a real browser against real footage: 228/228 frames detected, a drawn roto face, 37 backend and 234 frontend tests green. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
parent
686f897401
commit
83d106bbc5
14 changed files with 748 additions and 142 deletions
|
|
@ -70,9 +70,21 @@ class Extraction(models.Model):
|
|||
class Footage(models.Model):
|
||||
"""Tier 3: the frames and audio of one extraction, immutable.
|
||||
|
||||
`digest` is over the ordered frame digests plus the audio's, so two
|
||||
extractions of the same clip at the same rate are one footage and the same
|
||||
analysis can be reused across both.
|
||||
`digest` is over the PROXY VIDEO's digest plus the audio's and the rate, so
|
||||
two extractions of the same clip at the same settings are one footage and the
|
||||
same analysis can be reused across both.
|
||||
|
||||
THE PROXY IS THE ANALYSIS SOURCE AND THE FRAMES ARE NOT. `video` is one
|
||||
browser-safe H.264 file, and it is what the page seeks through to detect
|
||||
landmarks. `frame_set` is a JPEG per frame at tracing size: reference stills
|
||||
for the tracing editor, never the thing measured. The two are not
|
||||
interchangeable, and which one carries the pixels an analysis was computed
|
||||
from is the difference between a 6MB take and a 1.1GB one.
|
||||
|
||||
So the frame JPEGs are deliberately NOT in `digest`. They are a rendering of
|
||||
this footage for a human to trace over; re-rendering them at another size does
|
||||
not make it different footage, and putting them in the identity would throw
|
||||
away every analysis when the tracing size changed.
|
||||
|
||||
`feature_absence` is the manifest annotation step 8 introduced: known
|
||||
occlusion intervals, one-based and inclusive, expanded into presence tracks by
|
||||
|
|
@ -88,6 +100,10 @@ class Footage(models.Model):
|
|||
width = models.PositiveIntegerField()
|
||||
height = models.PositiveIntegerField()
|
||||
audio = models.ForeignKey(Blob, on_delete=models.PROTECT, related_name="audio_for")
|
||||
video = models.ForeignKey(
|
||||
Blob, null=True, blank=True, on_delete=models.PROTECT, related_name="video_for",
|
||||
help_text="the browser-safe proxy the page detects from; null on pre-proxy footage",
|
||||
)
|
||||
feature_absence = models.JSONField(default=dict, blank=True)
|
||||
created = models.DateTimeField(auto_now_add=True)
|
||||
|
||||
|
|
@ -99,12 +115,14 @@ class Footage(models.Model):
|
|||
|
||||
|
||||
class FootageFrame(models.Model):
|
||||
"""One source frame. A row rather than an entry in a JSON list, because a
|
||||
"""One tracing still. A row rather than an entry in a JSON list, because a
|
||||
frame is a thing the server serves, and because a blob's references have to be
|
||||
countable before anything can be collected."""
|
||||
countable before anything can be collected.
|
||||
|
||||
A REFERENCE IMAGE, NOT A MEASUREMENT INPUT. See `Footage.video`."""
|
||||
|
||||
footage = models.ForeignKey(Footage, on_delete=models.CASCADE, related_name="frame_set")
|
||||
index = models.PositiveIntegerField(help_text="0-based; source frame index + 1 is the PNG's name")
|
||||
index = models.PositiveIntegerField(help_text="0-based; source frame index + 1 is the JPEG's name")
|
||||
blob = models.ForeignKey(Blob, on_delete=models.PROTECT, related_name="frame_for")
|
||||
|
||||
class Meta:
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue