Measure the video, not a PNG per frame

Detection now walks a browser-seekable H.264 proxy in MediaPipe's VIDEO
running mode. The PNG sequence it replaces was 112MB for 7.6 seconds at
1440x1920 and 1.1GB at the 900-frame limit; the proxy is 6MB, and landmarks
detected off decoded H.264 rather than off the PNGs moved at most 0.0033 of
frame width.

Three things had to be true for video mode to work, and each was measured
against the same footage decoded to PNGs:

/blob/<digest> answers byte ranges. Django's FileResponse does no Range
handling, and a media element handed 200 with no Accept-Ranges reports an
empty `seekable`, no-ops every currentTime write, and detects frame one
ninety times without raising.

A seek aims at the MIDDLE of its frame. Aiming at i/fps sits on a frame
boundary and landed one frame early 31 times in 91; (i + 0.5)/fps was exact
on all 91.

Timestamps are strictly increasing footage milliseconds. Video mode is a
tracker: a repeat leaves the graph in an error state every later call
re-throws, so the landmarker is discarded on failure, and passing the frame
index instead of i*1000/fps moved landmarks six times further from the
per-frame answer.

Frames are verified rather than trusted. requestVideoFrameCallback states
which frame it handed over, the walker discards any other and fails loudly
if the one it asked for never arrives — a stale presentation from the tail
of a previous seek is what produced "asked for frame 1 and it presented
frame 2" on a video whose seeks were in fact exact.

The proxy is re-encoded even when the upload is already H.264: HEVC is not
decodable everywhere, and footage identity is the proxy's digest. The JPEG
stills beside it are tracing references, outside the footage digest because
re-rendering them at another size is not different footage.

Verified end to end in a real browser against real footage: 228/228 frames
detected, a drawn roto face, 37 backend and 234 frontend tests green.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
Olive Vaughn 2026-09-28 11:32:01 -04:00
parent 686f897401
commit 83d106bbc5
14 changed files with 748 additions and 142 deletions

View file

@ -16,9 +16,10 @@ modern conveniences belong in the workflow, not the output. See
## ClojureScript port
The active port plays the synthetic take, accepts video uploads, extracts their
frames and audio, analyzes real footage for mouth, eyes, brows and pixel-derived
teeth, and saves the project with reusable analysis data. The step 8 data model
The active port plays the synthetic take, accepts video uploads, transcodes them
to a browser-seekable proxy plus audio and tracing stills, analyzes real footage
for mouth, eyes, brows and pixel-derived teeth, and saves the project with
reusable analysis data. The step 8 data model
represents persistent feature IDs, eye pairs and feature-level observation gaps;
its controls are still pending. See the [port plan](docs/port-plan.md).
@ -63,7 +64,7 @@ Three tiers, cut by mutability and size — the full argument is in
| --- | --- | --- |
| 1 **authored** | the scene: nodes, channels, features, time maps | the database, as independently addressed leaves. Kilobytes |
| 2 **derived** | detected landmarks, raw mouth crops, and dense channel blocks | `var/blobs`, addressed by analysis and block inputs, including the detector version |
| 3 **source** | uploaded video, extracted frames, and audio | the same blob store, by the hash of their bytes |
| 3 **source** | the uploaded video, the H.264 proxy measured from it, its tracing stills, and audio | the same blob store, by the hash of their bytes |
Only tier 1 is the document. Tier 2 is a pure function of tiers 1 and 3, so a
saved project names its blocks rather than carrying them, and a knob change gives
@ -88,16 +89,27 @@ wasm, which is fetched from a CDN on first use.
For real footage:
```sh
./extract.sh /path/to/clip.mov # -> frames/*.png, audio.wav, manifest.json
```
Upload it in the app. `./extract.sh` still writes the old PNG-sequence bundle and
`ingest_bundle` still registers it, but footage ingested that way has no proxy and
the loader will say so — the measured pixels come out of the video now.
then **Load frames**. MediaPipe's wasm is fetched from jsdelivr on first use;
`face_landmarker.task` is local.
MediaPipe's wasm and `face_landmarker.task` are both local; nothing in detection
touches the network.
Frames are pre-extracted rather than decoded in the page because browser video
seeking is approximate and `requestVideoFrameCallback` only delivers frames at
playback speed — neither gives a deterministic per-frame pass.
Detection reads the VIDEO, not a frame per file. The page seeks the proxy to the
MIDDLE of each frame — `(i + 0.5) / fps` — and waits for
`requestVideoFrameCallback` to hand the frame over, then checks the `mediaTime` it
reports against the frame it asked for. Both halves are load-bearing and both were
measured against the same footage decoded to PNGs: aiming at `i / fps` sits on a
frame boundary and landed one frame early 31 times in 91, and aiming at the middle
was exact on all 91. A run that gets a frame it did not ask for stops and says so,
because a one-frame slip between the landmarks and the audio is not something
anyone finds by looking at the result.
This is what replaced the PNG sequence, which was 112MB for 7.6 seconds and would
be 1.1GB at the 900-frame limit. The proxy is 6MB, and the landmarks barely
notice: detected off decoded H.264 rather than off the PNGs, they moved at most
0.0033 of frame width.
`manifest.json` records the source rate. The extractor keeps every source frame;
the page reads that rate because a guessed fps desynchronises audio from picture.

View file

@ -1,4 +1,27 @@
"""Upload a video once, then decode it into the existing footage model."""
"""Upload a video once, then turn it into the two things the app actually reads.
WHAT CHANGED AND WHY. This used to decode one PNG per source frame and store every
one of them. A 7.6-second 1440x1920 take is 112MB that way, and the 900-frame limit
is 1.1GB — for pixels whose only consumer was a canvas that MediaPipe then read
once. The page now detects from the video itself (see `frontend/src/arthur/flow/
ingest.cljs`), so this produces:
THE PROXY. One browser-safe H.264/yuv420p MP4, CFR, `+faststart`. The same take
is 6MB. This is the analysis source, and it is re-encoded RATHER THAN KEPT AS
UPLOADED even when the upload is already H.264, for two reasons that are both
about not guessing: an iPhone's HEVC is not decodable in every browser, and the
footage's identity is the digest of this file — one produced by one ffmpeg
invocation, not one that depends on which branch the source happened to take.
THE TRACING STILLS. One JPEG per frame, long edge capped, for the tracing editor
to draw over. Reference images; nothing measures them. They are not in the
footage digest — see `models.Footage`.
The proxy is probed after it is written rather than before. `width`, `height` and
`frames` are properties of the file the browser will decode, and taking them from
the source instead is how a scaler or a dropped frame becomes a silent one-frame
offset between the landmarks and the audio.
"""
import hashlib
import json
@ -18,6 +41,14 @@ _active = set()
_lock = threading.Lock()
TIMEOUT = 3600
# Visually lossless enough that landmarks do not move: measured against the same
# frames as PNGs, IMAGE-mode landmarks shifted at most 0.0033 of frame width.
PROXY_CRF = "18"
# The long edge of a tracing still. The proxy keeps full resolution because the
# detector reads it; a still only has to be good enough to draw a cel over.
TRACING_EDGE = 1280
TRACING_QUALITY = "4"
def _command(args):
result = subprocess.run(args, capture_output=True, text=True, timeout=TIMEOUT)
@ -26,31 +57,34 @@ def _command(args):
return result.stdout
def _decode_frames(job, source_path, frames_dir, facts, root):
"""Decode one frame per source frame and publish ffmpeg's live frame count."""
progress_path = root / "frames.progress"
log_path = root / "frames.log"
total = facts.get("reported_frames") or round(facts["duration"] * facts["fps"])
def _run_with_progress(job, args, root, name, total, span):
"""Run one ffmpeg and publish its live frame count as `span` of the job.
ffmpeg's `-progress` file is the only honest source for this: parsing its
stderr means parsing a format that is explicitly not an interface, and a
spinner that is not attached to frames is a spinner that lies on a long take.
"""
progress_path = root / f"{name}.progress"
log_path = root / f"{name}.log"
first, last = span
args = ["ffmpeg", "-hide_banner", "-loglevel", "error", "-y",
"-stats_period", "0.25", "-progress", str(progress_path),
"-i", str(source_path), "-fps_mode", "passthrough",
str(frames_dir / "%04d.png")]
"-stats_period", "0.25", "-progress", str(progress_path)] + args
with open(log_path, "wb") as log:
proc = subprocess.Popen(args, stdout=log, stderr=subprocess.STDOUT)
deadline = time.monotonic() + TIMEOUT
try:
while proc.poll() is None:
if time.monotonic() >= deadline:
raise TimeoutError("video frame extraction timed out")
raise TimeoutError(f"{name} timed out")
if progress_path.exists():
lines = progress_path.read_text(errors="replace").splitlines()
count = next((int(line[6:].strip()) for line in reversed(lines)
if line.startswith("frame=") and
line[6:].strip().isdigit()), 0)
if count and total:
progress = min(59, int(60 * count / total))
if progress > job.progress:
job.progress = progress
reached = first + int((last - first) * min(1.0, count / total))
if reached > job.progress:
job.progress = reached
job.save(update_fields=["progress", "updated"])
time.sleep(0.2)
finally:
@ -58,8 +92,35 @@ def _decode_frames(job, source_path, frames_dir, facts, root):
proc.kill()
proc.wait()
if proc.returncode:
raise ValueError(log_path.read_text(errors="replace")[-1200:] or
"video frame extraction failed")
raise ValueError(log_path.read_text(errors="replace")[-1200:] or f"{name} failed")
def _encode_proxy(job, source_path, proxy_path, facts, root):
"""The uploaded video -> one H.264 file every browser can decode and seek."""
total = facts.get("reported_frames") or round(facts["duration"] * facts["fps"])
_run_with_progress(
job,
["-i", str(source_path), "-an",
# Constant frame rate at the source's own rate. `probe` has already
# refused VFR, so this asserts that rather than resampling.
"-fps_mode", "cfr", "-r", str(facts["fps"]),
"-c:v", "libx264", "-preset", "veryfast", "-crf", PROXY_CRF,
# yuv420p and an even frame size are what makes this playable everywhere
# rather than only in the browser that happened to be tested.
"-pix_fmt", "yuv420p", "-vf", "scale=trunc(iw/2)*2:trunc(ih/2)*2",
"-movflags", "+faststart", str(proxy_path)],
root, "proxy", total, (0, 55))
def _extract_stills(job, proxy_path, frames_dir, frames, root):
"""The proxy -> one tracing JPEG per frame, long edge capped."""
_run_with_progress(
job,
["-i", str(proxy_path), "-fps_mode", "passthrough",
"-vf", f"scale='if(gt(iw,ih),min({TRACING_EDGE},iw),-2)':"
f"'if(gt(iw,ih),-2,min({TRACING_EDGE},ih))'",
"-q:v", TRACING_QUALITY, str(frames_dir / "%04d.jpg")],
root, "stills", frames, (55, 85))
def probe(path):
@ -88,40 +149,58 @@ def probe(path):
"vfr": False}
def count_frames(path):
"""How many frames a file really holds, counted rather than reported.
`nb_frames` is a container's claim. This is the decoder's answer, and it is
what the page will get when it walks the proxy — so a disagreement between the
two has to be settled before the count reaches a manifest, not after it has
become a one-frame audio offset nobody can find.
"""
text = _command(["ffprobe", "-v", "error", "-select_streams", "v:0",
"-count_frames", "-show_entries", "stream=nb_read_frames",
"-of", "default=nokey=1:noprint_wrappers=1", str(path)])
counted = text.strip()
if not counted.isdigit():
raise ValueError("could not count the proxy's frames")
return int(counted)
def extraction_key(source, settings):
text = json.dumps({"scheme": 1, "source": source.blob_id, "settings": settings},
text = json.dumps({"scheme": 2, "source": source.blob_id, "settings": settings},
sort_keys=True, separators=(",", ":"))
return "sha256:" + hashlib.sha256(text.encode()).hexdigest()
def _register(job, frames, audio_path, facts):
width, height = blobs.png_size(frames[0])
frame_blobs = []
for index, path in enumerate(frames):
if blobs.png_size(path) != (width, height):
raise ValueError(f"decoded frame {index + 1} has different dimensions")
digest, size = blobs.adopt(path)
frame_blobs.append((index, digest, size))
def _register(job, proxy_path, stills, audio_path, facts):
proxy_digest, proxy_size = blobs.adopt(proxy_path)
audio_digest, audio_size = blobs.adopt(audio_path)
still_blobs = [(index, *blobs.adopt(path)) for index, path in enumerate(stills)]
width, height, fps, frames = facts["width"], facts["height"], facts["fps"], facts["frames"]
# The footage's own identity: the bytes the page will measure, the audio it
# will clock against, and the rate that ties them together. Scheme 2 — scheme
# 1 hashed a PNG per frame, and those footages name pixels this no longer has.
h = hashlib.sha256()
h.update(f"arthur-footage-1/{facts['fps']}/{len(frames)}/{width}x{height}\n".encode())
for _, digest, _ in frame_blobs:
h.update(digest.encode())
h.update(f"arthur-footage-2/{fps}/{frames}/{width}x{height}\n".encode())
h.update(proxy_digest.encode())
h.update(audio_digest.encode())
with transaction.atomic():
proxy_blob, _ = Blob.objects.get_or_create(
digest=proxy_digest, defaults={"size": proxy_size, "media_type": "video/mp4"})
audio_blob, _ = Blob.objects.get_or_create(
digest=audio_digest, defaults={"size": audio_size, "media_type": "audio/wav"})
footage, created = Footage.objects.get_or_create(
digest=h.hexdigest(),
defaults={"label": job.source.filename[:200], "source": job.source.filename[:200],
"fps": facts["fps"],
"frames": len(frames), "width": width,
"height": height, "audio": audio_blob})
"fps": fps, "frames": frames, "width": width, "height": height,
"audio": audio_blob, "video": proxy_blob})
if created:
rows = []
for index, digest, size in frame_blobs:
for index, digest, size in still_blobs:
blob, _ = Blob.objects.get_or_create(
digest=digest, defaults={"size": size, "media_type": "image/png"})
digest=digest, defaults={"size": size, "media_type": "image/jpeg"})
rows.append(FootageFrame(footage=footage, index=index, blob=blob))
FootageFrame.objects.bulk_create(rows)
return footage
@ -137,14 +216,29 @@ def run(key):
source_path = blobs.path_for(job.source.blob_id)
with tempfile.TemporaryDirectory(prefix="arthur-extract-") as directory:
root = Path(directory)
frames_dir = root / "frames"
frames_dir.mkdir()
_decode_frames(job, source_path, frames_dir, facts, root)
frames = sorted(frames_dir.glob("*.png"))
proxy_path = root / "proxy.mp4"
_encode_proxy(job, source_path, proxy_path, facts, root)
# Everything downstream describes the PROXY, not the upload.
proxy_facts = probe(proxy_path)
frames = count_frames(proxy_path)
if not 1 <= frames <= 900:
raise ValueError(f"the proxy holds {frames} frames; the limit is 1–900")
expected = facts.get("reported_frames")
if not frames or len(frames) > 900 or (expected and len(frames) != expected):
raise ValueError(f"decoded {len(frames)} frames; expected {expected or '1–900'}")
job.progress = 60
if expected and frames != expected:
raise ValueError(
f"the proxy holds {frames} frames and the upload reports {expected}; "
"refusing footage whose picture and audio would drift")
proxy_facts["frames"] = frames
frames_dir = root / "stills"
frames_dir.mkdir()
_extract_stills(job, proxy_path, frames_dir, frames, root)
stills = sorted(frames_dir.glob("*.jpg"))
if len(stills) != frames:
raise ValueError(f"wrote {len(stills)} tracing stills for {frames} frames")
job.progress = 85
job.save(update_fields=["progress", "updated"])
audio_path = root / "audio.wav"
if facts["has_audio"]:
@ -154,9 +248,9 @@ def run(key):
else:
_command(["ffmpeg", "-hide_banner", "-loglevel", "error", "-y",
"-f", "lavfi", "-i", "anullsrc=r=44100:cl=mono",
"-t", str(len(frames) / facts["fps"]), "-c:a", "pcm_s16le",
"-t", str(frames / proxy_facts["fps"]), "-c:a", "pcm_s16le",
str(audio_path)])
footage = _register(job, frames, audio_path, facts)
footage = _register(job, proxy_path, stills, audio_path, proxy_facts)
job.footage, job.state, job.progress = footage, "done", 100
job.save(update_fields=["footage", "state", "progress", "updated"])
except Exception as exc:

View file

@ -0,0 +1,24 @@
# Generated by Django 5.2.17 on 2026-09-28 15:10
import django.db.models.deletion
from django.db import migrations, models
class Migration(migrations.Migration):
dependencies = [
('clips', '0003_source_extraction'),
]
operations = [
migrations.AddField(
model_name='footage',
name='video',
field=models.ForeignKey(blank=True, help_text='the browser-safe proxy the page detects from; null on pre-proxy footage', null=True, on_delete=django.db.models.deletion.PROTECT, related_name='video_for', to='clips.blob'),
),
migrations.AlterField(
model_name='footageframe',
name='index',
field=models.PositiveIntegerField(help_text="0-based; source frame index + 1 is the JPEG's name"),
),
]

View file

@ -70,9 +70,21 @@ class Extraction(models.Model):
class Footage(models.Model):
"""Tier 3: the frames and audio of one extraction, immutable.
`digest` is over the ordered frame digests plus the audio's, so two
extractions of the same clip at the same rate are one footage and the same
analysis can be reused across both.
`digest` is over the PROXY VIDEO's digest plus the audio's and the rate, so
two extractions of the same clip at the same settings are one footage and the
same analysis can be reused across both.
THE PROXY IS THE ANALYSIS SOURCE AND THE FRAMES ARE NOT. `video` is one
browser-safe H.264 file, and it is what the page seeks through to detect
landmarks. `frame_set` is a JPEG per frame at tracing size: reference stills
for the tracing editor, never the thing measured. The two are not
interchangeable, and which one carries the pixels an analysis was computed
from is the difference between a 6MB take and a 1.1GB one.
So the frame JPEGs are deliberately NOT in `digest`. They are a rendering of
this footage for a human to trace over; re-rendering them at another size does
not make it different footage, and putting them in the identity would throw
away every analysis when the tracing size changed.
`feature_absence` is the manifest annotation step 8 introduced: known
occlusion intervals, one-based and inclusive, expanded into presence tracks by
@ -88,6 +100,10 @@ class Footage(models.Model):
width = models.PositiveIntegerField()
height = models.PositiveIntegerField()
audio = models.ForeignKey(Blob, on_delete=models.PROTECT, related_name="audio_for")
video = models.ForeignKey(
Blob, null=True, blank=True, on_delete=models.PROTECT, related_name="video_for",
help_text="the browser-safe proxy the page detects from; null on pre-proxy footage",
)
feature_absence = models.JSONField(default=dict, blank=True)
created = models.DateTimeField(auto_now_add=True)
@ -99,12 +115,14 @@ class Footage(models.Model):
class FootageFrame(models.Model):
"""One source frame. A row rather than an entry in a JSON list, because a
"""One tracing still. A row rather than an entry in a JSON list, because a
frame is a thing the server serves, and because a blob's references have to be
countable before anything can be collected."""
countable before anything can be collected.
A REFERENCE IMAGE, NOT A MEASUREMENT INPUT. See `Footage.video`."""
footage = models.ForeignKey(Footage, on_delete=models.CASCADE, related_name="frame_set")
index = models.PositiveIntegerField(help_text="0-based; source frame index + 1 is the PNG's name")
index = models.PositiveIntegerField(help_text="0-based; source frame index + 1 is the JPEG's name")
blob = models.ForeignKey(Blob, on_delete=models.PROTECT, related_name="frame_for")
class Meta:

View file

@ -105,6 +105,46 @@ class BlobStoreTests(TestCase):
self.assertEqual(404, self.client.get("/blob/" + "0" * 64).status_code)
self.assertEqual(404, self.client.get("/blob/nonsense").status_code)
def test_a_blob_serves_byte_ranges(self):
# NOT AN OPTIMISATION. A <video> that is handed 200 with no Accept-Ranges
# reports an empty `seekable`, every currentTime write is a no-op, and the
# detector then measures frame one over and over without anything raising.
# Django's FileResponse does no Range handling, so this is the whole of
# what makes the analysis source seekable.
digest, _ = blobs.write(b"0123456789")
Blob.objects.create(digest=digest, size=10, media_type="video/mp4")
whole = self.client.get(f"/blob/{digest}")
self.assertEqual(200, whole.status_code)
self.assertEqual("bytes", whole["Accept-Ranges"])
part = self.client.get(f"/blob/{digest}", headers={"range": "bytes=2-5"})
self.assertEqual(206, part.status_code)
self.assertEqual("bytes 2-5/10", part["Content-Range"])
self.assertEqual("4", part["Content-Length"])
self.assertEqual(b"2345", b"".join(part.streaming_content))
# An open end, which is what a media element actually sends first.
tail = self.client.get(f"/blob/{digest}", headers={"range": "bytes=7-"})
self.assertEqual(206, tail.status_code)
self.assertEqual("bytes 7-9/10", tail["Content-Range"])
self.assertEqual(b"789", b"".join(tail.streaming_content))
# A suffix range asks a different question: the LAST n bytes.
suffix = self.client.get(f"/blob/{digest}", headers={"range": "bytes=-3"})
self.assertEqual(206, suffix.status_code)
self.assertEqual("bytes 7-9/10", suffix["Content-Range"])
# Past the end is a 416 with the real length, so the client can recover.
over = self.client.get(f"/blob/{digest}", headers={"range": "bytes=50-60"})
self.assertEqual(416, over.status_code)
self.assertEqual("bytes */10", over["Content-Range"])
# Unparsable is not an error: RFC 9110 says ignore it and send it all.
junk = self.client.get(f"/blob/{digest}", headers={"range": "furlongs=1-2"})
self.assertEqual(200, junk.status_code)
self.assertEqual(b"0123456789", b"".join(junk.streaming_content))
@override_settings(BLOB_ROOT=BLOB_DIR)
class Tier2Tests(TestCase):
@ -509,11 +549,12 @@ class PageTests(TestCase):
@skipUnless(shutil.which("ffmpeg") and shutil.which("ffprobe"), "ffmpeg is required")
@override_settings(BLOB_ROOT=BLOB_DIR)
class UploadTests(TestCase):
def test_frame_decode_reports_live_progress(self):
def test_an_ffmpeg_stage_reports_live_progress_within_its_own_span(self):
# The job's percentage is shared between the encode and the stills, so a
# stage reports its own fraction of its own span rather than of the job.
# Half of the frames through a stage that owns 0-55 is 27.
with tempfile.TemporaryDirectory() as directory:
root = Path(directory)
frames = root / "frames"
frames.mkdir()
job = Mock(progress=0)
class FakeProcess:
@ -523,7 +564,7 @@ class UploadTests(TestCase):
def poll(self):
self.calls += 1
if self.calls == 1:
(root / "frames.progress").write_text("frame=2\nprogress=continue\n")
(root / "proxy.progress").write_text("frame=2\nprogress=continue\n")
return None
return 0
@ -532,10 +573,9 @@ class UploadTests(TestCase):
with patch("clips.extraction.subprocess.Popen", return_value=FakeProcess()), \
patch("clips.extraction.time.sleep"):
extraction._decode_frames(job, root / "source.mp4", frames,
{"reported_frames": 4, "duration": 1, "fps": 4},
root)
self.assertEqual(30, job.progress)
extraction._run_with_progress(job, ["-i", "in.mp4", "out.mp4"],
root, "proxy", 4, (0, 55))
self.assertEqual(27, job.progress)
job.save.assert_called_once_with(update_fields=["progress", "updated"])
def test_uploaded_video_extracts_to_reopenable_footage(self):
@ -568,6 +608,61 @@ class UploadTests(TestCase):
self.assertEqual("done", job["state"], job)
footage = self.client.get(f"/api/footage/{job['footage']}").json()
self.assertEqual((4, 64, 48), (footage["frames"], footage["width"], footage["height"]))
# THE PROXY IS THE ANALYSIS SOURCE. The page seeks this URL frame by
# frame, so it has to exist, be a video, and answer a Range request —
# without the last of those a media element cannot seek it at all.
self.assertTrue(footage["video"].startswith("/blob/"), footage)
proxy = self.client.get(footage["video"])
self.assertEqual(200, proxy.status_code)
self.assertEqual("video/mp4", proxy["Content-Type"])
self.assertEqual("bytes", proxy["Accept-Ranges"])
self.assertEqual(206, self.client.get(footage["video"],
headers={"range": "bytes=0-31"}).status_code)
# And the stills beside it are JPEGs for tracing, one per frame.
self.assertEqual(4, len(footage["urls"]))
self.assertEqual(200, self.client.get(footage["urls"][0]).status_code)
still = self.client.get(footage["urls"][0])
self.assertEqual(200, still.status_code)
self.assertEqual("image/jpeg", still["Content-Type"])
self.assertEqual(200, self.client.get(footage["audio"]).status_code)
def test_the_proxy_is_re_encoded_rather_than_the_upload_re_served(self):
# The footage's identity is the proxy's digest, and the proxy is produced
# by one ffmpeg invocation whatever the upload was. If the upload were
# passed through when it happened to be playable, identity would depend on
# which branch ran — and an HEVC upload would reach a browser that cannot
# decode it.
with tempfile.TemporaryDirectory() as directory:
path = Path(directory) / "already-h264.mp4"
subprocess.run([
"ffmpeg", "-hide_banner", "-loglevel", "error", "-y",
"-f", "lavfi", "-i", "testsrc=s=64x48:r=4:d=1",
"-c:v", "libx264", "-pix_fmt", "yuv420p", str(path),
], check=True, capture_output=True)
payload = path.read_bytes()
uploaded = self.client.post("/api/sources", {
"file": SimpleUploadedFile("already-h264.mp4", payload, content_type="video/mp4")})
with patch("clips.extraction.enqueue", side_effect=extraction.run):
queued = self.client.post("/api/extractions", json.dumps({
"source": uploaded.json()["id"], "settings": {},
}), content_type="application/json")
job = self.client.get(f"/api/extractions/{queued.json()['key']}").json()
self.assertEqual("done", job["state"], job)
footage = Footage.objects.get(id=job["footage"])
self.assertIsNotNone(footage.video)
self.assertNotEqual(Source.objects.get(id=uploaded.json()["id"]).blob_id,
footage.video_id)
def test_footage_without_a_proxy_says_so_rather_than_serving_nothing(self):
# Footage ingested before the proxy existed. The manifest reports a null
# video so the loader can name the fix; it does not omit the field and let
# the client discover it somewhere inside MediaPipe.
audio, size = blobs.write(b"RIFF....WAVEfmt ")
blob = Blob.objects.create(digest=audio, size=size, media_type="audio/wav")
footage = Footage.objects.create(
digest="e" * 64, fps=12, frames=3, width=8, height=6, audio=blob)
manifest = self.client.get(f"/api/footage/{footage.id}").json()
self.assertIsNone(manifest["video"])

View file

@ -25,6 +25,7 @@ tool got worse", with no event to attach it to.
"""
import hashlib
import json
import re
from functools import lru_cache
from pathlib import Path
from uuid import UUID
@ -235,6 +236,10 @@ def _footage_json(footage: Footage, urls=True):
"height": footage.height,
"footage": f"sha256:{footage.digest}",
"audio": f"/blob/{footage.audio.digest}",
# The analysis source. `null` on footage ingested before the proxy
# existed, which the loader reports as "re-extract this" rather than
# failing somewhere inside MediaPipe.
"video": f"/blob/{footage.video.digest}" if footage.video_id else None,
"feature-absence": footage.feature_absence or {},
}
if urls:
@ -252,26 +257,100 @@ def footage_list(request):
@require_http_methods(["GET"])
def footage_detail(request, footage_id):
try:
footage = Footage.objects.select_related("audio").get(id=footage_id)
footage = Footage.objects.select_related("audio", "video").get(id=footage_id)
except Footage.DoesNotExist:
return JsonResponse({"error": "no such footage"}, status=404)
return JsonResponse(_footage_json(footage))
_RANGE = re.compile(r"^bytes=(\d*)-(\d*)$")
class _Slice:
"""A file, readable only up to `remaining` bytes from where it was seeked."""
def __init__(self, handle, remaining):
self.handle, self.remaining = handle, remaining
def read(self, size=-1):
if self.remaining <= 0:
return b""
if size < 0 or size > self.remaining:
size = self.remaining
data = self.handle.read(size)
self.remaining -= len(data)
return data
def close(self):
self.handle.close()
def _byte_range(header, size):
"""One `Range` header -> (start, end) inclusive, or None for the whole blob.
A syntactically broken header is NOT an error: RFC 9110 says an unparsable
Range is ignored and the whole representation is sent, which is what a client
that meant nothing by it wants. `False` is the third answer — a range that
parses and cannot be satisfied — because that one is a 416.
"""
if not header:
return None
match = _RANGE.match(header.strip())
if not match or match.group(1) == "" and match.group(2) == "":
return None
first, last = match.group(1), match.group(2)
if first == "":
# `bytes=-500`: the LAST 500 bytes, which is a different question.
length = int(last)
if length == 0:
return False
return (max(0, size - length), size - 1)
start = int(first)
end = int(last) if last else size - 1
end = min(end, size - 1)
if start >= size or start > end:
return False
return (start, end)
@require_http_methods(["GET"])
def blob(request, digest):
"""Raw bytes, immutable.
"""Raw bytes, immutable, and serveable a slice at a time.
`immutable` is not optimism here, it is the definition: the name IS the hash of
the content, so a cached copy cannot be stale. That is what makes serving 600
frames out of this cheap enough to do on every load.
the content, so a cached copy cannot be stale. That is what makes serving a
take's frames out of this cheap enough to do on every load.
RANGE IS NOT AN OPTIMISATION HERE, IT IS THE FEATURE. Since the analysis source
became a video file, a `<video>` element seeks this URL, and a media element
that is handed 200OK with no `Accept-Ranges` cannot seek: it reports an empty
`seekable` range, every `currentTime` write is a no-op, and detection then runs
ninety times over frame one without anything raising. Django's `FileResponse`
does not do this for us — there is no Range handling anywhere in it — so the
absence of these thirty lines presents as "MediaPipe's video mode is broken".
"""
try:
row = Blob.objects.get(digest=digest)
path = blobs.path_for(digest)
except (Blob.DoesNotExist, ValueError):
return JsonResponse({"error": "no such blob"}, status=404)
response = FileResponse(open(path, "rb"), content_type=row.media_type)
size = path.stat().st_size
span = _byte_range(request.headers.get("Range"), size)
if span is False:
response = HttpResponse(status=416)
response["Content-Range"] = f"bytes */{size}"
elif span is None:
response = FileResponse(open(path, "rb"), content_type=row.media_type)
else:
start, end = span
handle = open(path, "rb")
handle.seek(start)
response = FileResponse(_Slice(handle, end - start + 1),
status=206, content_type=row.media_type)
response["Content-Range"] = f"bytes {start}-{end}/{size}"
response["Content-Length"] = str(end - start + 1)
response["Accept-Ranges"] = "bytes"
response["Cache-Control"] = "public, max-age=31536000, immutable"
response["ETag"] = f'"{digest}"'
return response

View file

@ -145,8 +145,8 @@ boundary.**
| # | Stage | In | Out | Cost |
| --- | --- | --- | --- | --- |
| 1 | **ingest** | video | footage: frames, audio, manifest | minutes, in-app |
| 2 | **detect** | footage | raw landmarks per frame | minutes, **cached** |
| 1 | **ingest** | video | footage: a seekable H.264 proxy, tracing stills, audio, manifest | minutes, in-app |
| 2 | **detect** | the proxy, walked one frame at a time | raw landmarks per frame | minutes, **cached** |
| 3 | **measure** | landmarks | anchor fit, residual, head-local rings, signals, interior pixels | seconds |
| 4 | **condition** | measurements | smoothed transforms and contours | milliseconds |
| 5 | **key** | conditioned signals + policy | channels: sparse keys, quantised holds, kept frames | milliseconds |
@ -575,7 +575,8 @@ PUT /api/analyses/<key> link dense landmarks, mask, crops
POST /api/blocks/missing {keys} -> {missing}
POST /api/blocks {key, descriptor, data, state}
GET /api/blocks/<key>
GET /api/footage/<id> the manifest, with a URL per frame
GET /api/footage/<id> the manifest: the proxy to measure, audio, a URL per tracing still
GET /blob/<digest> immutable bytes, and RANGE-capable so a <video> can seek one
POST /api/sources multipart video upload
POST /api/extractions idempotent decode job
GET /api/extractions/<key> job state and footage id
@ -600,10 +601,28 @@ separately.
**A manifest names frames; it does not locate them.** Until step 9 the client
fetched `manifest.json` and built `frames/0001.png` itself, which made the frame
layout a shared secret between a shell script and a ClojureScript namespace. The
manifest now carries a URL per frame, so the frames can live in the blob store —
or be produced by the app's video upload and server-side ffmpeg extraction. The
upload path needs no `manifest.json` file: the server builds the footage response
from the extracted frame records. The producer changes; the shape does not.
manifest now carries every URL, so the bytes can live in the blob store — or be
produced by the app's video upload and server-side ffmpeg extraction. The upload
path needs no `manifest.json` file: the server builds the footage response from
its own records. The producer changes; the shape does not.
**Tier 3 keeps a video, not a frame per file.** Stage 1 used to decode a PNG per
source frame: 112MB for 7.6 seconds at 1440x1920, and 1.1GB at the 900-frame
limit, for pixels whose only consumer was a canvas MediaPipe read once. It now
writes one browser-safe H.264 proxy — 6MB for the same take — and the page seeks
THAT, frame by frame, in MediaPipe's video running mode. The JPEG stills beside it
are reference images for tracing; nothing measures them, so they are deliberately
outside the footage digest and re-rendering them at another size does not
invalidate an analysis.
Two things make that trustworthy rather than merely smaller. `/blob/<digest>`
answers byte ranges, because a media element handed 200 with no `Accept-Ranges`
reports an empty `seekable` and silently refuses to move — Django's `FileResponse`
does no Range handling, so this is code we own. And every frame is CHECKED:
`requestVideoFrameCallback` states the `mediaTime` of the frame it hands over, the
walker compares it to the frame it asked for, and a mismatch ends the run. Content
addressing over landmarks whose frame alignment was assumed would be addressing a
guess.
**The document stores what a block IS, not what it holds.** A block's element type
is in its own descriptor, which is the only place it is written down: an

View file

@ -114,12 +114,29 @@ and real footage use `src/arthur/flow/take.cljs` for the measurement order and
### Real footage
Choose a video in the **footage** file input. The server probes it, extracts one
PNG per source frame and WAV audio, then makes the resulting footage selectable.
Click **load frames** to detect and freeze it. Extraction progress is currently
read from `/api/extractions/<key>`; a future WebSocket can push the same job state.
The uploaded bytes, extraction job, and decoded footage have separate records, so
the same uploaded video can be reopened without decoding it again.
Choose a video in the **footage** file input. The server probes it, re-encodes it
to a browser-seekable H.264 proxy, pulls WAV audio and one tracing JPEG per frame,
then makes the resulting footage selectable. Click **load frames** to detect and
freeze it. Extraction progress is currently read from `/api/extractions/<key>`; a
future WebSocket can push the same job state. The uploaded bytes, extraction job,
and decoded footage have separate records, so the same uploaded video can be
reopened without decoding it again.
**The proxy is what gets measured, and the stills are not.** `flow/ingest` steps
the proxy one frame at a time — seek to `(i + 0.5) / fps`, wait for
`requestVideoFrameCallback`, check the `mediaTime` it reports is the frame that
was asked for — and `flow/detect` hands each frame to MediaPipe in **VIDEO**
running mode at `i * 1000 / fps` milliseconds. That timestamp has to increase
strictly and has to be real footage time: video mode is a tracker, it reads the
gap between timestamps as motion, and a repeat leaves the graph in an error state
that every later call re-throws. The JPEGs beside the proxy are reference images
for the tracing editor and nothing measures them, so they are not in the footage
digest.
It is re-encoded even when the upload is already H.264, for two reasons: an
iPhone's HEVC is not decodable in every browser, and the footage's identity is the
proxy's digest — one produced by one ffmpeg invocation, not one that depends on
which branch the source happened to take.
The command-line route is also available for an existing extracted bundle:

View file

@ -16,47 +16,56 @@
[arthur.fx.http :as http]
[re-frame.core :as rf]))
(defn- detect-frames! [manifest model]
(let [canvas (.createElement js/document "canvas")
ctx (.getContext canvas "2d")
raw (atom [])
crops (atom [])
dims (atom nil)
total (:frames manifest)]
(js/Promise.
(fn [resolve reject]
(letfn [(next-frame [i]
(if (= i total)
(try
(resolve (assoc (detect/fill-gaps @raw)
:dimensions @dims :crops @crops))
(catch :default error (reject error)))
(-> (ingest/image! (ingest/frame-url manifest i))
(.then
(fn [image]
(let [wh [(.-naturalWidth image) (.-naturalHeight image)]]
(when (and @dims (not= @dims wh))
(throw (ex-info (str "frame " (inc i) " has different dimensions")
{:expected @dims :actual wh})))
(reset! dims wh)
(set! (.-width canvas) (first wh))
(set! (.-height canvas) (second wh))
(.drawImage ctx image 0 0)
(let [face (detect/detect! model canvas)
ring (when face (mapv #(nth face %) lm/LIPS-INNER))
box (when ring (interior/crop ring wh))
pixels (when box
(.-data (.getImageData ctx (:x box) (:y box)
(:w box) (:h box))))]
(swap! raw conj face)
(swap! crops conj (when box {:box box :data pixels})))
(when (or (zero? i) (zero? (mod (inc i) 4)) (= (inc i) total))
(rf/dispatch [::progress (str "detecting " (inc i) "/" total)]))
;; Let the status and the transport paint between sync
;; MediaPipe calls, and release each decoded PNG.
(js/setTimeout #(next-frame (inc i)) 0))))
(.catch reject))))]
(next-frame 0))))))
(defn- detect-frames!
"Walk the proxy once, forward, and measure every frame as it goes.
ONCE AND FORWARD IS A REQUIREMENT, NOT A STYLE. MediaPipe's video mode is a
tracker whose input stream refuses a timestamp that does not advance, and the
error it raises is terminal for the landmarker — so there is no re-reading a
frame, no retry of frame 40, and no second pass. Each frame is decoded, handed
to the detector at its own time in the take, and its mouth crop read off the
same canvas before the loop moves on."
[manifest model]
(let [[w h] [(:width manifest) (:height manifest)]
canvas (.createElement js/document "canvas")
ctx (.getContext canvas "2d" #js {:willReadFrequently true})
fps (:fps manifest)
raw (atom [])
crops (atom [])
total (:frames manifest)]
(set! (.-width canvas) w)
(set! (.-height canvas) h)
(-> (ingest/video! (ingest/video-url manifest) w h)
(.then
(fn [video]
(js/Promise.
(fn [resolve reject]
(letfn [(next-frame [i]
(if (= i total)
(try
(resolve (assoc (detect/fill-gaps @raw)
:dimensions [w h] :crops @crops))
(catch :default error (reject error)))
(-> (ingest/frame! video fps i)
(.then
(fn [_]
(.drawImage ctx video 0 0)
(let [face (detect/detect! model canvas
(ingest/frame-ms fps i))
ring (when face (mapv #(nth face %) lm/LIPS-INNER))
box (when ring (interior/crop ring [w h]))
pixels (when box
(.-data (.getImageData ctx (:x box) (:y box)
(:w box) (:h box))))]
(swap! raw conj face)
(swap! crops conj (when box {:box box :data pixels})))
(when (or (zero? i) (zero? (mod (inc i) 4)) (= (inc i) total))
(rf/dispatch [::progress (str "detecting " (inc i) "/" total)]))
;; Let the status paint between synchronous
;; MediaPipe calls.
(js/setTimeout #(next-frame (inc i)) 0)))
(.catch reject))))]
(next-frame 0)))))))))
(defn- build-clip [manifest detector
{:keys [dense detected dimensions crops missing first-real] :as source-inputs}]
@ -116,7 +125,7 @@
(do (rf/dispatch [::progress "loading MediaPipe…"])
(-> (detect/landmarker!)
(.then (fn [model]
(rf/dispatch [::progress "loading frames…"])
(rf/dispatch [::progress "opening the video…"])
(-> (detect-frames! manifest model)
(.then (fn [fresh]
(build-clip manifest detector fresh))))))))))))))
@ -125,6 +134,10 @@
(rf/dispatch [::loaded id (:summary entry)]))))
(.catch (fn [error]
(js/console.error error)
;; A run that ended badly may have ended on a MediaPipe graph
;; error, and a landmarker that has hit one throws the same error
;; for the rest of the page's life. Retrying has to get a new one.
(detect/discard!)
(rf/dispatch [::failed (or (ex-message error) (.-message error) (str error))]))))))
(rf/reg-fx

View file

@ -60,8 +60,15 @@
`:source` is in it and `:name` is NOT. A clip's name is a label a human types;
two clips of the same footage under different names are the same analysis and
must share it, which is the whole return on addressing."
[{:keys [detector version source footage frames fps aspect seed]}]
must share it, which is the whole return on addressing.
`:mode` is how the detector was RUN, and it belongs here for the same reason the
version does. MediaPipe's video mode is a tracker and its image mode is not:
over the same frames and the same model they disagree by up to 0.013 of frame
width, which is a visible difference on a mouth. Optional, because the synthetic
take has no running mode to declare and an absent field is how the other
optional inputs already say \"not applicable\"."
[{:keys [detector version source footage frames fps aspect seed mode]}]
(when-not (and (string? detector) (seq detector) (string? version) (seq version))
(throw (ex-info "an analysis names its detector and the detector's VERSION: an upgrade that silently reuses old landmarks is the failure content addressing exists to prevent"
{:detector detector :version version})))
@ -73,7 +80,8 @@
:aspect aspect}
source (assoc :source source)
footage (assoc :footage footage)
seed (assoc :seed seed))))
seed (assoc :seed seed)
mode (assoc :mode mode))))
(defn analysis
"An analysis record with its `:id` filled in. The record is tier 1 — it says

View file

@ -4,6 +4,21 @@
(defonce ^:private instance (atom nil))
(defonce ^:private pending (atom nil))
(defn discard!
"Throw away the cached landmarker so the next run builds a fresh one.
Because a MediaPipe graph error is PERMANENT for the instance that hit it. A
timestamp that did not advance leaves the graph in an error state, and every
later `detectForVideo` on that landmarker re-throws it — so a memoised instance
turns one bad run into a tool that is broken until the tab is reloaded. Called
from the failure path, not from the happy one: a landmarker costs 26MB of wasm
and a model parse, and that is worth keeping for a run that ended cleanly."
[]
(when-let [model @instance]
(try (.call (aget model "close") model) (catch :default _ nil)))
(reset! instance nil)
(reset! pending nil))
(defn landmarker!
"Initialize once, using the vendored wasm and the local model. CPU also works
in browsers where a GPU delegate initializes but fails on its first frame.
@ -27,7 +42,7 @@
#js {:baseOptions
#js {:modelAssetPath "/static/mediapipe/face_landmarker.task"
:delegate "CPU"}
:runningMode "IMAGE"
:runningMode "VIDEO"
:numFaces 1})))
(.then (fn [model]
(reset! instance model)
@ -40,9 +55,27 @@
(js/Promise.reject (js/Error. "local MediaPipe script did not load"))))))
(defn detect!
"Detect one already-drawn canvas frame. nil means no face was detected."
[model canvas]
(when-let [face (aget (aget (.call (aget model "detect") model canvas)
"Detect one already-drawn canvas frame at its own time in the take. nil means no
face was detected.
`at-ms` IS THE FRAME'S REAL PRESENTATION TIME, and all three words are
load-bearing. In VIDEO mode the graph is a tracker: it runs face DETECTION only
when it has lost the face, and otherwise follows the previous frame's region,
using the gap between timestamps as the motion it has to account for. So:
IT MUST INCREASE, STRICTLY. MediaPipe's input streams reject a timestamp that
does not advance — `Packet timestamp mismatch on a calculator receiving from
stream \"norm_rect\"` — and that error is not recoverable: the graph is left in an
error state and every later call on this landmarker throws the same thing. One
repeated frame kills the run, so the caller walks frames forward exactly once.
IT MUST BE MILLISECONDS OF FOOTAGE, not a frame counter. Feeding `i` instead of
`i * 1000 / fps` still runs — and measured over the same 91 frames it moved
landmarks six times further from the per-frame answer (0.079 of frame width
against 0.013), because a tracker told that every frame is 1ms apart expects a
face that has barely moved."
[model canvas at-ms]
(when-let [face (aget (aget (.call (aget model "detectForVideo") model canvas at-ms)
"faceLandmarks") 0)]
(mapv (fn [p] {:x (.-x p) :y (.-y p) :z (.-z p)}) face)))

View file

@ -1,8 +1,14 @@
(ns arthur.flow.ingest
"Read footage from the server. Its response gives source timing and one
content-addressed URL per frame. The upload path derives this response from
stored frame records; no `manifest.json` file is part of that path."
(:require [arthur.fx.http :as http]))
"Read footage from the server: source timing, one video to measure, and one
content-addressed URL per tracing still.
THE MEASURED PIXELS COME OUT OF A VIDEO NOW, not out of a PNG per frame. The old
arrangement stored 112MB for a 7.6-second take and 1.1GB at the 900-frame limit;
the same footage is a 6MB H.264 proxy the page steps through. What that costs is
the property a PNG sequence gave for free — that asking for frame 12 gets frame
12 — so `frame!` below buys it back explicitly, and refuses to guess."
(:require [arthur.fx.http :as http]
[clojure.string :as str]))
(defn feature-presence
"Expand one-based, inclusive absence intervals from a manifest into boolean
@ -40,6 +46,13 @@
(sequential? urls) (every? string? urls))
(throw (ex-info "a footage manifest needs fps, frames (1–900), audio and a url per frame"
{:manifest (dissoc m :urls)})))
;; Named as its own failure rather than folded into the check above, because
;; it has a specific cause and a specific fix: this footage was extracted
;; before the proxy existed, and its frames were stored as PNGs that the
;; measurement path no longer reads.
(when-not (and (string? (:video m)) (seq (:video m)))
(throw (ex-info "this footage has no video to measure — re-extract it from its source"
{:footage (:id m)})))
(when-not (= frames (count urls))
;; The count is the manifest's and the URLs are the manifest's, so a
;; disagreement between them is the server contradicting itself — and it
@ -76,7 +89,13 @@
(defn audio-url [manifest]
(:audio manifest))
(defn frame-url [manifest i]
(defn video-url [manifest]
(:video manifest))
(defn frame-url
"The tracing still for one frame. A reference image for drawing over — the
landmarks and the mouth crops come from the video, not from these."
[manifest i]
(nth (:urls manifest) i))
(defn image! [src]
@ -87,3 +106,126 @@
(set! (.-onerror image) #(reject (ex-info (str "frame did not load: " src)
{:src src})))
(set! (.-src image) src)))))
;; ---------------------------------------------------------------------------
;; walking the proxy, one frame at a time
(def ^:private seek-timeout-ms
"How long one frame may take to arrive before the run gives up.
Long, because the first seek of a take also opens the file and fills a buffer,
and short enough that a video the browser cannot decode fails with a sentence
instead of hanging with a spinner."
10000)
(defn video!
"Load the proxy as a decodable, seekable element.
`preload=auto` and nothing else: the element is never added to the document and
never played. It is a decoder with a seek function, and the only reason it is a
DOM element rather than a `VideoDecoder` is that a `VideoDecoder` needs the
container demuxed before it can be handed a single frame, and this does not."
[src width height]
(js/Promise.
(fn [resolve reject]
(let [video (.createElement js/document "video")]
(set! (.-muted video) true)
(set! (.-playsInline video) true)
(set! (.-preload video) "auto")
(set! (.-crossOrigin video) "anonymous")
(set! (.-onerror video)
(fn [_]
(reject (ex-info (str "the browser could not decode this footage's video"
(when-let [e (.-error video)]
(str " (" (.-message e) ")")))
{:src src}))))
(set! (.-onloadeddata video)
(fn [_]
(cond
(not (fn? (.-requestVideoFrameCallback video)))
(reject (ex-info (str "this browser has no requestVideoFrameCallback, so "
"which frame is on screen cannot be established")
{}))
(not= [(.-videoWidth video) (.-videoHeight video)] [width height])
(reject (ex-info "the video's size disagrees with the footage manifest"
{:manifest [width height]
:video [(.-videoWidth video) (.-videoHeight video)]}))
:else (resolve video))))
(set! (.-src video) src)))))
(defn seek-time
"When to ask the video for source frame `i`: the MIDDLE of the frame, not its
start.
A frame occupies the half-open interval `[i/fps, (i+1)/fps)`, so `i/fps` sits
exactly on a boundary — and a boundary is where a seek lands on whichever side
the container's timebase rounds to. Measured over 91 frames: seeking to `i/fps`
produced the previous frame 31 times, and seeking to the middle was exact on all
91. Half a frame of slack in both directions is the entire fix."
[fps i]
(/ (+ i 0.5) fps))
(defn presented-frame
"Which source frame a `mediaTime` from requestVideoFrameCallback refers to.
`mediaTime` is the presented frame's own START — measured on 91 frames at 12fps
it came back as an exact multiple of the frame duration — so this is the inverse
of `i/fps` and NOT of `seek-time`. The two are asymmetric on purpose: we aim
half a frame late because a seek target is fuzzy, and read back exactly because
a presentation timestamp is not."
[fps media-time]
(js/Math.round (* media-time fps)))
(defn frame-ms
"When source frame `i` happens, in milliseconds of footage.
What `flow/detect` hands MediaPipe as the frame's timestamp. Real elapsed time
rather than the frame number, because the tracker reads the gap between
timestamps as motion — see `detect/detect!`."
[fps i]
(/ (* i 1000) fps))
(defn frame!
"Seek to source frame `i` and resolve once the browser has PRESENTED it.
IT WAITS FOR THE FRAME IT ASKED FOR, rather than trusting the first callback.
`requestVideoFrameCallback` is a queue of presentations, not an answer to our
seek: a frame presented while the file was still opening, or the tail of the
previous seek, arrives on the next callback we happen to have registered. Taking
it at face value is what produced `asked the video for frame 1 and it presented
frame 2` on a video whose seeks were in fact exact. So a callback whose
`mediaTime` is not this frame's is DISCARDED and the wait re-armed — the
browser states which frame it handed over, and that is the only frame we let
through.
It still fails loudly. A silent one-frame slip between the landmarks and the
audio is not something anyone finds by looking at the result, so a frame that
never arrives inside `seek-timeout-ms` ends the run and says which one."
[^js video fps i]
(js/Promise.
(fn [resolve reject]
(let [settled (volatile! false)
seen (volatile! [])
finish (fn [f] (when-not @settled (vreset! settled true) (f)))
timer (js/setTimeout
#(finish
(fn []
(reject (ex-info (str "the video never presented frame " (inc i)
(when (seq @seen)
(str "; it offered "
(str/join ", " (map inc @seen)))))
{:frame i :offered @seen}))))
seek-timeout-ms)]
(letfn [(listen []
(.requestVideoFrameCallback
video
(fn [_now metadata]
(when-not @settled
(let [presented (presented-frame fps (.-mediaTime metadata))]
(if (= presented i)
(do (js/clearTimeout timer) (finish #(resolve video)))
(do (vswap! seen conj presented) (listen))))))))]
(listen))
(set! (.-currentTime video) (seek-time fps i))))))

View file

@ -24,7 +24,10 @@
:footage (:footage manifest)
:frames (:frames manifest)
:fps (:fps manifest)
:aspect aspect}))))
:aspect aspect
;; How the detector was run, not just which one it was. See
;; `flow/address/analysis-descriptor`.
:mode "video"}))))
(defn measure
"Condition the anchor before measuring rings through it."

View file

@ -1,5 +1,5 @@
(ns arthur.flow.ingest-test
(:require [cljs.test :refer [deftest is]]
(:require [cljs.test :refer [deftest is testing]]
[arthur.flow.ingest :as ingest]))
(deftest feature-absence-uses-the-source-frame-numbers
@ -13,3 +13,52 @@
{:eye-r [[6 9]]}
{:eye-r [[5 3]]}]]
(is (thrown? ExceptionInfo (ingest/feature-presence 8 bad)))))
;; ---------------------------------------------------------------------------
;; walking the proxy
;;
;; These three are arithmetic, and all three were wrong in a way that no error
;; reported. A seek that lands one frame early, a timestamp read back with the
;; wrong inverse, or a frame time given as an index all produce a take that runs
;; to completion and is quietly off — which is why the numbers are asserted here
;; rather than left inline where they read as obvious.
(deftest a-seek-aims-at-the-middle-of-its-frame
;; THE BUG THIS ENCODES: seeking to `i/fps` sits exactly on the boundary between
;; two frames, and over 91 frames of real footage it produced the PREVIOUS frame
;; 31 times. Every seek must land strictly inside its own frame's interval, with
;; half a frame of slack on each side.
(doseq [fps [12 23.976 24 25 29.97 30 60]]
(testing (str fps "fps")
(doseq [i [0 1 2 41 90 899]]
(let [t (ingest/seek-time fps i)]
(is (< (/ i fps) t (/ (inc i) fps))
(str "frame " i " at " fps "fps seeks to " t
", outside [" (/ i fps) ", " (/ (inc i) fps) ")"))
;; Within a rounding error of dead centre. Not `=`: `(/ i fps)` and
;; `(/ (+ i 0.5) fps)` are two different divisions, so their difference
;; carries the last bit of both and 0.4999999999999982 is a pass.
(is (< (abs (- 0.5 (* fps (- t (/ i fps))))) 1e-9)
"the aim point drifted off the middle of the frame"))))))
(deftest a-presented-media-time-reads-back-as-its-own-frame
;; The inverse of `i/fps` and NOT of `seek-time`: requestVideoFrameCallback
;; reports the frame's START. Rounding the wrong one is a silent off-by-one
;; between the landmarks and the audio.
(doseq [fps [12 24 29.97 30]]
(doseq [i [0 1 2 41 90]]
(is (= i (ingest/presented-frame fps (/ i fps)))
(str "frame " i " at " fps "fps did not read back as itself")))))
(deftest a-frame-time-is-footage-milliseconds-and-not-a-frame-number
;; MediaPipe's video mode is a tracker that reads the gap between timestamps as
;; motion. Handing it `i` still runs, and measurably moved landmarks six times
;; further from the per-frame answer. It must be real elapsed time.
(is (= 0.0 (ingest/frame-ms 12 0)))
(is (= 1000.0 (ingest/frame-ms 12 12)))
(is (= (/ 1000 30) (ingest/frame-ms 30 1)))
(testing "strictly increasing, which the input stream requires"
(doseq [fps [12 23.976 30 60 240]]
(let [times (mapv #(ingest/frame-ms fps %) (range 200))]
(is (every? pos? (map - (rest times) times))
(str "two frames at " fps "fps shared a timestamp"))))))