diff --git a/README.md b/README.md index 8e19db1..3bb4523 100644 --- a/README.md +++ b/README.md @@ -16,9 +16,10 @@ modern conveniences belong in the workflow, not the output. See ## ClojureScript port -The active port plays the synthetic take, accepts video uploads, extracts their -frames and audio, analyzes real footage for mouth, eyes, brows and pixel-derived -teeth, and saves the project with reusable analysis data. The step 8 data model +The active port plays the synthetic take, accepts video uploads, transcodes them +to a browser-seekable proxy plus audio and tracing stills, analyzes real footage +for mouth, eyes, brows and pixel-derived teeth, and saves the project with +reusable analysis data. The step 8 data model represents persistent feature IDs, eye pairs and feature-level observation gaps; its controls are still pending. See the [port plan](docs/port-plan.md). @@ -63,7 +64,7 @@ Three tiers, cut by mutability and size — the full argument is in | --- | --- | --- | | 1 **authored** | the scene: nodes, channels, features, time maps | the database, as independently addressed leaves. Kilobytes | | 2 **derived** | detected landmarks, raw mouth crops, and dense channel blocks | `var/blobs`, addressed by analysis and block inputs, including the detector version | -| 3 **source** | uploaded video, extracted frames, and audio | the same blob store, by the hash of their bytes | +| 3 **source** | the uploaded video, the H.264 proxy measured from it, its tracing stills, and audio | the same blob store, by the hash of their bytes | Only tier 1 is the document. Tier 2 is a pure function of tiers 1 and 3, so a saved project names its blocks rather than carrying them, and a knob change gives @@ -88,16 +89,27 @@ wasm, which is fetched from a CDN on first use. For real footage: -```sh -./extract.sh /path/to/clip.mov # -> frames/*.png, audio.wav, manifest.json -``` +Upload it in the app. `./extract.sh` still writes the old PNG-sequence bundle and +`ingest_bundle` still registers it, but footage ingested that way has no proxy and +the loader will say so — the measured pixels come out of the video now. -then **Load frames**. MediaPipe's wasm is fetched from jsdelivr on first use; -`face_landmarker.task` is local. +MediaPipe's wasm and `face_landmarker.task` are both local; nothing in detection +touches the network. -Frames are pre-extracted rather than decoded in the page because browser video -seeking is approximate and `requestVideoFrameCallback` only delivers frames at -playback speed — neither gives a deterministic per-frame pass. +Detection reads the VIDEO, not a frame per file. The page seeks the proxy to the +MIDDLE of each frame — `(i + 0.5) / fps` — and waits for +`requestVideoFrameCallback` to hand the frame over, then checks the `mediaTime` it +reports against the frame it asked for. Both halves are load-bearing and both were +measured against the same footage decoded to PNGs: aiming at `i / fps` sits on a +frame boundary and landed one frame early 31 times in 91, and aiming at the middle +was exact on all 91. A run that gets a frame it did not ask for stops and says so, +because a one-frame slip between the landmarks and the audio is not something +anyone finds by looking at the result. + +This is what replaced the PNG sequence, which was 112MB for 7.6 seconds and would +be 1.1GB at the 900-frame limit. The proxy is 6MB, and the landmarks barely +notice: detected off decoded H.264 rather than off the PNGs, they moved at most +0.0033 of frame width. `manifest.json` records the source rate. The extractor keeps every source frame; the page reads that rate because a guessed fps desynchronises audio from picture. diff --git a/clips/extraction.py b/clips/extraction.py index 8019103..25ea808 100644 --- a/clips/extraction.py +++ b/clips/extraction.py @@ -1,4 +1,27 @@ -"""Upload a video once, then decode it into the existing footage model.""" +"""Upload a video once, then turn it into the two things the app actually reads. + +WHAT CHANGED AND WHY. This used to decode one PNG per source frame and store every +one of them. A 7.6-second 1440x1920 take is 112MB that way, and the 900-frame limit +is 1.1GB — for pixels whose only consumer was a canvas that MediaPipe then read +once. The page now detects from the video itself (see `frontend/src/arthur/flow/ +ingest.cljs`), so this produces: + + THE PROXY. One browser-safe H.264/yuv420p MP4, CFR, `+faststart`. The same take + is 6MB. This is the analysis source, and it is re-encoded RATHER THAN KEPT AS + UPLOADED even when the upload is already H.264, for two reasons that are both + about not guessing: an iPhone's HEVC is not decodable in every browser, and the + footage's identity is the digest of this file — one produced by one ffmpeg + invocation, not one that depends on which branch the source happened to take. + + THE TRACING STILLS. One JPEG per frame, long edge capped, for the tracing editor + to draw over. Reference images; nothing measures them. They are not in the + footage digest — see `models.Footage`. + +The proxy is probed after it is written rather than before. `width`, `height` and +`frames` are properties of the file the browser will decode, and taking them from +the source instead is how a scaler or a dropped frame becomes a silent one-frame +offset between the landmarks and the audio. +""" import hashlib import json @@ -18,6 +41,14 @@ _active = set() _lock = threading.Lock() TIMEOUT = 3600 +# Visually lossless enough that landmarks do not move: measured against the same +# frames as PNGs, IMAGE-mode landmarks shifted at most 0.0033 of frame width. +PROXY_CRF = "18" +# The long edge of a tracing still. The proxy keeps full resolution because the +# detector reads it; a still only has to be good enough to draw a cel over. +TRACING_EDGE = 1280 +TRACING_QUALITY = "4" + def _command(args): result = subprocess.run(args, capture_output=True, text=True, timeout=TIMEOUT) @@ -26,31 +57,34 @@ def _command(args): return result.stdout -def _decode_frames(job, source_path, frames_dir, facts, root): - """Decode one frame per source frame and publish ffmpeg's live frame count.""" - progress_path = root / "frames.progress" - log_path = root / "frames.log" - total = facts.get("reported_frames") or round(facts["duration"] * facts["fps"]) +def _run_with_progress(job, args, root, name, total, span): + """Run one ffmpeg and publish its live frame count as `span` of the job. + + ffmpeg's `-progress` file is the only honest source for this: parsing its + stderr means parsing a format that is explicitly not an interface, and a + spinner that is not attached to frames is a spinner that lies on a long take. + """ + progress_path = root / f"{name}.progress" + log_path = root / f"{name}.log" + first, last = span args = ["ffmpeg", "-hide_banner", "-loglevel", "error", "-y", - "-stats_period", "0.25", "-progress", str(progress_path), - "-i", str(source_path), "-fps_mode", "passthrough", - str(frames_dir / "%04d.png")] + "-stats_period", "0.25", "-progress", str(progress_path)] + args with open(log_path, "wb") as log: proc = subprocess.Popen(args, stdout=log, stderr=subprocess.STDOUT) deadline = time.monotonic() + TIMEOUT try: while proc.poll() is None: if time.monotonic() >= deadline: - raise TimeoutError("video frame extraction timed out") + raise TimeoutError(f"{name} timed out") if progress_path.exists(): lines = progress_path.read_text(errors="replace").splitlines() count = next((int(line[6:].strip()) for line in reversed(lines) if line.startswith("frame=") and line[6:].strip().isdigit()), 0) if count and total: - progress = min(59, int(60 * count / total)) - if progress > job.progress: - job.progress = progress + reached = first + int((last - first) * min(1.0, count / total)) + if reached > job.progress: + job.progress = reached job.save(update_fields=["progress", "updated"]) time.sleep(0.2) finally: @@ -58,8 +92,35 @@ def _decode_frames(job, source_path, frames_dir, facts, root): proc.kill() proc.wait() if proc.returncode: - raise ValueError(log_path.read_text(errors="replace")[-1200:] or - "video frame extraction failed") + raise ValueError(log_path.read_text(errors="replace")[-1200:] or f"{name} failed") + + +def _encode_proxy(job, source_path, proxy_path, facts, root): + """The uploaded video -> one H.264 file every browser can decode and seek.""" + total = facts.get("reported_frames") or round(facts["duration"] * facts["fps"]) + _run_with_progress( + job, + ["-i", str(source_path), "-an", + # Constant frame rate at the source's own rate. `probe` has already + # refused VFR, so this asserts that rather than resampling. + "-fps_mode", "cfr", "-r", str(facts["fps"]), + "-c:v", "libx264", "-preset", "veryfast", "-crf", PROXY_CRF, + # yuv420p and an even frame size are what makes this playable everywhere + # rather than only in the browser that happened to be tested. + "-pix_fmt", "yuv420p", "-vf", "scale=trunc(iw/2)*2:trunc(ih/2)*2", + "-movflags", "+faststart", str(proxy_path)], + root, "proxy", total, (0, 55)) + + +def _extract_stills(job, proxy_path, frames_dir, frames, root): + """The proxy -> one tracing JPEG per frame, long edge capped.""" + _run_with_progress( + job, + ["-i", str(proxy_path), "-fps_mode", "passthrough", + "-vf", f"scale='if(gt(iw,ih),min({TRACING_EDGE},iw),-2)':" + f"'if(gt(iw,ih),-2,min({TRACING_EDGE},ih))'", + "-q:v", TRACING_QUALITY, str(frames_dir / "%04d.jpg")], + root, "stills", frames, (55, 85)) def probe(path): @@ -88,40 +149,58 @@ def probe(path): "vfr": False} +def count_frames(path): + """How many frames a file really holds, counted rather than reported. + + `nb_frames` is a container's claim. This is the decoder's answer, and it is + what the page will get when it walks the proxy — so a disagreement between the + two has to be settled before the count reaches a manifest, not after it has + become a one-frame audio offset nobody can find. + """ + text = _command(["ffprobe", "-v", "error", "-select_streams", "v:0", + "-count_frames", "-show_entries", "stream=nb_read_frames", + "-of", "default=nokey=1:noprint_wrappers=1", str(path)]) + counted = text.strip() + if not counted.isdigit(): + raise ValueError("could not count the proxy's frames") + return int(counted) + + def extraction_key(source, settings): - text = json.dumps({"scheme": 1, "source": source.blob_id, "settings": settings}, + text = json.dumps({"scheme": 2, "source": source.blob_id, "settings": settings}, sort_keys=True, separators=(",", ":")) return "sha256:" + hashlib.sha256(text.encode()).hexdigest() -def _register(job, frames, audio_path, facts): - width, height = blobs.png_size(frames[0]) - frame_blobs = [] - for index, path in enumerate(frames): - if blobs.png_size(path) != (width, height): - raise ValueError(f"decoded frame {index + 1} has different dimensions") - digest, size = blobs.adopt(path) - frame_blobs.append((index, digest, size)) +def _register(job, proxy_path, stills, audio_path, facts): + proxy_digest, proxy_size = blobs.adopt(proxy_path) audio_digest, audio_size = blobs.adopt(audio_path) + still_blobs = [(index, *blobs.adopt(path)) for index, path in enumerate(stills)] + width, height, fps, frames = facts["width"], facts["height"], facts["fps"], facts["frames"] + + # The footage's own identity: the bytes the page will measure, the audio it + # will clock against, and the rate that ties them together. Scheme 2 — scheme + # 1 hashed a PNG per frame, and those footages name pixels this no longer has. h = hashlib.sha256() - h.update(f"arthur-footage-1/{facts['fps']}/{len(frames)}/{width}x{height}\n".encode()) - for _, digest, _ in frame_blobs: - h.update(digest.encode()) + h.update(f"arthur-footage-2/{fps}/{frames}/{width}x{height}\n".encode()) + h.update(proxy_digest.encode()) h.update(audio_digest.encode()) + with transaction.atomic(): + proxy_blob, _ = Blob.objects.get_or_create( + digest=proxy_digest, defaults={"size": proxy_size, "media_type": "video/mp4"}) audio_blob, _ = Blob.objects.get_or_create( digest=audio_digest, defaults={"size": audio_size, "media_type": "audio/wav"}) footage, created = Footage.objects.get_or_create( digest=h.hexdigest(), defaults={"label": job.source.filename[:200], "source": job.source.filename[:200], - "fps": facts["fps"], - "frames": len(frames), "width": width, - "height": height, "audio": audio_blob}) + "fps": fps, "frames": frames, "width": width, "height": height, + "audio": audio_blob, "video": proxy_blob}) if created: rows = [] - for index, digest, size in frame_blobs: + for index, digest, size in still_blobs: blob, _ = Blob.objects.get_or_create( - digest=digest, defaults={"size": size, "media_type": "image/png"}) + digest=digest, defaults={"size": size, "media_type": "image/jpeg"}) rows.append(FootageFrame(footage=footage, index=index, blob=blob)) FootageFrame.objects.bulk_create(rows) return footage @@ -137,14 +216,29 @@ def run(key): source_path = blobs.path_for(job.source.blob_id) with tempfile.TemporaryDirectory(prefix="arthur-extract-") as directory: root = Path(directory) - frames_dir = root / "frames" - frames_dir.mkdir() - _decode_frames(job, source_path, frames_dir, facts, root) - frames = sorted(frames_dir.glob("*.png")) + proxy_path = root / "proxy.mp4" + _encode_proxy(job, source_path, proxy_path, facts, root) + + # Everything downstream describes the PROXY, not the upload. + proxy_facts = probe(proxy_path) + frames = count_frames(proxy_path) + if not 1 <= frames <= 900: + raise ValueError(f"the proxy holds {frames} frames; the limit is 1–900") expected = facts.get("reported_frames") - if not frames or len(frames) > 900 or (expected and len(frames) != expected): - raise ValueError(f"decoded {len(frames)} frames; expected {expected or '1–900'}") - job.progress = 60 + if expected and frames != expected: + raise ValueError( + f"the proxy holds {frames} frames and the upload reports {expected}; " + "refusing footage whose picture and audio would drift") + proxy_facts["frames"] = frames + + frames_dir = root / "stills" + frames_dir.mkdir() + _extract_stills(job, proxy_path, frames_dir, frames, root) + stills = sorted(frames_dir.glob("*.jpg")) + if len(stills) != frames: + raise ValueError(f"wrote {len(stills)} tracing stills for {frames} frames") + + job.progress = 85 job.save(update_fields=["progress", "updated"]) audio_path = root / "audio.wav" if facts["has_audio"]: @@ -154,9 +248,9 @@ def run(key): else: _command(["ffmpeg", "-hide_banner", "-loglevel", "error", "-y", "-f", "lavfi", "-i", "anullsrc=r=44100:cl=mono", - "-t", str(len(frames) / facts["fps"]), "-c:a", "pcm_s16le", + "-t", str(frames / proxy_facts["fps"]), "-c:a", "pcm_s16le", str(audio_path)]) - footage = _register(job, frames, audio_path, facts) + footage = _register(job, proxy_path, stills, audio_path, proxy_facts) job.footage, job.state, job.progress = footage, "done", 100 job.save(update_fields=["footage", "state", "progress", "updated"]) except Exception as exc: diff --git a/clips/migrations/0004_footage_video_alter_footageframe_index.py b/clips/migrations/0004_footage_video_alter_footageframe_index.py new file mode 100644 index 0000000..6e53732 --- /dev/null +++ b/clips/migrations/0004_footage_video_alter_footageframe_index.py @@ -0,0 +1,24 @@ +# Generated by Django 5.2.17 on 2026-09-28 15:10 + +import django.db.models.deletion +from django.db import migrations, models + + +class Migration(migrations.Migration): + + dependencies = [ + ('clips', '0003_source_extraction'), + ] + + operations = [ + migrations.AddField( + model_name='footage', + name='video', + field=models.ForeignKey(blank=True, help_text='the browser-safe proxy the page detects from; null on pre-proxy footage', null=True, on_delete=django.db.models.deletion.PROTECT, related_name='video_for', to='clips.blob'), + ), + migrations.AlterField( + model_name='footageframe', + name='index', + field=models.PositiveIntegerField(help_text="0-based; source frame index + 1 is the JPEG's name"), + ), + ] diff --git a/clips/models.py b/clips/models.py index daacd4c..031d5ae 100644 --- a/clips/models.py +++ b/clips/models.py @@ -70,9 +70,21 @@ class Extraction(models.Model): class Footage(models.Model): """Tier 3: the frames and audio of one extraction, immutable. - `digest` is over the ordered frame digests plus the audio's, so two - extractions of the same clip at the same rate are one footage and the same - analysis can be reused across both. + `digest` is over the PROXY VIDEO's digest plus the audio's and the rate, so + two extractions of the same clip at the same settings are one footage and the + same analysis can be reused across both. + + THE PROXY IS THE ANALYSIS SOURCE AND THE FRAMES ARE NOT. `video` is one + browser-safe H.264 file, and it is what the page seeks through to detect + landmarks. `frame_set` is a JPEG per frame at tracing size: reference stills + for the tracing editor, never the thing measured. The two are not + interchangeable, and which one carries the pixels an analysis was computed + from is the difference between a 6MB take and a 1.1GB one. + + So the frame JPEGs are deliberately NOT in `digest`. They are a rendering of + this footage for a human to trace over; re-rendering them at another size does + not make it different footage, and putting them in the identity would throw + away every analysis when the tracing size changed. `feature_absence` is the manifest annotation step 8 introduced: known occlusion intervals, one-based and inclusive, expanded into presence tracks by @@ -88,6 +100,10 @@ class Footage(models.Model): width = models.PositiveIntegerField() height = models.PositiveIntegerField() audio = models.ForeignKey(Blob, on_delete=models.PROTECT, related_name="audio_for") + video = models.ForeignKey( + Blob, null=True, blank=True, on_delete=models.PROTECT, related_name="video_for", + help_text="the browser-safe proxy the page detects from; null on pre-proxy footage", + ) feature_absence = models.JSONField(default=dict, blank=True) created = models.DateTimeField(auto_now_add=True) @@ -99,12 +115,14 @@ class Footage(models.Model): class FootageFrame(models.Model): - """One source frame. A row rather than an entry in a JSON list, because a + """One tracing still. A row rather than an entry in a JSON list, because a frame is a thing the server serves, and because a blob's references have to be - countable before anything can be collected.""" + countable before anything can be collected. + + A REFERENCE IMAGE, NOT A MEASUREMENT INPUT. See `Footage.video`.""" footage = models.ForeignKey(Footage, on_delete=models.CASCADE, related_name="frame_set") - index = models.PositiveIntegerField(help_text="0-based; source frame index + 1 is the PNG's name") + index = models.PositiveIntegerField(help_text="0-based; source frame index + 1 is the JPEG's name") blob = models.ForeignKey(Blob, on_delete=models.PROTECT, related_name="frame_for") class Meta: diff --git a/clips/tests/test_api.py b/clips/tests/test_api.py index 9954a4d..55a7b6b 100644 --- a/clips/tests/test_api.py +++ b/clips/tests/test_api.py @@ -105,6 +105,46 @@ class BlobStoreTests(TestCase): self.assertEqual(404, self.client.get("/blob/" + "0" * 64).status_code) self.assertEqual(404, self.client.get("/blob/nonsense").status_code) + def test_a_blob_serves_byte_ranges(self): + # NOT AN OPTIMISATION. A