Serve the document from a Django backend, split into three tiers

Step 9. The tier split was the work; Django was the easy half.

Tier 1 — the authored scene — is the document, and it is addressed as
independently versioned leaves rather than saved whole, so one vertex drag
cannot clobber a collaborator's keying. `domain/leaf` is the document as
path -> value; `domain/wire` puts it on the wire as transit, because JSON
has neither integer map keys nor keywords and a save would quietly turn
`{0 v}` into `{"0" v}`.

Tier 2 — the dense channel blocks — is content-addressed by a hash over
every input, with the detector version inside every key through the
analysis the block descriptor names. `flow/address`'s `block-knobs` is the
invalidation table, and `address-test` does not trust it: it re-freezes the
take once per knob and asserts the biconditional, that a block's bytes
changed if and only if its key changed. That found `brow-pos` not depending
on `contour-avg` — the brow ring is smoothed, the raise is not.

Tier 3 — frames and audio — is served by the hash of its bytes out of the
same store. A manifest now names frames and carries a URL for each, so the
frame layout stopped being a shared secret between a shell script and a
ClojureScript namespace, and the `?v=` cache-buster went with it: a blob's
name is the hash of its contents, so a stale copy is not a thing that can
happen. The synthetic take's `audio.wav` moved to `static/arthur/` — an
asset the project owns, not an extraction that churns.

The server verifies rather than trusting a name it was handed: it
recomputes every key from the descriptor stored beside it, refuses an
analysis that declares no detector version, and refuses a document naming
blocks it does not hold. It hashes the descriptor TEXT, because JS prints
an integral double as `1` and Python as `1.0`, and a scheme where both ends
re-render the numbers disagrees on the first parameter that happens to be
whole.

Two loose ends from step 8 closed on the way. `pack` no longer takes a
`(track, frame)` predicate whose call sites each re-derived a feature from
an index — every track names the feature it follows, which deleted five
hand-maintained mappings. And `:dev-http` is gone: Django serves the page,
shadow-cljs only builds into the staticfiles tree.

227 CLJS tests, 31 Django tests, green.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
Olive Vaughn 2026-09-28 01:11:41 -04:00
parent b6517f837a
commit 9cd5243983
61 changed files with 4694 additions and 269 deletions

View file

View file

View file

@ -0,0 +1,120 @@
"""Register an extracted bundle as tier 3.
python manage.py ingest_bundle # ./manifest.json
python manage.py ingest_bundle scratch/my-take # that bundle
WHAT THIS REPLACES. Until step 9 the page fetched `/manifest.json` and then built
`frames/0001.png` itself, with shadow-cljs's `:dev-http` serving the repo root. So
the frame layout was a shared secret between a shell script and a ClojureScript
namespace, and "where the frames are" was answered by a directory listing.
Now the server names every frame, and the client asks it. The frames go into the
content-addressed blob store — by hard link, so 112MB of PNGs is not copied — and
the manifest the client receives carries a URL per frame. That is the whole of what
makes the frames the backend's to serve, and it is what the in-browser wasm-ffmpeg
extraction docs/architecture.md describes will upload INTO, without the client
learning anything new when it arrives: the same blobs, the same manifest, a
different producer.
`extract.sh` still does the decoding. It is out of step 9's scope, it works, and it
is the only part of this that needs a terminal.
"""
import json
from pathlib import Path
from django.core.management.base import BaseCommand, CommandError
from django.db import transaction
from clips import blobs
from clips.models import Blob, Footage, FootageFrame
class Command(BaseCommand):
help = "Register an extracted frames+audio+manifest bundle as footage."
def add_arguments(self, parser):
parser.add_argument(
"bundle", nargs="?", default=".",
help="a directory holding manifest.json, or the manifest itself",
)
parser.add_argument("--label", default="", help="what to call it in the UI")
def handle(self, *args, **options):
manifest_path = Path(options["bundle"])
if manifest_path.is_dir():
manifest_path = manifest_path / "manifest.json"
if not manifest_path.exists():
raise CommandError(f"{manifest_path} does not exist — run ./extract.sh first")
manifest = json.loads(manifest_path.read_text())
root = manifest_path.parent
frames_dir = root / manifest["dir"]
audio_path = root / manifest["audio"]
count = int(manifest["frames"])
pngs = sorted(frames_dir.glob("*.png"))
if len(pngs) != count:
raise CommandError(
f"the manifest says {count} frames and {frames_dir} holds {len(pngs)}; "
"refusing an inaccurate footage"
)
if not audio_path.exists():
raise CommandError(f"{audio_path} does not exist")
width, height = blobs.png_size(pngs[0])
self.stdout.write(f"hashing {len(pngs)} frames…")
frame_blobs = []
for i, png in enumerate(pngs):
digest, size = blobs.adopt(png)
frame_blobs.append((i, digest, size))
if (i + 1) % 25 == 0 or i + 1 == len(pngs):
self.stdout.write(f" {i + 1}/{len(pngs)}")
audio_digest, audio_size = blobs.adopt(audio_path)
# The footage's own identity: every frame in order, plus the audio and the
# rate. Two extractions of one clip at one rate are one footage, so an
# analysis over it is reusable across both.
import hashlib
h = hashlib.sha256()
h.update(f"arthur-footage-1/{manifest['fps']}/{count}/{width}x{height}\n".encode())
for _, digest, _ in frame_blobs:
h.update(digest.encode())
h.update(audio_digest.encode())
footage_digest = h.hexdigest()
if existing := Footage.objects.filter(digest=footage_digest).first():
self.stdout.write(self.style.SUCCESS(f"already ingested: {existing.id}"))
return
with transaction.atomic():
audio_blob, _ = Blob.objects.get_or_create(
digest=audio_digest,
defaults={"size": audio_size, "media_type": "audio/wav"},
)
footage = Footage.objects.create(
digest=footage_digest,
label=options["label"] or manifest.get("source") or frames_dir.name,
source=manifest.get("source") or "",
fps=float(manifest["fps"]),
frames=count,
width=width,
height=height,
audio=audio_blob,
feature_absence=manifest.get("feature-absence") or {},
)
rows = []
for index, digest, size in frame_blobs:
blob, _ = Blob.objects.get_or_create(
digest=digest, defaults={"size": size, "media_type": "image/png"}
)
rows.append(FootageFrame(footage=footage, index=index, blob=blob))
FootageFrame.objects.bulk_create(rows)
self.stdout.write(
self.style.SUCCESS(
f"{count} frames at {manifest['fps']}fps, {width}x{height} -> footage {footage.id}"
)
)