Step 9. The tier split was the work; Django was the easy half.
Tier 1 — the authored scene — is the document, and it is addressed as
independently versioned leaves rather than saved whole, so one vertex drag
cannot clobber a collaborator's keying. `domain/leaf` is the document as
path -> value; `domain/wire` puts it on the wire as transit, because JSON
has neither integer map keys nor keywords and a save would quietly turn
`{0 v}` into `{"0" v}`.
Tier 2 — the dense channel blocks — is content-addressed by a hash over
every input, with the detector version inside every key through the
analysis the block descriptor names. `flow/address`'s `block-knobs` is the
invalidation table, and `address-test` does not trust it: it re-freezes the
take once per knob and asserts the biconditional, that a block's bytes
changed if and only if its key changed. That found `brow-pos` not depending
on `contour-avg` — the brow ring is smoothed, the raise is not.
Tier 3 — frames and audio — is served by the hash of its bytes out of the
same store. A manifest now names frames and carries a URL for each, so the
frame layout stopped being a shared secret between a shell script and a
ClojureScript namespace, and the `?v=` cache-buster went with it: a blob's
name is the hash of its contents, so a stale copy is not a thing that can
happen. The synthetic take's `audio.wav` moved to `static/arthur/` — an
asset the project owns, not an extraction that churns.
The server verifies rather than trusting a name it was handed: it
recomputes every key from the descriptor stored beside it, refuses an
analysis that declares no detector version, and refuses a document naming
blocks it does not hold. It hashes the descriptor TEXT, because JS prints
an integral double as `1` and Python as `1.0`, and a scheme where both ends
re-render the numbers disagrees on the first parameter that happens to be
whole.
Two loose ends from step 8 closed on the way. `pack` no longer takes a
`(track, frame)` predicate whose call sites each re-derived a feature from
an index — every track names the feature it follows, which deleted five
hand-maintained mappings. And `:dev-http` is gone: Django serves the page,
shadow-cljs only builds into the staticfiles tree.
227 CLJS tests, 31 Django tests, green.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
120 lines
4.9 KiB
Python
120 lines
4.9 KiB
Python
"""Register an extracted bundle as tier 3.
|
|
|
|
python manage.py ingest_bundle # ./manifest.json
|
|
python manage.py ingest_bundle scratch/my-take # that bundle
|
|
|
|
WHAT THIS REPLACES. Until step 9 the page fetched `/manifest.json` and then built
|
|
`frames/0001.png` itself, with shadow-cljs's `:dev-http` serving the repo root. So
|
|
the frame layout was a shared secret between a shell script and a ClojureScript
|
|
namespace, and "where the frames are" was answered by a directory listing.
|
|
|
|
Now the server names every frame, and the client asks it. The frames go into the
|
|
content-addressed blob store — by hard link, so 112MB of PNGs is not copied — and
|
|
the manifest the client receives carries a URL per frame. That is the whole of what
|
|
makes the frames the backend's to serve, and it is what the in-browser wasm-ffmpeg
|
|
extraction docs/architecture.md describes will upload INTO, without the client
|
|
learning anything new when it arrives: the same blobs, the same manifest, a
|
|
different producer.
|
|
|
|
`extract.sh` still does the decoding. It is out of step 9's scope, it works, and it
|
|
is the only part of this that needs a terminal.
|
|
"""
|
|
import json
|
|
from pathlib import Path
|
|
|
|
from django.core.management.base import BaseCommand, CommandError
|
|
from django.db import transaction
|
|
|
|
from clips import blobs
|
|
from clips.models import Blob, Footage, FootageFrame
|
|
|
|
|
|
class Command(BaseCommand):
|
|
help = "Register an extracted frames+audio+manifest bundle as footage."
|
|
|
|
def add_arguments(self, parser):
|
|
parser.add_argument(
|
|
"bundle", nargs="?", default=".",
|
|
help="a directory holding manifest.json, or the manifest itself",
|
|
)
|
|
parser.add_argument("--label", default="", help="what to call it in the UI")
|
|
|
|
def handle(self, *args, **options):
|
|
manifest_path = Path(options["bundle"])
|
|
if manifest_path.is_dir():
|
|
manifest_path = manifest_path / "manifest.json"
|
|
if not manifest_path.exists():
|
|
raise CommandError(f"{manifest_path} does not exist — run ./extract.sh first")
|
|
|
|
manifest = json.loads(manifest_path.read_text())
|
|
root = manifest_path.parent
|
|
frames_dir = root / manifest["dir"]
|
|
audio_path = root / manifest["audio"]
|
|
count = int(manifest["frames"])
|
|
|
|
pngs = sorted(frames_dir.glob("*.png"))
|
|
if len(pngs) != count:
|
|
raise CommandError(
|
|
f"the manifest says {count} frames and {frames_dir} holds {len(pngs)}; "
|
|
"refusing an inaccurate footage"
|
|
)
|
|
if not audio_path.exists():
|
|
raise CommandError(f"{audio_path} does not exist")
|
|
|
|
width, height = blobs.png_size(pngs[0])
|
|
|
|
self.stdout.write(f"hashing {len(pngs)} frames…")
|
|
frame_blobs = []
|
|
for i, png in enumerate(pngs):
|
|
digest, size = blobs.adopt(png)
|
|
frame_blobs.append((i, digest, size))
|
|
if (i + 1) % 25 == 0 or i + 1 == len(pngs):
|
|
self.stdout.write(f" {i + 1}/{len(pngs)}")
|
|
|
|
audio_digest, audio_size = blobs.adopt(audio_path)
|
|
|
|
# The footage's own identity: every frame in order, plus the audio and the
|
|
# rate. Two extractions of one clip at one rate are one footage, so an
|
|
# analysis over it is reusable across both.
|
|
import hashlib
|
|
|
|
h = hashlib.sha256()
|
|
h.update(f"arthur-footage-1/{manifest['fps']}/{count}/{width}x{height}\n".encode())
|
|
for _, digest, _ in frame_blobs:
|
|
h.update(digest.encode())
|
|
h.update(audio_digest.encode())
|
|
footage_digest = h.hexdigest()
|
|
|
|
if existing := Footage.objects.filter(digest=footage_digest).first():
|
|
self.stdout.write(self.style.SUCCESS(f"already ingested: {existing.id}"))
|
|
return
|
|
|
|
with transaction.atomic():
|
|
audio_blob, _ = Blob.objects.get_or_create(
|
|
digest=audio_digest,
|
|
defaults={"size": audio_size, "media_type": "audio/wav"},
|
|
)
|
|
footage = Footage.objects.create(
|
|
digest=footage_digest,
|
|
label=options["label"] or manifest.get("source") or frames_dir.name,
|
|
source=manifest.get("source") or "",
|
|
fps=float(manifest["fps"]),
|
|
frames=count,
|
|
width=width,
|
|
height=height,
|
|
audio=audio_blob,
|
|
feature_absence=manifest.get("feature-absence") or {},
|
|
)
|
|
rows = []
|
|
for index, digest, size in frame_blobs:
|
|
blob, _ = Blob.objects.get_or_create(
|
|
digest=digest, defaults={"size": size, "media_type": "image/png"}
|
|
)
|
|
rows.append(FootageFrame(footage=footage, index=index, blob=blob))
|
|
FootageFrame.objects.bulk_create(rows)
|
|
|
|
self.stdout.write(
|
|
self.style.SUCCESS(
|
|
f"{count} frames at {manifest['fps']}fps, {width}x{height} -> footage {footage.id}"
|
|
)
|
|
)
|