Serve the document from a Django backend, split into three tiers

Step 9. The tier split was the work; Django was the easy half.

Tier 1 — the authored scene — is the document, and it is addressed as
independently versioned leaves rather than saved whole, so one vertex drag
cannot clobber a collaborator's keying. `domain/leaf` is the document as
path -> value; `domain/wire` puts it on the wire as transit, because JSON
has neither integer map keys nor keywords and a save would quietly turn
`{0 v}` into `{"0" v}`.

Tier 2 — the dense channel blocks — is content-addressed by a hash over
every input, with the detector version inside every key through the
analysis the block descriptor names. `flow/address`'s `block-knobs` is the
invalidation table, and `address-test` does not trust it: it re-freezes the
take once per knob and asserts the biconditional, that a block's bytes
changed if and only if its key changed. That found `brow-pos` not depending
on `contour-avg` — the brow ring is smoothed, the raise is not.

Tier 3 — frames and audio — is served by the hash of its bytes out of the
same store. A manifest now names frames and carries a URL for each, so the
frame layout stopped being a shared secret between a shell script and a
ClojureScript namespace, and the `?v=` cache-buster went with it: a blob's
name is the hash of its contents, so a stale copy is not a thing that can
happen. The synthetic take's `audio.wav` moved to `static/arthur/` — an
asset the project owns, not an extraction that churns.

The server verifies rather than trusting a name it was handed: it
recomputes every key from the descriptor stored beside it, refuses an
analysis that declares no detector version, and refuses a document naming
blocks it does not hold. It hashes the descriptor TEXT, because JS prints
an integral double as `1` and Python as `1.0`, and a scheme where both ends
re-render the numbers disagrees on the first parameter that happens to be
whole.

Two loose ends from step 8 closed on the way. `pack` no longer takes a
`(track, frame)` predicate whose call sites each re-derived a feature from
an index — every track names the feature it follows, which deleted five
hand-maintained mappings. And `:dev-http` is gone: Django serves the page,
shadow-cljs only builds into the staticfiles tree.

227 CLJS tests, 31 Django tests, green.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
Olive Vaughn 2026-09-28 01:11:41 -04:00
parent b6517f837a
commit 9cd5243983
61 changed files with 4694 additions and 269 deletions

0
clips/__init__.py Normal file
View file

63
clips/admin.py Normal file
View file

@ -0,0 +1,63 @@
"""The admin, which is here for one reason: tier 1 is readable.
docs/architecture.md's argument against a CRDT is partly this — "the canonical
document moves into an opaque blob, and every server-side thing that reads the
document needs it materialised back out". A leaf is transit-as-JSON in a
JSONField, so it is legible here, and that is a property worth being able to see.
"""
from django.contrib import admin
from .models import Analysis, Block, Blob, Clip, Footage, FootageFrame, Leaf, Project, Revision
@admin.register(Project)
class ProjectAdmin(admin.ModelAdmin):
list_display = ("name", "id", "seq", "updated")
search_fields = ("name", "id")
@admin.register(Clip)
class ClipAdmin(admin.ModelAdmin):
list_display = ("cid", "project", "name", "footage", "analysis")
list_filter = ("project",)
@admin.register(Leaf)
class LeafAdmin(admin.ModelAdmin):
list_display = ("path", "project", "version", "updated")
list_filter = ("project",)
search_fields = ("path",)
@admin.register(Revision)
class RevisionAdmin(admin.ModelAdmin):
list_display = ("project", "seq", "summary", "author", "created")
@admin.register(Footage)
class FootageAdmin(admin.ModelAdmin):
list_display = ("label", "source", "fps", "frames", "width", "height", "created")
@admin.register(FootageFrame)
class FootageFrameAdmin(admin.ModelAdmin):
list_display = ("footage", "index", "blob")
list_filter = ("footage",)
@admin.register(Analysis)
class AnalysisAdmin(admin.ModelAdmin):
list_display = ("key", "detector", "version", "footage", "created")
search_fields = ("key", "detector", "version")
@admin.register(Block)
class BlockAdmin(admin.ModelAdmin):
list_display = ("key", "role", "analysis", "data", "state", "created")
list_filter = ("role",)
search_fields = ("key",)
@admin.register(Blob)
class BlobAdmin(admin.ModelAdmin):
list_display = ("digest", "media_type", "size", "created")

14
clips/apps.py Normal file
View file

@ -0,0 +1,14 @@
from django.apps import AppConfig
class ClipsConfig(AppConfig):
"""The one app.
`clips` because the CLIP is the entity the whole tool is about and the one the
prototype had exactly one of and never named — `state` in `js/app.js` is a clip
with its analysis inlined and its palette global. Project, Footage, Analysis,
Block, Leaf and Revision all hang off it.
"""
default_auto_field = "django.db.models.BigAutoField"
name = "clips"

104
clips/blobs.py Normal file
View file

@ -0,0 +1,104 @@
"""The content-addressed blob store: tiers 2 and 3 on disk.
One store for both, and docs/architecture.md says why in a sentence: once tier 3
is decoded by the app rather than by a shell script, frames and audio become "the
same kind of thing as tier 2 — a cache with a hash". So there is one place that
writes bytes, one that reads them, and one URL shape for both.
TWO KINDS OF HASH, AND THEY ARE NOT THE SAME HASH. A blob is named by the sha256
of its BYTES: that is what makes identical frames in two extractions one file. A
derived thing — an analysis artifact, a dense block — is named by a sha256 over
its INPUTS, which is what lets the client ask for the block the current settings
want before anything has computed it. So `Block.key` is an input hash and
`Block.data.digest` is a byte hash, and conflating them would break the half of
addressing that answers questions about work not yet done.
"""
import hashlib
import os
from pathlib import Path
from django.conf import settings
CHUNK = 1 << 20
def digest_bytes(data: bytes) -> str:
return hashlib.sha256(data).hexdigest()
def digest_file(path: Path) -> str:
h = hashlib.sha256()
with open(path, "rb") as fh:
while chunk := fh.read(CHUNK):
h.update(chunk)
return h.hexdigest()
def path_for(digest: str) -> Path:
"""Where a blob lives.
Fanned out two levels, so that a take's worth of frames does not put a hundred
thousand entries in one directory — which is slow on every filesystem and
unusable on some.
"""
if len(digest) != 64 or any(c not in "0123456789abcdef" for c in digest):
raise ValueError(f"not a sha256: {digest!r}")
return Path(settings.BLOB_ROOT) / digest[:2] / digest[2:4] / digest
def write(data: bytes) -> tuple[str, int]:
"""Store bytes, return (digest, size). Writing the same bytes twice is a
no-op, which is what content addressing is for."""
digest = digest_bytes(data)
dest = path_for(digest)
if not dest.exists():
dest.parent.mkdir(parents=True, exist_ok=True)
tmp = dest.with_suffix(".part")
with open(tmp, "wb") as fh:
fh.write(data)
os.replace(tmp, dest)
return digest, len(data)
def adopt(source: Path) -> tuple[str, int]:
"""Store a file already on disk, by hard link where the filesystem allows it.
112MB of PNGs is a normal extraction and copying them into a second place in
the tree for no reason is not. A hard link is exact — the blob is immutable, so
two names for one inode is the whole of what is wanted — and a copy is the
fallback when `extract.sh` wrote to another volume.
"""
digest = digest_file(source)
dest = path_for(digest)
size = source.stat().st_size
if not dest.exists():
dest.parent.mkdir(parents=True, exist_ok=True)
try:
os.link(source, dest)
except OSError:
tmp = dest.with_suffix(".part")
with open(source, "rb") as src, open(tmp, "wb") as out:
while chunk := src.read(CHUNK):
out.write(chunk)
os.replace(tmp, dest)
return digest, size
def read(digest: str) -> bytes:
with open(path_for(digest), "rb") as fh:
return fh.read()
def png_size(path: Path) -> tuple[int, int]:
"""A PNG's dimensions, out of its IHDR.
Twenty-four bytes rather than a dependency. The footage's width and height are
manifest data — docs/architecture.md's entity model puts them there — and
Pillow to read two integers out of a header that has held them in the same
place since 1996 is not a trade worth making.
"""
with open(path, "rb") as fh:
head = fh.read(24)
if head[:8] != b"\x89PNG\r\n\x1a\n" or head[12:16] != b"IHDR":
raise ValueError(f"{path} is not a PNG")
return int.from_bytes(head[16:20], "big"), int.from_bytes(head[20:24], "big")

View file

View file

View file

@ -0,0 +1,120 @@
"""Register an extracted bundle as tier 3.
python manage.py ingest_bundle # ./manifest.json
python manage.py ingest_bundle scratch/my-take # that bundle
WHAT THIS REPLACES. Until step 9 the page fetched `/manifest.json` and then built
`frames/0001.png` itself, with shadow-cljs's `:dev-http` serving the repo root. So
the frame layout was a shared secret between a shell script and a ClojureScript
namespace, and "where the frames are" was answered by a directory listing.
Now the server names every frame, and the client asks it. The frames go into the
content-addressed blob store — by hard link, so 112MB of PNGs is not copied — and
the manifest the client receives carries a URL per frame. That is the whole of what
makes the frames the backend's to serve, and it is what the in-browser wasm-ffmpeg
extraction docs/architecture.md describes will upload INTO, without the client
learning anything new when it arrives: the same blobs, the same manifest, a
different producer.
`extract.sh` still does the decoding. It is out of step 9's scope, it works, and it
is the only part of this that needs a terminal.
"""
import json
from pathlib import Path
from django.core.management.base import BaseCommand, CommandError
from django.db import transaction
from clips import blobs
from clips.models import Blob, Footage, FootageFrame
class Command(BaseCommand):
help = "Register an extracted frames+audio+manifest bundle as footage."
def add_arguments(self, parser):
parser.add_argument(
"bundle", nargs="?", default=".",
help="a directory holding manifest.json, or the manifest itself",
)
parser.add_argument("--label", default="", help="what to call it in the UI")
def handle(self, *args, **options):
manifest_path = Path(options["bundle"])
if manifest_path.is_dir():
manifest_path = manifest_path / "manifest.json"
if not manifest_path.exists():
raise CommandError(f"{manifest_path} does not exist — run ./extract.sh first")
manifest = json.loads(manifest_path.read_text())
root = manifest_path.parent
frames_dir = root / manifest["dir"]
audio_path = root / manifest["audio"]
count = int(manifest["frames"])
pngs = sorted(frames_dir.glob("*.png"))
if len(pngs) != count:
raise CommandError(
f"the manifest says {count} frames and {frames_dir} holds {len(pngs)}; "
"refusing an inaccurate footage"
)
if not audio_path.exists():
raise CommandError(f"{audio_path} does not exist")
width, height = blobs.png_size(pngs[0])
self.stdout.write(f"hashing {len(pngs)} frames…")
frame_blobs = []
for i, png in enumerate(pngs):
digest, size = blobs.adopt(png)
frame_blobs.append((i, digest, size))
if (i + 1) % 25 == 0 or i + 1 == len(pngs):
self.stdout.write(f" {i + 1}/{len(pngs)}")
audio_digest, audio_size = blobs.adopt(audio_path)
# The footage's own identity: every frame in order, plus the audio and the
# rate. Two extractions of one clip at one rate are one footage, so an
# analysis over it is reusable across both.
import hashlib
h = hashlib.sha256()
h.update(f"arthur-footage-1/{manifest['fps']}/{count}/{width}x{height}\n".encode())
for _, digest, _ in frame_blobs:
h.update(digest.encode())
h.update(audio_digest.encode())
footage_digest = h.hexdigest()
if existing := Footage.objects.filter(digest=footage_digest).first():
self.stdout.write(self.style.SUCCESS(f"already ingested: {existing.id}"))
return
with transaction.atomic():
audio_blob, _ = Blob.objects.get_or_create(
digest=audio_digest,
defaults={"size": audio_size, "media_type": "audio/wav"},
)
footage = Footage.objects.create(
digest=footage_digest,
label=options["label"] or manifest.get("source") or frames_dir.name,
source=manifest.get("source") or "",
fps=float(manifest["fps"]),
frames=count,
width=width,
height=height,
audio=audio_blob,
feature_absence=manifest.get("feature-absence") or {},
)
rows = []
for index, digest, size in frame_blobs:
blob, _ = Blob.objects.get_or_create(
digest=digest, defaults={"size": size, "media_type": "image/png"}
)
rows.append(FootageFrame(footage=footage, index=index, blob=blob))
FootageFrame.objects.bulk_create(rows)
self.stdout.write(
self.style.SUCCESS(
f"{count} frames at {manifest['fps']}fps, {width}x{height} -> footage {footage.id}"
)
)

View file

@ -0,0 +1,149 @@
# Generated by Django 5.2.17 on 2026-09-28 04:44
import django.db.models.deletion
import uuid
from django.db import migrations, models
class Migration(migrations.Migration):
initial = True
dependencies = [
]
operations = [
migrations.CreateModel(
name='Blob',
fields=[
('digest', models.CharField(max_length=64, primary_key=True, serialize=False)),
('media_type', models.CharField(default='application/octet-stream', max_length=100)),
('size', models.BigIntegerField()),
('created', models.DateTimeField(auto_now_add=True)),
],
),
migrations.CreateModel(
name='Project',
fields=[
('id', models.UUIDField(default=uuid.uuid4, editable=False, primary_key=True, serialize=False)),
('name', models.CharField(default='untitled', max_length=200)),
('seq', models.PositiveBigIntegerField(default=0)),
('palette', models.CharField(default='arthur/default', max_length=64)),
('created', models.DateTimeField(auto_now_add=True)),
('updated', models.DateTimeField(auto_now=True)),
],
options={
'ordering': ['-updated'],
},
),
migrations.CreateModel(
name='Analysis',
fields=[
('key', models.CharField(max_length=71, primary_key=True, serialize=False)),
('descriptor', models.TextField()),
('detector', models.CharField(max_length=64)),
('version', models.CharField(max_length=64)),
('created', models.DateTimeField(auto_now_add=True)),
('artifact', models.ForeignKey(blank=True, help_text='the dense landmark track, once bake A is uploaded', null=True, on_delete=django.db.models.deletion.SET_NULL, related_name='analysis_for', to='clips.blob')),
],
options={
'verbose_name_plural': 'analyses',
},
),
migrations.CreateModel(
name='Block',
fields=[
('key', models.CharField(max_length=71, primary_key=True, serialize=False)),
('descriptor', models.TextField()),
('role', models.CharField(max_length=32)),
('created', models.DateTimeField(auto_now_add=True)),
('analysis', models.ForeignKey(blank=True, null=True, on_delete=django.db.models.deletion.SET_NULL, related_name='blocks', to='clips.analysis')),
('data', models.ForeignKey(on_delete=django.db.models.deletion.PROTECT, related_name='block_data_for', to='clips.blob')),
('state', models.ForeignKey(blank=True, help_text='the per-track absence mask, when the take has one', null=True, on_delete=django.db.models.deletion.PROTECT, related_name='block_state_for', to='clips.blob')),
],
),
migrations.CreateModel(
name='Footage',
fields=[
('id', models.UUIDField(default=uuid.uuid4, editable=False, primary_key=True, serialize=False)),
('digest', models.CharField(max_length=64, unique=True)),
('label', models.CharField(blank=True, max_length=200)),
('source', models.CharField(blank=True, max_length=200)),
('fps', models.FloatField()),
('frames', models.PositiveIntegerField()),
('width', models.PositiveIntegerField()),
('height', models.PositiveIntegerField()),
('feature_absence', models.JSONField(blank=True, default=dict)),
('created', models.DateTimeField(auto_now_add=True)),
('audio', models.ForeignKey(on_delete=django.db.models.deletion.PROTECT, related_name='audio_for', to='clips.blob')),
],
options={
'ordering': ['-created'],
},
),
migrations.AddField(
model_name='analysis',
name='footage',
field=models.ForeignKey(blank=True, null=True, on_delete=django.db.models.deletion.SET_NULL, related_name='analyses', to='clips.footage'),
),
migrations.CreateModel(
name='Revision',
fields=[
('id', models.BigAutoField(auto_created=True, primary_key=True, serialize=False, verbose_name='ID')),
('seq', models.PositiveBigIntegerField()),
('author', models.CharField(blank=True, max_length=200)),
('summary', models.CharField(blank=True, max_length=500)),
('document', models.JSONField(help_text='every leaf of the project, by path')),
('created', models.DateTimeField(auto_now_add=True)),
('project', models.ForeignKey(on_delete=django.db.models.deletion.CASCADE, related_name='revisions', to='clips.project')),
],
options={
'ordering': ['-seq'],
},
),
migrations.CreateModel(
name='FootageFrame',
fields=[
('id', models.BigAutoField(auto_created=True, primary_key=True, serialize=False, verbose_name='ID')),
('index', models.PositiveIntegerField(help_text="0-based; source frame index + 1 is the PNG's name")),
('blob', models.ForeignKey(on_delete=django.db.models.deletion.PROTECT, related_name='frame_for', to='clips.blob')),
('footage', models.ForeignKey(on_delete=django.db.models.deletion.CASCADE, related_name='frame_set', to='clips.footage')),
],
options={
'ordering': ['index'],
'constraints': [models.UniqueConstraint(fields=('footage', 'index'), name='one_blob_per_frame')],
},
),
migrations.CreateModel(
name='Leaf',
fields=[
('id', models.BigAutoField(auto_created=True, primary_key=True, serialize=False, verbose_name='ID')),
('path', models.CharField(max_length=300)),
('value', models.JSONField()),
('version', models.PositiveBigIntegerField(default=1)),
('updated', models.DateTimeField(auto_now=True)),
('project', models.ForeignKey(on_delete=django.db.models.deletion.CASCADE, related_name='leaves', to='clips.project')),
],
options={
'ordering': ['path'],
'constraints': [models.UniqueConstraint(fields=('project', 'path'), name='one_leaf_per_path')],
},
),
migrations.CreateModel(
name='Clip',
fields=[
('id', models.BigAutoField(auto_created=True, primary_key=True, serialize=False, verbose_name='ID')),
('cid', models.SlugField(max_length=64)),
('name', models.CharField(blank=True, max_length=200)),
('order', models.IntegerField(default=0)),
('analysis', models.ForeignKey(blank=True, null=True, on_delete=django.db.models.deletion.SET_NULL, related_name='clips', to='clips.analysis')),
('blocks', models.ManyToManyField(blank=True, help_text="the tier-2 blocks this clip's channels name", related_name='clips', to='clips.block')),
('footage', models.ForeignKey(blank=True, null=True, on_delete=django.db.models.deletion.SET_NULL, related_name='clips', to='clips.footage')),
('project', models.ForeignKey(on_delete=django.db.models.deletion.CASCADE, related_name='clips', to='clips.project')),
],
options={
'ordering': ['order', 'cid'],
'constraints': [models.UniqueConstraint(fields=('project', 'cid'), name='one_cid_per_project')],
},
),
]

View file

261
clips/models.py Normal file
View file

@ -0,0 +1,261 @@
"""The entity model, as tables.
It follows docs/architecture.md's model exactly, and the one thing worth reading
it for is which tier each table is in, because that is what decides whether a row
is a document, a cache entry or a source.
TIER 1, the document. Project, Clip, Leaf, Revision. Kilobytes, authored,
versioned, and the only tier anything will ever sync.
TIER 2, derived. Analysis, Block. Content-addressed by a hash over every input
that produced them — including the detector version — so a stale bake is
unreachable rather than wrong, and a collaborator's bake is fetchable by the
same key.
TIER 3, source. Footage, FootageFrame. Immutable, by hash.
Blob is under all three of them: bytes, named by the sha256 of themselves.
WHAT IS DELIBERATELY NOT HERE. `Clip` does not store fps, frames, width or height.
They are in the document — the `timing` and `stage` leaves — and a copy of them in
a column is a copy that comes to disagree with the scene it describes. The columns
`Clip` does have are the ones the SERVER needs to answer a question about a clip
without parsing its leaves: which footage, which analysis, which blocks.
"""
import uuid
from django.db import models
class Blob(models.Model):
"""Bytes, named by the sha256 of themselves. The file is on disk under
`BLOB_ROOT`; this row is the index and the size."""
digest = models.CharField(primary_key=True, max_length=64)
media_type = models.CharField(max_length=100, default="application/octet-stream")
size = models.BigIntegerField()
created = models.DateTimeField(auto_now_add=True)
def __str__(self):
return f"{self.digest[:12]}… {self.size}B {self.media_type}"
class Footage(models.Model):
"""Tier 3: the frames and audio of one extraction, immutable.
`digest` is over the ordered frame digests plus the audio's, so two
extractions of the same clip at the same rate are one footage and the same
analysis can be reused across both.
`feature_absence` is the manifest annotation step 8 introduced: known
occlusion intervals, one-based and inclusive, expanded into presence tracks by
the loader. An input format, not a control UI.
"""
id = models.UUIDField(primary_key=True, default=uuid.uuid4, editable=False)
digest = models.CharField(max_length=64, unique=True)
label = models.CharField(max_length=200, blank=True)
source = models.CharField(max_length=200, blank=True)
fps = models.FloatField()
frames = models.PositiveIntegerField()
width = models.PositiveIntegerField()
height = models.PositiveIntegerField()
audio = models.ForeignKey(Blob, on_delete=models.PROTECT, related_name="audio_for")
feature_absence = models.JSONField(default=dict, blank=True)
created = models.DateTimeField(auto_now_add=True)
class Meta:
ordering = ["-created"]
def __str__(self):
return f"{self.label or self.source or self.id} ({self.frames}f @{self.fps})"
class FootageFrame(models.Model):
"""One source frame. A row rather than an entry in a JSON list, because a
frame is a thing the server serves, and because a blob's references have to be
countable before anything can be collected."""
footage = models.ForeignKey(Footage, on_delete=models.CASCADE, related_name="frame_set")
index = models.PositiveIntegerField(help_text="0-based; source frame index + 1 is the PNG's name")
blob = models.ForeignKey(Blob, on_delete=models.PROTECT, related_name="frame_for")
class Meta:
ordering = ["index"]
constraints = [
models.UniqueConstraint(fields=["footage", "index"], name="one_blob_per_frame"),
]
class Analysis(models.Model):
"""Tier 2: one detector, at one version, over one footage.
`key` is a content address over every input, and `descriptor` is the exact
canonical text that key is the sha256 of — sent by the client and stored, not
recomputed here. `clips/views.py` says why that is the honest arrangement: JS
prints an integral double as `1` and Python as `1.0`, so a scheme where both
sides re-render the numbers breaks on the first one of them.
`detector` and `version` are columns as well as descriptor fields so that the
question "which model produced this take" is answerable in the admin and in a
query, rather than only by parsing a hash's preimage.
"""
key = models.CharField(primary_key=True, max_length=71)
descriptor = models.TextField()
detector = models.CharField(max_length=64)
version = models.CharField(max_length=64)
footage = models.ForeignKey(
Footage, null=True, blank=True, on_delete=models.SET_NULL, related_name="analyses"
)
artifact = models.ForeignKey(
Blob, null=True, blank=True, on_delete=models.SET_NULL, related_name="analysis_for",
help_text="the dense landmark track, once bake A is uploaded",
)
created = models.DateTimeField(auto_now_add=True)
class Meta:
verbose_name_plural = "analyses"
def __str__(self):
return f"{self.detector} {self.version} → {self.key[7:19]}…"
class Block(models.Model):
"""Tier 2: one dense channel block.
Two hashes, and they are not the same hash. `key` is over the block's INPUTS,
which is what lets a client ask for the block its current settings want before
anything has computed it. `data.digest` is over the bytes. See clips/blobs.py.
"""
key = models.CharField(primary_key=True, max_length=71)
descriptor = models.TextField()
role = models.CharField(max_length=32)
analysis = models.ForeignKey(
Analysis, null=True, blank=True, on_delete=models.SET_NULL, related_name="blocks"
)
data = models.ForeignKey(Blob, on_delete=models.PROTECT, related_name="block_data_for")
state = models.ForeignKey(
Blob, null=True, blank=True, on_delete=models.PROTECT, related_name="block_state_for",
help_text="the per-track absence mask, when the take has one",
)
created = models.DateTimeField(auto_now_add=True)
def __str__(self):
return f"{self.role} {self.key[7:19]}…"
class Project(models.Model):
"""Tier 1: the document's root.
`seq` is the monotonic project version docs/architecture.md asks for. Every
write bumps it, and a client that sees `seq > local + 1` refetches — which is
what makes staleness self-healing rather than permanent once there is a
broadcast to miss.
"""
id = models.UUIDField(primary_key=True, default=uuid.uuid4, editable=False)
name = models.CharField(max_length=200, default="untitled")
seq = models.PositiveBigIntegerField(default=0)
palette = models.CharField(max_length=64, default="arthur/default")
created = models.DateTimeField(auto_now_add=True)
updated = models.DateTimeField(auto_now=True)
class Meta:
ordering = ["-updated"]
def __str__(self):
return f"{self.name} ({self.id})"
def bump(self):
self.seq += 1
self.save(update_fields=["seq", "updated"])
return self.seq
class Clip(models.Model):
"""Tier 1: the unit of work, and the thing leaf paths are scoped by.
`cid` is what appears in `clip/<cid>/...`, so it is the clip's identity as far
as addressing is concerned and it does not change.
"""
project = models.ForeignKey(Project, on_delete=models.CASCADE, related_name="clips")
cid = models.SlugField(max_length=64)
name = models.CharField(max_length=200, blank=True)
order = models.IntegerField(default=0)
footage = models.ForeignKey(
Footage, null=True, blank=True, on_delete=models.SET_NULL, related_name="clips"
)
analysis = models.ForeignKey(
Analysis, null=True, blank=True, on_delete=models.SET_NULL, related_name="clips"
)
blocks = models.ManyToManyField(
Block, blank=True, related_name="clips",
help_text="the tier-2 blocks this clip's channels name",
)
class Meta:
ordering = ["order", "cid"]
constraints = [
models.UniqueConstraint(fields=["project", "cid"], name="one_cid_per_project"),
]
def __str__(self):
return f"{self.cid} of {self.project.name}"
class Leaf(models.Model):
"""Tier 1: one independently addressed, independently versioned piece of the
document.
The value is transit-as-JSON in a JSONField, so the column holds JSON rather
than a string containing JSON: the admin can read a leaf, and the field-wise
merge of a channel leaf that docs/architecture.md describes as fifteen lines of
Python is possible over it. `version` is the entity tag a conditional write
compares — RFC 7232, not a bespoke invention.
"""
project = models.ForeignKey(Project, on_delete=models.CASCADE, related_name="leaves")
path = models.CharField(max_length=300)
value = models.JSONField()
version = models.PositiveBigIntegerField(default=1)
updated = models.DateTimeField(auto_now=True)
class Meta:
ordering = ["path"]
constraints = [
models.UniqueConstraint(fields=["project", "path"], name="one_leaf_per_path"),
]
@property
def etag(self):
return f'"{self.version}"'
def __str__(self):
return f"{self.path}@{self.version}"
class Revision(models.Model):
"""Tier 1: a snapshot of the authored layer, with a user and a summary.
ON AN EXPLICIT TRIGGER, not on every save. tl snapshots a small annotation
layer; arthur's tier 1 will contain cel polygons, so a snapshot per save bloats
the table — docs/architecture.md's "revisions need a coarser trigger". So this
is written by `POST /api/projects/<id>/revisions`, which is a "mark version"
button, and never by a save.
"""
project = models.ForeignKey(Project, on_delete=models.CASCADE, related_name="revisions")
seq = models.PositiveBigIntegerField()
author = models.CharField(max_length=200, blank=True)
summary = models.CharField(max_length=500, blank=True)
document = models.JSONField(help_text="every leaf of the project, by path")
created = models.DateTimeField(auto_now_add=True)
class Meta:
ordering = ["-seq"]
def __str__(self):
return f"{self.project.name} r{self.seq}: {self.summary}"

View file

@ -0,0 +1,60 @@
{% load static %}<!doctype html>
{% comment %}
The host page, served by Django since port-plan step 9.
It was `frontend/public/index.html`, served by shadow-cljs's `:dev-http`, and that
key is gone. The bundle is unchanged: shadow-cljs writes it into
`static/arthur/js` and staticfiles serves it from there, so `manage.py runserver`
and `shadow-cljs watch app` are the whole dev loop with nothing copying files
between them.
The CSRF token is rendered so that Django sets its cookie, which is what
`arthur.fx.http` reads to write the `X-CSRFToken` header. Saves are ordinary POSTs
and PUTs with ordinary CSRF protection — no endpoint in this app is exempt.
{% endcomment %}
<html lang="en">
<head>
<meta charset="utf-8">
<meta name="viewport" content="width=device-width, initial-scale=1">
<title>arthur</title>
<style>
:root { color-scheme: dark; --bg: #12141c; --fg: #c9c3b4; }
html, body { margin: 0; height: 100%; background: var(--bg); color: var(--fg); }
body { font: 14px/1.5 ui-monospace, SFMono-Regular, Menlo, monospace; }
main { padding: 24px; }
/* The preview is nearest-neighbour everywhere. A browser that smooths the
upscale would misrepresent the look the tool exists to judge. */
canvas { image-rendering: pixelated; }
h1 { font-size: 14px; font-weight: normal; opacity: .5; margin: 0 0 12px; }
.stage { display: block; background: #12141c; }
audio { display: none; }
.transport { margin-top: 12px; width: 640px; }
.transport .row { display: flex; flex-wrap: wrap; gap: 6px; align-items: center; }
.transport .gap { flex: 1; }
button {
font: inherit; color: var(--fg); background: #1c1f2b;
border: 1px solid #2b3040; padding: 3px 10px; cursor: pointer;
}
button:hover { background: #242836; }
button:disabled { opacity: .45; cursor: wait; }
button.on { background: #3a4258; border-color: #556080; }
.scrub { width: 100%; margin: 10px 0 6px; }
.readout { display: flex; gap: 18px; opacity: .55; font-size: 12px; }
.readout .warn { color: #d98f5a; opacity: 1; }
.picture-rate { display: flex; align-items: center; gap: 6px; margin-top: 7px;
font-size: 12px; }
.source-path { display: block; margin-top: 8px; font-size: 12px; opacity: .7; }
.source-path select { margin: 0 8px; padding: 3px 5px;
color: var(--fg); background: #1c1f2b; border: 1px solid #2b3040;
font: inherit; max-width: 360px; }
.load-status { margin-top: 6px; font-size: 12px; opacity: .75; }
.note { opacity: .35; font-size: 12px; max-width: 640px; }
</style>
</head>
<body>
{% csrf_token %}
<div id="app"></div>
<script src="{% static 'mediapipe/vision_bundle.js' %}"></script>
<script src="{% static 'arthur/js/main.js' %}"></script>
</body>
</html>

0
clips/tests/__init__.py Normal file
View file

476
clips/tests/test_api.py Normal file
View file

@ -0,0 +1,476 @@
"""What the server guarantees, as opposed to what the client intends.
The two interesting groups here are the ones that make the tier split a property
of the system rather than a convention in ClojureScript:
A KEY DESCRIBES ITS BYTES. The server recomputes every tier-2 key it is handed
and refuses a mismatch, so nothing can store a block under a name that is not
the hash of its own descriptor.
A BLOCK CAN NAME ITS DETECTOR VERSION. Every block names an analysis and every
analysis declares a detector and a version, both enforced here. That chain is
what docs/architecture.md asks for: without it, a model upgrade that silently
reuses old landmarks presents as "the tool got worse" with no event to attach it
to.
The rest is the load/save round trip, the conditional write, and the footage
manifest that makes the frames the backend's to serve.
"""
import hashlib
import json
import struct
import tempfile
import zlib
from pathlib import Path
from django.test import TestCase, override_settings
from clips import blobs
from clips.models import Analysis, Block, Blob, Clip, Footage, Leaf, Project, Revision
BLOB_DIR = tempfile.mkdtemp(prefix="arthur-test-blobs-")
def key_for(descriptor: str) -> str:
return "sha256:" + hashlib.sha256(descriptor.encode("utf-8")).hexdigest()
def analysis_descriptor(version="1.0.1"):
# Canonical JSON, written the way arthur.domain.canon writes it: sorted keys,
# no spaces, integral doubles with no point.
return ('{"aspect":1,"detector":"mediapipe","frames":48,"fps":30,"scheme":1,'
f'"version":"{version}"}}')
def block_descriptor(analysis_key, role="geom", anchor_avg=2):
return (f'{{"analysis":"{analysis_key}","features":["mouth"],'
f'"layout":{{"frames":48,"scale":16384,"stride":16,"tracks":1,"type":"int16"}},'
f'"observation":null,"params":{{"anchor-avg":{anchor_avg}}},'
f'"role":"{role}","scheme":1,"tracks":["outer"]}}')
def png(width=4, height=3):
"""The smallest valid PNG of a given size, written by hand.
So that `blobs.png_size` and the ingest path are exercised without Pillow. The
one thing the backend needs from a PNG is its IHDR, and this is a PNG with one.
"""
def chunk(kind, payload):
return (struct.pack(">I", len(payload)) + kind + payload
+ struct.pack(">I", zlib.crc32(kind + payload) & 0xFFFFFFFF))
ihdr = struct.pack(">IIBBBBB", width, height, 8, 2, 0, 0, 0)
raw = b"".join(b"\x00" + b"\x40\x40\x40" * width for _ in range(height))
return (b"\x89PNG\r\n\x1a\n" + chunk(b"IHDR", ihdr)
+ chunk(b"IDAT", zlib.compress(raw)) + chunk(b"IEND", b""))
@override_settings(BLOB_ROOT=BLOB_DIR)
class BlobStoreTests(TestCase):
def test_the_same_bytes_are_stored_once(self):
a, size = blobs.write(b"the same bytes")
b, _ = blobs.write(b"the same bytes")
self.assertEqual(a, b)
self.assertEqual(size, 14)
self.assertEqual(blobs.read(a), b"the same bytes")
def test_a_path_that_is_not_a_hash_is_refused(self):
# The blob route takes its digest from the URL, so this is the check that
# stops `/blob/../../etc/passwd` being a path at all.
with self.assertRaises(ValueError):
blobs.path_for("../../etc/passwd")
with self.assertRaises(ValueError):
blobs.path_for("deadbeef")
def test_a_png_reports_its_own_size(self):
with tempfile.NamedTemporaryFile(suffix=".png", delete=False) as fh:
fh.write(png(17, 5))
self.assertEqual((17, 5), blobs.png_size(Path(fh.name)))
def test_a_blob_is_served_immutable(self):
digest, size = blobs.write(b"bytes on the wire")
Blob.objects.create(digest=digest, size=size, media_type="application/octet-stream")
response = self.client.get(f"/blob/{digest}")
self.assertEqual(200, response.status_code)
self.assertIn("immutable", response["Cache-Control"])
self.assertEqual(f'"{digest}"', response["ETag"])
self.assertEqual(b"bytes on the wire", b"".join(response.streaming_content))
def test_an_unknown_blob_is_a_404_and_not_a_traceback(self):
self.assertEqual(404, self.client.get("/blob/" + "0" * 64).status_code)
self.assertEqual(404, self.client.get("/blob/nonsense").status_code)
@override_settings(BLOB_ROOT=BLOB_DIR)
class Tier2Tests(TestCase):
def post(self, url, payload):
return self.client.post(url, data=json.dumps(payload),
content_type="application/json")
def register_analysis(self, version="1.0.1"):
descriptor = analysis_descriptor(version)
key = key_for(descriptor)
response = self.post("/api/analyses", {
"key": key, "descriptor": descriptor,
"detector": "mediapipe", "version": version,
})
self.assertEqual(201, response.status_code, response.content)
return key
def test_an_analysis_is_its_own_descriptors_hash(self):
key = self.register_analysis()
row = Analysis.objects.get(key=key)
self.assertEqual("mediapipe", row.detector)
self.assertEqual("1.0.1", row.version)
# Idempotent: the same inputs are the same key are the same row.
again = self.post("/api/analyses", {
"key": key, "descriptor": analysis_descriptor(), "detector": "mediapipe",
"version": "1.0.1",
})
self.assertEqual(200, again.status_code)
self.assertEqual(1, Analysis.objects.count())
def test_a_key_that_is_not_the_hash_of_its_descriptor_is_refused(self):
response = self.post("/api/analyses", {
"key": "sha256:" + "0" * 64, "descriptor": analysis_descriptor(),
})
self.assertEqual(409, response.status_code)
self.assertIn("not the hash", response.json()["error"])
self.assertEqual(0, Analysis.objects.count())
def test_an_analysis_without_a_detector_version_is_refused(self):
# The rule docs/architecture.md is most insistent about, enforced where a
# client cannot forget it.
descriptor = '{"detector":"mediapipe","frames":48,"scheme":1}'
response = self.post("/api/analyses", {
"key": key_for(descriptor), "descriptor": descriptor,
})
self.assertEqual(400, response.status_code)
self.assertEqual("version", response.json()["missing"])
def test_a_block_is_stored_under_the_hash_of_its_inputs(self):
analysis = self.register_analysis()
descriptor = block_descriptor(analysis)
key = key_for(descriptor)
response = self.post("/api/blocks", {
"key": key, "descriptor": descriptor,
"data": "AAECAwQFBgc=", "state": "AAE=",
})
self.assertEqual(201, response.status_code, response.content)
row = Block.objects.get(key=key)
self.assertEqual("geom", row.role)
self.assertEqual(analysis, row.analysis_id)
# Two hashes, and they are not the same hash: the key is over the inputs,
# the blob's digest is over the bytes.
self.assertNotEqual(key[7:], row.data.digest)
self.assertEqual(8, row.data.size)
fetched = self.client.get(f"/api/blocks/{key}").json()
self.assertEqual("AAECAwQFBgc=", fetched["data"])
self.assertEqual("AAE=", fetched["state"])
self.assertEqual(descriptor, fetched["descriptor"])
def test_a_block_whose_analysis_is_unknown_is_refused(self):
descriptor = block_descriptor("sha256:" + "f" * 64)
response = self.post("/api/blocks", {
"key": key_for(descriptor), "descriptor": descriptor, "data": "AA==",
})
self.assertEqual(400, response.status_code)
self.assertIn("analysis the server does not know", response.json()["error"])
def test_a_block_that_does_not_say_what_its_elements_are_is_refused(self):
analysis = self.register_analysis()
descriptor = ('{"analysis":"%s","layout":{"frames":48},"role":"geom","scheme":1}'
% analysis)
response = self.post("/api/blocks", {
"key": key_for(descriptor), "descriptor": descriptor, "data": "AA==",
})
self.assertEqual(400, response.status_code)
self.assertIn("valid readings", response.json()["error"])
def test_only_the_missing_blocks_are_asked_for(self):
analysis = self.register_analysis()
here = key_for(block_descriptor(analysis))
self.post("/api/blocks", {
"key": here, "descriptor": block_descriptor(analysis), "data": "AA==",
})
elsewhere = key_for(block_descriptor(analysis, anchor_avg=3))
response = self.post("/api/blocks/missing", {"keys": [here, elsewhere]})
self.assertEqual([elsewhere], response.json()["missing"])
def test_a_detector_upgrade_gives_a_block_a_new_name(self):
# The end-to-end statement of the requirement: the same measurements under
# a new model version are a different, additional block, and the old one is
# unreachable from the new document rather than wrong.
old = self.register_analysis("1.0.1")
new_descriptor = analysis_descriptor("1.0.2")
self.post("/api/analyses", {"key": key_for(new_descriptor),
"descriptor": new_descriptor})
for analysis in (old, key_for(new_descriptor)):
descriptor = block_descriptor(analysis)
self.post("/api/blocks", {"key": key_for(descriptor),
"descriptor": descriptor, "data": "AAEC"})
self.assertEqual(2, Block.objects.count())
# One set of bytes, two names: the upgrade renamed the block and did not
# duplicate it on disk.
self.assertEqual(1, Blob.objects.filter(block_data_for__isnull=False).distinct().count())
@override_settings(BLOB_ROOT=BLOB_DIR)
class DocumentTests(TestCase):
"""Tier 1: load, save, and the conditional write."""
def setUp(self):
self.project = Project.objects.create(name="a project")
descriptor = analysis_descriptor()
self.analysis = key_for(descriptor)
self.client.post("/api/analyses", data=json.dumps(
{"key": self.analysis, "descriptor": descriptor}),
content_type="application/json")
block = block_descriptor(self.analysis)
self.block = key_for(block)
self.client.post("/api/blocks", data=json.dumps(
{"key": self.block, "descriptor": block, "data": "AAECAwQFBgc="}),
content_type="application/json")
def put(self, url, payload, **headers):
return self.client.put(url, data=json.dumps(payload),
content_type="application/json", **headers)
def leaves(self):
# Transit-shaped, because that is what a leaf actually holds: a map with a
# cache marker, keyword keys, and a frame-keyed inner map.
return {
"clip/c1/timing": ["^ ", "~:fps", 30, "~:frames", 48],
"clip/c1/node/mouth": ["^ ", "~:id", "~:mouth", "~:z", "a1"],
"clip/c1/channel/mouth/geom.pts": [
"^ ", "~:animated?", True, "~:dense",
["^ ", "~:store", self.block, "~:offset", 0, "~:stride", 16],
],
"clip/c1/channel/mouth-in/vis": [
"^ ", "~:animated?", True, "~:keys", ["^ ", "~i0", True, "~i12", False],
],
}
def save(self, leaves=None, blocks=None):
return self.put(f"/api/projects/{self.project.id}", {
"name": "a project",
"clips": [{"cid": "c1", "name": "take", "analysis": self.analysis,
"leaves": leaves if leaves is not None else self.leaves(),
"blocks": blocks if blocks is not None else [self.block]}],
})
def test_a_document_comes_back_exactly(self):
response = self.save()
self.assertEqual(200, response.status_code, response.content)
self.assertEqual(4, len(response.json()["written"]))
loaded = self.client.get(f"/api/projects/{self.project.id}").json()
self.assertEqual(1, len(loaded["clips"]))
clip = loaded["clips"][0]
self.assertEqual("c1", clip["cid"])
self.assertEqual([self.block], clip["blocks"])
self.assertEqual(self.analysis, clip["analysis"])
# The whole point: byte-identical values, including the integer frame keys
# transit writes as "~i0". A JSON round trip that stringified them would
# come back "0" and the part would hold its first pose forever.
self.assertEqual(self.leaves(), clip["leaves"])
def test_an_unchanged_leaf_keeps_its_version(self):
# What makes an entity tag worth having: a save where one channel moved
# invalidates one leaf's etag, not the whole document's.
self.save()
first = {leaf.path: leaf.version for leaf in Leaf.objects.all()}
moved = self.leaves()
moved["clip/c1/channel/mouth-in/vis"] = [
"^ ", "~:animated?", True, "~:keys", ["^ ", "~i0", False],
]
response = self.save(moved)
self.assertEqual(["clip/c1/channel/mouth-in/vis"], response.json()["written"])
self.assertEqual(3, response.json()["unchanged"])
after = {leaf.path: leaf.version for leaf in Leaf.objects.all()}
self.assertEqual(2, after["clip/c1/channel/mouth-in/vis"])
self.assertEqual(first["clip/c1/timing"], after["clip/c1/timing"])
def test_a_removed_node_removes_its_leaf(self):
self.save()
fewer = {k: v for k, v in self.leaves().items() if k != "clip/c1/node/mouth"}
response = self.save(fewer)
self.assertEqual(["clip/c1/node/mouth"], response.json()["removed"])
self.assertEqual(3, Leaf.objects.count())
def test_a_save_does_not_disturb_another_clip(self):
# A save is not the only way the document changes, so a save that cleared
# what it did not mention would undo a collaborator.
Leaf.objects.create(project=self.project, path="clip/c2/timing", value=["^ "])
self.save()
self.assertTrue(Leaf.objects.filter(path="clip/c2/timing").exists())
def test_a_leaf_addressed_to_another_clip_is_refused(self):
response = self.save({"clip/c9/timing": ["^ "]})
self.assertEqual(400, response.status_code)
self.assertIn("not addressed to clip", response.json()["error"])
self.assertEqual(0, Leaf.objects.count())
def test_a_document_naming_blocks_the_server_lacks_is_refused(self):
# Referential integrity across the tiers. Saved without this, the document
# loads into a blank stage on any other machine.
response = self.save(blocks=[self.block, "sha256:" + "a" * 64])
self.assertEqual(409, response.status_code)
self.assertEqual(["sha256:" + "a" * 64], response.json()["missing"])
self.assertEqual(0, Leaf.objects.count())
def test_every_write_bumps_the_projects_version(self):
before = Project.objects.get(id=self.project.id).seq
self.save()
self.assertEqual(before + 1, Project.objects.get(id=self.project.id).seq)
# --- the conditional write ---------------------------------------------
def test_a_leaf_write_carries_an_etag(self):
self.save()
url = f"/api/projects/{self.project.id}/leaves/clip/c1/node/mouth"
got = self.client.get(url)
self.assertEqual('"1"', got["ETag"])
ok = self.put(url, {"value": ["^ ", "~:id", "~:mouth", "~:z", "a2"]},
HTTP_IF_MATCH='"1"')
self.assertEqual(200, ok.status_code)
self.assertEqual('"2"', ok["ETag"])
self.assertEqual(["^ ", "~:id", "~:mouth", "~:z", "a2"],
self.client.get(url).json()["value"])
def test_a_stale_write_is_refused_and_says_what_is_there(self):
# 409 with the current value, so the client can offer keep-mine /
# take-theirs. A PUT that replaced unconditionally is the bug where the
# loser's work disappears silently.
self.save()
url = f"/api/projects/{self.project.id}/leaves/clip/c1/node/mouth"
self.put(url, {"value": ["^ ", "~:z", "a2"]}, HTTP_IF_MATCH='"1"')
stale = self.put(url, {"value": ["^ ", "~:z", "a3"]}, HTTP_IF_MATCH='"1"')
self.assertEqual(409, stale.status_code)
self.assertEqual(2, stale.json()["version"])
self.assertEqual(["^ ", "~:z", "a2"], stale.json()["value"])
# And the value on the server is the one that won, not the one refused.
self.assertEqual(["^ ", "~:z", "a2"], self.client.get(url).json()["value"])
def test_an_unconditional_write_still_works(self):
# Conditional writes are the protocol, not a requirement: the first write
# of a leaf has no etag to match.
url = f"/api/projects/{self.project.id}/leaves/clip/c1/stage"
response = self.put(url, {"value": ["^ ", "~:width", 320]})
self.assertEqual(200, response.status_code)
self.assertEqual('"1"', response["ETag"])
def test_if_match_star_requires_the_leaf_to_exist(self):
url = f"/api/projects/{self.project.id}/leaves/clip/c1/nothing"
self.assertEqual(409, self.put(url, {"value": []}, HTTP_IF_MATCH="*").status_code)
# --- revisions ---------------------------------------------------------
def test_a_revision_snapshots_the_authored_layer(self):
self.save()
response = self.client.post(
f"/api/projects/{self.project.id}/revisions",
data=json.dumps({"summary": "first pass", "author": "olive"}),
content_type="application/json",
)
self.assertEqual(201, response.status_code)
revision = Revision.objects.get()
self.assertEqual(4, len(revision.document))
self.assertEqual(self.leaves(), revision.document)
# Coarse on purpose: a save does not write one, because tier 1 will hold
# cel polygons and a snapshot per save bloats the table.
self.save()
self.assertEqual(1, Revision.objects.count())
@override_settings(BLOB_ROOT=BLOB_DIR)
class FootageTests(TestCase):
"""Tier 3, and the thing that makes the frames the backend's to serve: the
manifest names every frame by URL."""
def bundle(self, frames=3, absence=None):
root = Path(tempfile.mkdtemp(prefix="arthur-test-bundle-"))
(root / "frames").mkdir()
for i in range(frames):
(root / "frames" / f"{i + 1:04d}.png").write_bytes(png(8, 6) + bytes([i]))
(root / "audio.wav").write_bytes(b"RIFF....WAVEfmt ")
manifest = {"fps": 12, "frames": frames, "dir": "frames",
"audio": "audio.wav", "source": "IMG_8608.MOV"}
if absence:
manifest["feature-absence"] = absence
(root / "manifest.json").write_text(json.dumps(manifest))
return root
def ingest(self, root):
from django.core.management import call_command
from io import StringIO
call_command("ingest_bundle", str(root), stdout=StringIO())
return Footage.objects.get()
def test_a_bundle_becomes_footage_with_a_url_per_frame(self):
footage = self.ingest(self.bundle(frames=3, absence={"eye-r": [[1, 2]]}))
self.assertEqual(3, footage.frames)
self.assertEqual((8, 6), (footage.width, footage.height))
self.assertEqual(12, footage.fps)
self.assertEqual({"eye-r": [[1, 2]]}, footage.feature_absence)
manifest = self.client.get(f"/api/footage/{footage.id}").json()
self.assertEqual(3, len(manifest["urls"]))
self.assertTrue(all(url.startswith("/blob/") for url in manifest["urls"]))
self.assertEqual(f"sha256:{footage.digest}", manifest["footage"])
self.assertTrue(manifest["audio"].startswith("/blob/"))
self.assertEqual({"eye-r": [[1, 2]]}, manifest["feature-absence"])
# The frames are in order, and each one is fetchable.
first = self.client.get(manifest["urls"][0])
self.assertEqual(200, first.status_code)
self.assertEqual("image/png", first["Content-Type"])
def test_ingesting_the_same_bundle_twice_is_one_footage(self):
root = self.bundle()
self.ingest(root)
self.ingest(root)
self.assertEqual(1, Footage.objects.count())
def test_a_bundle_whose_count_disagrees_with_its_frames_is_refused(self):
from django.core.management import call_command
from django.core.management.base import CommandError
from io import StringIO
root = self.bundle(frames=3)
(root / "frames" / "0003.png").unlink()
with self.assertRaisesMessage(CommandError, "refusing an inaccurate footage"):
call_command("ingest_bundle", str(root), stdout=StringIO())
def test_the_footage_list_does_not_carry_every_url(self):
# A list of takes should not be a list of six hundred URLs each.
self.ingest(self.bundle())
listed = self.client.get("/api/footage").json()["footage"]
self.assertEqual(1, len(listed))
self.assertNotIn("urls", listed[0])
class PageTests(TestCase):
def test_django_serves_the_page_at_both_urls(self):
for url in ("/", "/index.html"):
response = self.client.get(url)
self.assertEqual(200, response.status_code, url)
body = response.content.decode()
self.assertIn("/static/arthur/js/main.js", body)
self.assertIn("/static/mediapipe/vision_bundle.js", body)
self.assertIn('id="app"', body)
# The token is rendered so Django sets its cookie, which is what the
# save path reads to write the X-CSRFToken header.
self.assertIn("csrfmiddlewaretoken", body)
def test_the_detector_reports_a_version_derived_from_the_model(self):
# The version is the package version plus the model asset's own hash,
# because a version string in the client is one somebody has to remember to
# bump, and the server is the thing that serves the model.
report = self.client.get("/api/detector").json()
self.assertEqual("mediapipe", report["detector"])
self.assertNotEqual("unknown", report["version"])
self.assertTrue(report["model"].startswith("sha256:"))
self.assertIn("+", report["version"])

30
clips/urls.py Normal file
View file

@ -0,0 +1,30 @@
"""The API, which is nine endpoints and no framework.
The shape is RFC 7232 over addressed resources: a leaf is a resource, its version
is an entity tag, and a conditional write answers 409. docs/architecture.md is
explicit that this part is not a bespoke invention — "optimistic concurrency
control over addressed resources with an entity tag" is what HTTP has done for
thirty years — so the plumbing here is deliberately boring.
WRITES ARE ON HTTP AND STAY THERE. When the websocket arrives it carries presence
and broadcasts, and not writes: auth, idempotency, status codes, retries and
conditional requests all come for free here, and a dropped socket cannot lose a
write.
"""
from django.urls import path
from . import views
urlpatterns = [
path("detector", views.detector),
path("footage", views.footage_list),
path("footage/<uuid:footage_id>", views.footage_detail),
path("projects", views.projects),
path("projects/<uuid:project_id>", views.project_detail),
path("projects/<uuid:project_id>/leaves/<path:leaf_path>", views.leaf_detail),
path("projects/<uuid:project_id>/revisions", views.revisions),
path("analyses", views.analyses),
path("blocks", views.blocks),
path("blocks/missing", views.blocks_missing),
path("blocks/<str:key>", views.block_detail),
]

578
clips/views.py Normal file
View file

@ -0,0 +1,578 @@
"""The API's implementation.
Two things in here are load-bearing and neither is Django.
THE SERVER VERIFIES EVERY TIER-2 KEY IT IS HANDED. A key is the sha256 of a
canonical descriptor, and this recomputes it and refuses a mismatch. That is what
makes content addressing a property of the system rather than a convention in the
client: nothing can store bytes under a name that does not describe them.
It hashes THE TEXT IT WAS SENT rather than re-rendering the descriptor from parsed
values, and that is the honest arrangement rather than a shortcut. JS prints an
integral double as `1` and Python prints `1.0`, so a scheme where both sides
re-render the numbers would disagree on the first parameter whose value happens to
be whole — and the failure would be an upload that 409s with nothing wrong. The
bytes are the contract; the schema on top of them is a convention, and the two
fields this file actually reads out of that schema are checked separately.
AND IT REFUSES A BLOCK WHOSE ANALYSIS IT DOES NOT KNOW. Every block descriptor
names an analysis, and every analysis declares a detector and a VERSION. So the
chain from a stored block to the model version that produced it cannot be broken
by a client that forgot a step — which is the whole point of
docs/architecture.md's insistence that the cache key include the detector version.
A model upgrade that silently reused old landmarks would otherwise present as "the
tool got worse", with no event to attach it to.
"""
import hashlib
import json
from functools import lru_cache
from pathlib import Path
from django.conf import settings
from django.db import transaction
from django.http import FileResponse, HttpResponse, JsonResponse
from django.shortcuts import render
from django.views.decorators.http import require_http_methods
from . import blobs
from .models import Analysis, Block, Blob, Clip, Footage, Leaf, Project, Revision
KEY_LENGTH = 71 # "sha256:" + 64 hex
# ---------------------------------------------------------------------------
# helpers
def _body(request):
try:
return json.loads(request.body or b"{}")
except json.JSONDecodeError as exc:
raise Bad(f"the request body is not JSON: {exc}") from exc
class Bad(Exception):
"""A 400 with a message, raised where the problem is noticed."""
def __init__(self, message, status=400, **detail):
super().__init__(message)
self.message = message
self.status = status
self.detail = detail
def _error(exc: Bad):
return JsonResponse({"error": exc.message, **exc.detail}, status=exc.status)
def _check_key(key, descriptor):
"""The verification. A key is the sha256 of the descriptor stored beside it."""
if not isinstance(key, str) or len(key) != KEY_LENGTH or not key.startswith("sha256:"):
raise Bad(f"not a content address: {key!r}")
if not isinstance(descriptor, str) or not descriptor:
raise Bad("a key without its descriptor addresses nothing")
actual = hashlib.sha256(descriptor.encode("utf-8")).hexdigest()
if actual != key[7:]:
raise Bad(
"the key is not the hash of its descriptor",
status=409,
expected=f"sha256:{actual}",
given=key,
)
try:
return json.loads(descriptor)
except json.JSONDecodeError as exc:
raise Bad(f"the descriptor is not canonical JSON: {exc}") from exc
def _blob(b64, media_type="application/octet-stream"):
import base64
digest, size = blobs.write(base64.b64decode(b64))
blob, _ = Blob.objects.get_or_create(
digest=digest, defaults={"size": size, "media_type": media_type}
)
return blob
# ---------------------------------------------------------------------------
# the page
def page(request):
"""The host page. This replaced `frontend/public/index.html` at step 9, and
`:dev-http` in shadow-cljs.edn went away with it."""
return render(request, "clips/index.html")
# ---------------------------------------------------------------------------
# the detector
#
# WHY THE SERVER ANSWERS THIS. The analysis key has to include the detector
# version, and a version string in the client is a string somebody has to remember
# to bump. The server serves the model, so it can hash the model — and then the
# version is a fact about the bytes that produced the landmarks rather than a
# claim about them.
@lru_cache(maxsize=4)
def _model_digest(path: str, mtime: float) -> str:
return blobs.digest_file(Path(path))
def _package_version() -> str:
pkg = Path(settings.BASE_DIR) / "frontend" / "package.json"
try:
deps = json.loads(pkg.read_text())["dependencies"]
return deps["@mediapipe/tasks-vision"].lstrip("^~")
except Exception:
return "unknown"
@require_http_methods(["GET"])
def detector(request):
model = Path(settings.BASE_DIR) / "frontend" / "public" / "mediapipe" / "face_landmarker.task"
if not model.exists():
# Honest rather than fatal: detection will fail at the MediaPipe boundary
# with a better message than this one could give, and an analysis stamped
# "unknown" is a take somebody can still look at and re-freeze later.
return JsonResponse({"detector": "mediapipe", "version": "unknown", "model": None})
digest = _model_digest(str(model), model.stat().st_mtime)
return JsonResponse(
{
"detector": "mediapipe",
# The package version AND the model's own hash. Either alone can change
# while the other does not, and both change the landmarks.
"version": f"{_package_version()}+{digest[:16]}",
"model": f"sha256:{digest}",
}
)
# ---------------------------------------------------------------------------
# tier 3: footage
#
# THE MANIFEST NOW CARRIES URLS. It used to carry a directory and the loader built
# `frames/0001.png` itself, which quietly made the frame layout a shared secret
# between a shell script and a ClojureScript namespace. The server names every
# frame instead, so the frames can move into the blob store — or later be uploaded
# from the browser by wasm ffmpeg — without the client learning anything new.
def _footage_json(footage: Footage, urls=True):
out = {
"id": str(footage.id),
"label": footage.label or footage.source,
"source": footage.source,
"fps": footage.fps,
"frames": footage.frames,
"width": footage.width,
"height": footage.height,
"footage": f"sha256:{footage.digest}",
"audio": f"/blob/{footage.audio.digest}",
"feature-absence": footage.feature_absence or {},
}
if urls:
out["urls"] = [f"/blob/{f.blob.digest}" for f in footage.frame_set.select_related("blob")]
return out
@require_http_methods(["GET"])
def footage_list(request):
return JsonResponse(
{"footage": [_footage_json(f, urls=False) for f in Footage.objects.all()]}
)
@require_http_methods(["GET"])
def footage_detail(request, footage_id):
try:
footage = Footage.objects.select_related("audio").get(id=footage_id)
except Footage.DoesNotExist:
return JsonResponse({"error": "no such footage"}, status=404)
return JsonResponse(_footage_json(footage))
@require_http_methods(["GET"])
def blob(request, digest):
"""Raw bytes, immutable.
`immutable` is not optimism here, it is the definition: the name IS the hash of
the content, so a cached copy cannot be stale. That is what makes serving 600
frames out of this cheap enough to do on every load.
"""
try:
row = Blob.objects.get(digest=digest)
path = blobs.path_for(digest)
except (Blob.DoesNotExist, ValueError):
return JsonResponse({"error": "no such blob"}, status=404)
response = FileResponse(open(path, "rb"), content_type=row.media_type)
response["Cache-Control"] = "public, max-age=31536000, immutable"
response["ETag"] = f'"{digest}"'
return response
# ---------------------------------------------------------------------------
# tier 2: analyses and blocks
@require_http_methods(["POST"])
def analyses(request):
"""Register an analysis artifact's identity. Idempotent: the same inputs are
the same key are the same row."""
try:
data = _body(request)
key = data.get("key")
descriptor = data.get("descriptor")
parsed = _check_key(key, descriptor)
for field in ("detector", "version"):
if not parsed.get(field):
raise Bad(
f"the descriptor does not declare a {field}: a cache key that "
"omits the detector version lets a model upgrade silently reuse "
"old landmarks",
missing=field,
)
footage = None
if data.get("footage"):
digest = str(data["footage"]).removeprefix("sha256:")
footage = Footage.objects.filter(digest=digest).first()
if footage is None:
raise Bad("the analysis names footage this server does not have",
footage=data["footage"])
row, created = Analysis.objects.get_or_create(
key=key,
defaults={
"descriptor": descriptor,
"detector": parsed["detector"],
"version": str(parsed["version"]),
"footage": footage,
},
)
return JsonResponse({"key": row.key, "created": created}, status=201 if created else 200)
except Bad as exc:
return _error(exc)
@require_http_methods(["POST"])
def blocks_missing(request):
"""Which of these keys the server does not have.
The return on content addressing, as one request: a save uploads the blocks
that are new and nothing else, so re-saving a document after a knob-free edit
moves kilobytes.
"""
try:
keys = _body(request).get("keys") or []
if not isinstance(keys, list):
raise Bad("keys must be a list")
have = set(Block.objects.filter(key__in=keys).values_list("key", flat=True))
return JsonResponse({"missing": [k for k in keys if k not in have]})
except Bad as exc:
return _error(exc)
@require_http_methods(["POST"])
def blocks(request):
"""Store one dense block: its bytes, its optional absence mask, and the
descriptor its key is the hash of."""
try:
data = _body(request)
key = data.get("key")
descriptor = data.get("descriptor")
parsed = _check_key(key, descriptor)
role = parsed.get("role")
if not role:
raise Bad("a block's descriptor names its role")
if not parsed.get("layout", {}).get("type"):
raise Bad(
"a block's descriptor must say what its elements are: an Int16Array "
"and a Float32Array over the same bytes are both valid readings and "
"only one of them is the block"
)
analysis_key = parsed.get("analysis")
analysis = Analysis.objects.filter(key=analysis_key).first()
if analysis is None:
raise Bad(
"this block names an analysis the server does not know; register the "
"analysis first, so that every stored block can name the detector "
"version that produced it",
analysis=analysis_key,
)
if not data.get("data"):
raise Bad("a block with no bytes")
with transaction.atomic():
row, created = Block.objects.get_or_create(
key=key,
defaults={
"descriptor": descriptor,
"role": role,
"analysis": analysis,
"data": _blob(data["data"]),
"state": _blob(data["state"]) if data.get("state") else None,
},
)
return JsonResponse({"key": row.key, "created": created}, status=201 if created else 200)
except Bad as exc:
return _error(exc)
@require_http_methods(["GET"])
def block_detail(request, key):
import base64
try:
row = Block.objects.select_related("data", "state").get(key=key)
except Block.DoesNotExist:
return JsonResponse({"error": "no such block"}, status=404)
out = {
"key": row.key,
"descriptor": row.descriptor,
"data": base64.b64encode(blobs.read(row.data.digest)).decode("ascii"),
}
if row.state_id:
out["state"] = base64.b64encode(blobs.read(row.state.digest)).decode("ascii")
response = JsonResponse(out)
response["Cache-Control"] = "public, max-age=31536000, immutable"
return response
# ---------------------------------------------------------------------------
# tier 1: projects, clips, leaves
def _project_json(project: Project):
leaves = list(project.leaves.all())
clips = []
for clip in project.clips.all():
prefix = f"clip/{clip.cid}/"
clips.append(
{
"cid": clip.cid,
"name": clip.name,
"footage": str(clip.footage_id) if clip.footage_id else None,
"analysis": clip.analysis_id,
"blocks": sorted(clip.blocks.values_list("key", flat=True)),
"leaves": {leaf.path: leaf.value for leaf in leaves if leaf.path.startswith(prefix)},
}
)
return {
"id": str(project.id),
"name": project.name,
"seq": project.seq,
"palette": project.palette,
"clips": clips,
}
@require_http_methods(["GET", "POST"])
def projects(request):
if request.method == "GET":
return JsonResponse(
{
"projects": [
{"id": str(p.id), "name": p.name, "seq": p.seq,
"updated": p.updated.isoformat()}
for p in Project.objects.all()[:100]
]
}
)
try:
data = _body(request)
project = Project.objects.create(name=data.get("name") or "untitled")
return JsonResponse(_project_json(project), status=201)
except Bad as exc:
return _error(exc)
@require_http_methods(["GET", "PUT"])
def project_detail(request, project_id):
try:
project = Project.objects.get(id=project_id)
except Project.DoesNotExist:
return JsonResponse({"error": "no such project"}, status=404)
if request.method == "GET":
return JsonResponse(_project_json(project))
try:
return _save(project, _body(request))
except Bad as exc:
return _error(exc)
@transaction.atomic
def _save(project: Project, data):
"""A whole-document save: one clip's leaves replace that clip's leaves.
SCOPED BY CLIP, not by project. A payload that carries clip `a` does not
disturb clip `b`'s leaves, because a save is not the only way the document
changes — a single-leaf conditional write is — and a save that cleared
everything it did not mention would be a save that undoes a collaborator.
A leaf whose value is unchanged keeps its VERSION. That is what makes the
entity tag mean something: a save of a document where one channel moved
invalidates one leaf's etag, not all four hundred.
"""
if data.get("name"):
project.name = data["name"]
if data.get("palette"):
project.palette = data["palette"]
written, removed, unchanged = [], [], []
for spec in data.get("clips") or []:
cid = spec.get("cid")
if not cid:
raise Bad("every clip in a save names its cid")
leaves = spec.get("leaves") or {}
prefix = f"clip/{cid}/"
for path in leaves:
if not path.startswith(prefix):
raise Bad(
f"leaf {path!r} is not addressed to clip {cid!r}",
clip=cid, path=path,
)
keys = spec.get("blocks") or []
have = set(Block.objects.filter(key__in=keys).values_list("key", flat=True))
if missing := [k for k in keys if k not in have]:
# Referential integrity across the tiers, enforced where it can be:
# a document that names blocks the server does not hold would load
# into a blank stage on any other machine.
raise Bad(
"this clip names tier-2 blocks the server does not have; upload them "
"before saving the document that points at them",
status=409, missing=missing,
)
analysis = Analysis.objects.filter(key=spec.get("analysis")).first()
footage = None
if spec.get("footage"):
footage = Footage.objects.filter(id=spec["footage"]).first()
clip, _ = Clip.objects.update_or_create(
project=project,
cid=cid,
defaults={"name": spec.get("name") or "", "analysis": analysis, "footage": footage},
)
clip.blocks.set(Block.objects.filter(key__in=keys))
existing = {leaf.path: leaf for leaf in project.leaves.filter(path__startswith=prefix)}
for path, value in leaves.items():
leaf = existing.get(path)
if leaf is None:
Leaf.objects.create(project=project, path=path, value=value)
written.append(path)
elif leaf.value != value:
leaf.value = value
leaf.version += 1
leaf.save(update_fields=["value", "version", "updated"])
written.append(path)
else:
unchanged.append(path)
for path, leaf in existing.items():
if path not in leaves:
leaf.delete()
removed.append(path)
seq = project.seq + 1
project.seq = seq
project.save()
return JsonResponse(
{
"id": str(project.id),
"seq": seq,
"written": sorted(written),
"removed": sorted(removed),
"unchanged": len(unchanged),
}
)
@require_http_methods(["GET", "PUT"])
def leaf_detail(request, project_id, leaf_path):
"""One leaf, conditionally.
`If-Match` and a 409 whose body carries the CURRENT value, so the client can
offer keep-mine / take-theirs. A PUT that replaced unconditionally is the bug
docs/architecture.md calls out in tl: the loser's work disappears silently, and
for a painted cel that is the class of bug that ends trust in a tool.
"""
try:
project = Project.objects.get(id=project_id)
except Project.DoesNotExist:
return JsonResponse({"error": "no such project"}, status=404)
leaf = project.leaves.filter(path=leaf_path).first()
if request.method == "GET":
if leaf is None:
return JsonResponse({"error": "no such leaf"}, status=404)
response = JsonResponse({"path": leaf.path, "value": leaf.value, "version": leaf.version})
response["ETag"] = leaf.etag
return response
try:
data = _body(request)
except Bad as exc:
return _error(exc)
if "value" not in data:
return _error(Bad("a leaf write carries a value"))
match = request.headers.get("If-Match")
if leaf is None:
# ANY `If-Match` on a leaf that does not exist is a failed precondition,
# `*` included: RFC 7232 gives `*` the meaning "the resource must already
# exist", which is exactly the write a client makes when it believes it is
# editing something. Creating it instead would turn "somebody deleted this
# node" into a silent resurrection.
if match:
return JsonResponse(
{"error": "no such leaf", "path": leaf_path}, status=409
)
leaf = Leaf.objects.create(project=project, path=leaf_path, value=data["value"])
else:
if match and match not in ("*", leaf.etag):
response = JsonResponse(
{
"error": "stale write",
"path": leaf.path,
"version": leaf.version,
"value": leaf.value,
},
status=409,
)
response["ETag"] = leaf.etag
return response
leaf.value = data["value"]
leaf.version += 1
leaf.save(update_fields=["value", "version", "updated"])
seq = project.bump()
response = JsonResponse({"path": leaf.path, "version": leaf.version, "seq": seq})
response["ETag"] = leaf.etag
return response
@require_http_methods(["GET", "POST"])
def revisions(request, project_id):
"""Mark a version: one snapshot of the authored layer, with a summary."""
try:
project = Project.objects.get(id=project_id)
except Project.DoesNotExist:
return JsonResponse({"error": "no such project"}, status=404)
if request.method == "GET":
return JsonResponse(
{
"revisions": [
{"seq": r.seq, "author": r.author, "summary": r.summary,
"created": r.created.isoformat(), "leaves": len(r.document)}
for r in project.revisions.all()[:100]
]
}
)
data = json.loads(request.body or b"{}")
revision = Revision.objects.create(
project=project,
seq=project.seq,
author=data.get("author") or "",
summary=data.get("summary") or "",
document={leaf.path: leaf.value for leaf in project.leaves.all()},
)
return JsonResponse({"seq": revision.seq, "leaves": len(revision.document)}, status=201)