Serve the document from a Django backend, split into three tiers
Step 9. The tier split was the work; Django was the easy half.
Tier 1 — the authored scene — is the document, and it is addressed as
independently versioned leaves rather than saved whole, so one vertex drag
cannot clobber a collaborator's keying. `domain/leaf` is the document as
path -> value; `domain/wire` puts it on the wire as transit, because JSON
has neither integer map keys nor keywords and a save would quietly turn
`{0 v}` into `{"0" v}`.
Tier 2 — the dense channel blocks — is content-addressed by a hash over
every input, with the detector version inside every key through the
analysis the block descriptor names. `flow/address`'s `block-knobs` is the
invalidation table, and `address-test` does not trust it: it re-freezes the
take once per knob and asserts the biconditional, that a block's bytes
changed if and only if its key changed. That found `brow-pos` not depending
on `contour-avg` — the brow ring is smoothed, the raise is not.
Tier 3 — frames and audio — is served by the hash of its bytes out of the
same store. A manifest now names frames and carries a URL for each, so the
frame layout stopped being a shared secret between a shell script and a
ClojureScript namespace, and the `?v=` cache-buster went with it: a blob's
name is the hash of its contents, so a stale copy is not a thing that can
happen. The synthetic take's `audio.wav` moved to `static/arthur/` — an
asset the project owns, not an extraction that churns.
The server verifies rather than trusting a name it was handed: it
recomputes every key from the descriptor stored beside it, refuses an
analysis that declares no detector version, and refuses a document naming
blocks it does not hold. It hashes the descriptor TEXT, because JS prints
an integral double as `1` and Python as `1.0`, and a scheme where both ends
re-render the numbers disagrees on the first parameter that happens to be
whole.
Two loose ends from step 8 closed on the way. `pack` no longer takes a
`(track, frame)` predicate whose call sites each re-derived a feature from
an index — every track names the feature it follows, which deleted five
hand-maintained mappings. And `:dev-http` is gone: Django serves the page,
shadow-cljs only builds into the staticfiles tree.
227 CLJS tests, 31 Django tests, green.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
parent
b6517f837a
commit
9cd5243983
61 changed files with 4694 additions and 269 deletions
261
clips/models.py
Normal file
261
clips/models.py
Normal file
|
|
@ -0,0 +1,261 @@
|
|||
"""The entity model, as tables.
|
||||
|
||||
It follows docs/architecture.md's model exactly, and the one thing worth reading
|
||||
it for is which tier each table is in, because that is what decides whether a row
|
||||
is a document, a cache entry or a source.
|
||||
|
||||
TIER 1, the document. Project, Clip, Leaf, Revision. Kilobytes, authored,
|
||||
versioned, and the only tier anything will ever sync.
|
||||
|
||||
TIER 2, derived. Analysis, Block. Content-addressed by a hash over every input
|
||||
that produced them — including the detector version — so a stale bake is
|
||||
unreachable rather than wrong, and a collaborator's bake is fetchable by the
|
||||
same key.
|
||||
|
||||
TIER 3, source. Footage, FootageFrame. Immutable, by hash.
|
||||
|
||||
Blob is under all three of them: bytes, named by the sha256 of themselves.
|
||||
|
||||
WHAT IS DELIBERATELY NOT HERE. `Clip` does not store fps, frames, width or height.
|
||||
They are in the document — the `timing` and `stage` leaves — and a copy of them in
|
||||
a column is a copy that comes to disagree with the scene it describes. The columns
|
||||
`Clip` does have are the ones the SERVER needs to answer a question about a clip
|
||||
without parsing its leaves: which footage, which analysis, which blocks.
|
||||
"""
|
||||
import uuid
|
||||
|
||||
from django.db import models
|
||||
|
||||
|
||||
class Blob(models.Model):
|
||||
"""Bytes, named by the sha256 of themselves. The file is on disk under
|
||||
`BLOB_ROOT`; this row is the index and the size."""
|
||||
|
||||
digest = models.CharField(primary_key=True, max_length=64)
|
||||
media_type = models.CharField(max_length=100, default="application/octet-stream")
|
||||
size = models.BigIntegerField()
|
||||
created = models.DateTimeField(auto_now_add=True)
|
||||
|
||||
def __str__(self):
|
||||
return f"{self.digest[:12]}… {self.size}B {self.media_type}"
|
||||
|
||||
|
||||
class Footage(models.Model):
|
||||
"""Tier 3: the frames and audio of one extraction, immutable.
|
||||
|
||||
`digest` is over the ordered frame digests plus the audio's, so two
|
||||
extractions of the same clip at the same rate are one footage and the same
|
||||
analysis can be reused across both.
|
||||
|
||||
`feature_absence` is the manifest annotation step 8 introduced: known
|
||||
occlusion intervals, one-based and inclusive, expanded into presence tracks by
|
||||
the loader. An input format, not a control UI.
|
||||
"""
|
||||
|
||||
id = models.UUIDField(primary_key=True, default=uuid.uuid4, editable=False)
|
||||
digest = models.CharField(max_length=64, unique=True)
|
||||
label = models.CharField(max_length=200, blank=True)
|
||||
source = models.CharField(max_length=200, blank=True)
|
||||
fps = models.FloatField()
|
||||
frames = models.PositiveIntegerField()
|
||||
width = models.PositiveIntegerField()
|
||||
height = models.PositiveIntegerField()
|
||||
audio = models.ForeignKey(Blob, on_delete=models.PROTECT, related_name="audio_for")
|
||||
feature_absence = models.JSONField(default=dict, blank=True)
|
||||
created = models.DateTimeField(auto_now_add=True)
|
||||
|
||||
class Meta:
|
||||
ordering = ["-created"]
|
||||
|
||||
def __str__(self):
|
||||
return f"{self.label or self.source or self.id} ({self.frames}f @{self.fps})"
|
||||
|
||||
|
||||
class FootageFrame(models.Model):
|
||||
"""One source frame. A row rather than an entry in a JSON list, because a
|
||||
frame is a thing the server serves, and because a blob's references have to be
|
||||
countable before anything can be collected."""
|
||||
|
||||
footage = models.ForeignKey(Footage, on_delete=models.CASCADE, related_name="frame_set")
|
||||
index = models.PositiveIntegerField(help_text="0-based; source frame index + 1 is the PNG's name")
|
||||
blob = models.ForeignKey(Blob, on_delete=models.PROTECT, related_name="frame_for")
|
||||
|
||||
class Meta:
|
||||
ordering = ["index"]
|
||||
constraints = [
|
||||
models.UniqueConstraint(fields=["footage", "index"], name="one_blob_per_frame"),
|
||||
]
|
||||
|
||||
|
||||
class Analysis(models.Model):
|
||||
"""Tier 2: one detector, at one version, over one footage.
|
||||
|
||||
`key` is a content address over every input, and `descriptor` is the exact
|
||||
canonical text that key is the sha256 of — sent by the client and stored, not
|
||||
recomputed here. `clips/views.py` says why that is the honest arrangement: JS
|
||||
prints an integral double as `1` and Python as `1.0`, so a scheme where both
|
||||
sides re-render the numbers breaks on the first one of them.
|
||||
|
||||
`detector` and `version` are columns as well as descriptor fields so that the
|
||||
question "which model produced this take" is answerable in the admin and in a
|
||||
query, rather than only by parsing a hash's preimage.
|
||||
"""
|
||||
|
||||
key = models.CharField(primary_key=True, max_length=71)
|
||||
descriptor = models.TextField()
|
||||
detector = models.CharField(max_length=64)
|
||||
version = models.CharField(max_length=64)
|
||||
footage = models.ForeignKey(
|
||||
Footage, null=True, blank=True, on_delete=models.SET_NULL, related_name="analyses"
|
||||
)
|
||||
artifact = models.ForeignKey(
|
||||
Blob, null=True, blank=True, on_delete=models.SET_NULL, related_name="analysis_for",
|
||||
help_text="the dense landmark track, once bake A is uploaded",
|
||||
)
|
||||
created = models.DateTimeField(auto_now_add=True)
|
||||
|
||||
class Meta:
|
||||
verbose_name_plural = "analyses"
|
||||
|
||||
def __str__(self):
|
||||
return f"{self.detector} {self.version} → {self.key[7:19]}…"
|
||||
|
||||
|
||||
class Block(models.Model):
|
||||
"""Tier 2: one dense channel block.
|
||||
|
||||
Two hashes, and they are not the same hash. `key` is over the block's INPUTS,
|
||||
which is what lets a client ask for the block its current settings want before
|
||||
anything has computed it. `data.digest` is over the bytes. See clips/blobs.py.
|
||||
"""
|
||||
|
||||
key = models.CharField(primary_key=True, max_length=71)
|
||||
descriptor = models.TextField()
|
||||
role = models.CharField(max_length=32)
|
||||
analysis = models.ForeignKey(
|
||||
Analysis, null=True, blank=True, on_delete=models.SET_NULL, related_name="blocks"
|
||||
)
|
||||
data = models.ForeignKey(Blob, on_delete=models.PROTECT, related_name="block_data_for")
|
||||
state = models.ForeignKey(
|
||||
Blob, null=True, blank=True, on_delete=models.PROTECT, related_name="block_state_for",
|
||||
help_text="the per-track absence mask, when the take has one",
|
||||
)
|
||||
created = models.DateTimeField(auto_now_add=True)
|
||||
|
||||
def __str__(self):
|
||||
return f"{self.role} {self.key[7:19]}…"
|
||||
|
||||
|
||||
class Project(models.Model):
|
||||
"""Tier 1: the document's root.
|
||||
|
||||
`seq` is the monotonic project version docs/architecture.md asks for. Every
|
||||
write bumps it, and a client that sees `seq > local + 1` refetches — which is
|
||||
what makes staleness self-healing rather than permanent once there is a
|
||||
broadcast to miss.
|
||||
"""
|
||||
|
||||
id = models.UUIDField(primary_key=True, default=uuid.uuid4, editable=False)
|
||||
name = models.CharField(max_length=200, default="untitled")
|
||||
seq = models.PositiveBigIntegerField(default=0)
|
||||
palette = models.CharField(max_length=64, default="arthur/default")
|
||||
created = models.DateTimeField(auto_now_add=True)
|
||||
updated = models.DateTimeField(auto_now=True)
|
||||
|
||||
class Meta:
|
||||
ordering = ["-updated"]
|
||||
|
||||
def __str__(self):
|
||||
return f"{self.name} ({self.id})"
|
||||
|
||||
def bump(self):
|
||||
self.seq += 1
|
||||
self.save(update_fields=["seq", "updated"])
|
||||
return self.seq
|
||||
|
||||
|
||||
class Clip(models.Model):
|
||||
"""Tier 1: the unit of work, and the thing leaf paths are scoped by.
|
||||
|
||||
`cid` is what appears in `clip/<cid>/...`, so it is the clip's identity as far
|
||||
as addressing is concerned and it does not change.
|
||||
"""
|
||||
|
||||
project = models.ForeignKey(Project, on_delete=models.CASCADE, related_name="clips")
|
||||
cid = models.SlugField(max_length=64)
|
||||
name = models.CharField(max_length=200, blank=True)
|
||||
order = models.IntegerField(default=0)
|
||||
footage = models.ForeignKey(
|
||||
Footage, null=True, blank=True, on_delete=models.SET_NULL, related_name="clips"
|
||||
)
|
||||
analysis = models.ForeignKey(
|
||||
Analysis, null=True, blank=True, on_delete=models.SET_NULL, related_name="clips"
|
||||
)
|
||||
blocks = models.ManyToManyField(
|
||||
Block, blank=True, related_name="clips",
|
||||
help_text="the tier-2 blocks this clip's channels name",
|
||||
)
|
||||
|
||||
class Meta:
|
||||
ordering = ["order", "cid"]
|
||||
constraints = [
|
||||
models.UniqueConstraint(fields=["project", "cid"], name="one_cid_per_project"),
|
||||
]
|
||||
|
||||
def __str__(self):
|
||||
return f"{self.cid} of {self.project.name}"
|
||||
|
||||
|
||||
class Leaf(models.Model):
|
||||
"""Tier 1: one independently addressed, independently versioned piece of the
|
||||
document.
|
||||
|
||||
The value is transit-as-JSON in a JSONField, so the column holds JSON rather
|
||||
than a string containing JSON: the admin can read a leaf, and the field-wise
|
||||
merge of a channel leaf that docs/architecture.md describes as fifteen lines of
|
||||
Python is possible over it. `version` is the entity tag a conditional write
|
||||
compares — RFC 7232, not a bespoke invention.
|
||||
"""
|
||||
|
||||
project = models.ForeignKey(Project, on_delete=models.CASCADE, related_name="leaves")
|
||||
path = models.CharField(max_length=300)
|
||||
value = models.JSONField()
|
||||
version = models.PositiveBigIntegerField(default=1)
|
||||
updated = models.DateTimeField(auto_now=True)
|
||||
|
||||
class Meta:
|
||||
ordering = ["path"]
|
||||
constraints = [
|
||||
models.UniqueConstraint(fields=["project", "path"], name="one_leaf_per_path"),
|
||||
]
|
||||
|
||||
@property
|
||||
def etag(self):
|
||||
return f'"{self.version}"'
|
||||
|
||||
def __str__(self):
|
||||
return f"{self.path}@{self.version}"
|
||||
|
||||
|
||||
class Revision(models.Model):
|
||||
"""Tier 1: a snapshot of the authored layer, with a user and a summary.
|
||||
|
||||
ON AN EXPLICIT TRIGGER, not on every save. tl snapshots a small annotation
|
||||
layer; arthur's tier 1 will contain cel polygons, so a snapshot per save bloats
|
||||
the table — docs/architecture.md's "revisions need a coarser trigger". So this
|
||||
is written by `POST /api/projects/<id>/revisions`, which is a "mark version"
|
||||
button, and never by a save.
|
||||
"""
|
||||
|
||||
project = models.ForeignKey(Project, on_delete=models.CASCADE, related_name="revisions")
|
||||
seq = models.PositiveBigIntegerField()
|
||||
author = models.CharField(max_length=200, blank=True)
|
||||
summary = models.CharField(max_length=500, blank=True)
|
||||
document = models.JSONField(help_text="every leaf of the project, by path")
|
||||
created = models.DateTimeField(auto_now_add=True)
|
||||
|
||||
class Meta:
|
||||
ordering = ["-seq"]
|
||||
|
||||
def __str__(self):
|
||||
return f"{self.project.name} r{self.seq}: {self.summary}"
|
||||
Loading…
Add table
Add a link
Reference in a new issue