Serve the document from a Django backend, split into three tiers

Step 9. The tier split was the work; Django was the easy half.

Tier 1 — the authored scene — is the document, and it is addressed as
independently versioned leaves rather than saved whole, so one vertex drag
cannot clobber a collaborator's keying. `domain/leaf` is the document as
path -> value; `domain/wire` puts it on the wire as transit, because JSON
has neither integer map keys nor keywords and a save would quietly turn
`{0 v}` into `{"0" v}`.

Tier 2 — the dense channel blocks — is content-addressed by a hash over
every input, with the detector version inside every key through the
analysis the block descriptor names. `flow/address`'s `block-knobs` is the
invalidation table, and `address-test` does not trust it: it re-freezes the
take once per knob and asserts the biconditional, that a block's bytes
changed if and only if its key changed. That found `brow-pos` not depending
on `contour-avg` — the brow ring is smoothed, the raise is not.

Tier 3 — frames and audio — is served by the hash of its bytes out of the
same store. A manifest now names frames and carries a URL for each, so the
frame layout stopped being a shared secret between a shell script and a
ClojureScript namespace, and the `?v=` cache-buster went with it: a blob's
name is the hash of its contents, so a stale copy is not a thing that can
happen. The synthetic take's `audio.wav` moved to `static/arthur/` — an
asset the project owns, not an extraction that churns.

The server verifies rather than trusting a name it was handed: it
recomputes every key from the descriptor stored beside it, refuses an
analysis that declares no detector version, and refuses a document naming
blocks it does not hold. It hashes the descriptor TEXT, because JS prints
an integral double as `1` and Python as `1.0`, and a scheme where both ends
re-render the numbers disagrees on the first parameter that happens to be
whole.

Two loose ends from step 8 closed on the way. `pack` no longer takes a
`(track, frame)` predicate whose call sites each re-derived a feature from
an index — every track names the feature it follows, which deleted five
hand-maintained mappings. And `:dev-http` is gone: Django serves the page,
shadow-cljs only builds into the staticfiles tree.

227 CLJS tests, 31 Django tests, green.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
Olive Vaughn 2026-09-28 01:11:41 -04:00
parent b6517f837a
commit 9cd5243983
61 changed files with 4694 additions and 269 deletions

View file

@ -527,6 +527,102 @@ becomes addressable as a path; and two people keying different frames of one par
merge field-wise with no merge algorithm at all. `selectKeys` returns an array
today — change it before anything depends on the order.
### Serving tiers 2 and 3
Built at step 9. Above this point the tiers are a rule about what is allowed on
the wire; this is the shape that enforces it.
**One store for both, named by the sha256 of the bytes.** Once tier 3 is decoded
by the app rather than by a shell script it becomes the same kind of thing as tier
2 — a cache with a hash — so there is one place that writes bytes, one that reads
them, and one URL shape:
```
GET /blob/<sha256> raw bytes, Cache-Control: immutable
```
`immutable` is not optimism there, it is the definition: the name IS the hash of
the content, so a cached copy cannot be stale. That is what makes serving six
hundred frames out of it cheap enough to do on every load.
**Two kinds of hash, and they are not the same hash.** A blob is named by the hash
of its BYTES, which is what makes an identical frame in two extractions one file.
A derived thing — an analysis artifact, a dense block — is named by a hash over its
INPUTS, which is what lets a client ask for the block the current settings want
*before* anything has computed it, and what makes a stale bake unreachable rather
than wrong. So a `Block` row has both: `key` over the inputs, and a foreign key to
the blob whose digest is over the bytes. Conflating them would break the half of
addressing that answers questions about work not yet done.
```
POST /api/analyses {key, descriptor} idempotent
POST /api/blocks/missing {keys} -> {missing}
POST /api/blocks {key, descriptor, data, state}
GET /api/blocks/<key>
GET /api/footage/<id> the manifest, with a URL per frame
```
**The server verifies, rather than trusting a name it was handed.** It recomputes
`sha256(descriptor)` for every key and refuses a mismatch; it refuses an analysis
whose descriptor does not declare a detector and a version; it refuses a block
whose analysis it does not know; and it refuses a document naming blocks it does
not hold. The chain from a stored block to the model version that produced it
therefore cannot be broken by a client that skipped a step — which is what the
cache-key rule above actually requires, as opposed to recommends.
It hashes the descriptor TEXT rather than re-rendering it from parsed values, and
that is not a shortcut. JS prints an integral double as `1` and Python prints
`1.0`, so a scheme where both ends re-render the numbers disagrees on the first
parameter whose value happens to be whole — and the failure is an upload that 409s
with nothing wrong. The bytes are the contract; the schema on top of them is a
convention, and the two fields the server reads out of that schema are checked
separately.
**A manifest names frames; it does not locate them.** Until step 9 the client
fetched `manifest.json` and built `frames/0001.png` itself, which made the frame
layout a shared secret between a shell script and a ClojureScript namespace. The
manifest now carries a URL per frame, so the frames can live in the blob store —
or, when wasm-ffmpeg extraction arrives, be uploaded into the same store by the
app — and the client learns nothing new when that happens. The producer changes;
the shape does not.
**The document stores what a block IS, not what it holds.** A block's element type
is in its own descriptor, which is the only place it is written down: an
`Int16Array` and a `Float32Array` over the same bytes are both valid readings, and
only one of them is the block. That makes the descriptor load-bearing rather than
documentation, which is the right way round for the thing a key is the hash of.
### Leaf paths, as built
The list under **Make the merge unit small instead of clever** is the design; this
is what step 9 implemented, for the subset that exists:
```
clip/<cid>/name clip/<cid>/subject/<sid>
clip/<cid>/timing clip/<cid>/feature/<fid>
clip/<cid>/stage clip/<cid>/group/<gid>
clip/<cid>/source clip/<cid>/node/<nid>
clip/<cid>/measured/<nid> clip/<cid>/channel/<nid>/<prop>
```
Two departures from the design above, both because step 8 moved settings.
`params/:area` is **not** a leaf. That path came from a draft where params were one
blob per clip, and two people tuning teeth and eyes collided on every slider move.
Settings now live on the subject, the feature and the group, and a feature has
exactly one area — so the feature leaf already *is* the area-scoped leaf, and
splitting it again would separate a feature's params from its identity.
`measured/<nid>` is one leaf holding several channels, which contradicts "every
channel gets its own". `:head`'s measured channels are not authored: a freeze
writes them together and a re-freeze replaces them together, and `head-mode` reads
them to write `:channels`. A leaf per measured channel would offer a write nobody
can make.
A leaf path is "/"-delimited and an id is one segment of it, so a namespaced id —
`:eye-r/iris`, as drawn under **The node, decomposed** — is written `eye-r~iris`,
and `~` is then refused inside a name. That is the whole of the escaping.
## Collaboration
`../tl` already has the model, and it is the right one to copy: