arthur/clips/blobs.py
Olive Vaughn 9cd5243983 Serve the document from a Django backend, split into three tiers
Step 9. The tier split was the work; Django was the easy half.

Tier 1 — the authored scene — is the document, and it is addressed as
independently versioned leaves rather than saved whole, so one vertex drag
cannot clobber a collaborator's keying. `domain/leaf` is the document as
path -> value; `domain/wire` puts it on the wire as transit, because JSON
has neither integer map keys nor keywords and a save would quietly turn
`{0 v}` into `{"0" v}`.

Tier 2 — the dense channel blocks — is content-addressed by a hash over
every input, with the detector version inside every key through the
analysis the block descriptor names. `flow/address`'s `block-knobs` is the
invalidation table, and `address-test` does not trust it: it re-freezes the
take once per knob and asserts the biconditional, that a block's bytes
changed if and only if its key changed. That found `brow-pos` not depending
on `contour-avg` — the brow ring is smoothed, the raise is not.

Tier 3 — frames and audio — is served by the hash of its bytes out of the
same store. A manifest now names frames and carries a URL for each, so the
frame layout stopped being a shared secret between a shell script and a
ClojureScript namespace, and the `?v=` cache-buster went with it: a blob's
name is the hash of its contents, so a stale copy is not a thing that can
happen. The synthetic take's `audio.wav` moved to `static/arthur/` — an
asset the project owns, not an extraction that churns.

The server verifies rather than trusting a name it was handed: it
recomputes every key from the descriptor stored beside it, refuses an
analysis that declares no detector version, and refuses a document naming
blocks it does not hold. It hashes the descriptor TEXT, because JS prints
an integral double as `1` and Python as `1.0`, and a scheme where both ends
re-render the numbers disagrees on the first parameter that happens to be
whole.

Two loose ends from step 8 closed on the way. `pack` no longer takes a
`(track, frame)` predicate whose call sites each re-derived a feature from
an index — every track names the feature it follows, which deleted five
hand-maintained mappings. And `:dev-http` is gone: Django serves the page,
shadow-cljs only builds into the staticfiles tree.

227 CLJS tests, 31 Django tests, green.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-28 01:11:41 -04:00

104 lines
3.8 KiB
Python

"""The content-addressed blob store: tiers 2 and 3 on disk.
One store for both, and docs/architecture.md says why in a sentence: once tier 3
is decoded by the app rather than by a shell script, frames and audio become "the
same kind of thing as tier 2 — a cache with a hash". So there is one place that
writes bytes, one that reads them, and one URL shape for both.
TWO KINDS OF HASH, AND THEY ARE NOT THE SAME HASH. A blob is named by the sha256
of its BYTES: that is what makes identical frames in two extractions one file. A
derived thing — an analysis artifact, a dense block — is named by a sha256 over
its INPUTS, which is what lets the client ask for the block the current settings
want before anything has computed it. So `Block.key` is an input hash and
`Block.data.digest` is a byte hash, and conflating them would break the half of
addressing that answers questions about work not yet done.
"""
import hashlib
import os
from pathlib import Path
from django.conf import settings
CHUNK = 1 << 20
def digest_bytes(data: bytes) -> str:
return hashlib.sha256(data).hexdigest()
def digest_file(path: Path) -> str:
h = hashlib.sha256()
with open(path, "rb") as fh:
while chunk := fh.read(CHUNK):
h.update(chunk)
return h.hexdigest()
def path_for(digest: str) -> Path:
"""Where a blob lives.
Fanned out two levels, so that a take's worth of frames does not put a hundred
thousand entries in one directory — which is slow on every filesystem and
unusable on some.
"""
if len(digest) != 64 or any(c not in "0123456789abcdef" for c in digest):
raise ValueError(f"not a sha256: {digest!r}")
return Path(settings.BLOB_ROOT) / digest[:2] / digest[2:4] / digest
def write(data: bytes) -> tuple[str, int]:
"""Store bytes, return (digest, size). Writing the same bytes twice is a
no-op, which is what content addressing is for."""
digest = digest_bytes(data)
dest = path_for(digest)
if not dest.exists():
dest.parent.mkdir(parents=True, exist_ok=True)
tmp = dest.with_suffix(".part")
with open(tmp, "wb") as fh:
fh.write(data)
os.replace(tmp, dest)
return digest, len(data)
def adopt(source: Path) -> tuple[str, int]:
"""Store a file already on disk, by hard link where the filesystem allows it.
112MB of PNGs is a normal extraction and copying them into a second place in
the tree for no reason is not. A hard link is exact — the blob is immutable, so
two names for one inode is the whole of what is wanted — and a copy is the
fallback when `extract.sh` wrote to another volume.
"""
digest = digest_file(source)
dest = path_for(digest)
size = source.stat().st_size
if not dest.exists():
dest.parent.mkdir(parents=True, exist_ok=True)
try:
os.link(source, dest)
except OSError:
tmp = dest.with_suffix(".part")
with open(source, "rb") as src, open(tmp, "wb") as out:
while chunk := src.read(CHUNK):
out.write(chunk)
os.replace(tmp, dest)
return digest, size
def read(digest: str) -> bytes:
with open(path_for(digest), "rb") as fh:
return fh.read()
def png_size(path: Path) -> tuple[int, int]:
"""A PNG's dimensions, out of its IHDR.
Twenty-four bytes rather than a dependency. The footage's width and height are
manifest data — docs/architecture.md's entity model puts them there — and
Pillow to read two integers out of a header that has held them in the same
place since 1996 is not a trade worth making.
"""
with open(path, "rb") as fh:
head = fh.read(24)
if head[:8] != b"\x89PNG\r\n\x1a\n" or head[12:16] != b"IHDR":
raise ValueError(f"{path} is not a PNG")
return int.from_bytes(head[16:20], "big"), int.from_bytes(head[20:24], "big")