Step 9. The tier split was the work; Django was the easy half.
Tier 1 — the authored scene — is the document, and it is addressed as
independently versioned leaves rather than saved whole, so one vertex drag
cannot clobber a collaborator's keying. `domain/leaf` is the document as
path -> value; `domain/wire` puts it on the wire as transit, because JSON
has neither integer map keys nor keywords and a save would quietly turn
`{0 v}` into `{"0" v}`.
Tier 2 — the dense channel blocks — is content-addressed by a hash over
every input, with the detector version inside every key through the
analysis the block descriptor names. `flow/address`'s `block-knobs` is the
invalidation table, and `address-test` does not trust it: it re-freezes the
take once per knob and asserts the biconditional, that a block's bytes
changed if and only if its key changed. That found `brow-pos` not depending
on `contour-avg` — the brow ring is smoothed, the raise is not.
Tier 3 — frames and audio — is served by the hash of its bytes out of the
same store. A manifest now names frames and carries a URL for each, so the
frame layout stopped being a shared secret between a shell script and a
ClojureScript namespace, and the `?v=` cache-buster went with it: a blob's
name is the hash of its contents, so a stale copy is not a thing that can
happen. The synthetic take's `audio.wav` moved to `static/arthur/` — an
asset the project owns, not an extraction that churns.
The server verifies rather than trusting a name it was handed: it
recomputes every key from the descriptor stored beside it, refuses an
analysis that declares no detector version, and refuses a document naming
blocks it does not hold. It hashes the descriptor TEXT, because JS prints
an integral double as `1` and Python as `1.0`, and a scheme where both ends
re-render the numbers disagrees on the first parameter that happens to be
whole.
Two loose ends from step 8 closed on the way. `pack` no longer takes a
`(track, frame)` predicate whose call sites each re-derived a feature from
an index — every track names the feature it follows, which deleted five
hand-maintained mappings. And `:dev-http` is gone: Django serves the page,
shadow-cljs only builds into the staticfiles tree.
227 CLJS tests, 31 Django tests, green.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
104 lines
3.8 KiB
Python
104 lines
3.8 KiB
Python
"""The content-addressed blob store: tiers 2 and 3 on disk.
|
|
|
|
One store for both, and docs/architecture.md says why in a sentence: once tier 3
|
|
is decoded by the app rather than by a shell script, frames and audio become "the
|
|
same kind of thing as tier 2 — a cache with a hash". So there is one place that
|
|
writes bytes, one that reads them, and one URL shape for both.
|
|
|
|
TWO KINDS OF HASH, AND THEY ARE NOT THE SAME HASH. A blob is named by the sha256
|
|
of its BYTES: that is what makes identical frames in two extractions one file. A
|
|
derived thing — an analysis artifact, a dense block — is named by a sha256 over
|
|
its INPUTS, which is what lets the client ask for the block the current settings
|
|
want before anything has computed it. So `Block.key` is an input hash and
|
|
`Block.data.digest` is a byte hash, and conflating them would break the half of
|
|
addressing that answers questions about work not yet done.
|
|
"""
|
|
import hashlib
|
|
import os
|
|
from pathlib import Path
|
|
|
|
from django.conf import settings
|
|
|
|
CHUNK = 1 << 20
|
|
|
|
|
|
def digest_bytes(data: bytes) -> str:
|
|
return hashlib.sha256(data).hexdigest()
|
|
|
|
|
|
def digest_file(path: Path) -> str:
|
|
h = hashlib.sha256()
|
|
with open(path, "rb") as fh:
|
|
while chunk := fh.read(CHUNK):
|
|
h.update(chunk)
|
|
return h.hexdigest()
|
|
|
|
|
|
def path_for(digest: str) -> Path:
|
|
"""Where a blob lives.
|
|
|
|
Fanned out two levels, so that a take's worth of frames does not put a hundred
|
|
thousand entries in one directory — which is slow on every filesystem and
|
|
unusable on some.
|
|
"""
|
|
if len(digest) != 64 or any(c not in "0123456789abcdef" for c in digest):
|
|
raise ValueError(f"not a sha256: {digest!r}")
|
|
return Path(settings.BLOB_ROOT) / digest[:2] / digest[2:4] / digest
|
|
|
|
|
|
def write(data: bytes) -> tuple[str, int]:
|
|
"""Store bytes, return (digest, size). Writing the same bytes twice is a
|
|
no-op, which is what content addressing is for."""
|
|
digest = digest_bytes(data)
|
|
dest = path_for(digest)
|
|
if not dest.exists():
|
|
dest.parent.mkdir(parents=True, exist_ok=True)
|
|
tmp = dest.with_suffix(".part")
|
|
with open(tmp, "wb") as fh:
|
|
fh.write(data)
|
|
os.replace(tmp, dest)
|
|
return digest, len(data)
|
|
|
|
|
|
def adopt(source: Path) -> tuple[str, int]:
|
|
"""Store a file already on disk, by hard link where the filesystem allows it.
|
|
|
|
112MB of PNGs is a normal extraction and copying them into a second place in
|
|
the tree for no reason is not. A hard link is exact — the blob is immutable, so
|
|
two names for one inode is the whole of what is wanted — and a copy is the
|
|
fallback when `extract.sh` wrote to another volume.
|
|
"""
|
|
digest = digest_file(source)
|
|
dest = path_for(digest)
|
|
size = source.stat().st_size
|
|
if not dest.exists():
|
|
dest.parent.mkdir(parents=True, exist_ok=True)
|
|
try:
|
|
os.link(source, dest)
|
|
except OSError:
|
|
tmp = dest.with_suffix(".part")
|
|
with open(source, "rb") as src, open(tmp, "wb") as out:
|
|
while chunk := src.read(CHUNK):
|
|
out.write(chunk)
|
|
os.replace(tmp, dest)
|
|
return digest, size
|
|
|
|
|
|
def read(digest: str) -> bytes:
|
|
with open(path_for(digest), "rb") as fh:
|
|
return fh.read()
|
|
|
|
|
|
def png_size(path: Path) -> tuple[int, int]:
|
|
"""A PNG's dimensions, out of its IHDR.
|
|
|
|
Twenty-four bytes rather than a dependency. The footage's width and height are
|
|
manifest data — docs/architecture.md's entity model puts them there — and
|
|
Pillow to read two integers out of a header that has held them in the same
|
|
place since 1996 is not a trade worth making.
|
|
"""
|
|
with open(path, "rb") as fh:
|
|
head = fh.read(24)
|
|
if head[:8] != b"\x89PNG\r\n\x1a\n" or head[12:16] != b"IHDR":
|
|
raise ValueError(f"{path} is not a PNG")
|
|
return int.from_bytes(head[16:20], "big"), int.from_bytes(head[20:24], "big")
|