53 KiB
arthur — structure
How the suite is laid out once it is a suite: ClojureScript and re-frame, a clip as the unit of work, keying as its own step, and painting as a peer of rotoscoping rather than a sketch bolted to the side.
docs/design.md says what the thing is and why. This says where the code goes.
docs/animation-model.md specifies the data both of them are about — nodes,
channels, symbols and time maps — and supersedes this document wherever the two
describe the same type.
The newer Lane Model takes precedence for occurrence ownership,
playback semantics, shared editing operations, and multi-view UX. It explicitly
allows replacing the current format without backward compatibility.
Nothing here revises an aesthetic decision; several things here split a
decision that is currently made in two places at once.
What is actually wrong with the current shape
The pure modules are fine. pipeline.js, mathutil.js, landmarks.js,
interior.js and raster.js are functions over data with the
reasoning written down next to them, and they port nearly verbatim.
app.js is the whole problem, and not because it is long. It is long because
six unrelated jobs are braided together in it:
- reading knob values out of the DOM (
opts()), - deriving everything from them (
rebuild,buildEyes,buildBrows,resolveTeeth), - resolving what is on screen at a frame (
plateIndex,perfIndex,leadIndex), - rasterising (
renderFrame,compositeRender), - driving the clock (
tick), - persisting (
saveCels,loadCels).
Every one of those is a different layer, and rebuild recomputes all of them
whenever any knob moves. That is affordable at one clip and 72 frames and is
not affordable at a project. The re-frame port is worth doing chiefly because
the subscription graph is the staged dataflow docs/design.md already
describes — "three artifacts, not two" is a dependency graph written in prose.
Three frame spaces, named
Most of the future bugs live here, so name them before anything else.
| Space | Symbol | Meaning |
|---|---|---|
| source frame | sf |
Index into the footage and the dense track. What detection produced. |
| clip frame | cf |
0-based within a clip. cf = sf - (:in clip). |
| sequence frame | qf |
Position on the timeline. qf = (:at clip) + cf. |
Nothing authored is ever stored in sequence-frame space. Keys, cels, overrides and kept-frame sets are all in clip-frame space, so a clip slides on the timeline without a single stored number changing. Exposure and lead are transforms within clip space:
Detection retains every source frame. Symbols carry their native fps; project fps selects the output grid without changing source data or audio speed. See Time selection for boundary sampling and frame units. Tracing references remain an independent selection.
(defn pose-frame [clip cf]
(-> cf (expose (:exposure clip)) (shift (:lead clip) (count-frames clip))))
which is the composition app.js currently performs in leadIndex and
perfIndex with the order implicit. Expose first, then lead: exposure floors
onto a grid and lead deliberately reads the future, so leading first and then
flooring would quietly discard the lead on most frames.
The entity model
project
├── palette authored ramp; indices, never RGB
├── character plate library, per-plate slots
├── footage[] frames + audio + manifest (fps, w, h)
├── analysis[] dense track, cached, keyed by footage+model+version
├── clip[] the unit of work
│ ├── source footage id, analysis id, in/out in sf
│ ├── exposure, lead timing, in cf
│ ├── scene node tree
│ ├── channels node+property -> keyframe stream, in cf
│ ├── overrides (node, property, cf) -> value, applied last
│ └── cels painted vector layers, keyed by chosen clip frames
└── sequence[] clip placements: {clip-id, at, in, out}
clip is the entity the current app has exactly one of and never names. Giving
it a name is most of the work: state in app.js is a clip with its analysis
inlined and its palette global.
Cel keys select where drawings begin and how long they hold. The source frames shown beneath a cel while tracing are chosen independently. Project fps controls which native frames can appear on the output grid.
Two things called "track"
docs/design.md says "dense track" for the landmark stream. A timeline also
wants tracks. Pick now: the landmark stream is analysis, a keyframe stream
is a channel, and the word track is reserved for a row on the timeline.
Renaming this later costs a day.
The node, decomposed
The prototype today gives a part a parent and a clip, and parent is doing
nothing except documenting intent — mouth_in is already in the same
head-local raster space as mouth, so composing its transform would be
composing identity. The moment a painted cel is attached to a head plate that
moves, transform composition becomes real, and the three ideas currently sharing
two fields have to come apart:
| Field | What it does | Wrong to conflate because |
|---|---|---|
:parent |
Transform composition. Child geometry is in parent's local space. | A node can be drawn over its parent without inheriting its motion. |
:stencil |
Colour-key clip: write only where the buffer already holds that index. | The iris is stencilled by the sclera and parented to the lid ring; those are different nodes. |
:z |
Draw order. | Order is authored per scene, not implied by the tree. |
A node is then:
{:id :eye-r/iris
:source {:kind :primitive ...} ; the five kinds from design.md
:parent :eye-r/lid
:stencil :eye-r/sclera
:z 42
:color :iris ; palette key, never a hex
:interp :hold}
:source carries the taxonomy docs/design.md already has — :plate,
:feature, :interior, :primitive, :scalar — plus :cel for painted
vectors and :group for a pure transform node. That is the decomposition that
makes painting a peer rather than an annex: paint.js is not a special case,
it is a source kind whose channels happen to be authored by hand instead of
measured. docs/design.md already promises "pluggable sources"; this is the
data shape that keeps the promise.
Each node also gets an optional local transform channel — translate, rotate,
scale, squash. That is the thing the current tool cannot express and that the
slot_mouth / scale / rot / squash fields in the prototype are reaching for.
The flow, in seven stages
The user-facing flow is frames -> analysis -> keying -> geometry -> palette.
Analysis is three stages, not one, and the split is the most load-bearing
decision in this document, because exactly one boundary is a cache
boundary.
| # | Stage | In | Out | Cost |
|---|---|---|---|---|
| 1 | ingest | video | footage: an H.264 proxy and raw stream, tracing stills, audio, manifest | minutes, in-app |
| 2 | detect | the raw stream, decoded one frame at a time | raw landmarks per frame | minutes, cached |
| 3 | measure | landmarks | anchor fit, residual, head-local rings, signals, interior pixels | seconds |
| 4 | condition | measurements | smoothed transforms and contours | milliseconds |
| 5 | key | conditioned signals + policy | channels: sparse keys, quantised holds, kept frames | milliseconds |
| 6 | resolve | channels + scene + exposure + lead + overrides | geometry per output frame | per frame |
| 7 | palette | geometry | indexed raster | per frame |
Stage 2 is the only artifact worth persisting and the only one worth a progress
bar. Stage 3 reads pixels — teeth extraction is here, and that is why it is
here rather than in keying: after stage 3, nothing downstream may look at a
source pixel. Stage 4 is separated from 3 only because contour avg and
anchor avg are knobs and the rest of stage 3 is not; splitting them means
dragging that slider does not re-run the interior extraction.
That last clause is the guarantee, and it is narrower than the table's ordering
looks. Stage 3 is not one pass that finishes before stage 4 begins. The anchor fit
is knob-free; conditioning smooths its four parameters; the head-local rings are
then measured through the conditioned transform — so anchor avg does re-run
the ring mapping, which is a few hundred frames of twenty points and free. What it
must not re-run is the part that reads a source pixel, and that part takes the
landmarks and the frames and never the transform, so it does not. Built as
measure/anchor → condition/anchor → measure/mouth → condition/contours,
and stage 7's eyes, brows and interior follow the same shape.
Where the current code lands
| Now | Stage |
|---|---|
stabilize |
3 measure |
eyeSignals, browSignals, pairIrises, pairBrows |
3 measure |
extractTeeth |
3 measure — it reads pixels |
smoothTransforms, smoothContours |
4 condition |
selectKeys, suggestPlateFrames |
5 key |
quantizeSnap, resolveBlink |
5 key — a held value is a key |
exposeIndex, shiftIndex, activeKey, heldFrame |
6 resolve |
toRasterRing, iris placement at ring slots, offsetRing |
6 resolve |
IndexedRaster, drawCel |
7 palette |
drawRegistered, posterizeInto |
reference only, off to the side |
Three things in app.js currently straddle a boundary, and each straddle is a
bug waiting for a bigger project:
resolveTeethextracts from pixels and applies contrast/dwell/blob knobs. Extraction is stage 3 and cached with the track; resolution is stage 5. Today, nudgingteeth dwellre-runs Otsu over every frame.buildEyesmeasures, quantises, and places the iris on the subsampled lid ring. Those are stages 3, 5 and 6. The placement argument indocs/design.md— that the iris is read off the already-smoothed ring — is a stage-6 fact and stays true; it just needs to be in stage 6.buildBrowslikewise: measure the height out of the ring (3), quantise it (5), put it back (6). The doc's warning about moving the brow twice is precisely a warning about these being one function.
What each knob invalidates
This table is the argument for the split, and it should be derivable from the sub graph rather than maintained by hand.
As built, half of it is, by a different route. flow/address/block-knobs is this
table for tier 2 — per block, the settings its bytes depend on — and
address-test asserts it by biconditional rather than deriving it, which is
stronger than a derivation would have been: a derived table is only as right as
the graph it reads. What is not written down anywhere is the rest of the column.
Stages 3 to 5 are one call chain in flow/take/measure rather than a chain of
subs, so there is no graph above ::base-scene to read a dependency off, and the
tier-1 half of a re-freeze is therefore whole-clip.
An earlier arrangement had a second table in domain/params, per feature area,
saying which areas a knob affected. It disagreed with the tested one on the first
entry where the two granularities part company and it had no caller; see that
namespace for why one checked table beats two.
| Knob | Invalidates from |
|---|---|
| model / footage | 2 detect |
teeth contrast, cavity erode, blob grow, tongue reject, prefer upper |
3 measure (mouth crop only) |
contour avg, anchor avg |
4 condition |
teeth dwell, blink cut/hold/dwell, gaze step/dwell, brow step/dwell, suggest tolerance |
5 key |
vertices, eye vertices, brow vertices, teeth vertices |
5 key |
exposure, mouth lead, kept-frame edits, overrides |
6 resolve |
lash line, iris size, pupil, brow weight |
6 resolve |
| palette edits, colour assignment | 7 palette |
Namespaces
src/arthur/
domain/ pure data, specs, ops. No re-frame, no DOM.
palette.cljs ramp; index lookup; the no-RGB rule as a spec
landmarks.cljs index tables (port verbatim)
ring.cljs ordered traversal: subsample, offset, simplicity
geom.cljs similarity fit, procrustes, moving average
node.cljs scene node: source, parent, stencil, z
timeline.cljs node tree: topo order, transform composition
channel.cljs keyframe stream: active-key-at, hold semantics
clip.cljs clip entity; the frame-space conversions
cel.cljs painted vector layers
timeline.cljs sequence: clip placement, qf <-> cf
flow/
ingest.cljs
detect.cljs the MediaPipe boundary, and the only one
measure/
anchor.cljs stabilise: procrustes, similarity, residual
mouth.cljs
eyes.cljs openness, gaze, iris pairing vote
brows.cljs raise/tilt, ring and end pairing votes
interior.cljs otsu, morphology, components, radial contour
condition.cljs
key.cljs velocity minima, dwell, quantise, blink, decimate
resolve.cljs scene eval at a frame; exposure, lead, overrides
raster.cljs indexed scanline fill, stencil, disc, rect
reference.cljs registered underlay, posterise
db.cljs app-db schema + spec
store.cljs handles for things too big for app-db
clock.cljs audio clock, outside app-db
events/ one ns per domain area
subs/ one ns per stage
ui/
shell.cljs
mode/ one ns per tool mode
panel/ strip, worksheet, readouts, palette, layers
canvas.cljs the one imperative sink
fx/ mediapipe, files, audio, persistence
Two rules about this tree. domain/ may not require flow/, and neither may
require re-frame. And flow/ namespaces take explicit arguments — no
namespace below subs/ ever calls subscribe.
app-db, and what is not in it
app-db holds authored data and ids. Nothing derived, and nothing large.
That sounds like ordinary hygiene and it is not: it is the precondition for two features that are otherwise unbuildable.
- Spec validation on every event. Worth having, affordable only over authored data.
- Cheap writes. Every edit
assoces into app-db, and every mounted layer-2 sub compares the result. Both are structural-sharing operations and stay cheap only while the map is small.
Undo is not on that list, and an earlier draft said it was. See Collaboration: snapshot-based undo is right single-player and wrong once two people edit.
So the dense track and the decoded frames live in arthur.store, a defonce
map of id to JS object, and app-db holds {:analysis/id "sha-..."}. Landmarks
stay as they arrive — arrays of {x, y, z}, or better, a flat Float32Array —
and are not converted to Clojure maps. 478 points × 600 frames is 286,800
maps, allocated for nothing: no sub ever needs to diff them, and every stage
that touches them walks all of them anyway.
The playhead, and where per-frame work goes
A reaction propagates downstream only when its own output value changes, not
when one of its inputs notifies. So with layer-2 extractors and layer-3
computations, a playhead tick costs one cheap extractor run per mounted layer-2
sub, each of which returns the same value for every subtree the tick did not
touch and therefore notifies nobody. ::measurements, ::channels and
::resolver do not re-run. The :<- graph is what buys that, and it buys it
whether or not the playhead is in app-db.
So the playhead lives in app-db, like everything else. An earlier draft of this document put it in a standalone ratom to avoid an invalidation storm that does not happen. Two independent reasons it belongs in db:
[:playback/seek ...]in the event log is how scrubbing becomes inspectable in re-frame-10x.- A collaborator's playhead is a feature. Putting it outside app-db puts it outside the machinery that shares it. See Collaboration.
What is true is narrower, and is about sub authoring and interceptors rather than about the playhead:
- Expensive work must not live in a layer-2 sub. A sub that derefs app-db directly re-runs on every db change whatever changed. That is the actual content of the folklore about re-frame and canvas.
- Do not put a per-frame event behind a global interceptor that walks the
whole db — a spec-validating
after, orstd-interceptors/debugand itsclojure.data/diff. At 30fps that is thirty full-db traversals a second. The transport event carries its own interceptor chain and is excluded from the global ones. - The render sink is not a Reagent component. Stage 7 writes bytes into a canvas from an rAF loop; it is not a view that re-renders.
Stage 6 is still arranged so the per-frame path is a lookup, not a computation:
;; recomputes when channels, scene, exposure, lead or overrides change
(rf/reg-sub ::resolver ...) ;; => (fn [cf] geometry), or an index into a bake
;; the rAF loop: reads, blits, dispatches nothing
(defn tick []
(let [cf (clock/clip-frame)]
(raster/paint! (@resolver cf))
(canvas/blit!)))
The sub produces a resolver; the loop applies it. A knob change costs one sub recomputation; a frame costs a lookup and a blit.
Events and subs
Events are named for intent, and carry the clip they apply to:
[:clip/set-exposure clip-id 2]
[:clip/toggle-kept clip-id cf]
[:keying/re-key clip-id] ; not a setter; policy unchanged, redo the work
[:keying/suggest clip-id]
[:paint/commit-shape clip-id node-id cel-frame pts]
[:node/set-parent clip-id node-id parent-id]
[:override/set clip-id node-id :pts cf value]
[:timeline/move-clip seq-id clip-id qf]
Setters are fine where the intent is the value — set-exposure is honest.
Where they differ, name the intent: :keying/re-key takes no value and is not
:clip/set-keys.
Subs mirror the stages one-to-one, so the graph and the table above are the same object:
::footage -> ::landmarks -> ::measurements -> ::conditioned -> ::channels -> ::resolved -> ::raster
Each is a layer-3 sub over the previous plus the parameters for its stage only.
That is what makes "drag exposure" recompute ::resolved and nothing above it.
Tool modes
A mode is a namespace, not a branch. Each exposes a map:
{:id :paint/pen
:cursor "crosshair"
:keymap {"Enter" [:paint/close-shape] "Escape" [:paint/abort]}
:on-pointer (fn [ev ctx] ...)
:overlay (fn [g ctx] ...) ; draws handles, never output pixels
:enter :exit (fn [ctx] ...)}
with two rules worth stating because the current PaintUI breaks both and gets
away with it at this size:
- In-progress gesture state is mode-local, in a ratom the mode owns — not in
app-db. An unclosed polygon must not appear in the undo history, and a drag
must produce one undo entry rather than one per pointermove. The mode
dispatches a single
:paint/commit-shapeon release. - Overlays draw to a separate canvas. Handles, vertex boxes and the green close-target are not indexed pixels and must never be in the raster. The current code composites them into the same context as the output, which is fine for a preview and lies to you the moment you want to judge the look.
The palette-index constraint and grid snapping stay where they are: they belong
to domain/cel and domain/palette, enforced at commit, so no tool can bypass
them.
Where things live
Superseded in detail by Serialization and Baking below; the local picture is:
| Thing | Where | Keyed by |
|---|---|---|
| project (tier 1) | server, plus a local autosave copy | project id, then leaf path |
| footage frames, audio (tier 3) | on disk / OPFS, untouched | content hash |
| analysis artifact, geometry bakes (tier 2) | OPFS, with IndexedDB for the index | content hash over every input |
| render output | never persisted | — |
The cache key on the analysis artifact must include the detector version. A model upgrade that silently reuses old landmarks presents as "the tool got worse" with no event to attach it to — which is the argument for content-addressing rather than version-numbering everything derived.
Baking
"No analysis during playback" is docs/design.md's three-artifact rule with a
number attached. The expensive stages are 2 (detect), 3 (measure) and, over a
long take, 4 (condition). Stages 5–7 are arithmetic over small arrays. So there
are two bakes, not one, and they answer different questions.
Bake A — the analysis artifact. Mandatory.
Stages 2–4 for one footage at one set of conditioning parameters. This is what must never run while the transport is moving.
| Field | Shape |
|---|---|
| landmarks | Float32Array[n_frames × 478 × 2] |
| transforms | Float32Array[n_frames × 4] — s, θ, tx, ty |
| residual | Float32Array[n_frames] |
| head-local rings | Int16Array per ring table, isotropic fixed point |
| signals | openness, gaze, brow ends, aperture — Float32Array[n_frames] each |
| mouth crops | Uint8Array, one small crop per frame |
The mouth crops are the non-obvious entry, and they are what makes remote work
possible at all. extractTeeth reads source pixels, so without them a
collaborator holding the analysis but not the source video cannot touch a
single teeth knob without decoding video again. These are RGBA crops: a 40×30
crop is 4.8KB raw, and current 200-pixel-wide crops can total over 12MB for a
take. The blob store compresses source/crops losslessly with zlib; block reads
return the original pixels. Run python manage.py compress_crop_blocks once to
convert existing raw crop blobs and remove their unreferenced copies.
Bake B — resolved geometry. For scale.
Stage 6 output per node. Not needed for one clip — resolving a frame is a key lookup and a transform — and needed the moment six roto tracks are live at 30fps.
Fixed topology is what makes this flat. Because every key of a part carries the same vertex count with the same vertex meanings, a node's whole bake is a rectangular array with no per-frame header and no indirection:
pts Int16Array[n_frames × n_verts × 2] raster space, already grid-snapped
state Uint8Array[n_frames] hidden flag + palette index
Frame f of a node is the subarray at f * n_verts * 2. 600 frames × 20 verts
is 48KB; a twelve-node clip about 600KB; six roto tracks about 3.5MB. So this is
the payoff on the aesthetic constraint rather than a cost of it — a variable
vertex count would force a per-frame offset table and a scan.
Rasters are not baked
Stage 7 is the cheapest stage and the one most worth keeping live: palette and colour assignment are the last decisions and the ones you want to change in motion. Baking to pixels would freeze exactly the loop that should be instant. Small raster thumbnails for the strip and the timeline are a separate artifact with a separate cache.
Unbaking is not an inverse
The bake never replaces its input. Authored state — channels, params, overrides — is always retained and always authoritative, so "unbake this roto and edit it" is a boolean about which side of the cache the renderer reads, not a computation. Instant by construction. Three rules keep it that way:
- No destructive bake. Nothing is discarded when something is baked.
- Hand corrections never go into the bake. They are
:overrides, applied at stage 6 after the bake is read, so a correction survives a rebake. This isdocs/design.md's "hand corrections must not live in the take", and it is the rule that stops baking from becoming a trap. - Bakes are content-addressed by a hash over every input that produced them: footage hash, detector version, and the parameters of each stage at or above the bake line. A stale bake is then unreachable rather than wrong, and a collaborator's bake is fetchable by the same key — see Collaboration.
Partial bakes
Bake per node, per frame range. Scrubbing into an unbaked range resolves on demand and fills the cache; a worker bakes ahead of the playhead. Show it: a bake-state bar under the timeline is an affordance every NLE has already taught people to read, and it is honest about what is ready.
Serialization, in three tiers
Cut by mutability and size. The cut is what makes collaboration and baking both tractable, because it decides what is allowed on the wire.
| Tier | What | Size | Synced | Undoable |
|---|---|---|---|---|
| 1 authored | palette, clips, scene graphs, channels, cels, overrides, sequences | KB | yes — it is the document | yes |
| 2 derived | analysis artifacts, geometry bakes, thumbnails | MB | no — content-addressed blobs, fetched | no |
| 3 source | frames, audio | MB–GB | no — immutable, by hash | no |
Tier 3 is produced by the app, not by extract.sh: wasm ffmpeg decodes the
clip to frames and audio in the browser, and they are uploaded and served as
content-addressed blobs. That makes tier 3 the same kind of thing as tier 2 — a
cache with a hash — and it removes the one step that currently needs a terminal.
Two consequences worth planning for: the extraction rate is recorded by the code
that chose it rather than by a manifest.json a human might edit, and offline
work needs the frames already fetched, so the local blob cache is what makes a
plane usable. Decoding server-side instead is the same data model with a
different worker; the format does not care which side runs it.
Only tier 1 is the document. A bake inside the shared document is a system that puts 48KB on the wire per vertex drag, and it is unnecessary, because tier 2 is a pure function of tiers 1 and 3.
Tier 1 stays small enough to read. The only vertex data in it is what a human placed by hand: cel polygons and overrides. Everything traced is tier 2.
Keys are a map, not a vector
;; not this
{:keys [{:f 0 :pts [...]} {:f 4 :pts [...]}]}
;; this
{:keys {0 {:pts [...]}, 4 {:pts [...]}}}
activeKey becomes a lookup in a sorted map rather than a scan; a single key
becomes addressable as a path; and two people keying different frames of one part
merge field-wise with no merge algorithm at all. selectKeys returns an array
today — change it before anything depends on the order.
Serving tiers 2 and 3
Built at step 9. Above this point the tiers are a rule about what is allowed on the wire; this is the shape that enforces it.
One store for both, named by the sha256 of the bytes. Once tier 3 is decoded by the app rather than by a shell script it becomes the same kind of thing as tier 2 — a cache with a hash — so there is one place that writes bytes, one that reads them, and one URL shape:
GET /blob/<sha256> raw bytes, Cache-Control: immutable
immutable is not optimism there, it is the definition: the name IS the hash of
the content, so a cached copy cannot be stale. That is what makes serving six
hundred frames out of it cheap enough to do on every load.
Two kinds of hash, and they are not the same hash. A blob is named by the hash
of its BYTES, which is what makes an identical frame in two extractions one file.
A derived thing — an analysis artifact, a dense block — is named by a hash over its
INPUTS, which is what lets a client ask for the block the current settings want
before anything has computed it, and what makes a stale bake unreachable rather
than wrong. So a Block row has both: key over the inputs, and a foreign key to
the blob whose digest is over the bytes. Conflating them would break the half of
addressing that answers questions about work not yet done.
POST /api/analyses {key, descriptor} idempotent
GET /api/analyses/<key> metadata + source block keys
PUT /api/analyses/<key> link dense landmarks, mask, crops
POST /api/blocks/missing {keys} -> {missing}
POST /api/blocks multipart: key, descriptor, data file, optional state file (JSON also accepted)
GET /api/blocks/<key>
GET /api/footage/<id> the manifest: video and stream URLs, audio, a URL per tracing still
GET /blob/<digest> immutable bytes, with byte ranges for video playback
POST /api/sources multipart video upload
POST /api/extractions idempotent decode job
GET /api/extractions/<key> job state and footage id
The server verifies, rather than trusting a name it was handed. It recomputes
sha256(descriptor) for every key and refuses a mismatch; it refuses an analysis
whose descriptor does not declare a detector and a version; it refuses a block
whose analysis it does not know; and it refuses a document naming blocks it does
not hold. The chain from a stored block to the model version that produced it
therefore cannot be broken by a client that skipped a step — which is what the
cache-key rule above actually requires, as opposed to recommends.
It hashes the descriptor TEXT rather than re-rendering it from parsed values, and
that is not a shortcut. JS prints an integral double as 1 and Python prints
1.0, so a scheme where both ends re-render the numbers disagrees on the first
parameter whose value happens to be whole — and the failure is an upload that 409s
with nothing wrong. The bytes are the contract; the schema on top of them is a
convention, and the two fields the server reads out of that schema are checked
separately.
A manifest names frames; it does not locate them. Until step 9 the client
fetched manifest.json and built frames/0001.png itself, which made the frame
layout a shared secret between a shell script and a ClojureScript namespace. The
manifest now carries every URL, so the bytes can live in the blob store — or be
produced by the app's video upload and server-side ffmpeg extraction. The upload
path needs no manifest.json file: the server builds the footage response from
its own records. The producer changes; the shape does not.
Tier 3 keeps a video, not a frame per file. Stage 1 used to decode a PNG per source frame: 112MB for 7.6 seconds at 1440x1920, and 1.1GB at the 900-frame limit, for pixels whose only consumer was a canvas MediaPipe read once. It now writes a browser-safe H.264 proxy — 6MB for the same take — and copies its coded frames into an Annex-B stream. WebCodecs decodes that stream in order, and the page gives each frame to MediaPipe in video running mode. The JPEG stills beside it are reference images for tracing; nothing measures them, so they are deliberately outside the footage digest and re-rendering them at another size does not invalidate an analysis.
The proxy is encoded without B-frames, so decode order matches presentation
order. The client checks that the stream has exactly the manifest's frame count
before detection. /blob/<digest> also answers byte ranges for ordinary video
playback; Django's FileResponse does no Range handling, so this is code we own.
The document stores what a block IS, not what it holds. A block's element type
is in its own descriptor, which is the only place it is written down: an
Int16Array and a Float32Array over the same bytes are both valid readings, and
only one of them is the block. That makes the descriptor load-bearing rather than
documentation, which is the right way round for the thing a key is the hash of.
Leaf paths, as built
The list under Make the merge unit small instead of clever is the design; this is what step 9 implemented, for the subset that exists:
clip/<cid>/name clip/<cid>/subject/<sid>
clip/<cid>/timing clip/<cid>/feature/<fid>
clip/<cid>/stage clip/<cid>/group/<gid>
clip/<cid>/source clip/<cid>/symbol/<sid>
clip/<cid>/symbol/<sid>/node/<nid>
clip/<cid>/symbol/<sid>/measured/<nid>
clip/<cid>/symbol/<sid>/channel/<nid>/<prop>
Settings live on subject, feature and group leaves. Each feature has one area, so these leaves give settings their own address without separating them from the identity they describe.
measured/<nid> holds several channels together. :head's measured channels are
written and replaced together by a freeze; head-mode reads them to write authored
:channels.
A leaf path is "/"-delimited and an id is one segment of it, so a namespaced id —
:eye-r/iris, as drawn under The node, decomposed — is written eye-r~iris,
and ~ is then refused inside a name. That is the whole of the escaping.
Collaboration
../tl already has the model, and it is the right one to copy:
- An authoritative server, and one broadcast source. Edits go over HTTP; the server persists and then broadcasts the delta to the room. The socket is read-only for document state. No OT, no CRDT.
- Presence gossiped between peers on the same socket, with the server doing exactly two things: handing out a connection id, and stamping the sender's server-side identity onto every message so nobody can post as somebody else.
- Parties as a pure function of the roster, so every peer computes the same answer and the server knows nothing about them.
- Revisions: one snapshot of the authored layer per save, with a user and a summary.
That maps onto the tiers exactly — the server stores tier 1 and snapshots tier 1, with tiers 2 and 3 as content-addressed blobs beside it.
As built
- Addresses.
/is the index of the projects you own or edit. A project is only ever at/p/<uuid>/<slug>: the id finds it, the slug is its name and follows a rename without a history entry. "new" makes the project on the server first and opens it; a built-in example opened in the editor is saved at once as a project of its own. There is no bare project. The address, the title and the socket follow[:project :id]through one global interceptor (events/collab). - Ownership. Every project has an owner (
Project.owner, not nullable) andeditors. Anyone with the link reads; the owner and editors write. Making a project needs you signed in. A reader can make a copy of their own./api/{me,login,signup,logout},/api/projects/<id>/editors[/<username>]. - Every edit saves. No save button. The same interceptor sees
:paint/revisionmove and saves; one request is in flight at a time, and an edit made meanwhile goes when it lands — so a drag reaches the room as fast as the round trip allows. A save sends only the leaves that differ from what was last synced, and a clean document sends nothing. Analyses and blocks already put on the server are not asked about again. - Saves are patches. With
base(the seq last caught up to) a save names only its changed and removed leaves. A named leaf somebody else changed afterbase, to something else, fails the whole save with 409. - The first write wins. A remote write to a leaf with a local change not
yet sent waits in the entry's
:behind; the next save, or a 409, puts theirs on screen over ours and says so. Ours stays in the undo list. - The socket (
/ws/projects/<uuid>, channels + daphne) carries presence, the delta each committed write broadcasts, andaccesswhen the editor list changes. A gap inseq, a welcome, or a 409 refetches the document. - Undo is per person, recorded in
events/editas leaf befores and afters (domain/history), and applied as an ordinary edit. A step undoes only if every leaf it touched still holds what it left there: somebody else's edit since refuses it rather than being undone with it. Edits to the same leaves within a second, each starting where the last left off, are one step. - Snapshots are named revisions (
/api/projects/<id>/revisions), with each clip's block keys. Restoring one is an ordinary write, broadcast like any.
Not yet: follow mode, frame/selection in presence, the advisory :editing
lease, the durable outbox.
Why this model, and not a CRDT
The usual reason to reach for Yjs or Automerge is automatic convergence without conflict dialogs. In this domain the primary contested data type is a vector path, and automatic convergence of a vector path is undesirable — two people dragging vertices on one polygon merge into a shape neither drew. So the CRDT's headline benefit is neutralised on exactly the data that needs help, and the parts where it would work fine (scalars, maps keyed by frame) are the parts that LWW already handles.
What a CRDT would additionally cost here is not the library, it is the second
source of truth: the canonical document moves into an opaque blob, and every
server-side thing that reads the document — revisions, collaborator lists,
thumbnail generation, the admin, any query — needs it materialised back out with
y-py or a Node sidecar. That is real operational weight bought for a feature
this domain does not want.
So: authoritative server, HTTP writes, LWW on addressed leaves, presence on the side. And writes stay on HTTP rather than moving to the socket, which is tl's choice and the right one: auth, idempotency, status codes, retries and conditional requests all come for free, and a dropped socket cannot lose a write. Latency is not the reason to move them, because latency is handled on the client (below) and never by waiting for a round trip.
The model is standard; only tl's plumbing is hand-rolled
Worth separating. Optimistic concurrency control over addressed resources with
an entity tag is RFC 7232 — ETag, If-Match, 412/409 — and it is what
databases and HTTP APIs have done for thirty years. It is not a bespoke
invention, and every HTTP client implements its half. The hand-rolled part of tl
is the plumbing around it, so the move is to keep the model and replace the
plumbing with the boring standard wherever one exists.
Four things to add
tl's implementation is missing two pieces that hand-rolled sync usually omits, and arthur needs two more that tl does not because its write rate is low.
-
Conditional writes.
PUT /project/:id/leaf/<path>withIf-Match: <etag>, answering409on a stale write. tl's PUT replaces, so the loser's work disappears silently. For a knob that is survivable; for a painted cel it is the class of bug that ends trust in a tool. The 409 body carries the current value so the client can offer keep-mine / take-theirs. -
A monotonic project version, and gap detection. The broadcast is fire-and-forget, so a peer that misses one delta — queue pressure, a reconnect gap — is silently stale forever. Every delta carries
seq; a client that seesseq > local + 1refetches. This is what makes staleness self-healing instead of permanent, and it is about twenty lines. -
Gesture coalescing, client side. This is the one real load difference between tl and arthur: tl's user annotates a clip every few seconds, arthur's user drags a vertex. A drag is one write on release, not one per pointermove. With the gesture held in mode-local state and committed once (see Tool modes), arthur's write rate is comparable to tl's and the model holds unchanged. Without it, no sync model survives.
-
An outbox, so nothing ever waits on the network. "No lag" is satisfied entirely on the client: apply locally, enqueue, reconcile on ack.
{:pending {"clip/7/cel/bg/12" {:value … :etag … :tries 0}}}With one rule that is easy to get wrong: while a leaf has a pending local write, incoming broadcasts for that leaf are ignored. Otherwise a remote delta lands mid-gesture and the shape snaps back and forth between the two values until the ack arrives.
Two operational notes
- tl's
ROOMSis process-local, as its own comment says. That is a single-worker ceiling, not a design flaw, and it has to move to Redis before a second worker. - Revisions need a coarser trigger than every save. tl snapshots a small annotation layer; arthur's tier 1 contains cel polygons, so a snapshot per save will bloat the table. Snapshot on an explicit "mark version", or time-boxed.
When this model is wrong, and what replaces it
State the tripwire rather than trusting the judgement: if 409s on cel leaves become common in real use, the merge unit is too coarse. The answer then is a finer unit, not a CRDT, and there is a planned progression:
| Step | Merge unit | Gets you |
|---|---|---|
| now | the cel (cel/:node/:frame) |
two people on different frames |
| next | the layer | two people on different layers of one cel |
| last | the stroke — an append-only list of immutable stroke records | two people painting one layer at once |
The third step is worth seeing now because it is the escape hatch that keeps this decision from being a dead end. Once a drawing is a list of immutable records rather than a mutable point array, concurrent painting merges by construction — two people appending different records cannot conflict — and that is a domain-specific version of what a CRDT would have given, without the second source of truth. It is also a plausible thing to want for undo granularity anyway.
Make the merge unit small instead of clever
Last-writer-wins clobbers only when its unit is too big. tl can save a scene; arthur cannot save a project, because one vertex drag would then clobber a collaborator's keying. The fix is addressing, not an algorithm:
palette
lane/:sid
clip/:cid/timing clip rate
clip/:cid/subject/:sid tracked subject and settings
clip/:cid/feature/:fid tracked feature and settings
clip/:cid/group/:gid shared settings for an eye pair
clip/:cid/symbol/:sid frame count, palette
clip/:cid/symbol/:sid/node/:nid one node: parent, stencil, z, time
clip/:cid/symbol/:sid/channel/:nid/:prop
clip/:cid/symbol/:sid/measured/:nid
clip/:cid/symbol/:sid/cel/:nid/:frame
clip/:cid/symbol/:sid/overrides/:nid/:prop
Each feature and node has its own leaf, so tuning separate features and adding
separate nodes use separate addresses. Fractional :z keeps draw order on the
node leaf.
Each path is a leaf: independently addressed, independently versioned, LWW
with an If-Match on its version. The boundaries are chosen so the things people
actually do simultaneously land on different leaves — two people painting
different cels, or keying different parts, never meet. Within a channel leaf the
server merges field-wise by frame, which is the return on keys-as-a-map and
about fifteen lines of Python.
Where it stays honest about conflict
| Data | Concurrency | Resolution |
|---|---|---|
| palette, clip params, knobs | scalar | LWW on the leaf. Two people tuning one clip will fight — that is a real thing to surface, not to merge. |
| channels | map by frame | field-wise union, LWW per frame |
| kept-frame set | add / remove | store {frame → bool}, not a set, so removals merge instead of vanishing |
cel layer order, :z |
list insert | fractional indices, not integers |
| cel polygon points | sequence | LWW on the whole layer, plus an advisory lease |
| parentage | tree | LWW per node, with a cycle check that rejects the losing move |
| playhead, selection, tool | ephemeral | presence only — never in the document, never in history |
Two of those are cheap now and expensive later.
Fractional indices for anything ordered. Integer :z and integer layer
positions do not survive concurrent insertion: you get duplicates and gaps. A
fractional key between neighbours makes concurrent inserts commute. Adopt it in
domain/node and domain/cel from the first commit.
Do not merge one polygon between two people. Two simultaneous edits to the same path have no meaningful merge, and the result is a shape nobody drew. The layer is the unit, LWW, and presence shows who is on it before the collision. A coarse merge unit that is visible beats a fine one that invents geometry.
Presence carries more than a cursor
tl's roster entry, extended with what arthur's views need:
{:cid "…" :user "…" :party nil :joinable true
:clip cid :frame cf :playing? false :rate 1.0
:node :mouth :tool :paint/pen
:editing "clip/7/cel/bg/12"} ; advisory lease on a leaf
:editing is what makes leaf-level LWW liveable — the collision becomes visible
before it happens. Advisory only: no server enforcement, so there is no lock to
leak when a tab closes.
Follow mode, and why frames never go on the wire
tl's parties become arthur's follow mode nearly unchanged: mirror :clip,
:frame, :playing? and the plate mode.
Broadcast transport state, never frames. On play, send
{:clip :frame :playing? :rate} once; each peer's own audio element then becomes
its own clock and the picture follows it exactly as it does locally. Per-frame
messages would put thirty packets a second on the wire to reproduce something
every peer can compute. A peer whose bake is cold shows a cold indicator and
catches up rather than stalling the party.
Going offline
Arthur is offline-capable almost by accident, and it is worth noticing why: tier
1 is small enough to hold entirely in app-db, tier 2 is a local content-addressed
cache, and tier 3 is a directory of PNGs that extract.sh put on your own disk.
So when the network drops, everything needed to scrub, key, paint, resolve and
render is already local. Nothing about playback or editing touches the server.
What stops is only the three things that are inherently remote: your writes stop acking, others' deltas stop arriving, and presence goes dark.
Two preconditions, or the above is a lie:
- The outbox must be durable. In memory, a closed tab loses the session. It goes in IndexedDB, written in the same task as the optimistic local apply — a few hundred microseconds, invisible, and the difference between "offline is fine" and "offline is a trap".
- Vendor MediaPipe's wasm.
README.mdnotes thatface_landmarker.taskis local but the wasm bundle is fetched from jsdelivr on first use. That makes stage 2 the one thing that silently requires a network, and it will be discovered on a plane. Vendor it or precache it in a service worker.
Reconnecting
Refetch the document first, then drain the outbox — in that order, so conflicts
are detected against current state rather than against a stale etag. Since the
document is kilobytes, a full refetch is cheaper than reasoning about a long
seq gap.
Then each pending write lands in one of four cases, and the design goal is that only the last one reaches a human:
| Case | Resolution |
|---|---|
| nobody touched your leaves | etags still match, outbox drains, converged. The common case. |
| a channel leaf collided | auto-merges field-wise by frame — union the frames, LWW per frame. No human. |
| a scalar leaf collided (knob, timing) | a compact list with one "keep all mine / take all theirs", because a knob is one number and the value is visible in the output anyway |
| a cel leaf collided | see below — never a dialog |
The channel row is the third time keys-as-a-map pays for itself, and it is the reason a long offline session usually reconciles with no interaction at all.
A colliding cel forks into a layer
Do not ask an artist to choose between two drawings in a modal. Keep both: their cel stays where it is, and yours is added on top as a layer named for the session that made it. Nothing is lost, nothing is silently clobbered, and the conflict is resolved by looking at it and deleting a layer — which is an operation the paint tool already has, and which is what an artist would do anyway.
This is the general principle for this whole class of conflict: when the merge unit is coarse, resolve by stacking rather than by choosing. A stack is reversible and legible; a choice made in a dialog against two thumbnails is neither.
The awkward cases, named
- A write to a leaf whose parent was deleted — a node or clip removed while
you were away. The server answers
404; the client parks it in an orphaned-edits tray rather than dropping it. Rare, and the alternative is silent loss. - A very long offline session against a busy project. Merging may be the wrong frame entirely. Offer the escape hatch revisions already make possible: save the offline session as a revision, take theirs, and reconcile by hand from two named versions. Better than a hundred-row conflict list.
- Two peers both run detection offline. Both compute the same analysis artifact and both try to upload it. Content addressing makes that idempotent — same hash, same bytes — so let it race rather than coordinating a claim.
Undo is per-user — and this revises an earlier claim
An earlier section of this document said undo was free if app-db held only
authored data, via day8.re-frame/undo. That is true single-player and wrong the
moment two people edit: a snapshot of app-db reverts their work along with yours.
So undo is a per-user stack of inverse leaf writes, applied as ordinary edits. Your undo writes a leaf back to the value you last saw, and it conflicts with a concurrent editor in the same visible way any other write does. Keeping app-db small is still right — for validation cost and allocation — but it is no longer the undo story.
The 30fps budget
At 320×200 stage 7 is not the problem. An even-odd scanline fill over 64,000 pixels with a dozen mostly-small parts is well under 200,000 byte writes a frame: six million a second at 30fps. Six roto tracks off baked geometry adds a handful of subarray reads.
Two things in the current code will miss 30fps, and neither is the rasteriser:
posterizeIntodoes agetImageDataevery frame. A GPU→CPU readback stalls the pipeline, and it is followed by 64,000 linear scans over the palette to find the nearest colour — around a million distance computations a frame. Both are avoidable: posterise once per source frame into a cachedUint8Array(it is a pure function of the frame, the transform and the palette), and replace the nearest-colour scan with a 5-5-5 RGB lookup table — 32KB, built once per palette.drawRegistereddraws a full-resolution photo every frame. Correct as a drawing reference on a held frame; not something to do thirty times a second. Pre-scale each source frame to raster size once, at load.
The rule both are instances of: nothing that reads source pixels may run on the transport path. That is the same line the bake boundary draws, which is why drawing it once, in stage 3, pays twice.
Porting order
Tests first, and not as discipline — as an oracle.
- Port
synth.jsandselftest.jsfirst, tocljs.test. The synthetic track is the only fixture with ground truth, and the 105 assertions encode invariants that are invisible to inspection: the ring-simplicity check, the bowtie, the iris-pairing vote fed a deliberately swapped track. - Run both implementations on the same synthetic track and diff numerically
while porting each pure module.
fitSimilarityandprocrustesMeanshould agree to 1e-9; a disagreement is a port bug, not float noise. - Port
landmarks,geom,ring,raster,take— mechanical, and they are where the invariant comments live. Carry the comments across verbatim. They are the most valuable text in the repo and every one of them is a bug that already happened. - Port
flow/measure/*by lifting out ofpipeline.js, then splitconditionandkeyout of it. - Build
db,store,clock, and stage 6 + 7 against a single hardcoded clip. At this point the synthetic take should render, with no UI beyond a canvas and a play button. - Only then the UI, the modes, the timeline.
- Add the ring-simplicity and no-RGB checks as specs, so they fire on authored data too and not only in tests.
Three things are cheap in step 3 and expensive after step 6, so they go in early
even though nothing needs them yet: keys as a map keyed by frame, fractional
indices on :z and cel layer order, and leaf addressing for tier 1 — the
paths under Collaboration, used as the shape of app-db even while single-player.
None of them cost anything on day one and all three are retrofits that touch
every namespace.
The wiring cross-check in selftest.js — every el('id') must exist in
index.html — becomes unnecessary and should be deleted rather than ported. It
is a test for a failure mode that Reagent does not have.
Decisions I would defer
- Whether
rasterstays JS. 320×200 is 64,000 pixels and CLJS over aUint8Arraywith no seq allocation handles it comfortably; port it and measure. Do not port it into idiomaticmap/reduce. - Interpolation.
:holdis the only interp the aesthetic wants, and the channel model should still carry:interpper node, because a parented transform on a painted cel is the one place a tween might be right. Do not implement it until something needs it. - Multi-character scenes.
characteris in the entity model as one field and should stay a stub until there are two. - Audio per clip vs per sequence. One master track is right until it is not.
- Whether bake B exists at all in v1. One clip does not need it; the flat typed-array layout should be designed now and built when the second roto track appears.
- Sharing tier 2. Content addressing means a collaborator can fetch a bake instead of recomputing, and also means nothing breaks if they do not. Ship the cache-miss path first and treat blob sharing as an optimisation — the analysis artifact is the only one where it clearly pays, because it is the only one that costs minutes.