There were two answers in the tree to "which stored bytes stop being valid when
this knob moves", and only one of them was checked.
`flow/address/block-knobs` is per block and asserted by biconditional —
`address-test` re-freezes the take once per knob and requires that the bytes
changed if and only if the key did. `domain/params`'s `:affects` was per area,
had no caller but a test asserting it returned what it was written as, and was
already wrong in both directions on the one entry where the two granularities
disagree: `:aperture-cut` claimed `#{:mouth}`, where it reaches no block, and
omitted the teeth, whose contour bytes it genuinely moves by gating
`condition/interior`'s smoothing. `:blink-cut` claimed `#{:eye}` and reaches no
block either, because a blink is `[:vis]` keys in tier 1.
So `:affects` and `affected-areas` are gone, and `address/knob-roles` is the
derived inverse of the table that is asserted — which is what a parameter panel
actually wants to ask. A knob absent from it invalidates no block, and that is
an answer rather than a gap.
Two new assertions keep the derivation from rotting at either edge: every role
in the table is reachable from some knob, and every knob a block declares is one
the registry defines. The second closes a real hole — `block-descriptor` checks
only that a knob was PASSED, and the freeze's `merge take/knobs` makes that true
of anything spelled like a keyword, so a typo in `block-knobs` would have named a
setting no slider can move.
228 CLJS tests, green.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
49 KiB
arthur — structure
How the suite is laid out once it is a suite: ClojureScript and re-frame, a clip as the unit of work, keying as its own step, and painting as a peer of rotoscoping rather than a sketch bolted to the side.
docs/design.md says what the thing is and why. This says where the code goes.
docs/animation-model.md specifies the data both of them are about — nodes,
channels, symbols and time maps — and supersedes this document wherever the two
describe the same type.
Nothing here revises an aesthetic decision; several things here split a
decision that is currently made in two places at once.
What is actually wrong with the current shape
The pure modules are fine. pipeline.js, mathutil.js, landmarks.js,
interior.js and raster.js are functions over data with the
reasoning written down next to them, and they port nearly verbatim.
app.js is the whole problem, and not because it is long. It is long because
six unrelated jobs are braided together in it:
- reading knob values out of the DOM (
opts()), - deriving everything from them (
rebuild,buildEyes,buildBrows,resolveTeeth), - resolving what is on screen at a frame (
plateIndex,perfIndex,leadIndex), - rasterising (
renderFrame,compositeRender), - driving the clock (
tick), - persisting (
saveCels,loadCels).
Every one of those is a different layer, and rebuild recomputes all of them
whenever any knob moves. That is affordable at one clip and 72 frames and is
not affordable at a project. The re-frame port is worth doing chiefly because
the subscription graph is the staged dataflow docs/design.md already
describes — "three artifacts, not two" is a dependency graph written in prose.
Three frame spaces, named
Most of the future bugs live here, so name them before anything else.
| Space | Symbol | Meaning |
|---|---|---|
| source frame | sf |
Index into the footage and the dense track. What detection produced. |
| clip frame | cf |
0-based within a clip. cf = sf - (:in clip). |
| sequence frame | qf |
Position on the timeline. qf = (:at clip) + cf. |
Nothing authored is ever stored in sequence-frame space. Keys, cels, overrides and kept-frame sets are all in clip-frame space, so a clip slides on the timeline without a single stored number changing. Exposure and lead are transforms within clip space:
Detection retains every source frame. A chosen picture fps samples the frozen roto in clip time; it changes neither source-frame count nor the audio clock. The set of source frames an artist uses as cel tracing references is another selection, independent of the picture fps.
(defn pose-frame [clip cf]
(-> cf (expose (:exposure clip)) (shift (:lead clip) (count-frames clip))))
which is the composition app.js currently performs in leadIndex and
perfIndex with the order implicit. Expose first, then lead: exposure floors
onto a grid and lead deliberately reads the future, so leading first and then
flooring would quietly discard the lead on most frames.
The entity model
project
├── palette authored ramp; indices, never RGB
├── character plate library, per-plate slots
├── footage[] frames + audio + manifest (fps, w, h)
├── analysis[] dense track, cached, keyed by footage+model+version
├── clip[] the unit of work
│ ├── source footage id, analysis id, in/out in sf
│ ├── exposure, lead timing, in cf
│ ├── scene node tree
│ ├── channels node+property -> keyframe stream, in cf
│ ├── overrides (node, property, cf) -> value, applied last
│ └── cels painted vector layers, keyed by chosen clip frames
└── sequence[] clip placements: {clip-id, at, in, out}
clip is the entity the current app has exactly one of and never names. Giving
it a name is most of the work: state in app.js is a clip with its analysis
inlined and its palette global.
Cel keys select where drawings begin and how long they hold. The source frames shown beneath a cel while tracing are chosen independently, and picture fps only controls which analyzed pose the finished roto displays at a given time.
Two things called "track"
docs/design.md says "dense track" for the landmark stream. A timeline also
wants tracks. Pick now: the landmark stream is analysis, a keyframe stream
is a channel, and the word track is reserved for a row on the timeline.
Renaming this later costs a day.
The node, decomposed
The prototype today gives a part a parent and a clip, and parent is doing
nothing except documenting intent — mouth_in is already in the same
head-local raster space as mouth, so composing its transform would be
composing identity. The moment a painted cel is attached to a head plate that
moves, transform composition becomes real, and the three ideas currently sharing
two fields have to come apart:
| Field | What it does | Wrong to conflate because |
|---|---|---|
:parent |
Transform composition. Child geometry is in parent's local space. | A node can be drawn over its parent without inheriting its motion. |
:stencil |
Colour-key clip: write only where the buffer already holds that index. | The iris is stencilled by the sclera and parented to the lid ring; those are different nodes. |
:z |
Draw order. | Order is authored per scene, not implied by the tree. |
A node is then:
{:id :eye-r/iris
:source {:kind :primitive ...} ; the five kinds from design.md
:parent :eye-r/lid
:stencil :eye-r/sclera
:z 42
:color :iris ; palette key, never a hex
:interp :hold}
:source carries the taxonomy docs/design.md already has — :plate,
:feature, :interior, :primitive, :scalar — plus :cel for painted
vectors and :group for a pure transform node. That is the decomposition that
makes painting a peer rather than an annex: paint.js is not a special case,
it is a source kind whose channels happen to be authored by hand instead of
measured. docs/design.md already promises "pluggable sources"; this is the
data shape that keeps the promise.
Each node also gets an optional local transform channel — translate, rotate,
scale, squash. That is the thing the current tool cannot express and that the
slot_mouth / scale / rot / squash fields in the prototype are reaching for.
The flow, in seven stages
The user-facing flow is frames -> analysis -> keying -> geometry -> palette.
Analysis is three stages, not one, and the split is the most load-bearing
decision in this document, because exactly one boundary is a cache
boundary.
| # | Stage | In | Out | Cost |
|---|---|---|---|---|
| 1 | ingest | video | footage: frames, audio, manifest | minutes, in-app |
| 2 | detect | footage | raw landmarks per frame | minutes, cached |
| 3 | measure | landmarks | anchor fit, residual, head-local rings, signals, interior pixels | seconds |
| 4 | condition | measurements | smoothed transforms and contours | milliseconds |
| 5 | key | conditioned signals + policy | channels: sparse keys, quantised holds, kept frames | milliseconds |
| 6 | resolve | channels + scene + exposure + lead + overrides | geometry per output frame | per frame |
| 7 | palette | geometry | indexed raster | per frame |
Stage 2 is the only artifact worth persisting and the only one worth a progress
bar. Stage 3 reads pixels — teeth extraction is here, and that is why it is
here rather than in keying: after stage 3, nothing downstream may look at a
source pixel. Stage 4 is separated from 3 only because contour avg and
anchor avg are knobs and the rest of stage 3 is not; splitting them means
dragging that slider does not re-run the interior extraction.
That last clause is the guarantee, and it is narrower than the table's ordering
looks. Stage 3 is not one pass that finishes before stage 4 begins. The anchor fit
is knob-free; conditioning smooths its four parameters; the head-local rings are
then measured through the conditioned transform — so anchor avg does re-run
the ring mapping, which is a few hundred frames of twenty points and free. What it
must not re-run is the part that reads a source pixel, and that part takes the
landmarks and the frames and never the transform, so it does not. Built as
measure/anchor → condition/anchor → measure/mouth → condition/contours,
and stage 7's eyes, brows and interior follow the same shape.
Where the current code lands
| Now | Stage |
|---|---|
stabilize |
3 measure |
eyeSignals, browSignals, pairIrises, pairBrows |
3 measure |
extractTeeth |
3 measure — it reads pixels |
smoothTransforms, smoothContours |
4 condition |
selectKeys, suggestPlateFrames |
5 key |
quantizeSnap, resolveBlink |
5 key — a held value is a key |
exposeIndex, shiftIndex, activeKey, heldFrame |
6 resolve |
toRasterRing, iris placement at ring slots, offsetRing |
6 resolve |
IndexedRaster, drawCel |
7 palette |
drawRegistered, posterizeInto |
reference only, off to the side |
Three things in app.js currently straddle a boundary, and each straddle is a
bug waiting for a bigger project:
resolveTeethextracts from pixels and applies contrast/dwell/blob knobs. Extraction is stage 3 and cached with the track; resolution is stage 5. Today, nudgingteeth dwellre-runs Otsu over every frame.buildEyesmeasures, quantises, and places the iris on the subsampled lid ring. Those are stages 3, 5 and 6. The placement argument indocs/design.md— that the iris is read off the already-smoothed ring — is a stage-6 fact and stays true; it just needs to be in stage 6.buildBrowslikewise: measure the height out of the ring (3), quantise it (5), put it back (6). The doc's warning about moving the brow twice is precisely a warning about these being one function.
What each knob invalidates
This table is the argument for the split, and it should be derivable from the sub graph rather than maintained by hand.
As built, half of it is, by a different route. flow/address/block-knobs is this
table for tier 2 — per block, the settings its bytes depend on — and
address-test asserts it by biconditional rather than deriving it, which is
stronger than a derivation would have been: a derived table is only as right as
the graph it reads. What is not written down anywhere is the rest of the column.
Stages 3 to 5 are one call chain in flow/take/measure rather than a chain of
subs, so there is no graph above ::base-scene to read a dependency off, and the
tier-1 half of a re-freeze is therefore whole-clip.
An earlier arrangement had a second table in domain/params, per feature area,
saying which areas a knob affected. It disagreed with the tested one on the first
entry where the two granularities part company and it had no caller; see that
namespace for why one checked table beats two.
| Knob | Invalidates from |
|---|---|
| model / footage | 2 detect |
teeth contrast, cavity erode, blob grow, tongue reject, prefer upper |
3 measure (mouth crop only) |
contour avg, anchor avg |
4 condition |
teeth dwell, blink cut/hold/dwell, gaze step/dwell, brow step/dwell, suggest tolerance |
5 key |
vertices, eye vertices, brow vertices, teeth vertices |
5 key |
exposure, mouth lead, kept-frame edits, overrides |
6 resolve |
lash line, iris size, pupil, brow weight |
6 resolve |
| palette edits, colour assignment | 7 palette |
Namespaces
src/arthur/
domain/ pure data, specs, ops. No re-frame, no DOM.
palette.cljs ramp; index lookup; the no-RGB rule as a spec
landmarks.cljs index tables (port verbatim)
ring.cljs ordered traversal: subsample, offset, simplicity
geom.cljs similarity fit, procrustes, moving average
node.cljs scene node: source, parent, stencil, z
scene.cljs node tree: topo order, transform composition
channel.cljs keyframe stream: active-key-at, hold semantics
clip.cljs clip entity; the frame-space conversions
cel.cljs painted vector layers
timeline.cljs sequence: clip placement, qf <-> cf
flow/
ingest.cljs
detect.cljs the MediaPipe boundary, and the only one
measure/
anchor.cljs stabilise: procrustes, similarity, residual
mouth.cljs
eyes.cljs openness, gaze, iris pairing vote
brows.cljs raise/tilt, ring and end pairing votes
interior.cljs otsu, morphology, components, radial contour
condition.cljs
key.cljs velocity minima, dwell, quantise, blink, decimate
resolve.cljs scene eval at a frame; exposure, lead, overrides
raster.cljs indexed scanline fill, stencil, disc, rect
reference.cljs registered underlay, posterise
db.cljs app-db schema + spec
store.cljs handles for things too big for app-db
clock.cljs audio clock, outside app-db
events/ one ns per domain area
subs/ one ns per stage
ui/
shell.cljs
mode/ one ns per tool mode
panel/ strip, worksheet, readouts, palette, layers
canvas.cljs the one imperative sink
fx/ mediapipe, files, audio, persistence
Two rules about this tree. domain/ may not require flow/, and neither may
require re-frame. And flow/ namespaces take explicit arguments — no
namespace below subs/ ever calls subscribe.
app-db, and what is not in it
app-db holds authored data and ids. Nothing derived, and nothing large.
That sounds like ordinary hygiene and it is not: it is the precondition for two features that are otherwise unbuildable.
- Spec validation on every event. Worth having, affordable only over authored data.
- Cheap writes. Every edit
assoces into app-db, and every mounted layer-2 sub compares the result. Both are structural-sharing operations and stay cheap only while the map is small.
Undo is not on that list, and an earlier draft said it was. See Collaboration: snapshot-based undo is right single-player and wrong once two people edit.
So the dense track and the decoded frames live in arthur.store, a defonce
map of id to JS object, and app-db holds {:analysis/id "sha-..."}. Landmarks
stay as they arrive — arrays of {x, y, z}, or better, a flat Float32Array —
and are not converted to Clojure maps. 478 points × 600 frames is 286,800
maps, allocated for nothing: no sub ever needs to diff them, and every stage
that touches them walks all of them anyway.
The playhead, and where per-frame work goes
A reaction propagates downstream only when its own output value changes, not
when one of its inputs notifies. So with layer-2 extractors and layer-3
computations, a playhead tick costs one cheap extractor run per mounted layer-2
sub, each of which returns the same value for every subtree the tick did not
touch and therefore notifies nobody. ::measurements, ::channels and
::resolver do not re-run. The :<- graph is what buys that, and it buys it
whether or not the playhead is in app-db.
So the playhead lives in app-db, like everything else. An earlier draft of this document put it in a standalone ratom to avoid an invalidation storm that does not happen. Two independent reasons it belongs in db:
[:playback/seek ...]in the event log is how scrubbing becomes inspectable in re-frame-10x.- A collaborator's playhead is a feature. Putting it outside app-db puts it outside the machinery that shares it. See Collaboration.
What is true is narrower, and is about sub authoring and interceptors rather than about the playhead:
- Expensive work must not live in a layer-2 sub. A sub that derefs app-db directly re-runs on every db change whatever changed. That is the actual content of the folklore about re-frame and canvas.
- Do not put a per-frame event behind a global interceptor that walks the
whole db — a spec-validating
after, orstd-interceptors/debugand itsclojure.data/diff. At 30fps that is thirty full-db traversals a second. The transport event carries its own interceptor chain and is excluded from the global ones. - The render sink is not a Reagent component. Stage 7 writes bytes into a canvas from an rAF loop; it is not a view that re-renders.
Stage 6 is still arranged so the per-frame path is a lookup, not a computation:
;; recomputes when channels, scene, exposure, lead or overrides change
(rf/reg-sub ::resolver ...) ;; => (fn [cf] geometry), or an index into a bake
;; the rAF loop: reads, blits, dispatches nothing
(defn tick []
(let [cf (clock/clip-frame)]
(raster/paint! (@resolver cf))
(canvas/blit!)))
The sub produces a resolver; the loop applies it. A knob change costs one sub recomputation; a frame costs a lookup and a blit.
Events and subs
Events are named for intent, and carry the clip they apply to:
[:clip/set-exposure clip-id 2]
[:clip/toggle-kept clip-id cf]
[:keying/re-key clip-id] ; not a setter; policy unchanged, redo the work
[:keying/suggest clip-id]
[:paint/commit-shape clip-id node-id cel-frame pts]
[:node/set-parent clip-id node-id parent-id]
[:override/set clip-id node-id :pts cf value]
[:timeline/move-clip seq-id clip-id qf]
Setters are fine where the intent is the value — set-exposure is honest.
Where they differ, name the intent: :keying/re-key takes no value and is not
:clip/set-keys.
Subs mirror the stages one-to-one, so the graph and the table above are the same object:
::footage -> ::landmarks -> ::measurements -> ::conditioned -> ::channels -> ::resolved -> ::raster
Each is a layer-3 sub over the previous plus the parameters for its stage only.
That is what makes "drag exposure" recompute ::resolved and nothing above it.
Tool modes
A mode is a namespace, not a branch. Each exposes a map:
{:id :paint/pen
:cursor "crosshair"
:keymap {"Enter" [:paint/close-shape] "Escape" [:paint/abort]}
:on-pointer (fn [ev ctx] ...)
:overlay (fn [g ctx] ...) ; draws handles, never output pixels
:enter :exit (fn [ctx] ...)}
with two rules worth stating because the current PaintUI breaks both and gets
away with it at this size:
- In-progress gesture state is mode-local, in a ratom the mode owns — not in
app-db. An unclosed polygon must not appear in the undo history, and a drag
must produce one undo entry rather than one per pointermove. The mode
dispatches a single
:paint/commit-shapeon release. - Overlays draw to a separate canvas. Handles, vertex boxes and the green close-target are not indexed pixels and must never be in the raster. The current code composites them into the same context as the output, which is fine for a preview and lies to you the moment you want to judge the look.
The palette-index constraint and grid snapping stay where they are: they belong
to domain/cel and domain/palette, enforced at commit, so no tool can bypass
them.
Where things live
Superseded in detail by Serialization and Baking below; the local picture is:
| Thing | Where | Keyed by |
|---|---|---|
| project (tier 1) | server, plus a local autosave copy | project id, then leaf path |
| footage frames, audio (tier 3) | on disk / OPFS, untouched | content hash |
| analysis artifact, geometry bakes (tier 2) | OPFS, with IndexedDB for the index | content hash over every input |
| render output | never persisted | — |
The cache key on the analysis artifact must include the detector version. A model upgrade that silently reuses old landmarks presents as "the tool got worse" with no event to attach it to — which is the argument for content-addressing rather than version-numbering everything derived.
Baking
"No analysis during playback" is docs/design.md's three-artifact rule with a
number attached. The expensive stages are 2 (detect), 3 (measure) and, over a
long take, 4 (condition). Stages 5–7 are arithmetic over small arrays. So there
are two bakes, not one, and they answer different questions.
Bake A — the analysis artifact. Mandatory.
Stages 2–4 for one footage at one set of conditioning parameters. This is what must never run while the transport is moving.
| Field | Shape |
|---|---|
| landmarks | Float32Array[n_frames × 478 × 2] |
| transforms | Float32Array[n_frames × 4] — s, θ, tx, ty |
| residual | Float32Array[n_frames] |
| head-local rings | Int16Array per ring table, isotropic fixed point |
| signals | openness, gaze, brow ends, aperture — Float32Array[n_frames] each |
| mouth crops | Uint8Array, one small crop per frame |
The mouth crops are the non-obvious entry, and they are what makes remote work
possible at all. extractTeeth reads source pixels, so without them a
collaborator holding the analysis but not the 600 source PNGs cannot touch a
single teeth knob. A 40×30 crop is about 1.2KB; a 600-frame take is under a
megabyte against hundreds for the footage.
Bake B — resolved geometry. For scale.
Stage 6 output per node. Not needed for one clip — resolving a frame is a key lookup and a transform — and needed the moment six roto tracks are live at 30fps.
Fixed topology is what makes this flat. Because every key of a part carries the same vertex count with the same vertex meanings, a node's whole bake is a rectangular array with no per-frame header and no indirection:
pts Int16Array[n_frames × n_verts × 2] raster space, already grid-snapped
state Uint8Array[n_frames] hidden flag + palette index
Frame f of a node is the subarray at f * n_verts * 2. 600 frames × 20 verts
is 48KB; a twelve-node clip about 600KB; six roto tracks about 3.5MB. So this is
the payoff on the aesthetic constraint rather than a cost of it — a variable
vertex count would force a per-frame offset table and a scan.
Rasters are not baked
Stage 7 is the cheapest stage and the one most worth keeping live: palette and colour assignment are the last decisions and the ones you want to change in motion. Baking to pixels would freeze exactly the loop that should be instant. Small raster thumbnails for the strip and the timeline are a separate artifact with a separate cache.
Unbaking is not an inverse
The bake never replaces its input. Authored state — channels, params, overrides — is always retained and always authoritative, so "unbake this roto and edit it" is a boolean about which side of the cache the renderer reads, not a computation. Instant by construction. Three rules keep it that way:
- No destructive bake. Nothing is discarded when something is baked.
- Hand corrections never go into the bake. They are
:overrides, applied at stage 6 after the bake is read, so a correction survives a rebake. This isdocs/design.md's "hand corrections must not live in the take", and it is the rule that stops baking from becoming a trap. - Bakes are content-addressed by a hash over every input that produced them: footage hash, detector version, and the parameters of each stage at or above the bake line. A stale bake is then unreachable rather than wrong, and a collaborator's bake is fetchable by the same key — see Collaboration.
Partial bakes
Bake per node, per frame range. Scrubbing into an unbaked range resolves on demand and fills the cache; a worker bakes ahead of the playhead. Show it: a bake-state bar under the timeline is an affordance every NLE has already taught people to read, and it is honest about what is ready.
Serialization, in three tiers
Cut by mutability and size. The cut is what makes collaboration and baking both tractable, because it decides what is allowed on the wire.
| Tier | What | Size | Synced | Undoable |
|---|---|---|---|---|
| 1 authored | palette, clips, scene graphs, channels, cels, overrides, sequences | KB | yes — it is the document | yes |
| 2 derived | analysis artifacts, geometry bakes, thumbnails | MB | no — content-addressed blobs, fetched | no |
| 3 source | frames, audio | MB–GB | no — immutable, by hash | no |
Tier 3 is produced by the app, not by extract.sh: wasm ffmpeg decodes the
clip to frames and audio in the browser, and they are uploaded and served as
content-addressed blobs. That makes tier 3 the same kind of thing as tier 2 — a
cache with a hash — and it removes the one step that currently needs a terminal.
Two consequences worth planning for: the extraction rate is recorded by the code
that chose it rather than by a manifest.json a human might edit, and offline
work needs the frames already fetched, so the local blob cache is what makes a
plane usable. Decoding server-side instead is the same data model with a
different worker; the format does not care which side runs it.
Only tier 1 is the document. A bake inside the shared document is a system that puts 48KB on the wire per vertex drag, and it is unnecessary, because tier 2 is a pure function of tiers 1 and 3.
Tier 1 stays small enough to read. The only vertex data in it is what a human placed by hand: cel polygons and overrides. Everything traced is tier 2.
Keys are a map, not a vector
;; not this
{:keys [{:f 0 :pts [...]} {:f 4 :pts [...]}]}
;; this
{:keys {0 {:pts [...]}, 4 {:pts [...]}}}
activeKey becomes a lookup in a sorted map rather than a scan; a single key
becomes addressable as a path; and two people keying different frames of one part
merge field-wise with no merge algorithm at all. selectKeys returns an array
today — change it before anything depends on the order.
Serving tiers 2 and 3
Built at step 9. Above this point the tiers are a rule about what is allowed on the wire; this is the shape that enforces it.
One store for both, named by the sha256 of the bytes. Once tier 3 is decoded by the app rather than by a shell script it becomes the same kind of thing as tier 2 — a cache with a hash — so there is one place that writes bytes, one that reads them, and one URL shape:
GET /blob/<sha256> raw bytes, Cache-Control: immutable
immutable is not optimism there, it is the definition: the name IS the hash of
the content, so a cached copy cannot be stale. That is what makes serving six
hundred frames out of it cheap enough to do on every load.
Two kinds of hash, and they are not the same hash. A blob is named by the hash
of its BYTES, which is what makes an identical frame in two extractions one file.
A derived thing — an analysis artifact, a dense block — is named by a hash over its
INPUTS, which is what lets a client ask for the block the current settings want
before anything has computed it, and what makes a stale bake unreachable rather
than wrong. So a Block row has both: key over the inputs, and a foreign key to
the blob whose digest is over the bytes. Conflating them would break the half of
addressing that answers questions about work not yet done.
POST /api/analyses {key, descriptor} idempotent
POST /api/blocks/missing {keys} -> {missing}
POST /api/blocks {key, descriptor, data, state}
GET /api/blocks/<key>
GET /api/footage/<id> the manifest, with a URL per frame
The server verifies, rather than trusting a name it was handed. It recomputes
sha256(descriptor) for every key and refuses a mismatch; it refuses an analysis
whose descriptor does not declare a detector and a version; it refuses a block
whose analysis it does not know; and it refuses a document naming blocks it does
not hold. The chain from a stored block to the model version that produced it
therefore cannot be broken by a client that skipped a step — which is what the
cache-key rule above actually requires, as opposed to recommends.
It hashes the descriptor TEXT rather than re-rendering it from parsed values, and
that is not a shortcut. JS prints an integral double as 1 and Python prints
1.0, so a scheme where both ends re-render the numbers disagrees on the first
parameter whose value happens to be whole — and the failure is an upload that 409s
with nothing wrong. The bytes are the contract; the schema on top of them is a
convention, and the two fields the server reads out of that schema are checked
separately.
A manifest names frames; it does not locate them. Until step 9 the client
fetched manifest.json and built frames/0001.png itself, which made the frame
layout a shared secret between a shell script and a ClojureScript namespace. The
manifest now carries a URL per frame, so the frames can live in the blob store —
or, when wasm-ffmpeg extraction arrives, be uploaded into the same store by the
app — and the client learns nothing new when that happens. The producer changes;
the shape does not.
The document stores what a block IS, not what it holds. A block's element type
is in its own descriptor, which is the only place it is written down: an
Int16Array and a Float32Array over the same bytes are both valid readings, and
only one of them is the block. That makes the descriptor load-bearing rather than
documentation, which is the right way round for the thing a key is the hash of.
Leaf paths, as built
The list under Make the merge unit small instead of clever is the design; this is what step 9 implemented, for the subset that exists:
clip/<cid>/name clip/<cid>/subject/<sid>
clip/<cid>/timing clip/<cid>/feature/<fid>
clip/<cid>/stage clip/<cid>/group/<gid>
clip/<cid>/source clip/<cid>/node/<nid>
clip/<cid>/measured/<nid> clip/<cid>/channel/<nid>/<prop>
Two departures from the design above, both because step 8 moved settings.
params/:area is not a leaf. That path came from a draft where params were one
blob per clip, and two people tuning teeth and eyes collided on every slider move.
Settings now live on the subject, the feature and the group, and a feature has
exactly one area — so the feature leaf already is the area-scoped leaf, and
splitting it again would separate a feature's params from its identity.
measured/<nid> is one leaf holding several channels, which contradicts "every
channel gets its own". :head's measured channels are not authored: a freeze
writes them together and a re-freeze replaces them together, and head-mode reads
them to write :channels. A leaf per measured channel would offer a write nobody
can make.
A leaf path is "/"-delimited and an id is one segment of it, so a namespaced id —
:eye-r/iris, as drawn under The node, decomposed — is written eye-r~iris,
and ~ is then refused inside a name. That is the whole of the escaping.
Collaboration
../tl already has the model, and it is the right one to copy:
- An authoritative server, and one broadcast source. Edits go over HTTP; the server persists and then broadcasts the delta to the room. The socket is read-only for document state. No OT, no CRDT.
- Presence gossiped between peers on the same socket, with the server doing exactly two things: handing out a connection id, and stamping the sender's server-side identity onto every message so nobody can post as somebody else.
- Parties as a pure function of the roster, so every peer computes the same answer and the server knows nothing about them.
- Revisions: one snapshot of the authored layer per save, with a user and a summary.
That maps onto the tiers exactly — the server stores tier 1 and snapshots tier 1, with tiers 2 and 3 as content-addressed blobs beside it.
Why this model, and not a CRDT
The usual reason to reach for Yjs or Automerge is automatic convergence without conflict dialogs. In this domain the primary contested data type is a vector path, and automatic convergence of a vector path is undesirable — two people dragging vertices on one polygon merge into a shape neither drew. So the CRDT's headline benefit is neutralised on exactly the data that needs help, and the parts where it would work fine (scalars, maps keyed by frame) are the parts that LWW already handles.
What a CRDT would additionally cost here is not the library, it is the second
source of truth: the canonical document moves into an opaque blob, and every
server-side thing that reads the document — revisions, collaborator lists,
thumbnail generation, the admin, any query — needs it materialised back out with
y-py or a Node sidecar. That is real operational weight bought for a feature
this domain does not want.
So: authoritative server, HTTP writes, LWW on addressed leaves, presence on the side. And writes stay on HTTP rather than moving to the socket, which is tl's choice and the right one: auth, idempotency, status codes, retries and conditional requests all come for free, and a dropped socket cannot lose a write. Latency is not the reason to move them, because latency is handled on the client (below) and never by waiting for a round trip.
The model is standard; only tl's plumbing is hand-rolled
Worth separating. Optimistic concurrency control over addressed resources with
an entity tag is RFC 7232 — ETag, If-Match, 412/409 — and it is what
databases and HTTP APIs have done for thirty years. It is not a bespoke
invention, and every HTTP client implements its half. The hand-rolled part of tl
is the plumbing around it, so the move is to keep the model and replace the
plumbing with the boring standard wherever one exists.
Four things to add
tl's implementation is missing two pieces that hand-rolled sync usually omits, and arthur needs two more that tl does not because its write rate is low.
-
Conditional writes.
PUT /project/:id/leaf/<path>withIf-Match: <etag>, answering409on a stale write. tl's PUT replaces, so the loser's work disappears silently. For a knob that is survivable; for a painted cel it is the class of bug that ends trust in a tool. The 409 body carries the current value so the client can offer keep-mine / take-theirs. -
A monotonic project version, and gap detection. The broadcast is fire-and-forget, so a peer that misses one delta — queue pressure, a reconnect gap — is silently stale forever. Every delta carries
seq; a client that seesseq > local + 1refetches. This is what makes staleness self-healing instead of permanent, and it is about twenty lines. -
Gesture coalescing, client side. This is the one real load difference between tl and arthur: tl's user annotates a clip every few seconds, arthur's user drags a vertex. A drag is one write on release, not one per pointermove. With the gesture held in mode-local state and committed once (see Tool modes), arthur's write rate is comparable to tl's and the model holds unchanged. Without it, no sync model survives.
-
An outbox, so nothing ever waits on the network. "No lag" is satisfied entirely on the client: apply locally, enqueue, reconcile on ack.
{:pending {"clip/7/cel/bg/12" {:value … :etag … :tries 0}}}With one rule that is easy to get wrong: while a leaf has a pending local write, incoming broadcasts for that leaf are ignored. Otherwise a remote delta lands mid-gesture and the shape snaps back and forth between the two values until the ack arrives.
Two operational notes
- tl's
ROOMSis process-local, as its own comment says. That is a single-worker ceiling, not a design flaw, and it has to move to Redis before a second worker. - Revisions need a coarser trigger than every save. tl snapshots a small annotation layer; arthur's tier 1 contains cel polygons, so a snapshot per save will bloat the table. Snapshot on an explicit "mark version", or time-boxed.
When this model is wrong, and what replaces it
State the tripwire rather than trusting the judgement: if 409s on cel leaves become common in real use, the merge unit is too coarse. The answer then is a finer unit, not a CRDT, and there is a planned progression:
| Step | Merge unit | Gets you |
|---|---|---|
| now | the cel (cel/:node/:frame) |
two people on different frames |
| next | the layer | two people on different layers of one cel |
| last | the stroke — an append-only list of immutable stroke records | two people painting one layer at once |
The third step is worth seeing now because it is the escape hatch that keeps this decision from being a dead end. Once a drawing is a list of immutable records rather than a mutable point array, concurrent painting merges by construction — two people appending different records cannot conflict — and that is a domain-specific version of what a CRDT would have given, without the second source of truth. It is also a plausible thing to want for undo granularity anyway.
Make the merge unit small instead of clever
Last-writer-wins clobbers only when its unit is too big. tl can save a scene; arthur cannot save a project, because one vertex drag would then clobber a collaborator's keying. The fix is addressing, not an algorithm:
palette
sequence/:sid
clip/:cid/timing exposure, lead, kept frames
clip/:cid/params/:area teeth | eyes | brows | mouth | plate
clip/:cid/node/:nid one node: source, parent, stencil, z, colour
clip/:cid/channel/:nid/:prop
clip/:cid/cel/:nid/:frame
clip/:cid/overrides/:nid/:prop
An earlier draft of this list had params and scene as one leaf each, and both
were too coarse: one person tuning teeth while another tunes eyes would have
collided on every slider move, and two people adding nodes would have collided
always. Split params by feature area and give every node its own leaf. With
fractional :z there is no separate order leaf to contend on, which is the
second thing fractional indices buy.
Each path is a leaf: independently addressed, independently versioned, LWW
with an If-Match on its version. The boundaries are chosen so the things people
actually do simultaneously land on different leaves — two people painting
different cels, or keying different parts, never meet. Within a channel leaf the
server merges field-wise by frame, which is the return on keys-as-a-map and
about fifteen lines of Python.
Where it stays honest about conflict
| Data | Concurrency | Resolution |
|---|---|---|
| palette, clip params, knobs | scalar | LWW on the leaf. Two people tuning one clip will fight — that is a real thing to surface, not to merge. |
| channels | map by frame | field-wise union, LWW per frame |
| kept-frame set | add / remove | store {frame → bool}, not a set, so removals merge instead of vanishing |
cel layer order, :z |
list insert | fractional indices, not integers |
| cel polygon points | sequence | LWW on the whole layer, plus an advisory lease |
| parentage | tree | LWW per node, with a cycle check that rejects the losing move |
| playhead, selection, tool | ephemeral | presence only — never in the document, never in history |
Two of those are cheap now and expensive later.
Fractional indices for anything ordered. Integer :z and integer layer
positions do not survive concurrent insertion: you get duplicates and gaps. A
fractional key between neighbours makes concurrent inserts commute. Adopt it in
domain/node and domain/cel from the first commit.
Do not merge one polygon between two people. Two simultaneous edits to the same path have no meaningful merge, and the result is a shape nobody drew. The layer is the unit, LWW, and presence shows who is on it before the collision. A coarse merge unit that is visible beats a fine one that invents geometry.
Presence carries more than a cursor
tl's roster entry, extended with what arthur's views need:
{:cid "…" :user "…" :party nil :joinable true
:clip cid :frame cf :playing? false :rate 1.0
:node :mouth :tool :paint/pen
:editing "clip/7/cel/bg/12"} ; advisory lease on a leaf
:editing is what makes leaf-level LWW liveable — the collision becomes visible
before it happens. Advisory only: no server enforcement, so there is no lock to
leak when a tab closes.
Follow mode, and why frames never go on the wire
tl's parties become arthur's follow mode nearly unchanged: mirror :clip,
:frame, :playing? and the plate mode.
Broadcast transport state, never frames. On play, send
{:clip :frame :playing? :rate} once; each peer's own audio element then becomes
its own clock and the picture follows it exactly as it does locally. Per-frame
messages would put thirty packets a second on the wire to reproduce something
every peer can compute. A peer whose bake is cold shows a cold indicator and
catches up rather than stalling the party.
Going offline
Arthur is offline-capable almost by accident, and it is worth noticing why: tier
1 is small enough to hold entirely in app-db, tier 2 is a local content-addressed
cache, and tier 3 is a directory of PNGs that extract.sh put on your own disk.
So when the network drops, everything needed to scrub, key, paint, resolve and
render is already local. Nothing about playback or editing touches the server.
What stops is only the three things that are inherently remote: your writes stop acking, others' deltas stop arriving, and presence goes dark.
Two preconditions, or the above is a lie:
- The outbox must be durable. In memory, a closed tab loses the session. It goes in IndexedDB, written in the same task as the optimistic local apply — a few hundred microseconds, invisible, and the difference between "offline is fine" and "offline is a trap".
- Vendor MediaPipe's wasm.
README.mdnotes thatface_landmarker.taskis local but the wasm bundle is fetched from jsdelivr on first use. That makes stage 2 the one thing that silently requires a network, and it will be discovered on a plane. Vendor it or precache it in a service worker.
Reconnecting
Refetch the document first, then drain the outbox — in that order, so conflicts
are detected against current state rather than against a stale etag. Since the
document is kilobytes, a full refetch is cheaper than reasoning about a long
seq gap.
Then each pending write lands in one of four cases, and the design goal is that only the last one reaches a human:
| Case | Resolution |
|---|---|
| nobody touched your leaves | etags still match, outbox drains, converged. The common case. |
| a channel leaf collided | auto-merges field-wise by frame — union the frames, LWW per frame. No human. |
| a scalar leaf collided (knob, timing) | a compact list with one "keep all mine / take all theirs", because a knob is one number and the value is visible in the output anyway |
| a cel leaf collided | see below — never a dialog |
The channel row is the third time keys-as-a-map pays for itself, and it is the reason a long offline session usually reconciles with no interaction at all.
A colliding cel forks into a layer
Do not ask an artist to choose between two drawings in a modal. Keep both: their cel stays where it is, and yours is added on top as a layer named for the session that made it. Nothing is lost, nothing is silently clobbered, and the conflict is resolved by looking at it and deleting a layer — which is an operation the paint tool already has, and which is what an artist would do anyway.
This is the general principle for this whole class of conflict: when the merge unit is coarse, resolve by stacking rather than by choosing. A stack is reversible and legible; a choice made in a dialog against two thumbnails is neither.
The awkward cases, named
- A write to a leaf whose parent was deleted — a node or clip removed while
you were away. The server answers
404; the client parks it in an orphaned-edits tray rather than dropping it. Rare, and the alternative is silent loss. - A very long offline session against a busy project. Merging may be the wrong frame entirely. Offer the escape hatch revisions already make possible: save the offline session as a revision, take theirs, and reconcile by hand from two named versions. Better than a hundred-row conflict list.
- Two peers both run detection offline. Both compute the same analysis artifact and both try to upload it. Content addressing makes that idempotent — same hash, same bytes — so let it race rather than coordinating a claim.
Undo is per-user — and this revises an earlier claim
An earlier section of this document said undo was free if app-db held only
authored data, via day8.re-frame/undo. That is true single-player and wrong the
moment two people edit: a snapshot of app-db reverts their work along with yours.
So undo is a per-user stack of inverse leaf writes, applied as ordinary edits. Your undo writes a leaf back to the value you last saw, and it conflicts with a concurrent editor in the same visible way any other write does. Keeping app-db small is still right — for validation cost and allocation — but it is no longer the undo story.
The 30fps budget
At 320×200 stage 7 is not the problem. An even-odd scanline fill over 64,000 pixels with a dozen mostly-small parts is well under 200,000 byte writes a frame: six million a second at 30fps. Six roto tracks off baked geometry adds a handful of subarray reads.
Two things in the current code will miss 30fps, and neither is the rasteriser:
posterizeIntodoes agetImageDataevery frame. A GPU→CPU readback stalls the pipeline, and it is followed by 64,000 linear scans over the palette to find the nearest colour — around a million distance computations a frame. Both are avoidable: posterise once per source frame into a cachedUint8Array(it is a pure function of the frame, the transform and the palette), and replace the nearest-colour scan with a 5-5-5 RGB lookup table — 32KB, built once per palette.drawRegistereddraws a full-resolution photo every frame. Correct as a drawing reference on a held frame; not something to do thirty times a second. Pre-scale each source frame to raster size once, at load.
The rule both are instances of: nothing that reads source pixels may run on the transport path. That is the same line the bake boundary draws, which is why drawing it once, in stage 3, pays twice.
Porting order
Tests first, and not as discipline — as an oracle.
- Port
synth.jsandselftest.jsfirst, tocljs.test. The synthetic track is the only fixture with ground truth, and the 105 assertions encode invariants that are invisible to inspection: the ring-simplicity check, the bowtie, the iris-pairing vote fed a deliberately swapped track. - Run both implementations on the same synthetic track and diff numerically
while porting each pure module.
fitSimilarityandprocrustesMeanshould agree to 1e-9; a disagreement is a port bug, not float noise. - Port
landmarks,geom,ring,raster,take— mechanical, and they are where the invariant comments live. Carry the comments across verbatim. They are the most valuable text in the repo and every one of them is a bug that already happened. - Port
flow/measure/*by lifting out ofpipeline.js, then splitconditionandkeyout of it. - Build
db,store,clock, and stage 6 + 7 against a single hardcoded clip. At this point the synthetic take should render, with no UI beyond a canvas and a play button. - Only then the UI, the modes, the timeline.
- Add the ring-simplicity and no-RGB checks as specs, so they fire on authored data too and not only in tests.
Three things are cheap in step 3 and expensive after step 6, so they go in early
even though nothing needs them yet: keys as a map keyed by frame, fractional
indices on :z and cel layer order, and leaf addressing for tier 1 — the
paths under Collaboration, used as the shape of app-db even while single-player.
None of them cost anything on day one and all three are retrofits that touch
every namespace.
The wiring cross-check in selftest.js — every el('id') must exist in
index.html — becomes unnecessary and should be deleted rather than ported. It
is a test for a failure mode that Reagent does not have.
Decisions I would defer
- Whether
rasterstays JS. 320×200 is 64,000 pixels and CLJS over aUint8Arraywith no seq allocation handles it comfortably; port it and measure. Do not port it into idiomaticmap/reduce. - Interpolation.
:holdis the only interp the aesthetic wants, and the channel model should still carry:interpper node, because a parented transform on a painted cel is the one place a tween might be right. Do not implement it until something needs it. - Multi-character scenes.
characteris in the entity model as one field and should stay a stub until there are two. - Audio per clip vs per sequence. One master track is right until it is not.
- Whether bake B exists at all in v1. One clip does not need it; the flat typed-array layout should be designed now and built when the second roto track appears.
- Sharing tier 2. Content addressing means a collaborator can fetch a bake instead of recomputing, and also means nothing breaks if they do not. Ship the cache-miss path first and treat blob sharing as an optimisation — the analysis artifact is the only one where it clearly pays, because it is the only one that costs minutes.