arthur/docs/architecture.md
Olive Vaughn 27bfe18bee Keep one invalidation table, and derive the inverse a UI wants
There were two answers in the tree to "which stored bytes stop being valid when
this knob moves", and only one of them was checked.

`flow/address/block-knobs` is per block and asserted by biconditional —
`address-test` re-freezes the take once per knob and requires that the bytes
changed if and only if the key did. `domain/params`'s `:affects` was per area,
had no caller but a test asserting it returned what it was written as, and was
already wrong in both directions on the one entry where the two granularities
disagree: `:aperture-cut` claimed `#{:mouth}`, where it reaches no block, and
omitted the teeth, whose contour bytes it genuinely moves by gating
`condition/interior`'s smoothing. `:blink-cut` claimed `#{:eye}` and reaches no
block either, because a blink is `[:vis]` keys in tier 1.

So `:affects` and `affected-areas` are gone, and `address/knob-roles` is the
derived inverse of the table that is asserted — which is what a parameter panel
actually wants to ask. A knob absent from it invalidates no block, and that is
an answer rather than a gap.

Two new assertions keep the derivation from rotting at either edge: every role
in the table is reachable from some knob, and every knob a block declares is one
the registry defines. The second closes a real hole — `block-descriptor` checks
only that a knob was PASSED, and the freeze's `merge take/knobs` makes that true
of anything spelled like a keyword, so a typo in `block-knobs` would have named a
setting no slider can move.

228 CLJS tests, green.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-28 01:23:26 -04:00

49 KiB
Raw Blame History

arthur — structure

How the suite is laid out once it is a suite: ClojureScript and re-frame, a clip as the unit of work, keying as its own step, and painting as a peer of rotoscoping rather than a sketch bolted to the side.

docs/design.md says what the thing is and why. This says where the code goes. docs/animation-model.md specifies the data both of them are about — nodes, channels, symbols and time maps — and supersedes this document wherever the two describe the same type. Nothing here revises an aesthetic decision; several things here split a decision that is currently made in two places at once.

What is actually wrong with the current shape

The pure modules are fine. pipeline.js, mathutil.js, landmarks.js, interior.js and raster.js are functions over data with the reasoning written down next to them, and they port nearly verbatim.

app.js is the whole problem, and not because it is long. It is long because six unrelated jobs are braided together in it:

  • reading knob values out of the DOM (opts()),
  • deriving everything from them (rebuild, buildEyes, buildBrows, resolveTeeth),
  • resolving what is on screen at a frame (plateIndex, perfIndex, leadIndex),
  • rasterising (renderFrame, compositeRender),
  • driving the clock (tick),
  • persisting (saveCels, loadCels).

Every one of those is a different layer, and rebuild recomputes all of them whenever any knob moves. That is affordable at one clip and 72 frames and is not affordable at a project. The re-frame port is worth doing chiefly because the subscription graph is the staged dataflow docs/design.md already describes — "three artifacts, not two" is a dependency graph written in prose.

Three frame spaces, named

Most of the future bugs live here, so name them before anything else.

Space Symbol Meaning
source frame sf Index into the footage and the dense track. What detection produced.
clip frame cf 0-based within a clip. cf = sf - (:in clip).
sequence frame qf Position on the timeline. qf = (:at clip) + cf.

Nothing authored is ever stored in sequence-frame space. Keys, cels, overrides and kept-frame sets are all in clip-frame space, so a clip slides on the timeline without a single stored number changing. Exposure and lead are transforms within clip space:

Detection retains every source frame. A chosen picture fps samples the frozen roto in clip time; it changes neither source-frame count nor the audio clock. The set of source frames an artist uses as cel tracing references is another selection, independent of the picture fps.

(defn pose-frame [clip cf]
  (-> cf (expose (:exposure clip)) (shift (:lead clip) (count-frames clip))))

which is the composition app.js currently performs in leadIndex and perfIndex with the order implicit. Expose first, then lead: exposure floors onto a grid and lead deliberately reads the future, so leading first and then flooring would quietly discard the lead on most frames.

The entity model

project
├── palette                  authored ramp; indices, never RGB
├── character                plate library, per-plate slots
├── footage[]                frames + audio + manifest (fps, w, h)
├── analysis[]               dense track, cached, keyed by footage+model+version
├── clip[]                   the unit of work
│   ├── source               footage id, analysis id, in/out in sf
│   ├── exposure, lead       timing, in cf
│   ├── scene                node tree
│   ├── channels             node+property -> keyframe stream, in cf
│   ├── overrides            (node, property, cf) -> value, applied last
│   └── cels                 painted vector layers, keyed by chosen clip frames
└── sequence[]               clip placements: {clip-id, at, in, out}

clip is the entity the current app has exactly one of and never names. Giving it a name is most of the work: state in app.js is a clip with its analysis inlined and its palette global.

Cel keys select where drawings begin and how long they hold. The source frames shown beneath a cel while tracing are chosen independently, and picture fps only controls which analyzed pose the finished roto displays at a given time.

Two things called "track"

docs/design.md says "dense track" for the landmark stream. A timeline also wants tracks. Pick now: the landmark stream is analysis, a keyframe stream is a channel, and the word track is reserved for a row on the timeline. Renaming this later costs a day.

The node, decomposed

The prototype today gives a part a parent and a clip, and parent is doing nothing except documenting intent — mouth_in is already in the same head-local raster space as mouth, so composing its transform would be composing identity. The moment a painted cel is attached to a head plate that moves, transform composition becomes real, and the three ideas currently sharing two fields have to come apart:

Field What it does Wrong to conflate because
:parent Transform composition. Child geometry is in parent's local space. A node can be drawn over its parent without inheriting its motion.
:stencil Colour-key clip: write only where the buffer already holds that index. The iris is stencilled by the sclera and parented to the lid ring; those are different nodes.
:z Draw order. Order is authored per scene, not implied by the tree.

A node is then:

{:id      :eye-r/iris
 :source  {:kind :primitive ...}     ; the five kinds from design.md
 :parent  :eye-r/lid
 :stencil :eye-r/sclera
 :z       42
 :color   :iris                      ; palette key, never a hex
 :interp  :hold}

:source carries the taxonomy docs/design.md already has — :plate, :feature, :interior, :primitive, :scalar — plus :cel for painted vectors and :group for a pure transform node. That is the decomposition that makes painting a peer rather than an annex: paint.js is not a special case, it is a source kind whose channels happen to be authored by hand instead of measured. docs/design.md already promises "pluggable sources"; this is the data shape that keeps the promise.

Each node also gets an optional local transform channel — translate, rotate, scale, squash. That is the thing the current tool cannot express and that the slot_mouth / scale / rot / squash fields in the prototype are reaching for.

The flow, in seven stages

The user-facing flow is frames -> analysis -> keying -> geometry -> palette. Analysis is three stages, not one, and the split is the most load-bearing decision in this document, because exactly one boundary is a cache boundary.

# Stage In Out Cost
1 ingest video footage: frames, audio, manifest minutes, in-app
2 detect footage raw landmarks per frame minutes, cached
3 measure landmarks anchor fit, residual, head-local rings, signals, interior pixels seconds
4 condition measurements smoothed transforms and contours milliseconds
5 key conditioned signals + policy channels: sparse keys, quantised holds, kept frames milliseconds
6 resolve channels + scene + exposure + lead + overrides geometry per output frame per frame
7 palette geometry indexed raster per frame

Stage 2 is the only artifact worth persisting and the only one worth a progress bar. Stage 3 reads pixels — teeth extraction is here, and that is why it is here rather than in keying: after stage 3, nothing downstream may look at a source pixel. Stage 4 is separated from 3 only because contour avg and anchor avg are knobs and the rest of stage 3 is not; splitting them means dragging that slider does not re-run the interior extraction.

That last clause is the guarantee, and it is narrower than the table's ordering looks. Stage 3 is not one pass that finishes before stage 4 begins. The anchor fit is knob-free; conditioning smooths its four parameters; the head-local rings are then measured through the conditioned transform — so anchor avg does re-run the ring mapping, which is a few hundred frames of twenty points and free. What it must not re-run is the part that reads a source pixel, and that part takes the landmarks and the frames and never the transform, so it does not. Built as measure/anchor → condition/anchor → measure/mouth → condition/contours, and stage 7's eyes, brows and interior follow the same shape.

Where the current code lands

Now Stage
stabilize 3 measure
eyeSignals, browSignals, pairIrises, pairBrows 3 measure
extractTeeth 3 measure — it reads pixels
smoothTransforms, smoothContours 4 condition
selectKeys, suggestPlateFrames 5 key
quantizeSnap, resolveBlink 5 key — a held value is a key
exposeIndex, shiftIndex, activeKey, heldFrame 6 resolve
toRasterRing, iris placement at ring slots, offsetRing 6 resolve
IndexedRaster, drawCel 7 palette
drawRegistered, posterizeInto reference only, off to the side

Three things in app.js currently straddle a boundary, and each straddle is a bug waiting for a bigger project:

  • resolveTeeth extracts from pixels and applies contrast/dwell/blob knobs. Extraction is stage 3 and cached with the track; resolution is stage 5. Today, nudging teeth dwell re-runs Otsu over every frame.
  • buildEyes measures, quantises, and places the iris on the subsampled lid ring. Those are stages 3, 5 and 6. The placement argument in docs/design.md — that the iris is read off the already-smoothed ring — is a stage-6 fact and stays true; it just needs to be in stage 6.
  • buildBrows likewise: measure the height out of the ring (3), quantise it (5), put it back (6). The doc's warning about moving the brow twice is precisely a warning about these being one function.

What each knob invalidates

This table is the argument for the split, and it should be derivable from the sub graph rather than maintained by hand.

As built, half of it is, by a different route. flow/address/block-knobs is this table for tier 2 — per block, the settings its bytes depend on — and address-test asserts it by biconditional rather than deriving it, which is stronger than a derivation would have been: a derived table is only as right as the graph it reads. What is not written down anywhere is the rest of the column. Stages 3 to 5 are one call chain in flow/take/measure rather than a chain of subs, so there is no graph above ::base-scene to read a dependency off, and the tier-1 half of a re-freeze is therefore whole-clip.

An earlier arrangement had a second table in domain/params, per feature area, saying which areas a knob affected. It disagreed with the tested one on the first entry where the two granularities part company and it had no caller; see that namespace for why one checked table beats two.

Knob Invalidates from
model / footage 2 detect
teeth contrast, cavity erode, blob grow, tongue reject, prefer upper 3 measure (mouth crop only)
contour avg, anchor avg 4 condition
teeth dwell, blink cut/hold/dwell, gaze step/dwell, brow step/dwell, suggest tolerance 5 key
vertices, eye vertices, brow vertices, teeth vertices 5 key
exposure, mouth lead, kept-frame edits, overrides 6 resolve
lash line, iris size, pupil, brow weight 6 resolve
palette edits, colour assignment 7 palette

Namespaces

src/arthur/
  domain/                     pure data, specs, ops. No re-frame, no DOM.
    palette.cljs              ramp; index lookup; the no-RGB rule as a spec
    landmarks.cljs            index tables (port verbatim)
    ring.cljs                 ordered traversal: subsample, offset, simplicity
    geom.cljs                 similarity fit, procrustes, moving average
    node.cljs                 scene node: source, parent, stencil, z
    scene.cljs                node tree: topo order, transform composition
    channel.cljs              keyframe stream: active-key-at, hold semantics
    clip.cljs                 clip entity; the frame-space conversions
    cel.cljs                  painted vector layers
    timeline.cljs             sequence: clip placement, qf <-> cf
  flow/
    ingest.cljs
    detect.cljs               the MediaPipe boundary, and the only one
    measure/
      anchor.cljs             stabilise: procrustes, similarity, residual
      mouth.cljs
      eyes.cljs               openness, gaze, iris pairing vote
      brows.cljs              raise/tilt, ring and end pairing votes
      interior.cljs           otsu, morphology, components, radial contour
    condition.cljs
    key.cljs                  velocity minima, dwell, quantise, blink, decimate
    resolve.cljs              scene eval at a frame; exposure, lead, overrides
    raster.cljs               indexed scanline fill, stencil, disc, rect
    reference.cljs            registered underlay, posterise
  db.cljs                     app-db schema + spec
  store.cljs                  handles for things too big for app-db
  clock.cljs                  audio clock, outside app-db
  events/                     one ns per domain area
  subs/                       one ns per stage
  ui/
    shell.cljs
    mode/                     one ns per tool mode
    panel/                    strip, worksheet, readouts, palette, layers
    canvas.cljs               the one imperative sink
  fx/                         mediapipe, files, audio, persistence

Two rules about this tree. domain/ may not require flow/, and neither may require re-frame. And flow/ namespaces take explicit arguments — no namespace below subs/ ever calls subscribe.

app-db, and what is not in it

app-db holds authored data and ids. Nothing derived, and nothing large.

That sounds like ordinary hygiene and it is not: it is the precondition for two features that are otherwise unbuildable.

  • Spec validation on every event. Worth having, affordable only over authored data.
  • Cheap writes. Every edit assoces into app-db, and every mounted layer-2 sub compares the result. Both are structural-sharing operations and stay cheap only while the map is small.

Undo is not on that list, and an earlier draft said it was. See Collaboration: snapshot-based undo is right single-player and wrong once two people edit.

So the dense track and the decoded frames live in arthur.store, a defonce map of id to JS object, and app-db holds {:analysis/id "sha-..."}. Landmarks stay as they arrive — arrays of {x, y, z}, or better, a flat Float32Array — and are not converted to Clojure maps. 478 points × 600 frames is 286,800 maps, allocated for nothing: no sub ever needs to diff them, and every stage that touches them walks all of them anyway.

The playhead, and where per-frame work goes

A reaction propagates downstream only when its own output value changes, not when one of its inputs notifies. So with layer-2 extractors and layer-3 computations, a playhead tick costs one cheap extractor run per mounted layer-2 sub, each of which returns the same value for every subtree the tick did not touch and therefore notifies nobody. ::measurements, ::channels and ::resolver do not re-run. The :<- graph is what buys that, and it buys it whether or not the playhead is in app-db.

So the playhead lives in app-db, like everything else. An earlier draft of this document put it in a standalone ratom to avoid an invalidation storm that does not happen. Two independent reasons it belongs in db:

  • [:playback/seek ...] in the event log is how scrubbing becomes inspectable in re-frame-10x.
  • A collaborator's playhead is a feature. Putting it outside app-db puts it outside the machinery that shares it. See Collaboration.

What is true is narrower, and is about sub authoring and interceptors rather than about the playhead:

  • Expensive work must not live in a layer-2 sub. A sub that derefs app-db directly re-runs on every db change whatever changed. That is the actual content of the folklore about re-frame and canvas.
  • Do not put a per-frame event behind a global interceptor that walks the whole db — a spec-validating after, or std-interceptors/debug and its clojure.data/diff. At 30fps that is thirty full-db traversals a second. The transport event carries its own interceptor chain and is excluded from the global ones.
  • The render sink is not a Reagent component. Stage 7 writes bytes into a canvas from an rAF loop; it is not a view that re-renders.

Stage 6 is still arranged so the per-frame path is a lookup, not a computation:

;; recomputes when channels, scene, exposure, lead or overrides change
(rf/reg-sub ::resolver ...)   ;; => (fn [cf] geometry), or an index into a bake

;; the rAF loop: reads, blits, dispatches nothing
(defn tick []
  (let [cf (clock/clip-frame)]
    (raster/paint! (@resolver cf))
    (canvas/blit!)))

The sub produces a resolver; the loop applies it. A knob change costs one sub recomputation; a frame costs a lookup and a blit.

Events and subs

Events are named for intent, and carry the clip they apply to:

[:clip/set-exposure  clip-id 2]
[:clip/toggle-kept   clip-id cf]
[:keying/re-key      clip-id]          ; not a setter; policy unchanged, redo the work
[:keying/suggest     clip-id]
[:paint/commit-shape clip-id node-id cel-frame pts]
[:node/set-parent    clip-id node-id parent-id]
[:override/set       clip-id node-id :pts cf value]
[:timeline/move-clip seq-id clip-id qf]

Setters are fine where the intent is the value — set-exposure is honest. Where they differ, name the intent: :keying/re-key takes no value and is not :clip/set-keys.

Subs mirror the stages one-to-one, so the graph and the table above are the same object:

::footage -> ::landmarks -> ::measurements -> ::conditioned -> ::channels -> ::resolved -> ::raster

Each is a layer-3 sub over the previous plus the parameters for its stage only. That is what makes "drag exposure" recompute ::resolved and nothing above it.

Tool modes

A mode is a namespace, not a branch. Each exposes a map:

{:id :paint/pen
 :cursor      "crosshair"
 :keymap      {"Enter" [:paint/close-shape] "Escape" [:paint/abort]}
 :on-pointer  (fn [ev ctx] ...)
 :overlay     (fn [g ctx] ...)          ; draws handles, never output pixels
 :enter :exit (fn [ctx] ...)}

with two rules worth stating because the current PaintUI breaks both and gets away with it at this size:

  • In-progress gesture state is mode-local, in a ratom the mode owns — not in app-db. An unclosed polygon must not appear in the undo history, and a drag must produce one undo entry rather than one per pointermove. The mode dispatches a single :paint/commit-shape on release.
  • Overlays draw to a separate canvas. Handles, vertex boxes and the green close-target are not indexed pixels and must never be in the raster. The current code composites them into the same context as the output, which is fine for a preview and lies to you the moment you want to judge the look.

The palette-index constraint and grid snapping stay where they are: they belong to domain/cel and domain/palette, enforced at commit, so no tool can bypass them.

Where things live

Superseded in detail by Serialization and Baking below; the local picture is:

Thing Where Keyed by
project (tier 1) server, plus a local autosave copy project id, then leaf path
footage frames, audio (tier 3) on disk / OPFS, untouched content hash
analysis artifact, geometry bakes (tier 2) OPFS, with IndexedDB for the index content hash over every input
render output never persisted —

The cache key on the analysis artifact must include the detector version. A model upgrade that silently reuses old landmarks presents as "the tool got worse" with no event to attach it to — which is the argument for content-addressing rather than version-numbering everything derived.

Baking

"No analysis during playback" is docs/design.md's three-artifact rule with a number attached. The expensive stages are 2 (detect), 3 (measure) and, over a long take, 4 (condition). Stages 5–7 are arithmetic over small arrays. So there are two bakes, not one, and they answer different questions.

Bake A — the analysis artifact. Mandatory.

Stages 2–4 for one footage at one set of conditioning parameters. This is what must never run while the transport is moving.

Field Shape
landmarks Float32Array[n_frames × 478 × 2]
transforms Float32Array[n_frames × 4] — s, θ, tx, ty
residual Float32Array[n_frames]
head-local rings Int16Array per ring table, isotropic fixed point
signals openness, gaze, brow ends, aperture — Float32Array[n_frames] each
mouth crops Uint8Array, one small crop per frame

The mouth crops are the non-obvious entry, and they are what makes remote work possible at all. extractTeeth reads source pixels, so without them a collaborator holding the analysis but not the 600 source PNGs cannot touch a single teeth knob. A 40×30 crop is about 1.2KB; a 600-frame take is under a megabyte against hundreds for the footage.

Bake B — resolved geometry. For scale.

Stage 6 output per node. Not needed for one clip — resolving a frame is a key lookup and a transform — and needed the moment six roto tracks are live at 30fps.

Fixed topology is what makes this flat. Because every key of a part carries the same vertex count with the same vertex meanings, a node's whole bake is a rectangular array with no per-frame header and no indirection:

pts     Int16Array[n_frames × n_verts × 2]    raster space, already grid-snapped
state   Uint8Array[n_frames]                  hidden flag + palette index

Frame f of a node is the subarray at f * n_verts * 2. 600 frames × 20 verts is 48KB; a twelve-node clip about 600KB; six roto tracks about 3.5MB. So this is the payoff on the aesthetic constraint rather than a cost of it — a variable vertex count would force a per-frame offset table and a scan.

Rasters are not baked

Stage 7 is the cheapest stage and the one most worth keeping live: palette and colour assignment are the last decisions and the ones you want to change in motion. Baking to pixels would freeze exactly the loop that should be instant. Small raster thumbnails for the strip and the timeline are a separate artifact with a separate cache.

Unbaking is not an inverse

The bake never replaces its input. Authored state — channels, params, overrides — is always retained and always authoritative, so "unbake this roto and edit it" is a boolean about which side of the cache the renderer reads, not a computation. Instant by construction. Three rules keep it that way:

  • No destructive bake. Nothing is discarded when something is baked.
  • Hand corrections never go into the bake. They are :overrides, applied at stage 6 after the bake is read, so a correction survives a rebake. This is docs/design.md's "hand corrections must not live in the take", and it is the rule that stops baking from becoming a trap.
  • Bakes are content-addressed by a hash over every input that produced them: footage hash, detector version, and the parameters of each stage at or above the bake line. A stale bake is then unreachable rather than wrong, and a collaborator's bake is fetchable by the same key — see Collaboration.

Partial bakes

Bake per node, per frame range. Scrubbing into an unbaked range resolves on demand and fills the cache; a worker bakes ahead of the playhead. Show it: a bake-state bar under the timeline is an affordance every NLE has already taught people to read, and it is honest about what is ready.

Serialization, in three tiers

Cut by mutability and size. The cut is what makes collaboration and baking both tractable, because it decides what is allowed on the wire.

Tier What Size Synced Undoable
1 authored palette, clips, scene graphs, channels, cels, overrides, sequences KB yes — it is the document yes
2 derived analysis artifacts, geometry bakes, thumbnails MB no — content-addressed blobs, fetched no
3 source frames, audio MB–GB no — immutable, by hash no

Tier 3 is produced by the app, not by extract.sh: wasm ffmpeg decodes the clip to frames and audio in the browser, and they are uploaded and served as content-addressed blobs. That makes tier 3 the same kind of thing as tier 2 — a cache with a hash — and it removes the one step that currently needs a terminal. Two consequences worth planning for: the extraction rate is recorded by the code that chose it rather than by a manifest.json a human might edit, and offline work needs the frames already fetched, so the local blob cache is what makes a plane usable. Decoding server-side instead is the same data model with a different worker; the format does not care which side runs it.

Only tier 1 is the document. A bake inside the shared document is a system that puts 48KB on the wire per vertex drag, and it is unnecessary, because tier 2 is a pure function of tiers 1 and 3.

Tier 1 stays small enough to read. The only vertex data in it is what a human placed by hand: cel polygons and overrides. Everything traced is tier 2.

Keys are a map, not a vector

;; not this
{:keys [{:f 0 :pts [...]} {:f 4 :pts [...]}]}
;; this
{:keys {0 {:pts [...]}, 4 {:pts [...]}}}

activeKey becomes a lookup in a sorted map rather than a scan; a single key becomes addressable as a path; and two people keying different frames of one part merge field-wise with no merge algorithm at all. selectKeys returns an array today — change it before anything depends on the order.

Serving tiers 2 and 3

Built at step 9. Above this point the tiers are a rule about what is allowed on the wire; this is the shape that enforces it.

One store for both, named by the sha256 of the bytes. Once tier 3 is decoded by the app rather than by a shell script it becomes the same kind of thing as tier 2 — a cache with a hash — so there is one place that writes bytes, one that reads them, and one URL shape:

GET /blob/<sha256>              raw bytes, Cache-Control: immutable

immutable is not optimism there, it is the definition: the name IS the hash of the content, so a cached copy cannot be stale. That is what makes serving six hundred frames out of it cheap enough to do on every load.

Two kinds of hash, and they are not the same hash. A blob is named by the hash of its BYTES, which is what makes an identical frame in two extractions one file. A derived thing — an analysis artifact, a dense block — is named by a hash over its INPUTS, which is what lets a client ask for the block the current settings want before anything has computed it, and what makes a stale bake unreachable rather than wrong. So a Block row has both: key over the inputs, and a foreign key to the blob whose digest is over the bytes. Conflating them would break the half of addressing that answers questions about work not yet done.

POST /api/analyses              {key, descriptor}        idempotent
POST /api/blocks/missing        {keys} -> {missing}
POST /api/blocks                {key, descriptor, data, state}
GET  /api/blocks/<key>
GET  /api/footage/<id>          the manifest, with a URL per frame

The server verifies, rather than trusting a name it was handed. It recomputes sha256(descriptor) for every key and refuses a mismatch; it refuses an analysis whose descriptor does not declare a detector and a version; it refuses a block whose analysis it does not know; and it refuses a document naming blocks it does not hold. The chain from a stored block to the model version that produced it therefore cannot be broken by a client that skipped a step — which is what the cache-key rule above actually requires, as opposed to recommends.

It hashes the descriptor TEXT rather than re-rendering it from parsed values, and that is not a shortcut. JS prints an integral double as 1 and Python prints 1.0, so a scheme where both ends re-render the numbers disagrees on the first parameter whose value happens to be whole — and the failure is an upload that 409s with nothing wrong. The bytes are the contract; the schema on top of them is a convention, and the two fields the server reads out of that schema are checked separately.

A manifest names frames; it does not locate them. Until step 9 the client fetched manifest.json and built frames/0001.png itself, which made the frame layout a shared secret between a shell script and a ClojureScript namespace. The manifest now carries a URL per frame, so the frames can live in the blob store — or, when wasm-ffmpeg extraction arrives, be uploaded into the same store by the app — and the client learns nothing new when that happens. The producer changes; the shape does not.

The document stores what a block IS, not what it holds. A block's element type is in its own descriptor, which is the only place it is written down: an Int16Array and a Float32Array over the same bytes are both valid readings, and only one of them is the block. That makes the descriptor load-bearing rather than documentation, which is the right way round for the thing a key is the hash of.

Leaf paths, as built

The list under Make the merge unit small instead of clever is the design; this is what step 9 implemented, for the subset that exists:

clip/<cid>/name                     clip/<cid>/subject/<sid>
clip/<cid>/timing                   clip/<cid>/feature/<fid>
clip/<cid>/stage                    clip/<cid>/group/<gid>
clip/<cid>/source                   clip/<cid>/node/<nid>
clip/<cid>/measured/<nid>           clip/<cid>/channel/<nid>/<prop>

Two departures from the design above, both because step 8 moved settings.

params/:area is not a leaf. That path came from a draft where params were one blob per clip, and two people tuning teeth and eyes collided on every slider move. Settings now live on the subject, the feature and the group, and a feature has exactly one area — so the feature leaf already is the area-scoped leaf, and splitting it again would separate a feature's params from its identity.

measured/<nid> is one leaf holding several channels, which contradicts "every channel gets its own". :head's measured channels are not authored: a freeze writes them together and a re-freeze replaces them together, and head-mode reads them to write :channels. A leaf per measured channel would offer a write nobody can make.

A leaf path is "/"-delimited and an id is one segment of it, so a namespaced id — :eye-r/iris, as drawn under The node, decomposed — is written eye-r~iris, and ~ is then refused inside a name. That is the whole of the escaping.

Collaboration

../tl already has the model, and it is the right one to copy:

  • An authoritative server, and one broadcast source. Edits go over HTTP; the server persists and then broadcasts the delta to the room. The socket is read-only for document state. No OT, no CRDT.
  • Presence gossiped between peers on the same socket, with the server doing exactly two things: handing out a connection id, and stamping the sender's server-side identity onto every message so nobody can post as somebody else.
  • Parties as a pure function of the roster, so every peer computes the same answer and the server knows nothing about them.
  • Revisions: one snapshot of the authored layer per save, with a user and a summary.

That maps onto the tiers exactly — the server stores tier 1 and snapshots tier 1, with tiers 2 and 3 as content-addressed blobs beside it.

Why this model, and not a CRDT

The usual reason to reach for Yjs or Automerge is automatic convergence without conflict dialogs. In this domain the primary contested data type is a vector path, and automatic convergence of a vector path is undesirable — two people dragging vertices on one polygon merge into a shape neither drew. So the CRDT's headline benefit is neutralised on exactly the data that needs help, and the parts where it would work fine (scalars, maps keyed by frame) are the parts that LWW already handles.

What a CRDT would additionally cost here is not the library, it is the second source of truth: the canonical document moves into an opaque blob, and every server-side thing that reads the document — revisions, collaborator lists, thumbnail generation, the admin, any query — needs it materialised back out with y-py or a Node sidecar. That is real operational weight bought for a feature this domain does not want.

So: authoritative server, HTTP writes, LWW on addressed leaves, presence on the side. And writes stay on HTTP rather than moving to the socket, which is tl's choice and the right one: auth, idempotency, status codes, retries and conditional requests all come for free, and a dropped socket cannot lose a write. Latency is not the reason to move them, because latency is handled on the client (below) and never by waiting for a round trip.

The model is standard; only tl's plumbing is hand-rolled

Worth separating. Optimistic concurrency control over addressed resources with an entity tag is RFC 7232 — ETag, If-Match, 412/409 — and it is what databases and HTTP APIs have done for thirty years. It is not a bespoke invention, and every HTTP client implements its half. The hand-rolled part of tl is the plumbing around it, so the move is to keep the model and replace the plumbing with the boring standard wherever one exists.

Four things to add

tl's implementation is missing two pieces that hand-rolled sync usually omits, and arthur needs two more that tl does not because its write rate is low.

  1. Conditional writes. PUT /project/:id/leaf/<path> with If-Match: <etag>, answering 409 on a stale write. tl's PUT replaces, so the loser's work disappears silently. For a knob that is survivable; for a painted cel it is the class of bug that ends trust in a tool. The 409 body carries the current value so the client can offer keep-mine / take-theirs.

  2. A monotonic project version, and gap detection. The broadcast is fire-and-forget, so a peer that misses one delta — queue pressure, a reconnect gap — is silently stale forever. Every delta carries seq; a client that sees seq > local + 1 refetches. This is what makes staleness self-healing instead of permanent, and it is about twenty lines.

  3. Gesture coalescing, client side. This is the one real load difference between tl and arthur: tl's user annotates a clip every few seconds, arthur's user drags a vertex. A drag is one write on release, not one per pointermove. With the gesture held in mode-local state and committed once (see Tool modes), arthur's write rate is comparable to tl's and the model holds unchanged. Without it, no sync model survives.

  4. An outbox, so nothing ever waits on the network. "No lag" is satisfied entirely on the client: apply locally, enqueue, reconcile on ack.

    {:pending {"clip/7/cel/bg/12" {:value … :etag … :tries 0}}}
    

    With one rule that is easy to get wrong: while a leaf has a pending local write, incoming broadcasts for that leaf are ignored. Otherwise a remote delta lands mid-gesture and the shape snaps back and forth between the two values until the ack arrives.

Two operational notes

  • tl's ROOMS is process-local, as its own comment says. That is a single-worker ceiling, not a design flaw, and it has to move to Redis before a second worker.
  • Revisions need a coarser trigger than every save. tl snapshots a small annotation layer; arthur's tier 1 contains cel polygons, so a snapshot per save will bloat the table. Snapshot on an explicit "mark version", or time-boxed.

When this model is wrong, and what replaces it

State the tripwire rather than trusting the judgement: if 409s on cel leaves become common in real use, the merge unit is too coarse. The answer then is a finer unit, not a CRDT, and there is a planned progression:

Step Merge unit Gets you
now the cel (cel/:node/:frame) two people on different frames
next the layer two people on different layers of one cel
last the stroke — an append-only list of immutable stroke records two people painting one layer at once

The third step is worth seeing now because it is the escape hatch that keeps this decision from being a dead end. Once a drawing is a list of immutable records rather than a mutable point array, concurrent painting merges by construction — two people appending different records cannot conflict — and that is a domain-specific version of what a CRDT would have given, without the second source of truth. It is also a plausible thing to want for undo granularity anyway.

Make the merge unit small instead of clever

Last-writer-wins clobbers only when its unit is too big. tl can save a scene; arthur cannot save a project, because one vertex drag would then clobber a collaborator's keying. The fix is addressing, not an algorithm:

palette
sequence/:sid
clip/:cid/timing                    exposure, lead, kept frames
clip/:cid/params/:area              teeth | eyes | brows | mouth | plate
clip/:cid/node/:nid                 one node: source, parent, stencil, z, colour
clip/:cid/channel/:nid/:prop
clip/:cid/cel/:nid/:frame
clip/:cid/overrides/:nid/:prop

An earlier draft of this list had params and scene as one leaf each, and both were too coarse: one person tuning teeth while another tunes eyes would have collided on every slider move, and two people adding nodes would have collided always. Split params by feature area and give every node its own leaf. With fractional :z there is no separate order leaf to contend on, which is the second thing fractional indices buy.

Each path is a leaf: independently addressed, independently versioned, LWW with an If-Match on its version. The boundaries are chosen so the things people actually do simultaneously land on different leaves — two people painting different cels, or keying different parts, never meet. Within a channel leaf the server merges field-wise by frame, which is the return on keys-as-a-map and about fifteen lines of Python.

Where it stays honest about conflict

Data Concurrency Resolution
palette, clip params, knobs scalar LWW on the leaf. Two people tuning one clip will fight — that is a real thing to surface, not to merge.
channels map by frame field-wise union, LWW per frame
kept-frame set add / remove store {frame → bool}, not a set, so removals merge instead of vanishing
cel layer order, :z list insert fractional indices, not integers
cel polygon points sequence LWW on the whole layer, plus an advisory lease
parentage tree LWW per node, with a cycle check that rejects the losing move
playhead, selection, tool ephemeral presence only — never in the document, never in history

Two of those are cheap now and expensive later.

Fractional indices for anything ordered. Integer :z and integer layer positions do not survive concurrent insertion: you get duplicates and gaps. A fractional key between neighbours makes concurrent inserts commute. Adopt it in domain/node and domain/cel from the first commit.

Do not merge one polygon between two people. Two simultaneous edits to the same path have no meaningful merge, and the result is a shape nobody drew. The layer is the unit, LWW, and presence shows who is on it before the collision. A coarse merge unit that is visible beats a fine one that invents geometry.

Presence carries more than a cursor

tl's roster entry, extended with what arthur's views need:

{:cid "…" :user "…" :party nil :joinable true
 :clip cid :frame cf :playing? false :rate 1.0
 :node :mouth :tool :paint/pen
 :editing "clip/7/cel/bg/12"}          ; advisory lease on a leaf

:editing is what makes leaf-level LWW liveable — the collision becomes visible before it happens. Advisory only: no server enforcement, so there is no lock to leak when a tab closes.

Follow mode, and why frames never go on the wire

tl's parties become arthur's follow mode nearly unchanged: mirror :clip, :frame, :playing? and the plate mode.

Broadcast transport state, never frames. On play, send {:clip :frame :playing? :rate} once; each peer's own audio element then becomes its own clock and the picture follows it exactly as it does locally. Per-frame messages would put thirty packets a second on the wire to reproduce something every peer can compute. A peer whose bake is cold shows a cold indicator and catches up rather than stalling the party.

Going offline

Arthur is offline-capable almost by accident, and it is worth noticing why: tier 1 is small enough to hold entirely in app-db, tier 2 is a local content-addressed cache, and tier 3 is a directory of PNGs that extract.sh put on your own disk. So when the network drops, everything needed to scrub, key, paint, resolve and render is already local. Nothing about playback or editing touches the server.

What stops is only the three things that are inherently remote: your writes stop acking, others' deltas stop arriving, and presence goes dark.

Two preconditions, or the above is a lie:

  • The outbox must be durable. In memory, a closed tab loses the session. It goes in IndexedDB, written in the same task as the optimistic local apply — a few hundred microseconds, invisible, and the difference between "offline is fine" and "offline is a trap".
  • Vendor MediaPipe's wasm. README.md notes that face_landmarker.task is local but the wasm bundle is fetched from jsdelivr on first use. That makes stage 2 the one thing that silently requires a network, and it will be discovered on a plane. Vendor it or precache it in a service worker.

Reconnecting

Refetch the document first, then drain the outbox — in that order, so conflicts are detected against current state rather than against a stale etag. Since the document is kilobytes, a full refetch is cheaper than reasoning about a long seq gap.

Then each pending write lands in one of four cases, and the design goal is that only the last one reaches a human:

Case Resolution
nobody touched your leaves etags still match, outbox drains, converged. The common case.
a channel leaf collided auto-merges field-wise by frame — union the frames, LWW per frame. No human.
a scalar leaf collided (knob, timing) a compact list with one "keep all mine / take all theirs", because a knob is one number and the value is visible in the output anyway
a cel leaf collided see below — never a dialog

The channel row is the third time keys-as-a-map pays for itself, and it is the reason a long offline session usually reconciles with no interaction at all.

A colliding cel forks into a layer

Do not ask an artist to choose between two drawings in a modal. Keep both: their cel stays where it is, and yours is added on top as a layer named for the session that made it. Nothing is lost, nothing is silently clobbered, and the conflict is resolved by looking at it and deleting a layer — which is an operation the paint tool already has, and which is what an artist would do anyway.

This is the general principle for this whole class of conflict: when the merge unit is coarse, resolve by stacking rather than by choosing. A stack is reversible and legible; a choice made in a dialog against two thumbnails is neither.

The awkward cases, named

  • A write to a leaf whose parent was deleted — a node or clip removed while you were away. The server answers 404; the client parks it in an orphaned-edits tray rather than dropping it. Rare, and the alternative is silent loss.
  • A very long offline session against a busy project. Merging may be the wrong frame entirely. Offer the escape hatch revisions already make possible: save the offline session as a revision, take theirs, and reconcile by hand from two named versions. Better than a hundred-row conflict list.
  • Two peers both run detection offline. Both compute the same analysis artifact and both try to upload it. Content addressing makes that idempotent — same hash, same bytes — so let it race rather than coordinating a claim.

Undo is per-user — and this revises an earlier claim

An earlier section of this document said undo was free if app-db held only authored data, via day8.re-frame/undo. That is true single-player and wrong the moment two people edit: a snapshot of app-db reverts their work along with yours.

So undo is a per-user stack of inverse leaf writes, applied as ordinary edits. Your undo writes a leaf back to the value you last saw, and it conflicts with a concurrent editor in the same visible way any other write does. Keeping app-db small is still right — for validation cost and allocation — but it is no longer the undo story.

The 30fps budget

At 320×200 stage 7 is not the problem. An even-odd scanline fill over 64,000 pixels with a dozen mostly-small parts is well under 200,000 byte writes a frame: six million a second at 30fps. Six roto tracks off baked geometry adds a handful of subarray reads.

Two things in the current code will miss 30fps, and neither is the rasteriser:

  • posterizeInto does a getImageData every frame. A GPU→CPU readback stalls the pipeline, and it is followed by 64,000 linear scans over the palette to find the nearest colour — around a million distance computations a frame. Both are avoidable: posterise once per source frame into a cached Uint8Array (it is a pure function of the frame, the transform and the palette), and replace the nearest-colour scan with a 5-5-5 RGB lookup table — 32KB, built once per palette.
  • drawRegistered draws a full-resolution photo every frame. Correct as a drawing reference on a held frame; not something to do thirty times a second. Pre-scale each source frame to raster size once, at load.

The rule both are instances of: nothing that reads source pixels may run on the transport path. That is the same line the bake boundary draws, which is why drawing it once, in stage 3, pays twice.

Porting order

Tests first, and not as discipline — as an oracle.

  1. Port synth.js and selftest.js first, to cljs.test. The synthetic track is the only fixture with ground truth, and the 105 assertions encode invariants that are invisible to inspection: the ring-simplicity check, the bowtie, the iris-pairing vote fed a deliberately swapped track.
  2. Run both implementations on the same synthetic track and diff numerically while porting each pure module. fitSimilarity and procrustesMean should agree to 1e-9; a disagreement is a port bug, not float noise.
  3. Port landmarks, geom, ring, raster, take — mechanical, and they are where the invariant comments live. Carry the comments across verbatim. They are the most valuable text in the repo and every one of them is a bug that already happened.
  4. Port flow/measure/* by lifting out of pipeline.js, then split condition and key out of it.
  5. Build db, store, clock, and stage 6 + 7 against a single hardcoded clip. At this point the synthetic take should render, with no UI beyond a canvas and a play button.
  6. Only then the UI, the modes, the timeline.
  7. Add the ring-simplicity and no-RGB checks as specs, so they fire on authored data too and not only in tests.

Three things are cheap in step 3 and expensive after step 6, so they go in early even though nothing needs them yet: keys as a map keyed by frame, fractional indices on :z and cel layer order, and leaf addressing for tier 1 — the paths under Collaboration, used as the shape of app-db even while single-player. None of them cost anything on day one and all three are retrofits that touch every namespace.

The wiring cross-check in selftest.js — every el('id') must exist in index.html — becomes unnecessary and should be deleted rather than ported. It is a test for a failure mode that Reagent does not have.

Decisions I would defer

  • Whether raster stays JS. 320×200 is 64,000 pixels and CLJS over a Uint8Array with no seq allocation handles it comfortably; port it and measure. Do not port it into idiomatic map/reduce.
  • Interpolation. :hold is the only interp the aesthetic wants, and the channel model should still carry :interp per node, because a parented transform on a painted cel is the one place a tween might be right. Do not implement it until something needs it.
  • Multi-character scenes. character is in the entity model as one field and should stay a stub until there are two.
  • Audio per clip vs per sequence. One master track is right until it is not.
  • Whether bake B exists at all in v1. One clip does not need it; the flat typed-array layout should be designed now and built when the second roto track appears.
  • Sharing tier 2. Content addressing means a collaborator can fetch a bake instead of recomputing, and also means nothing breaks if they do not. Ship the cache-miss path first and treat blob sharing as an optimisation — the analysis artifact is the only one where it clearly pays, because it is the only one that costs minutes.