Step 9. The tier split was the work; Django was the easy half.
Tier 1 — the authored scene — is the document, and it is addressed as
independently versioned leaves rather than saved whole, so one vertex drag
cannot clobber a collaborator's keying. `domain/leaf` is the document as
path -> value; `domain/wire` puts it on the wire as transit, because JSON
has neither integer map keys nor keywords and a save would quietly turn
`{0 v}` into `{"0" v}`.
Tier 2 — the dense channel blocks — is content-addressed by a hash over
every input, with the detector version inside every key through the
analysis the block descriptor names. `flow/address`'s `block-knobs` is the
invalidation table, and `address-test` does not trust it: it re-freezes the
take once per knob and asserts the biconditional, that a block's bytes
changed if and only if its key changed. That found `brow-pos` not depending
on `contour-avg` — the brow ring is smoothed, the raise is not.
Tier 3 — frames and audio — is served by the hash of its bytes out of the
same store. A manifest now names frames and carries a URL for each, so the
frame layout stopped being a shared secret between a shell script and a
ClojureScript namespace, and the `?v=` cache-buster went with it: a blob's
name is the hash of its contents, so a stale copy is not a thing that can
happen. The synthetic take's `audio.wav` moved to `static/arthur/` — an
asset the project owns, not an extraction that churns.
The server verifies rather than trusting a name it was handed: it
recomputes every key from the descriptor stored beside it, refuses an
analysis that declares no detector version, and refuses a document naming
blocks it does not hold. It hashes the descriptor TEXT, because JS prints
an integral double as `1` and Python as `1.0`, and a scheme where both ends
re-render the numbers disagrees on the first parameter that happens to be
whole.
Two loose ends from step 8 closed on the way. `pack` no longer takes a
`(track, frame)` predicate whose call sites each re-derived a feature from
an index — every track names the feature it follows, which deleted five
hand-maintained mappings. And `:dev-http` is gone: Django serves the page,
shadow-cljs only builds into the staticfiles tree.
227 CLJS tests, 31 Django tests, green.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
36 KiB
arthur — the animation model
The data that describes a moving picture: what the primitives are, how they nest, how they change over time, and how rotoscoped and hand-authored work end up being the same thing with one flag between them.
docs/design.md is the aesthetic argument. docs/architecture.md is where the
code goes. This is the type that both of them are about.
What this replaces
docs/design.md has a table of five kinds of part — plate, feature, interior,
primitive, scalar — each with its own source, vocabulary and interpolation. That
table is a good description of where data comes from and a bad description of
what data is, and the current code follows it too literally: eyes, brows,
teeth and mouth each get their own build function, their own key shape and their
own path through the prototype's writer.
They are all one thing. A part is a node with channels, and the five kinds collapse into differences of which channels exist and who filled them in.
Prior art, and what each one gets right
| System | The idea worth taking |
|---|---|
| Flash / SWF | A library of symbols and a timeline of instances at depths. "Framed" content that simply exists on a frame, versus tweened content. DefineMorphShape requires matching vertex counts — the fixed-topology rule, arrived at from the other direction. |
| Blender | Animation is addressed by path into the data (location[0]), not stored as fields on the object. An Action is a bag of F-Curves. That decoupling is what makes the dope sheet, the graph editor and the NLA three views of one dataset. Also: parenting captures a parent_inverse so the child does not jump. |
| After Effects | Every leaf property is animatable, uniformly. Property groups form a tree. Pre-comps nest arbitrarily and a pre-comp is just a layer. |
| Lottie | The uniform property shape: {a: 0, k: <value>} or {a: 1, k: [<keys>]}. One representation for static and animated, which is exactly "framed or keyframed". |
| Grease Pencil | A 2D layer holds frames at frame numbers, and a frame holds until the next one. Hold is the default, not a special case. |
What none of them get right for this project: colour. All four store RGB on the
shape. docs/design.md forbids that, so colour is a palette index here and it is
a channel like any other.
The one idea
Analysis is a channel generator. It does not produce a different kind of data; it produces keys, densely, on the same channels a hand would fill in sparsely. So:
footage ──▶ analysis ──▶ FREEZE ──▶ channels on nodes ──▶ evaluate ──▶ raster
▲
hand authoring ──┘
Freezing is not a conversion into a second format. There is one format, and freezing fills it in. That is what makes "the only difference is a special flag" literally true: the flag is provenance on a channel, and nothing in the renderer reads it.
Node
A node is an instance in the scene. The tree is stored flat, with parent pointers — never as nested maps.
{:id :mouth
:name "mouth"
:kind :poly ; :poly :disc :rect :group :bitmap :symbol
:parent :head ; nil at the root
:z "a3" ; fractional index, ordered among all siblings
:symbol nil ; or :sym/blink — see Symbols
:stencil :mouth-in ; colour-key clip; structural, not a channel
:span [0 240] ; in/out in the parent's frame space
:pinv [1 0 0 1 0 0] ; parent-inverse, captured when parented
:channels {...}}
Flat with pointers, for four reasons that all point the same way: any node is addressable without a walk; reparenting is a one-field write rather than a subtree move; an edit to a leaf does not change the identity of its ancestors, so re-frame's structural sharing keeps ancestor subs from invalidating; and it is what lets every node be its own sync leaf. Flash, Blender and AE all store it this way.
:span is Lottie's ip/op and Flash's PlaceObject/RemoveObject: the range
over which the node exists at all. Distinct from a [:vis] channel, which
blinks an existing node on and off.
Subjects and tracked features
Scene nodes describe drawings, not tracking identity. A scene may also carry a
flat :features map. A feature ID stays stable for the whole clip, including
frames where that feature is occluded and later reappears:
:subjects {:face-1 {:id :face-1}}
:features
{:eye-r {:id :eye-r :subject :face-1 :area :eye
:nodes [:eye-r :eye-r-in :iris-r :pupil-r] :params {}}
:eye-l {:id :eye-l :subject :face-1 :area :eye
:nodes [:eye-l :eye-l-in :iris-l :pupil-l] :params {}}
:mouth {:id :mouth :subject :face-1 :area :mouth
:nodes [:mouth :mouth-in] :params {}}
;; The teeth are their OWN feature and not three nodes of the mouth. A feature
;; carries the params of exactly one area, and the teeth have an `:area :teeth`
;; of their own — the otsu threshold, the tongue rejection, the radial contour's
;; vertex budget — which could not be reached if they were part of `:mouth`.
;; The coupling that made them look like the mouth's is real and is enforced
;; elsewhere: `:teeth` is STENCILLED by `:mouth-in`, and a node whose stencil drew
;; nothing is dropped, so an absent mouth takes the teeth with it without either
;; of them sharing an absence mask. An earlier draft of this block listed them
;; together; the code is right and this document was wrong.
:teeth {:id :teeth :subject :face-1 :area :teeth
:nodes [:teeth] :params {}}}
:groups
{:eyes-1 {:id :eyes-1 :kind :eye-pair :subject :face-1
:members [:eye-r :eye-l] :params {}}}
An eye pair is an explicit relationship between one or two eyes of the same subject. It may have one member when only one eye has been identified; it does not invent a second eye. Five subjects with nine identified eyes can have four two-member pairs and one one-member pair. Each eye still has its own feature ID and presence track. A group is a settings association, not a scene parent or a tracking ID. Membership lives only on the group, avoiding a second pointer on the feature that could disagree with it.
Each feature resolves settings from its area's definitions, then its group,
then its own :params. An eye can therefore inherit a pair setting or override
it without changing its partner. Removing it from a pair copies its effective
values into the feature first, so the result does not jump. Feature identity
and pair membership are clip-wide; a future parameter track can vary values
over time without splitting a feature at an observation gap.
Parameter definitions live in one registry: key, default, applicable area, value constraints and affected areas. The registry supplies the take's defaults today. The parameter UI and regeneration from edited values are later work.
Dense channel state records whether a measurement exists for that feature on
that frame. Occlusion means absent data on that frame, not a false [:vis]
value and not the end of the feature's identity. A full-face detection failure
makes all its features absent. A single occluded eye need only make that eye's
channels absent. Footage can carry explicit feature absence intervals in its
manifest, with one-based inclusive source frame numbers, for example
"feature-absence": {"eye-r": [[10, 14]]}. The loader expands these into
per-frame observation tracks before measurement. Unobserved landmarks may fill
rectangular numeric buffers, but they cannot contribute to an eye's contour,
blink or shared gaze. When one eye is absent, gaze uses the observed eye.
Until a detector supplies feature-level confidence, footage without annotations
uses the full-face detection mask as the fallback; it must not claim to detect
individual occlusions that it cannot see.
Channel
Every animatable property is a channel, and channels are addressed by path:
:channels
{[:xform :pos] {:animated? false :value [0.0 0.0]}
[:xform :rot] {:animated? false :value 0.0}
[:xform :scale] {:animated? false :value [1.0 1.0]}
[:xform :skew] {:animated? false :value [0.0 0.0]}
[:xform :anchor]{:animated? false :value [0.0 0.0]}
[:geom :pts] {:animated? true :interp :hold :dense {...} :generated {...}}
[:style :color] {:animated? false :value :skin-dark}
[:vis] {:animated? true :interp :hold :keys {0 true, 37 false}}}
A path is a vector, not a string — CLJS maps take vectors as keys natively,
so Blender's data_path idea arrives with no parsing. The set of valid paths for
a node follows from its :kind, and that is a spec, not a schema migration.
Three channel shapes, and the uniformity across them is the point:
;; FRAMED — one static thing. No animation, no vertex correspondence to worry
;; about. A painted background cel is this.
{:animated? false :value v}
;; KEYED — sparse, authored, in the document. Undoable and syncable.
{:animated? true :interp :hold :keys {0 v, 4 v, 12 v}}
;; DENSE — generated, one value per frame, held in tier 2 as a typed array.
{:animated? true :interp :hold
:dense {:store "sha256:…" :offset 0 :stride 40 :frames 600}
:generated {...}}
:interp defaults to :hold, which docs/design.md requires of every cut part.
A key may carry its own :interp to override the channel's, which is how Lottie
and Blender both do per-key easing; nothing uses it yet and the door is cheap to
leave open.
Keys are a map by frame, not a list
Already argued in docs/architecture.md for merge reasons; here it also gives
"the most recent key at or before f" as a rsubseq on a sorted map instead of
a scan. Store a plain map in the document — transit and JSON both lose
sortedness — and build the sorted index in the resolver.
The flag lives on the channel, not the node
:generated {:by :roto/lips-outer
:analysis "sha256:…" ; which analysis artifact
:params {:verts 8 :contour-avg 1 :aperture-cut 0.004}}
Present means the UI offers a parameter panel and a re-freeze button. Absent means the UI offers the keys directly. The renderer never reads it.
It belongs on the channel rather than the node because a node routinely wants
both at once: a mouth whose [:geom :pts] is rotoscoped and whose [:xform :pos]
is hand-animated to sit on a plate. Putting the flag on the node would forbid the
most useful thing in the model.
Channels are layered
A channel is a base plus optional override layers, and a layer declares how it combines:
{:animated? true :interp :hold
:dense {...} :generated {...}
:over [{:blend :offset :keys {88 [2 0], 96 [0 0]}}
{:blend :replace :keys {104 [[3 7] [4 7] …]}}]}
:offsetadds a delta to the base. "Nudge the mouth two pixels right for ten frames" survives a re-freeze at different parameters, because it was never a position — it was a correction.:replacewins outright. For the frame where detection simply failed.
This is what docs/design.md means by an override layer, and it is why
re-freezing is safe: the base is regenerated, the layers are untouched. It is
Blender's NLA blending and AE's effect stack at one property.
Layers are what "set it by hand" means for anything measured, and the measured
channel does not need to know. A hand-set gaze is an :over on
[:xform :pos] of the iris; a hand-set mouth shape is an :over on
[:geom :pts]. Turning the gaze-step or gaze-dwell knob regenerates the base and
leaves the correction alone, which is the entire reason a correction is stored as
a layer rather than written into the track.
A :replace layer overrides absence, an :offset layer does not. Sampling a
channel is: read the base, then apply the layers — and the base coming back
absent does not short-circuit that. :replace is explicitly for the frame
where detection failed, so it has to be able to supply a value where there is
none; :offset is a delta, and there is nothing to nudge, so an offset over an
absent base stays absent. Implemented the obvious way — bail out on absence
before reaching the layers — the one case the feature exists for is the one case
it would not cover.
One signal, two nodes
Gaze is deliberately one measurement shared by both eyes: at this size the
per-eye difference is noise, and independent noise reads as wall-eyed
immediately, which is the most expensive artefact on a face. But it is stored as
[:xform :pos] on :iris-r and on :iris-l, which are two channels on two
nodes with two different parents — so the invariant lives in measure and
nothing in the document enforces it.
That matters as soon as either one can be overridden by hand, because an :over
on one iris alone reproduces exactly the artefact the shared measurement exists
to prevent. Until drivers exist, the override is on both or on neither, and
that is a rule the UI has to keep rather than one the data can.
This is the case that will eventually justify drivers — one value, evaluated
once, feeding several channels — which is why gaze is named in Deferred as the
obvious first one. Nothing here forecloses it: a driver needs a place in the
document and a :driven-by on a channel, both of which are additive, and an
absent key means "not driven". So it stays deferred, and the shape does not have
to change to allow it.
Transform: decomposed, never a matrix
{:pos [x y] :rot θ :scale [sx sy] :skew [kx ky] :anchor [ax ay]}
Stored decomposed for two reasons. Each component has to be independently keyframable, which is the entire point of channels. And interpolating matrix entries is meaningless — a rotation tweened through its matrix shears on the way.
Composition, per node:
local = T(pos) · T(anchor) · R(rot) · K(skew) · S(scale) · T(-anchor)
world = world(parent) · pinv · local
:anchor is Flash's registration point and Blender's origin: rotation and scale
happen about it, and getting it wrong is why hand-placed parts swing rather than
turn.
:pinv is Blender's parent_inverse, captured at the moment of parenting so the
child does not jump when it acquires a parent. Small, and its absence is the kind
of thing that makes a parenting feature feel broken.
The similarity fit already produces a decomposition. fitSimilarity returns
{s θ tx ty}, which drops straight into [:xform :scale], [:xform :rot] and
[:xform :pos] with no conversion. The analysis output and the animation model
meet without an adapter, which is a sign the decomposition is the right one.
What space geometry is in
[:geom :pts] is always in the node's own local space, and the transform
chain says what that means. There is no global geometry space and no decision
to make about one.
| Node | Its local space | Why that one |
|---|---|---|
| a rotoscoped feature | head-local, isotropic, unit = one image height | what the anchor fit already produces; the xform to raster is not applied and not stored |
| a painted cel | the stage, in pixels, grid-snapped | the artist is placing pixels, so the pixel grid is the thing being authored |
| a primitive under a feature | its parent's | the iris is positioned on the lid ring, not on the stage |
This looks like a small clarification and it removes a whole class of argument.
The prototype bakes the framing into the numbers: toRasterRing applies
makeXform, which centres on the face oval's bounding box and zooms until the
face is 80% of the raster height, so every stored vertex carries a cropping
decision that was made once, at analysis time, from one frame's landmarks.
Dropping that step is a deletion, not a feature, and after it the framing is
simply a transform on a node.
Grid snapping belongs to the cel and not to the roto, for the same reason: a cel is authored on the grid and a traced contour is not. So it is a property of a node's space rather than a rule about all geometry, and the tension between "integer polygons" and "arbitrary placement" was never real.
Each dense block therefore carries its own fixed-point scale in its header,
because a block in image-height units and a block in stage pixels need different
ones to fill an Int16 usefully.
There is no camera node
A camera is a global transform over everything, and nothing here wants one.
Placing the face on the stage is a transform on a node, which already exists;
what is not on the stage hangs off the edges and the canvas clips it. Every fill
in domain/raster already clamps rather than assuming it is inside, so drawing
past the edge is not a feature to add.
Project dimensions are therefore independent of the footage. A 1440x1920 portrait clip composited onto a 320x200 stage is not a problem to solve — the head is placed and scaled where it belongs and the rest of the frame is simply not on stage. The full frame stays available for tracing without being visible, and those are different requirements.
The anchor: stabilisation is a channel, not a mode
stabilize produces {s, θ, tx, ty} per frame, which is exactly
[:xform :scale], [:xform :rot] and [:xform :pos]. So removing the head's
motion is not a pipeline setting — it is a question of which node holds that
motion, and the answer is one channel definition:
;; locked: the head sits still, for tracing and for judging articulation
[:xform :pos] {:animated? false :value [0.0 0.0]}
;; as filmed: the head moves around the stage
[:xform :pos] {:animated? true :interp :hold
:dense {:store "sha256:…" :stride 2 :frames 600}
:generated {:by :anchor/similarity}}
;; per plate: the head snaps at each selected frame and holds
[:xform :pos] {:animated? true :interp :hold :keys {0 […], 12 […], 23 […]}}
The three modes are the three channel shapes, on one channel, on one node. The
third is the one a plate strip wants — the head pose is stable for exactly as
long as a drawing is on screen — and it costs nothing because :keys already
exists. Its frame set is the kept-frame set, which is suggestPlateFrames in the
prototype and belongs to painting rather than to measurement.
Always measure, always store factored, toggle the parent. The fit is computed and the geometry is stored head-local in every mode, and only the parent's channel changes. Two things downstream require it, and both would be lost by making this an analysis-time switch:
- Smoothing. "Smooth the transform, never the contour" only means anything while the two are separate.
- Key selection. A velocity minimum is "articulation paused" in head-local space and "the head happened to be still" in image space.
It also makes the toggle an edit to the document rather than a reason to re-analyse: tier 1, undoable, syncable, and instant.
Two nodes, because two different things want that transform
:face group — AUTHORED. where the face sits on the stage, and how big.
:head group — MEASURED. the head's motion, or identity.
:mouth :mouth-in :teeth :lid-r :lid-l :brow-r :brow-l …
Switching modes rewrites :head and never touches :face, so it cannot move
something that was placed by hand. A group node is free, and keeping the authored
and the measured transform apart is the whole reason the transform is decomposed
in the first place.
Time maps — exposure, lead and symbol timing are one thing
Every node may map the frame it is evaluated at:
:time {:mode :inherit} ; the default, and almost always right
:time {:mode :map :expose 2 :offset -1 :rate 1.0 :loop? false}
:time {:mode :map :source-fps 30 :sample-fps 12} ; root: lower picture cadence
Three features that look unrelated are this one mechanism:
- exposure is
⌊f/n⌋·n, - picture fps quantises source time to a chosen picture grid, then reads the latest source pose at or before that time; source analysis and audio keep their original cadence,
- mouth lead is
f + k, - a symbol instance's timing is
(f - at)·rate + in, with optional looping.
Composed along the nesting chain, outermost first. Two rules follow, and they are different rules:
- Exposure inherits strictly.
docs/design.mdis emphatic that everything rides one grid, because a head cutting on odd frames against a mouth cutting on even ones reads as two performances. The model permits a per-node grid; the default must be:inherit, and setting it lower is a deliberate act the UI should make feel like one. - Offset is per-node by design. Mouth lead applies to performance nodes and not to the plate, which is the whole point of it — so the offset genuinely belongs at the node, not the clip.
Timelines, and why a scene is one
A timeline is an ordered bag of nodes in its own frame space:
{:frames 91
:palette {...} ; see Palettes
:nodes {id -> node}}
That is the whole type, and everything that holds nodes is one of these:
- a clip's scene is its root timeline,
- a symbol in the library is a timeline,
- a node with
:kind :symbolis an instance of one.
An earlier draft of this document had a scene and a :kind :timeline symbol as
two structures with the same fields and never said they were the same thing.
They are. Flash's _root is a MovieClip; After Effects' "a pre-comp is just a
layer" is already in the prior-art table above. Collapsing them is what makes
nesting arbitrary and free, rather than a feature to be added.
Two axes of nesting, and they are different
This is the distinction the flat-storage rule is about, and conflating the two is why "nested" and "flat with parent pointers" sound contradictory when they are not:
| Axis | What nests | How it is stored |
|---|---|---|
| parent / child | transform composition within one timeline | flat, with parent pointers — never nested maps |
| instance | a timeline inside another timeline | by reference into the library |
Each timeline is flat. Timelines nest. Every argument for flat storage — addressability, one-field reparenting, structural sharing, per-node sync leaves — is about the first axis and is untouched by the second.
The instance boundary is also the only place the frame space changes. Within
a timeline, :time is exposure and lead: a shift inside one space. At an
instance it is (f - at)·rate + in, into a different one. That is why :rate is
meaningless on an ordinary node and why sampling one must fail loudly rather than
be ignored.
What is scoped to a timeline
Three fields on a node only have meaning relative to a timeline, and the answer for all three is the same — their own:
:zorders among siblings; a node cannot interleave with nodes inside a nested instance. The instance occupies one position in its parent's order and its contents sort beneath it, which the z path gives for free by being a vector.:stencilnames a node in the same timeline. A colour key does not naturally respect a boundary — it is just pixels — so this is a rule rather than a consequence, and it is Flash's rule for masks.:spanis in the parent node's frame space.
Instances
A node with :kind :symbol and :of :sym/blink places one. Its own channels
compose over the symbol's, so one definition is placed many times and tinted,
offset or retimed at each placement — that is how a three-frame blink is reused
at frames 40, 88 and 200 without copying it.
This is also where docs/design.md's "closed vocabulary is right for the head"
lands: a plate library is a set of :sym/head-* timelines, and the strip chooses
which is instanced on which frame.
Cursors and point buffers are per-instance, not per-node. Two instances of
one symbol sit at different frames in their own space, so they cannot share a
reading head over the same channel. The resolver keys its caches by the instance
path, not by node id — which is a detail of Making it fast below, and the one
place symbol nesting is not free.
Evaluating a frame
(defn eval-frame
"Scene at clip frame f -> draw ops in z order. Pure."
[scene f] ...)
- Walk nodes in topological order by parent depth (cached; recompute only when parentage changes).
- Skip nodes outside
:span. - Apply the node's time map to get its own local frame
fn. - Sample each channel at
fn: a map lookup for framed, a sorted-index lookup for keyed, an array read for dense. Then apply:overlayers. - Compose
worldfrom the parent's. - Transform geometry into raster space, writing into a preallocated buffer owned by the node.
- Emit
{:kind :poly :pts buf :n 20 :color idx :stencil id}. - Sort by resolved
z.
The op list is the boundary with stage 7 in docs/architecture.md: the
rasteriser takes ops and knows nothing about nodes, channels or time.
A photographic underlay is not an op. The registered source frame that an
animator traces over is a reference, not output, and it may not enter the indexed
buffer — the same rule docs/architecture.md already sets for handles and
vertex boxes. It is a drawImage at an affine on a separate canvas, which clips
at the canvas edge for free, and the only thing it needs from the model is the
world transform of the node it rides:
(world-of resolver :head) ;; -> Float64Array[6]
Composed with image-pixels-to-local — both axes divided by imgH, never by
their own dimension — the photo is registered with the shapes by construction,
and an unregistered underlay is merely decorative. The tracing editor chooses
which source frame to show under a cel. That reference choice is independent of
the finished picture fps and does not change the dense analysis track. A cel can
therefore use any useful source frame as its drawing reference, even when that
frame is not one of the displayed picture poses.
A photo that has to sit between two drawn layers is the case that would make it
a :bitmap node with an op of its own. Nothing wants that yet: a reference is
either under everything or over everything at low alpha.
Making it fast in CLJS
Three things, and only these three matter:
-
Decomposed and persistent for storage; flat and mutable for evaluation. Composed transforms are 6-element
Float64Arrays, not maps. Every renderer does this; the storage form and the evaluation form are allowed to differ. -
A cursor per channel. Playback is sequential, so "most recent key at or before
f" is an advance of a saved index, O(1) amortised. Binary search only on a seek. This is the difference between arsubseqallocation per channel per frame and none. -
Preallocated point buffers per node. Fixed topology means the size is known at freeze time, so the vertices — the overwhelming majority of the per-frame bytes — are written into a buffer the node already owns. A frame still allocates its op maps and the sorted op vector; that is a dozen small objects against hundreds of points, and pooling them would buy nothing and cost the ability to pass an op list around as plain data. At 30fps, per-vertex allocation is the thing that will make this stutter.
Because the buffers are reused, ops must be consumed before the next frame is asked for. That is the contract the rAF loop wants anyway: it reads, blits, and dispatches nothing.
What is in app-db, and what is not
| In app-db (tier 1) | In tier 2, behind a handle |
|---|---|
| nodes, parentage, z, spans, stencils | dense channel blocks |
channel definitions, :interp, :generated |
analysis artifacts |
framed values, keyed keys, :over layers |
preallocated eval buffers |
| library / symbol definitions | composed transform scratch |
The rule: anything a human placed is in the document; anything a generator
produced is a handle. Which is the same line docs/architecture.md draws for
sync and baking, arrived at again from the renderer's side.
The current parts, in this model
Proof that it covers what exists, not just what is wanted:
| Now | Becomes |
|---|---|
mouth outer ring, every frame |
node :mouth, [:geom :pts] dense, :generated {:by :roto/lips-outer} |
mouth_in, hidden below aperture |
node :mouth-in, parent :mouth, [:geom :pts] dense + [:vis] dense |
teeth from image content |
node :teeth, stencil :mouth-in, [:geom :pts] dense, :generated {:by :interior/teeth} |
| lid rings | nodes :lid-r/-l, [:geom :pts] dense |
lash line (offsetRing) |
not data — a stage-6 parameter on the node, {:grow px} |
| iris disc | node :iris-r, :kind :disc, parent :lid-r, stencil :sclera-r, [:xform :pos] dense (quantised at freeze), radius framed |
| square pupil | node :pupil-r, :kind :rect, parent :iris-r, stencil :iris-r |
| brow ring + quantised raise | node :brow-r, [:geom :pts] dense (the traced ring with height removed), [:xform :pos] dense (the quantised raise). The decomposition design.md insists on is two channels. |
| head plate, kept frames | node :head, :symbol per instance, keys on [:symbol] at kept frames |
makeXform face-oval crop |
gone. Placement is [:xform :*] on :face; the stage clips |
stabilize transforms |
[:xform :*] on :head — framed identity, dense, or keyed at kept frames |
| registered underlay | not data — a UI layer riding (world-of resolver :head) |
| painted background cel | node per layer, [:geom :pts] framed, [:style :color] framed |
mouth lead |
:time {:offset k} on performance nodes only |
exposure |
:time {:expose n} on the clip root, inherited |
| picture fps | :time {:source-fps s :sample-fps p} on the clip root, applied after analysis |
| hand correction | an :over layer, :offset or :replace |
The brow row is the one worth looking at twice. docs/design.md argues at length
that the traced ring already contains the height, so the quantised raise must be
measured out and put back or the brow moves twice. In this model that is not
an argument to remember — it is two channels on one node, and getting it wrong
would mean writing the height into both.
Palettes
Three levels, and keeping them apart is what makes a palette swap a reinterpretation rather than an edit:
| Level | Holds | Lives on |
|---|---|---|
| tone | which mark this is — :skin-dark |
[:style :color], a channel on the node |
| ramp | what that tone looks like here | :palette, a channel on the timeline |
| the ramps | every named palette | the project |
A node names a tone, never a colour and never a ramp. Which ramp the tone is
read in is decided by the timeline the node is in. So the same drawing reads day
or night without one stored value changing — which is the entire payoff of
indexed colour, and is why docs/design.md forbids sampled RGB: once a shape
holds a measured colour there is nothing left to reinterpret.
Named palettes are variants over one tone vocabulary, not arbitrary colour
lists. :day and :night both define :skin-dark; that is what keeps a swap
total and keeps docs/design.md's closed vocabulary closed. A tone the ramp in
scope does not define resolves to the loud magenta, like any other missing index.
The scope rule
:palette on a timeline is a channel like any other:
{:frames 91
:palette {:animated? true :interp :hold :keys {0 :day, 48 :dusk, 72 :night}}
:nodes {...}}
Absent means inherit from the instancing context. Present means this
timeline's content is read in that ramp, and it travels with the timeline — a
symbol authored against :night stays night wherever it is placed. That is
lexical scope, and deliberately: a character with their own palette is a
character, not a decoration of whichever scene they were dropped into.
Composition is the same walk as :time — down the instance chain, innermost
set palette wins. An enclosing timeline's palette therefore applies to
everything inside it that does not set its own, which is adjustment-layer
behaviour with no adjustment layer in it. It is just scope.
And because it is an ordinary channel, a project switches palette over time with keys on the root timeline, a child timeline switches on its own, and neither knows about the other.
One index space, partitioned by palette
A raster is one Uint8Array and an index means one colour in it, so two ramps in
one frame cannot both own index 2. The resolution: the output index space is
the concatenation of the named palettes, and a tone resolves to
palette-base + tone-index.
Everything downstream is then unchanged — one buffer, one flat table for
->rgba, no per-frame palette construction, and an index does not change meaning
between frames, so bakes and thumbnails stay valid.
Two consequences worth stating rather than discovering:
- The limit is real and reachable. 256 indices over a nine-tone vocabulary is twenty-eight palettes. Detect it and say so; do not let it arrive as wrapped colour.
- It makes the stencil sharper. A stencil is a colour key, so two nodes sharing a tone share a stencil — a genuine weakness of the technique. Partitioning the index space by palette means two nodes in different palettes no longer collide at all, and the resolved stencil picks up whichever index the stencil node actually drew in.
Where it is resolved
At the op boundary, and nowhere else. [:style :color] holds a keyword all the
way through evaluation; the walk carries the palette in scope the same way it
carries the parent transform and the local frame; the op carries a resolved
index. The rasteriser never sees a tone name and the node never sees an index.
This also means the palette is a parameter of evaluation, not a global. The resolver takes it alongside the store.
Format on disk and on the wire
Tier 1 is EDN/transit: the node tree, channel definitions, framed values, keys, layers, library. Kilobytes, human-readable, diffable, and leaf-addressable for sync.
Dense blocks are separate content-addressed binaries — Int16Array for
geometry, Float32Array for transforms — with a small header naming the channel
path, frame count, stride, and the fixed-point scale of the node-local space
the block is in. Geometry is stored in the node's own space, not in raster space;
see "What space geometry is in".
Not Lottie internally, despite the property shape being borrowed from it. Lottie has no palette-indexed colour, its shapes are bezier with in/out tangents where these are integer polygons, and its interpolation defaults are the opposite of what is wanted. It is a fine thing to write out one day and a bad thing to store.
Output is deliberately not specified here. The target is encoding video in the browser, which touches the op list and nothing above it — a writer consumes frames, and frames are what stage 7 already produces.
Deferred
- Per-key easing. The structure allows it; nothing should use it until a parented transform on a painted cel asks for it.
- More than two channel layers. The
:overvector is already a list; a real blend stack with weights is the NLA, and it is not needed to fix a bad frame. - Skew beyond the field.
[:xform :skew]is in the transform and in the composition order from the start, because adding a component to a decomposition later means migrating every stored transform. - Instance channel overrides on symbols. Compose-over is specified; only colour and transform need it at first.
- Constraints and drivers. Blender's other half. A gaze that aims at a null object is the obvious first one, and it is a long way off. Until then the one gaze shared by two iris nodes is a UI rule, not a stored relationship — see "One signal, two nodes".