arthur/docs/animation-model.md
Your Name e7f5f82845 A drawing's origin is the middle of what it draws
A shape keyed from the bottom left to the top centre with a 360° turn on the way
left the stage completely in the middle of the spin and came back. The keys were
right and every frame between them was wrong, which is the signature of a wrong
pivot and a full turn: 0° and 360° are the only two frames where a wrong pivot
cannot be seen at all.

`paint/new-shape` stored a stroke EXACTLY AS DRAWN, in the containing symbol's
coordinates, and wrote no `pos`. So a drawing's origin was the SYMBOL's origin —
on the stage, its top-left corner. `node/local!` turns and scales about the
node's own origin and nothing else, deliberately, since `[:xform :anchor]` was
deleted in 925c12f. A node whose origin is nowhere near its content therefore
turns about nowhere near its content: the reported shape orbited at a radius of
126 px on a 320x200 stage.

It could not be seen while a drag was the only way to turn something, because
`gesture/about` solves for the `pos` that holds the chosen pivot still and
`turn` wrote it alongside the rotation — exactly right on the frame of the drag.
But that solution is `p' = c + R(θ)(p − c)`, an ARC, and `pos` interpolates along
the CHORD. Right on a drag, right on a key, wrong on every frame between two.

So `paint/centred` splits a stroke into a ring about its own middle and the `pos`
that puts it back, and `new-shape` is the one place every drawing is born — the
pen, the brush, and each piece the eraser leaves. The pivot rule is unchanged,
the middle of what the node draws; for a drawing that point is now its ORIGIN, so
`gesture/at-origin?` holds, `turn` writes `rot` alone, `scale` writes `scale`
alone, and a keyed turn is right on every frame. `pos` goes back to being the
motion path it reads as. Hand-authored scenes were always written this way:
`demo/scene.edn`'s card is `[-44 -30 44 -30 44 30 -44 30]` with its place in
`pos`.

NOT the universal rule, and `a-face-part-scales-about-its-own-middle` is why. A
measured part's points and position are dense tier-2 geometry in the footage's
space and cannot be re-originated, so its pivot is not its origin and `about` is
the only thing that will hold it; the same is true of an instance, whose origin
IS its symbol's coordinate system. Both still drag correctly about their middle,
both are inexact if that drag is keyed, and for both a pivot that has to persist
or be keyed is a peg — which is a node, so its pivot is its own origin, so it
collapses again one level up.

`cut/erase` took EVERY leftover piece back through `world⁻¹`, the cut shape's own
coordinates, and handed the offcuts to `new-shape`, which gives them a fresh
identity transform. Those two spaces coincide only while a shape has `pos [0 0]`,
which was every shape, so erasing anything that had been moved already scattered
its offcuts, silently. The kept piece comes back through `world⁻¹` and the new
ones through `parent⁻¹`, the space a node's `pos` lives in.

And `::adjust-last` re-traces the same stroke from stage pixels, so it goes
through `paint/place-points` rather than writing symbol-space points into a node
that now has a position of its own.

No schema change: the same fields, better values. An existing document keeps
evaluating exactly as it does now.

`a-keyed-turn-holds-its-pivot-between-its-keys` checks all 31 frames of the
tween. Checking the keys is what let this through. 604 CLJS tests pass.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-10-06 10:37:31 -04:00

904 lines
46 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# arthur — the animation model
The revised target for lanes, occurrences, source playback, shared editing, and
multi-view UX is [The Lane Model](lane-model.md). It supersedes conflicting
proposals below. Backward compatibility is not required; this document still
contains descriptions of earlier shapes and planned features.
The data that describes a moving picture: what the primitives are, how they
nest, how they change over time, and how rotoscoped and hand-authored work end
up being the same thing with one flag between them.
`docs/design.md` is the aesthetic argument. `docs/architecture.md` is where the
code goes. This is the type that both of them are about.
## What this replaces
`docs/design.md` has a table of five kinds of part — plate, feature, interior,
primitive, scalar — each with its own source, vocabulary and interpolation. That
table is a good description of **where data comes from** and a bad description of
**what data is**, and the current code follows it too literally: eyes, brows,
teeth and mouth each get their own build function, their own key shape and their
own path through the prototype's writer.
They are all one thing. A part is a **node** with **channels**, and the five
kinds collapse into differences of which channels exist and who filled them in.
## Prior art, and what each one gets right
| System | The idea worth taking |
| --- | --- |
| **Flash / SWF** | A **library of symbols** and a timeline of **instances** at depths. "Framed" content that simply exists on a frame, versus tweened content. `DefineMorphShape` requires matching vertex counts — the fixed-topology rule, arrived at from the other direction. |
| **Blender** | Animation is **addressed by path into the data** (`location[0]`), not stored as fields on the object. An Action is a bag of F-Curves. That decoupling is what makes the dope sheet, the graph editor and the NLA three views of one dataset. Also: parenting captures a `parent_inverse` so the child does not jump. |
| **After Effects** | Every leaf property is animatable, uniformly. Property groups form a tree. Pre-comps nest arbitrarily and a pre-comp is just a layer. |
| **Lottie** | The uniform property shape: `{a: 0, k: <value>}` or `{a: 1, k: [<keys>]}`. One representation for static and animated, which is exactly "framed or keyframed". |
| **Grease Pencil** | A 2D layer holds frames at frame numbers, and a frame **holds until the next one**. Hold is the default, not a special case. |
What none of them get right for this project: colour. All four store RGB on the
shape. `docs/design.md` forbids that, so colour is a palette index here and it is
a channel like any other.
## The one idea
**Analysis is a channel generator.** It does not produce a different kind of
data; it produces keys, densely, on the same channels a hand would fill in
sparsely. So:
```
footage ──▶ analysis ──▶ FREEZE ──▶ channels on nodes ──▶ evaluate ──▶ raster
▲
hand authoring ──┘
```
Freezing is not a conversion into a second format. There is one format, and
freezing fills it in. That is what makes "the only difference is a special flag"
literally true: the flag is provenance on a channel, and nothing in the renderer
reads it.
## Node
A node is an instance in the scene. The tree is stored **flat, with parent
pointers** — never as nested maps.
```clojure
{:id :mouth
:name "mouth"
:kind :poly ; :poly :disc :rect :group :bitmap :symbol
:parent :head ; nil at the root
:z "a3" ; fractional index, ordered among all siblings
:symbol nil ; or :sym/blink — see Symbols
:stencil :mouth-in ; colour-key clip; structural, not a channel
:span [0 240] ; in/out in the parent's frame space
:pinv [1 0 0 1 0 0] ; parent-inverse, captured when parented
:channels {...}}
```
Flat with pointers, for four reasons that all point the same way: any node is
addressable without a walk; reparenting is a one-field write rather than a
subtree move; an edit to a leaf does not change the identity of its ancestors, so
re-frame's structural sharing keeps ancestor subs from invalidating; and it is
what lets every node be its own sync leaf. Flash, Blender and AE all store it
this way.
`:span` is Lottie's `ip`/`op` and Flash's `PlaceObject`/`RemoveObject`: the range
over which the node exists at all. Distinct from a `[:vis]` channel, which
blinks an existing node on and off.
**Every node has the same two maps into its parent**, whatever kind it is:
- **space** — the matrix its transform channels compose to, times a `:pinv` if
it has been moved in from elsewhere;
- **time** — `local = rate · (parent − at)`, from `:time :at` and `:rate`,
identity when absent. `:span` and every key are in the node's **own** frames.
A move keeps a node's world maps and re-expresses them under its new parent:
the matrix becomes a `:pinv`, the time becomes a new `:at` and `:rate`, and its
channels, keys and span are not touched. Both maps are affine, so any depth of
nesting is one map and every move is one inverse. `node/time-of`,
`node/then-time` and `node/placed-span` are the time half; `clip/move-node` and
`clip/group` are the move.
### Subjects and tracked features
Scene nodes describe drawings, not tracking identity. A scene may also carry a
flat `:features` map. A feature ID stays stable for the whole clip, including
frames where that feature is occluded and later reappears:
```clojure
:subjects {:face-1 {:id :face-1}}
:features
{:eye-r {:id :eye-r :subject :face-1 :area :eye
:nodes [:eye-r :eye-r-in :iris-r :pupil-r] :params {}}
:eye-l {:id :eye-l :subject :face-1 :area :eye
:nodes [:eye-l :eye-l-in :iris-l :pupil-l] :params {}}
:mouth {:id :mouth :subject :face-1 :area :mouth
:nodes [:mouth :mouth-in] :params {}}
;; The teeth are their OWN feature and not three nodes of the mouth. A feature
;; carries the params of exactly one area, and the teeth have an `:area :teeth`
;; of their own — the otsu threshold, the tongue rejection, the radial contour's
;; vertex budget — which could not be reached if they were part of `:mouth`.
;; The coupling that made them look like the mouth's is real and is enforced
;; elsewhere: `:teeth` is STENCILLED by `:mouth-in`, and a node whose stencil drew
;; nothing is dropped, so an absent mouth takes the teeth with it without either
;; of them sharing an absence mask. An earlier draft of this block listed them
;; together; the code is right and this document was wrong.
:teeth {:id :teeth :subject :face-1 :area :teeth
:nodes [:teeth] :params {}}}
:groups
{:eyes-1 {:id :eyes-1 :kind :eye-pair :subject :face-1
:members [:eye-r :eye-l] :params {}}}
```
An eye pair is an explicit relationship between one or two eyes of the **same
subject**. It may have one member when only one eye has been identified; it does
not invent a second eye. Five subjects with nine identified eyes can have four
two-member pairs and one one-member pair. Each eye still has its own feature ID
and presence track. A group is a settings association, not a scene parent or a
tracking ID. Membership lives only on the group, avoiding a second pointer on
the feature that could disagree with it.
Each feature resolves settings from its area's definitions, then its group,
then its own `:params`. An eye can therefore inherit a pair setting or override
it without changing its partner. Removing it from a pair copies its effective
values into the feature first, so the result does not jump. Feature identity
and pair membership are clip-wide; a future parameter track can vary values
over time without splitting a feature at an observation gap.
Parameter definitions live in one registry: key, default, applicable area,
value constraints and affected areas. The registry supplies the take's defaults
today. The parameter UI and regeneration from edited values are later work.
Dense channel state records whether a measurement exists **for that feature on
that frame**. Occlusion means absent data on that frame, not a false `[:vis]`
value and not the end of the feature's identity. A full-face detection failure
makes all its features absent. A single occluded eye need only make that eye's
channels absent. Footage can carry explicit feature absence intervals in its
manifest, with one-based inclusive source frame numbers, for example
`"feature-absence": {"eye-r": [[10, 14]]}`. The loader expands these into
per-frame observation tracks before measurement. Unobserved landmarks may fill
rectangular numeric buffers, but they cannot contribute to an eye's contour,
blink or shared gaze. When one eye is absent, gaze uses the observed eye.
Until a detector supplies feature-level confidence, footage without annotations
uses the full-face detection mask as the fallback; it must not claim to detect
individual occlusions that it cannot see.
## Channel
Every animatable property is a channel, and channels are addressed **by path**:
```clojure
:channels
{[:xform :pos] {:animated? false :value [0.0 0.0]}
[:xform :rot] {:animated? false :value 0.0}
[:xform :scale] {:animated? false :value [1.0 1.0]}
[:xform :skew] {:animated? false :value [0.0 0.0]}
[:geom :pts] {:animated? true :interp :hold :dense {...} :generated {...}}
[:style :color] {:animated? false :value :skin-dark}
[:vis] {:animated? true :interp :hold :keys {0 true, 37 false}}}
```
A path is a **vector**, not a string — CLJS maps take vectors as keys natively,
so Blender's `data_path` idea arrives with no parsing. The set of valid paths for
a node follows from its `:kind`, and that is a spec, not a schema migration.
Three channel shapes, and the uniformity across them is the point:
```clojure
;; FRAMED — one static thing. No animation, no vertex correspondence to worry
;; about. A painted background cel is this.
{:animated? false :value v}
;; KEYED — sparse, authored, in the document. Undoable and syncable.
{:animated? true :interp :hold :keys {0 v, 4 v, 12 v}}
;; DENSE — generated, one value per frame, held in tier 2 as a typed array.
{:animated? true :interp :hold
:dense {:store "sha256:…" :offset 0 :stride 40 :frames 600}
:generated {...}}
```
`:interp` defaults to `:hold`, which `docs/design.md` requires of every cut part.
An authored keyed channel may also carry `:segments {8 :linear}`: the key at 8
tweens toward the next key, while other gaps use the channel default. The
transition belongs to the gap starting at a key, so a shape can cut into one
drawing and tween out of it. Per-key easing beyond hold and linear is deferred.
### Keys are a map by frame, not a list
Already argued in `docs/architecture.md` for merge reasons; here it also gives
"the most recent key at or before `f`" as a `rsubseq` on a sorted map instead of
a scan. **Store a plain map** in the document — transit and JSON both lose
sortedness — and build the sorted index in the resolver.
### The flag lives on the channel, not the node
```clojure
:generated {:by :roto/lips-outer
:analysis "sha256:…" ; which analysis artifact
:params {:verts 8 :contour-avg 1 :aperture-cut 0.004}}
```
Present means the UI offers a parameter panel and a re-freeze button. Absent
means the UI offers the keys directly. **The renderer never reads it.**
It belongs on the channel rather than the node because a node routinely wants
both at once: a mouth whose `[:geom :pts]` is rotoscoped and whose `[:xform :pos]`
is hand-animated to sit on a plate. Putting the flag on the node would forbid the
most useful thing in the model.
### Channels are layered
A channel is a base plus optional override layers, and a layer declares how it
combines:
```clojure
{:animated? true :interp :hold
:dense {...} :generated {...}
:over [{:id :nudge :support [88 98] :op :offset
:values {:animated? true :interp :linear :keys {88 [2 0], 96 [0 0]}}}
{:id :redraw :support [104 105] :op :replace
:values {:animated? false :value [[3 7] [4 7] …]}}]}
```
A LAYER'S VALUES ARE A CHANNEL, which is what keeps a constant adjustment, a
ramp and a return motion from being three mechanisms: a framed one says the same
thing on every frame it covers, a keyed one moves. They read through `value-at`
and `cursor` like any channel, one reading head each, so the specification and
the playback path share their blending and differ only in how they read — and a
layer's values may not carry layers of their own, which the stack already
orders.
`:support` is half-open and explicit, `[in out)`. Outside it a layer is inactive
and the base evaluates exactly as it did before, which is the difference between
a bounded correction and inserting boundary keys — the latter alters the
neighbouring segments. And a layer has NO TIME SPACE of its own: its support and
its values' keys are in the frames the base channel's keys are in, the node's
own. A correction on a lane is therefore in lane frames and reaches across the
drawings exposed under it; one on a single occurrence is in that occurrence's
frames and travels with it when the exposure moves. Ownership had already
answered the question, so there is no field to disagree with.
- **`:offset`** adds a delta to the base. "Nudge the mouth two pixels right for
ten frames" survives a re-freeze at different parameters, because it was never
a position — it was a correction.
- **`:replace`** wins outright. For the frame where detection simply failed.
This is what `docs/design.md` means by an override layer, and it is why
re-freezing is safe: the base is regenerated, the layers are untouched. It is
Blender's NLA blending and AE's effect stack at one property.
WHEN THE BASE OUTGROWS A CORRECTION it is a CONFLICT, which is neither a dropped
layer nor an applied one. Turning `:verts` gives the mouth a different number of
points, and an `:offset` is a row of components that has to match: so the
regeneration records `:conflict` on the layer, the layer stays in the document,
the picture is the base meanwhile, and `clip/conflicts` is the list a view
offers to resolve. Deliberately not `problems` — the document loads and saves
fine, it just contains a decision nobody has made yet. A later regeneration
that restores the shape clears the mark. Only `:offset` can conflict; `:replace`
states a whole value and has nothing to agree with.
A correction is NOT a hand placement. `regenerate-head` leaves the head's
authored channels alone once somebody has placed it by hand, and it compares the
channels WITHOUT their layers to decide: otherwise the first correction anyone
made would stop the head following re-measurement forever, which is the opposite
of what a layer is for.
Layers are what "set it by hand" means for anything measured, and the measured
channel does not need to know. A hand-set gaze is an `:over` on
`[:xform :pos]` of the iris; a hand-set mouth shape is an `:over` on
`[:geom :pts]`. Turning the gaze-step or gaze-dwell knob regenerates the base and
leaves the correction alone, which is the entire reason a correction is stored as
a layer rather than written into the track.
**A `:replace` layer overrides absence, an `:offset` layer does not.** Sampling a
channel is: read the base, then apply the layers — and the base coming back
`absent` does not short-circuit that. `:replace` is explicitly for the frame
where detection failed, so it has to be able to supply a value where there is
none; `:offset` is a delta, and there is nothing to nudge, so an offset over an
absent base stays absent. Implemented the obvious way — bail out on absence
before reaching the layers — the one case the feature exists for is the one case
it would not cover.
### One signal, two nodes
Gaze is deliberately **one measurement shared by both eyes**: at this size the
per-eye difference is noise, and independent noise reads as wall-eyed
immediately, which is the most expensive artefact on a face. But it is stored as
`[:xform :pos]` on `:iris-r` and on `:iris-l`, which are two channels on two
nodes with two different parents — so the invariant lives in `measure` and
nothing in the document enforces it.
That matters as soon as either one can be overridden by hand, because an `:over`
on one iris alone reproduces exactly the artefact the shared measurement exists
to prevent. Until drivers exist, **the override is on both or on neither**, and
that is a rule the UI has to keep rather than one the data can.
This is the case that will eventually justify **drivers** — one value, evaluated
once, feeding several channels — which is why gaze is named in Deferred as the
obvious first one. Nothing here forecloses it: a driver needs a place in the
document and a `:driven-by` on a channel, both of which are additive, and an
absent key means "not driven". So it stays deferred, and the shape does not have
to change to allow it.
## Transform: decomposed, never a matrix
```clojure
{:pos [x y] :rot θ :scale [sx sy] :skew [kx ky]}
```
Stored decomposed for two reasons. Each component has to be independently
keyframable, which is the entire point of channels. And interpolating matrix
entries is meaningless — a rotation tweened through its matrix shears on the way.
Composition, per node:
```
local = T(pos) · R(rot) · K(skew) · S(scale)
world = world(parent) · pinv · local
```
`:pinv` is Blender's `parent_inverse`, captured at the moment of parenting so the
child does not jump when it acquires a parent. Small, and its absence is the kind
of thing that makes a parenting feature feel broken.
### There is no `:anchor`, because an anchor is a peg
Rotation and scale happen about the node's **own origin**. There is no
registration point in the decomposition, and that is a deletion rather than a
gap, because
```
T(pos) · T(a) · R·K·S · T(-a) ≡ peg at pos+a carrying R·K·S, child at -a
```
to the last bit of the mantissa — `node-test` asserts it. `T(a)·M·T(-a)` is `M`
conjugated by a translation, which is "do `M` in a frame shifted by `a`", and a
**parent already is a shifted frame**. So an anchor was a peg that could not be
selected, could not be keyed, could not be shared between nodes, and could not be
put above a measured channel. Same expressive content, strictly less reach.
What it did, two mechanisms now do, split along who owns the pivot:
**A pivot nobody chose is derived per drag and never stored.** `gesture/pivot` is
the middle of what the node draws — `pick/bounds-of`, the same call the stage
draws the selection box from, on the same frame — or the node's own origin when it
draws nothing. `gesture/about` then solves for the position that holds that point
still:
```
q = M⁻¹(c − p) the material point under c
p' = c − M'·q = c − M'·M⁻¹(c − p)
```
so a turn about a point that is **not** the node's origin writes `pos` as well as
`rot`. Nothing is cached, so nothing can go stale: the stored anchor was the
centre of what the node drew, captured once at creation, while the box beside it
was recomputed every render — so on anything edited since it was made, the cross
and the box visibly disagreed and the pivot was wrong. A symbol with more than one
node diverged on the first edit.
**And a drawing's origin is the middle of what it draws**, from the moment it is
drawn — `paint/centred`, which splits a stroke into a ring about its own middle
and the `pos` that puts it back. This is what keeps the paragraph above from being
the whole story, because the `pos` that `about` solves for is an **arc** in the
angle and `pos` interpolates along the **chord**:
| | pivot = origin | pivot ≠ origin |
| --- | --- | --- |
| one drag | right | right |
| between two keys | right | **wrong**, by the sagitta of the arc |
A 360° turn is where that is unmissable and was first seen: 0° and 360° are the
only two frames where a wrong pivot cannot be seen at all, so the keys looked
right and every frame between them was wrong — a shape keyed bottom-left to
top-centre with one full turn on the way left the stage completely in the middle
of the spin, orbiting its origin at a radius of 126 px on a 320×200 stage, because
a stroke used to be stored exactly as drawn and its origin was therefore the
**symbol's** origin, the top-left corner of the stage.
With the origin on the content there is nothing to solve: `gesture/at-origin?`
holds, `turn` writes `rot` alone, `scale` writes `scale` alone, and a keyed turn is
right on every frame. `about` is then needed only where the pivot genuinely is not
any node's origin — a multi-selection about its shared box, a measured part, or a
drawing whose points have been edited away from their own middle — and in each of
those a pivot that has to be **keyed** is a peg, below. Hand-authored scenes were
always written this way: `demo/scene.edn`'s card is
`[-44 -30 44 -30 44 30 -44 30]` with its place in `pos`.
**A pivot somebody chose is a peg** — an ordinary `:group` parent, `nest/peg`,
with `:pinv` captured so nothing moves when it appears. Toon Boom's peg, Fusion's
separate Transform node, Harmony's peg-over-the-drawing. It is the answer to the
three things a derived pivot cannot do:
| want | why a derived pivot cannot | what the peg does |
| --- | --- | --- |
| a pivot that persists — an arm turning about its shoulder | a gesture's pivot is the middle of the drawing and lives for one drag, and the drawing's own origin cannot be moved there without moving its points out from under everything that reads them | the peg's `pos`, static, nowhere near the middle |
| a pivot that travels — a foot roll | an anchor could only be keyed against `pos`, interpolated in the same breath, the two obliged to agree frame for frame | the peg's `pos` is an ordinary channel, so key it |
| a hand transform over a **measured** one | impossible: `local`'s translation is `pos − M·a`, and under a measured `M` writing `a` moves the thing it was meant to leave alone | the peg's channels are its own, so the hand transform composes outside the measurement, which stays regenerable |
That last row is why the rotoscoped parts were worst. `flow/freeze` used to run a
`pivoted` pass writing a default anchor onto everything it had made, and it
**skipped every `node/measured?` node** — correctly, for the reason in the table.
So the traced mouth, lids and brows got no pivot at all and turned about the
origin of head-local space, which is the top-left corner of the *footage*: on a
320×200 stage the mouth pivoted about (−234, −395), off the stage by more than a
stage. The pass is gone; there is no node a derived pivot can be missing from.
A peg is also how `demo/stage` places its seven faces, and that is the case that
makes the pair necessary rather than tidy: `:scale` is **keyed** — the faces pulse
— and the source's middle has to stay on its authored centre throughout. A static
`pos` cannot do it alone, since `T(pos)·S(k(f))` moves that point whenever `k`
changes. `T(center)·S(k(f))·T(-origin)` does, for every `k`, with nothing keyed
that was not keyed before.
`gesture/refusal` still turns a hand edit on a measured channel away, because the
next regenerate would discard it — but it can now name a way through, and the way
is a peg.
**The similarity fit already produces a decomposition.** `fitSimilarity` returns
`{s θ tx ty}`, which drops straight into `[:xform :scale]`, `[:xform :rot]` and
`[:xform :pos]` with no conversion. The analysis output and the animation model
meet without an adapter, which is a sign the decomposition is the right one.
## What space geometry is in
**`[:geom :pts]` is always in the node's own local space, and the transform
chain says what that means.** There is no global geometry space and no decision
to make about one.
| Node | Its local space | Why that one |
| --- | --- | --- |
| a rotoscoped feature | head-local, isotropic, unit = one image height | what the anchor fit already produces; the `xform` to raster is not applied and not stored |
| a painted cel | the stage, in pixels, grid-snapped | the artist is placing pixels, so the pixel grid is the thing being authored |
| a primitive under a feature | its parent's | the iris is positioned on the lid ring, not on the stage |
This looks like a small clarification and it removes a whole class of argument.
The prototype bakes the framing into the numbers: `toRasterRing` applies
`makeXform`, which centres on the face oval's bounding box and zooms until the
face is 80% of the raster height, so **every stored vertex carries a cropping
decision** that was made once, at analysis time, from one frame's landmarks.
Dropping that step is a deletion, not a feature, and after it the framing is
simply a transform on a node.
Grid snapping belongs to the cel and not to the roto, for the same reason: a cel
is authored on the grid and a traced contour is not. So it is a property of a
node's space rather than a rule about all geometry, and the tension between
"integer polygons" and "arbitrary placement" was never real.
Each dense block therefore carries its own **fixed-point scale** in its header,
because a block in image-height units and a block in stage pixels need different
ones to fill an `Int16` usefully.
### There is no camera node
A camera is a global transform over everything, and nothing here wants one.
Placing the face on the stage is a transform on a node, which already exists;
what is not on the stage hangs off the edges and the canvas clips it. Every fill
in `domain/raster` already clamps rather than assuming it is inside, so drawing
past the edge is not a feature to add.
Project dimensions are therefore **independent of the footage**. A 1440x1920
portrait clip composited onto a 320x200 stage is not a problem to solve — the
head is placed and scaled where it belongs and the rest of the frame is simply
not on stage. The full frame stays *available* for tracing without being
*visible*, and those are different requirements.
## Head motion: free or anchored to measured frames
`stabilize` produces `{s, θ, tx, ty}` per source frame. Its inverse is stored
densely on `:head`'s position, rotation and scale channels. The same measured
track serves every placement choice:
```clojure
;; no :anchors — free: read the measured transform at the current frame
;; one key — lock to a chosen measured frame throughout
:anchors {0 12}
;; several keys — cut to another measured head transform at frame 40
:anchors {0 12, 40 42}
```
The map is `local change frame -> measured source frame`. A single lock is a
one-key map. Position, rotation and scale read the same held source frame. The
frame set belongs to head placement, independently of plate drawings and stage
pose cuts. No measured block is copied into authored transform keys.
**Always measure, always store factored.** The fit is computed and the geometry
is stored head-local in every mode. Only the frame address used to read the
head's measured transform changes. Two things downstream require that split:
- *Smoothing.* "Smooth the transform, never the contour" only means anything
while the two are separate.
- *Key selection.* A velocity minimum is "articulation paused" in head-local
space and "the head happened to be still" in image space.
This is a document edit, not a reason to re-analyse. A registered tracing photo
will use its own source frame's stabilising transform followed by the same
selected head placement, so it aligns with the vectors drawn over it.
### Two nodes, because two different things want that transform
```
:face group — AUTHORED. where the face sits on the stage, and how big.
:head group — MEASURED. dense head motion read at the selected frame.
:mouth :mouth-in :teeth :lid-r :lid-l :brow-r :brow-l …
```
Changing anchor keys edits `:head` and never touches `:place`, so it cannot move
something that was placed by hand. A group node is free, and keeping the authored
and the measured transform apart is the whole reason the transform is decomposed
in the first place.
## Time maps — exposure, lead and symbol timing are one thing
Every node may map the frame it is evaluated at:
```clojure
:time {:mode :inherit} ; the default, and almost always right
:time {:mode :map :expose 2 :offset -1 :rate 1.0 :loop? false}
:time {:mode :map :source-fps 30 :sample-fps 12} ; root: lower picture cadence
```
Three features that look unrelated are this one mechanism:
- **exposure** is `⌊f/n⌋·n`,
- **picture fps** quantises source time to a chosen picture grid, then reads the
latest source pose at or before that time; source analysis and audio keep their
original cadence,
- **mouth lead** is `f + k`,
- **a symbol instance's timing** is `(f - at)·rate + in`, with optional looping.
Composed along the nesting chain, outermost first. Two rules follow, and they are
different rules:
- **Exposure inherits strictly.** `docs/design.md` is emphatic that everything
rides one grid, because a head cutting on odd frames against a mouth cutting on
even ones reads as two performances. The model permits a per-node grid; the
default must be `:inherit`, and setting it lower is a deliberate act the UI
should make feel like one.
- **Offset is per-node by design.** Mouth lead applies to performance nodes and
*not* to the plate, which is the whole point of it — so the offset genuinely
belongs at the node, not the clip.
## Symbols, and why a scene is one
A **symbol** is an ordered bag of nodes in its own frame space. (Earlier drafts
and code called this a *timeline*; that word now means only the UI pane that
shows one.)
```clojure
{:frames 91
:palette {...} ; see Palettes
:nodes {id -> node}}
```
That is the whole type, and **everything that holds nodes is one of these**:
- what a document opens on is a symbol, and **no symbol is reserved** — a new
document's is called `main` only because it has to be called something,
- anything placed inside another symbol is a symbol,
- a node with `:kind :instance` is an **instance** of one.
An earlier draft of this document had a scene and a `:kind :timeline` symbol as
two structures with the same fields and never said they were the same thing.
They are. Flash's `_root` is a MovieClip; After Effects' "a pre-comp is just a
layer" is already in the prior-art table above. Collapsing them is what makes
nesting arbitrary and free, rather than a feature to be added.
### Two axes of nesting, and they are different
This is the distinction the flat-storage rule is about, and conflating the two is
why "nested" and "flat with parent pointers" sound contradictory when they are
not:
| Axis | What nests | How it is stored |
| --- | --- | --- |
| **parent / child** | transform composition within one timeline | **flat, with parent pointers** — never nested maps |
| **instance** | a timeline inside another timeline | by reference into the library |
Each timeline is flat. Timelines nest. Every argument for flat storage —
addressability, one-field reparenting, structural sharing, per-node sync leaves —
is about the first axis and is untouched by the second.
The instance boundary is also **the only place the frame space changes.** Within
a timeline, `:time` is exposure and lead: a shift inside one space. At an
instance it is `(f - at)·rate + in`, into a different one. That is why `:rate` is
meaningless on an ordinary node and why sampling one must fail loudly rather than
be ignored.
### What is scoped to a timeline
Three fields on a node only have meaning relative to a timeline, and the answer
for all three is the same — **their own**:
- **`:z`** orders among siblings; a node cannot interleave with nodes inside a
nested instance. The instance occupies one position in its parent's order and
its contents sort beneath it, which the z path gives for free by being a
vector.
- **`:stencil`** names a node in the same timeline. A colour key does not
naturally respect a boundary — it is just pixels — so this is a rule rather
than a consequence, and it is Flash's rule for masks.
- **`:span`** is in the parent node's frame space.
### Instances
A node with `:kind :instance` and `:source {:symbol :sym/blink}` places one, and
its `:playback` says how time runs inside it — which drawing is used and how it
is played are separate facts, per [the lane model](lane-model.md). Its own channels
compose *over* the symbol's, so one definition is placed many times and tinted,
offset or retimed at each placement — that is how a three-frame blink is reused
at frames 40, 88 and 200 without copying it.
This is also where `docs/design.md`'s "closed vocabulary is right for the head"
lands: a plate library is a set of `:sym/head-*` timelines, and the strip chooses
which is instanced on which frame.
**Cursors and point buffers are per-instance, not per-node.** Two instances of
one symbol sit at different frames in their own space, so they cannot share a
reading head over the same channel. The resolver keys its caches by the instance
path, not by node id — which is a detail of `Making it fast` below, and the one
place symbol nesting is not free.
### Audio placements and controls
Sound is placed on a timeline as a separate `:audio` node. It uses the same
`:span`, `:time`, and channel representation as a drawn node. A `:linked-to` id
records which picture instance it was placed with; it does not force the two
spans or source in-points to match.
```clojure
{:id :voice-right :kind :audio :parent :root :z "a4"
:linked-to :right
:source {:footage "f8cace9e-..."}
:span [48 260]
:time {:mode :map :at 48 :in 0 :rate 1}
:channels {[:audio :gain]
{:animated? true :interp :linear
:keys {48 0.0, 60 1.0, 245 1.0, 259 0.0} :over []}}}
```
`[:audio :gain]`, `[:audio :pan]`, and `[:audio :rate]` are ordinary scalar
channels. They may be framed, keyed, or dense; numeric keyed channels can ramp
linearly. The time map sets the placement's base source rate, and
`[:audio :rate]` multiplies it. Audio is mixed from the referenced immutable
footage when the clip opens. The mix is derived output; the saved document holds
the nodes and channel keys, not another audio file. One audio element plays that
mix and remains the clock for both sound and picture.
This is also the boundary for a future control surface. A control has a stable
target, such as a feature's `:verts` setting or an audio node's
`[:audio :gain]` channel. The UI and a MIDI binding can address both through the
same control interface. Their update costs differ: gain can be keyed over time;
changing the number of lip vertices changes topology and must regenerate its
dense geometry. A topology setting cannot be treated as a per-frame gain curve.
## Evaluating a frame
```clojure
(defn eval-frame
"Scene at clip frame f -> draw ops in z order. Pure."
[scene f] ...)
```
1. Walk nodes in **topological order** by parent depth (cached; recompute only
when parentage changes).
2. Skip nodes outside `:span`.
3. Apply the node's time map to get its own local frame `fn`.
4. **Sample** each channel at `fn`: a map lookup for framed, a sorted-index
lookup for keyed, an array read for dense. Then apply `:over` layers.
5. Compose `world` from the parent's.
6. Transform geometry into raster space, writing into a **preallocated buffer**
owned by the node.
7. Emit `{:kind :poly :pts buf :n 20 :color idx :stencil id}`.
8. Sort by resolved `z`.
The op list is the boundary with stage 7 in `docs/architecture.md`: the
rasteriser takes ops and knows nothing about nodes, channels or time.
**A tracing layer is an op that never reaches the raster.** Footage or a still
to draw over is a symbol with `:type :trace` and a `:media`, placed by an ordinary
instance — so it is moved, scaled, trimmed, held and put in a lane like anything
else — and it resolves to one `:trace` op: `{:kind :trace :node :layer :media
:frame :size :m}`. The raster refuses that kind, the player hands it to a
`drawImage` on a separate canvas over the picture, and `clip/resolver` makes one
only when asked with `:tracing?`, which only the stage does. An export, a
symbol's centre and a thumbnail never ask, so a reference cannot reach the
picture by any path that forgets to filter it. See `docs/tracing-symbol-plan.md`.
A face's footage is one of these, placed as `:plate` under `:head` with the
anchor fit itself as its measured transform — the inverse of the head's, over
image height. Its world is `head · fit · 1/H`, so on a frame where the head and
the plate read the same measured frame the two cancel and the photo sits where
the face was filmed; on any other frame it rides the head. Registration is the
ordinary walk, not a matrix built beside it.
Which frame it shows is the placement's: `:time {:holds [...]}` holds it on
chosen frames, and a head with `:reads {:holds-of :plate}` jumps to the same
ones. Whether it is showing at all is the editor's, `[:ui :tracing]`.
### Making it fast in CLJS
Three things, and only these three matter:
- **Decomposed and persistent for storage; flat and mutable for evaluation.**
Composed transforms are 6-element `Float64Array`s, not maps. Every renderer
does this; the storage form and the evaluation form are allowed to differ.
- **A cursor per channel.** Playback is sequential, so "most recent key at or
before `f`" is an advance of a saved index, O(1) amortised. Binary search only
on a seek. This is the difference between a `rsubseq` allocation per channel per
frame and none.
- **Preallocated point buffers per node.** Fixed topology means the size is known
at freeze time, so the vertices — the overwhelming majority of the per-frame
bytes — are written into a buffer the node already owns. A frame still
allocates its op maps and the sorted op vector; that is a dozen small objects
against hundreds of points, and pooling them would buy nothing and cost the
ability to pass an op list around as plain data. At 30fps, per-vertex
allocation is the thing that will make this stutter.
Because the buffers are reused, **ops must be consumed before the next frame is
asked for.** That is the contract the rAF loop wants anyway: it reads, blits,
and dispatches nothing.
### What is in app-db, and what is not
| In app-db (tier 1) | In tier 2, behind a handle |
| --- | --- |
| nodes, parentage, z, spans, stencils | dense channel blocks |
| channel definitions, `:interp`, `:generated` | analysis artifacts |
| **framed** values, **keyed** keys, `:over` layers | preallocated eval buffers |
| library / symbol definitions | composed transform scratch |
The rule: **anything a human placed is in the document; anything a generator
produced is a handle.** Which is the same line `docs/architecture.md` draws for
sync and baking, arrived at again from the renderer's side.
## The current parts, in this model
Proof that it covers what exists, not just what is wanted:
| Now | Becomes |
| --- | --- |
| `mouth` outer ring, every frame | node `:mouth`, `[:geom :pts]` dense, `:generated {:by :roto/lips-outer}` |
| `mouth_in`, hidden below aperture | node `:mouth-in`, parent `:mouth`, `[:geom :pts]` dense + `[:vis]` dense |
| `teeth` from image content | node `:teeth`, stencil `:mouth-in`, `[:geom :pts]` dense, `:generated {:by :interior/teeth}` |
| lid rings | nodes `:lid-r/-l`, `[:geom :pts]` dense |
| lash line (`offsetRing`) | not data — a stage-6 parameter on the node, `{:grow px}` |
| iris disc | node `:iris-r`, `:kind :disc`, parent `:lid-r`, stencil `:sclera-r`, `[:xform :pos]` dense (quantised at freeze), radius framed |
| square pupil | node `:pupil-r`, `:kind :rect`, parent `:iris-r`, stencil `:iris-r` |
| brow ring + quantised raise | node `:brow-r`, `[:geom :pts]` dense (the traced ring with height removed), `[:xform :pos]` dense (the quantised raise). **The decomposition design.md insists on is two channels.** |
| head plate, kept frames | node `:head`, `:symbol` per instance, keys on `[:symbol]` at kept frames |
| `makeXform` face-oval crop | **gone.** Placement is `[:xform :*]` on the face's own `:place`; the stage clips |
| `stabilize` transforms | dense `[:xform :*]` on `:head` (the inverse fit) and on its `:plate` (the fit), read where `:reads` and `:time :holds` say |
| registered underlay | the face's `:plate`, an instance of the footage's tracing symbol under `:head`; a `:trace` op the raster never sees |
| painted background cel | node per layer, `[:geom :pts]` **framed**, `[:style :color]` framed |
| `mouth lead` | `:time {:offset k}` on performance nodes only |
| `exposure` | `:time {:expose n}` on the clip root, inherited |
| picture fps | resolver samples marked generated channels at the picture rate; authored keys keep their own time |
| hand correction | an `:over` layer, `:offset` or `:replace` |
The brow row is the one worth looking at twice. `docs/design.md` argues at length
that the traced ring already contains the height, so the quantised raise must be
measured *out* and put *back* or the brow moves twice. In this model that is not
an argument to remember — it is two channels on one node, and getting it wrong
would mean writing the height into both.
## Palettes
Three levels, and keeping them apart is what makes a palette swap a
**reinterpretation** rather than an edit:
| Level | Holds | Lives on |
| --- | --- | --- |
| **tone** | which mark this is — `:skin-dark` | `[:style :color]`, a channel on the node |
| **ramp** | what that tone looks like *here* | `:palette`, a channel on the timeline |
| **the ramps** | every named palette | the project |
A node names a **tone**, never a colour and never a ramp. Which ramp the tone is
read in is decided by the timeline the node is in. So the same drawing reads day
or night without one stored value changing — which is the entire payoff of
indexed colour, and is why `docs/design.md` forbids sampled RGB: once a shape
holds a measured colour there is nothing left to reinterpret.
Named palettes are **variants over one tone vocabulary**, not arbitrary colour
lists. `:day` and `:night` both define `:skin-dark`; that is what keeps a swap
total and keeps `docs/design.md`'s closed vocabulary closed. A tone the ramp in
scope does not define resolves to the loud magenta, like any other missing index.
### The scope rule
`:palette-track` points to an ordinary lane symbol. Its clips are instances of
restricted palette symbols: a palette symbol owns no nodes and points at exactly
one project palette.
```clojure
{:id :shot :frames 91 :palette :day :palette-track :shot-palettes :nodes {...}}
{:id :shot-palettes :type :palette-track :display :lane :frames 91
:nodes {:day-clip {:kind :instance :source {:symbol :day-palette} ...}
:dusk-clip {:kind :instance :source {:symbol :dusk-palette} ...}}}
{:id :day-palette :type :palette :palette-ref :day :frames 1 :nodes {}}
```
`:palette` is the symbol's authoring/preview palette. It seeds evaluation only
when that symbol is the viewed root; nested symbols do not replace the root's
choice merely because they were authored under another ramp. When absent, the
project default seeds evaluation.
Covered clips of the viewed root's palette track override that seed. An
uncovered lane interval is a genuine gap, restoring the authoring palette or
project default. Palette clips use the same trim, roll, slide, claim-time and
undo commands as visual clips; palette code does not duplicate those edits.
Thus palette-track coverage, authoring preview, and project fallback are
separate facts rather than three accidental meanings of one field. There is no
second keyed palette control on symbols or instances: time-varying palette
changes are authored only as clips in the palette lane.
### One index space, partitioned by palette
A raster is one `Uint8Array` and an index means one colour in it, so two ramps in
one frame cannot both own index 2. The resolution: **the output index space is
the concatenation of the named palettes**, and a tone resolves to
`palette-base + tone-index`.
Everything downstream is then unchanged — one buffer, one flat table for
`->rgba`, no per-frame palette construction, and an index does not change meaning
between frames, so bakes and thumbnails stay valid.
Two consequences worth stating rather than discovering:
- **The limit is real and reachable.** 256 indices over a nine-tone vocabulary is
twenty-eight palettes. Detect it and say so; do not let it arrive as wrapped
colour.
- **It makes the stencil sharper.** A stencil is a colour key, so two nodes
sharing a tone share a stencil — a genuine weakness of the technique.
Partitioning the index space by palette means two nodes in *different* palettes
no longer collide at all, and the resolved stencil picks up whichever index the
stencil node actually drew in.
### Where it is resolved
At the op boundary, and nowhere else. `[:style :color]` holds a keyword all the
way through evaluation; the walk carries the palette in scope the same way it
carries the parent transform and the local frame; the op carries a resolved
index. The rasteriser never sees a tone name and the node never sees an index.
This also means the palette is a **parameter of evaluation**, not a global. The
resolver takes it alongside the store.
## Format on disk and on the wire
Tier 1 is EDN/transit: the node tree, channel definitions, framed values, keys,
layers, library. Kilobytes, human-readable, diffable, and leaf-addressable for
sync.
Dense blocks are separate content-addressed binaries — `Int16Array` for
geometry, `Float32Array` for transforms — with a small header naming the channel
path, frame count, stride, and the **fixed-point scale** of the node-local space
the block is in. Geometry is stored in the node's own space, not in raster space;
see "What space geometry is in".
**Not Lottie internally**, despite the property shape being borrowed from it.
Lottie has no palette-indexed colour, its shapes are bezier with in/out tangents
where these are integer polygons, and its interpolation defaults are the opposite
of what is wanted. It is a fine thing to write out one day and a bad thing to
store.
Output is deliberately not specified here. The target is encoding video in the
browser, which touches the op list and nothing above it — a writer consumes
frames, and frames are what stage 7 already produces.
## Deferred
- **Per-key easing.** The structure allows it; nothing should use it until a
parented transform on a painted cel asks for it.
- **More than two channel layers.** The `:over` vector is already a list; a real
blend stack with weights is the NLA, and it is not needed to fix a bad frame.
- **Skew beyond the field.** `[:xform :skew]` is in the transform and in the
composition order from the start, because adding a component to a decomposition
later means migrating every stored transform.
- **Instance channel overrides on symbols.** Compose-over is specified; only
colour and transform need it at first.
- **Constraints and drivers.** Blender's other half. A gaze that aims at a null
object is the obvious first one, and it is a long way off. Until then the one
gaze shared by two iris nodes is a UI rule, not a stored relationship — see
"One signal, two nodes".