The source-to-stage mapping moves off :main's :face group and onto each face's own :place, above its head. `face-placement` computes exactly what it computed before, over every subject together, so two faces filmed side by side keep their filmed relation — it is written into each face instead of onto a group above them all. Same transform, same subtree, one level lower, and the composite is identical to the pixel: a digest over every op :main emits across the whole take is unchanged either way. THE OWNER IS THE POINT. A face carrying its own mapping is the right size wherever it is put — dropped into another symbol, or opened in its own tab to be drawn over — and the take that holds it needs to know nothing. On a group above the instances the scale belonged to the take, so a face taken out of it had no size at all and drew at a fraction of a pixel. The pool's thumbnails drop the workaround that knew about this: a symbol is rendered rooted at itself again, because a face now carries the placement that makes that honest, so the pool needs to know nothing about where a symbol happens to be used. `domain/node` and `arthur.export` leave its requires with it. The tests here were reading the placement off :main. The photo registration test changes shape rather than location: its premise was that face-1's head is its own root, so a photo sitting where it was filmed was image pixels over image height and nothing else. The head still cancels — that is what the test is about — but it now cancels against the face's own placement, which is why the photo comes with the face into its own tab instead of sitting at a fraction of a pixel beside it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
812 lines
41 KiB
Markdown
812 lines
41 KiB
Markdown
# arthur — the animation model
|
||
|
||
The revised target for lanes, occurrences, source playback, shared editing, and
|
||
multi-view UX is [The Lane Model](lane-model.md). It supersedes conflicting
|
||
proposals below. Backward compatibility is not required; this document still
|
||
contains descriptions of earlier shapes and planned features.
|
||
|
||
The data that describes a moving picture: what the primitives are, how they
|
||
nest, how they change over time, and how rotoscoped and hand-authored work end
|
||
up being the same thing with one flag between them.
|
||
|
||
`docs/design.md` is the aesthetic argument. `docs/architecture.md` is where the
|
||
code goes. This is the type that both of them are about.
|
||
|
||
## What this replaces
|
||
|
||
`docs/design.md` has a table of five kinds of part — plate, feature, interior,
|
||
primitive, scalar — each with its own source, vocabulary and interpolation. That
|
||
table is a good description of **where data comes from** and a bad description of
|
||
**what data is**, and the current code follows it too literally: eyes, brows,
|
||
teeth and mouth each get their own build function, their own key shape and their
|
||
own path through the prototype's writer.
|
||
|
||
They are all one thing. A part is a **node** with **channels**, and the five
|
||
kinds collapse into differences of which channels exist and who filled them in.
|
||
|
||
## Prior art, and what each one gets right
|
||
|
||
| System | The idea worth taking |
|
||
| --- | --- |
|
||
| **Flash / SWF** | A **library of symbols** and a timeline of **instances** at depths. "Framed" content that simply exists on a frame, versus tweened content. `DefineMorphShape` requires matching vertex counts — the fixed-topology rule, arrived at from the other direction. |
|
||
| **Blender** | Animation is **addressed by path into the data** (`location[0]`), not stored as fields on the object. An Action is a bag of F-Curves. That decoupling is what makes the dope sheet, the graph editor and the NLA three views of one dataset. Also: parenting captures a `parent_inverse` so the child does not jump. |
|
||
| **After Effects** | Every leaf property is animatable, uniformly. Property groups form a tree. Pre-comps nest arbitrarily and a pre-comp is just a layer. |
|
||
| **Lottie** | The uniform property shape: `{a: 0, k: <value>}` or `{a: 1, k: [<keys>]}`. One representation for static and animated, which is exactly "framed or keyframed". |
|
||
| **Grease Pencil** | A 2D layer holds frames at frame numbers, and a frame **holds until the next one**. Hold is the default, not a special case. |
|
||
|
||
What none of them get right for this project: colour. All four store RGB on the
|
||
shape. `docs/design.md` forbids that, so colour is a palette index here and it is
|
||
a channel like any other.
|
||
|
||
## The one idea
|
||
|
||
**Analysis is a channel generator.** It does not produce a different kind of
|
||
data; it produces keys, densely, on the same channels a hand would fill in
|
||
sparsely. So:
|
||
|
||
```
|
||
footage ──▶ analysis ──▶ FREEZE ──▶ channels on nodes ──▶ evaluate ──▶ raster
|
||
▲
|
||
hand authoring ──┘
|
||
```
|
||
|
||
Freezing is not a conversion into a second format. There is one format, and
|
||
freezing fills it in. That is what makes "the only difference is a special flag"
|
||
literally true: the flag is provenance on a channel, and nothing in the renderer
|
||
reads it.
|
||
|
||
## Node
|
||
|
||
A node is an instance in the scene. The tree is stored **flat, with parent
|
||
pointers** — never as nested maps.
|
||
|
||
```clojure
|
||
{:id :mouth
|
||
:name "mouth"
|
||
:kind :poly ; :poly :disc :rect :group :bitmap :symbol
|
||
:parent :head ; nil at the root
|
||
:z "a3" ; fractional index, ordered among all siblings
|
||
:symbol nil ; or :sym/blink — see Symbols
|
||
:stencil :mouth-in ; colour-key clip; structural, not a channel
|
||
:span [0 240] ; in/out in the parent's frame space
|
||
:pinv [1 0 0 1 0 0] ; parent-inverse, captured when parented
|
||
:channels {...}}
|
||
```
|
||
|
||
Flat with pointers, for four reasons that all point the same way: any node is
|
||
addressable without a walk; reparenting is a one-field write rather than a
|
||
subtree move; an edit to a leaf does not change the identity of its ancestors, so
|
||
re-frame's structural sharing keeps ancestor subs from invalidating; and it is
|
||
what lets every node be its own sync leaf. Flash, Blender and AE all store it
|
||
this way.
|
||
|
||
`:span` is Lottie's `ip`/`op` and Flash's `PlaceObject`/`RemoveObject`: the range
|
||
over which the node exists at all. Distinct from a `[:vis]` channel, which
|
||
blinks an existing node on and off.
|
||
|
||
**Every node has the same two maps into its parent**, whatever kind it is:
|
||
|
||
- **space** — the matrix its transform channels compose to, times a `:pinv` if
|
||
it has been moved in from elsewhere;
|
||
- **time** — `local = rate · (parent − at)`, from `:time :at` and `:rate`,
|
||
identity when absent. `:span` and every key are in the node's **own** frames.
|
||
|
||
A move keeps a node's world maps and re-expresses them under its new parent:
|
||
the matrix becomes a `:pinv`, the time becomes a new `:at` and `:rate`, and its
|
||
channels, keys and span are not touched. Both maps are affine, so any depth of
|
||
nesting is one map and every move is one inverse. `node/time-of`,
|
||
`node/then-time` and `node/placed-span` are the time half; `clip/move-node` and
|
||
`clip/group` are the move.
|
||
|
||
### Subjects and tracked features
|
||
|
||
Scene nodes describe drawings, not tracking identity. A scene may also carry a
|
||
flat `:features` map. A feature ID stays stable for the whole clip, including
|
||
frames where that feature is occluded and later reappears:
|
||
|
||
```clojure
|
||
:subjects {:face-1 {:id :face-1}}
|
||
:features
|
||
{:eye-r {:id :eye-r :subject :face-1 :area :eye
|
||
:nodes [:eye-r :eye-r-in :iris-r :pupil-r] :params {}}
|
||
:eye-l {:id :eye-l :subject :face-1 :area :eye
|
||
:nodes [:eye-l :eye-l-in :iris-l :pupil-l] :params {}}
|
||
:mouth {:id :mouth :subject :face-1 :area :mouth
|
||
:nodes [:mouth :mouth-in] :params {}}
|
||
;; The teeth are their OWN feature and not three nodes of the mouth. A feature
|
||
;; carries the params of exactly one area, and the teeth have an `:area :teeth`
|
||
;; of their own — the otsu threshold, the tongue rejection, the radial contour's
|
||
;; vertex budget — which could not be reached if they were part of `:mouth`.
|
||
;; The coupling that made them look like the mouth's is real and is enforced
|
||
;; elsewhere: `:teeth` is STENCILLED by `:mouth-in`, and a node whose stencil drew
|
||
;; nothing is dropped, so an absent mouth takes the teeth with it without either
|
||
;; of them sharing an absence mask. An earlier draft of this block listed them
|
||
;; together; the code is right and this document was wrong.
|
||
:teeth {:id :teeth :subject :face-1 :area :teeth
|
||
:nodes [:teeth] :params {}}}
|
||
:groups
|
||
{:eyes-1 {:id :eyes-1 :kind :eye-pair :subject :face-1
|
||
:members [:eye-r :eye-l] :params {}}}
|
||
```
|
||
|
||
An eye pair is an explicit relationship between one or two eyes of the **same
|
||
subject**. It may have one member when only one eye has been identified; it does
|
||
not invent a second eye. Five subjects with nine identified eyes can have four
|
||
two-member pairs and one one-member pair. Each eye still has its own feature ID
|
||
and presence track. A group is a settings association, not a scene parent or a
|
||
tracking ID. Membership lives only on the group, avoiding a second pointer on
|
||
the feature that could disagree with it.
|
||
|
||
Each feature resolves settings from its area's definitions, then its group,
|
||
then its own `:params`. An eye can therefore inherit a pair setting or override
|
||
it without changing its partner. Removing it from a pair copies its effective
|
||
values into the feature first, so the result does not jump. Feature identity
|
||
and pair membership are clip-wide; a future parameter track can vary values
|
||
over time without splitting a feature at an observation gap.
|
||
|
||
Parameter definitions live in one registry: key, default, applicable area,
|
||
value constraints and affected areas. The registry supplies the take's defaults
|
||
today. The parameter UI and regeneration from edited values are later work.
|
||
|
||
Dense channel state records whether a measurement exists **for that feature on
|
||
that frame**. Occlusion means absent data on that frame, not a false `[:vis]`
|
||
value and not the end of the feature's identity. A full-face detection failure
|
||
makes all its features absent. A single occluded eye need only make that eye's
|
||
channels absent. Footage can carry explicit feature absence intervals in its
|
||
manifest, with one-based inclusive source frame numbers, for example
|
||
`"feature-absence": {"eye-r": [[10, 14]]}`. The loader expands these into
|
||
per-frame observation tracks before measurement. Unobserved landmarks may fill
|
||
rectangular numeric buffers, but they cannot contribute to an eye's contour,
|
||
blink or shared gaze. When one eye is absent, gaze uses the observed eye.
|
||
Until a detector supplies feature-level confidence, footage without annotations
|
||
uses the full-face detection mask as the fallback; it must not claim to detect
|
||
individual occlusions that it cannot see.
|
||
|
||
## Channel
|
||
|
||
Every animatable property is a channel, and channels are addressed **by path**:
|
||
|
||
```clojure
|
||
:channels
|
||
{[:xform :pos] {:animated? false :value [0.0 0.0]}
|
||
[:xform :rot] {:animated? false :value 0.0}
|
||
[:xform :scale] {:animated? false :value [1.0 1.0]}
|
||
[:xform :skew] {:animated? false :value [0.0 0.0]}
|
||
[:xform :anchor]{:animated? false :value [0.0 0.0]}
|
||
[:geom :pts] {:animated? true :interp :hold :dense {...} :generated {...}}
|
||
[:style :color] {:animated? false :value :skin-dark}
|
||
[:vis] {:animated? true :interp :hold :keys {0 true, 37 false}}}
|
||
```
|
||
|
||
A path is a **vector**, not a string — CLJS maps take vectors as keys natively,
|
||
so Blender's `data_path` idea arrives with no parsing. The set of valid paths for
|
||
a node follows from its `:kind`, and that is a spec, not a schema migration.
|
||
|
||
Three channel shapes, and the uniformity across them is the point:
|
||
|
||
```clojure
|
||
;; FRAMED — one static thing. No animation, no vertex correspondence to worry
|
||
;; about. A painted background cel is this.
|
||
{:animated? false :value v}
|
||
|
||
;; KEYED — sparse, authored, in the document. Undoable and syncable.
|
||
{:animated? true :interp :hold :keys {0 v, 4 v, 12 v}}
|
||
|
||
;; DENSE — generated, one value per frame, held in tier 2 as a typed array.
|
||
{:animated? true :interp :hold
|
||
:dense {:store "sha256:…" :offset 0 :stride 40 :frames 600}
|
||
:generated {...}}
|
||
```
|
||
|
||
`:interp` defaults to `:hold`, which `docs/design.md` requires of every cut part.
|
||
An authored keyed channel may also carry `:segments {8 :linear}`: the key at 8
|
||
tweens toward the next key, while other gaps use the channel default. The
|
||
transition belongs to the gap starting at a key, so a shape can cut into one
|
||
drawing and tween out of it. Per-key easing beyond hold and linear is deferred.
|
||
|
||
### Keys are a map by frame, not a list
|
||
|
||
Already argued in `docs/architecture.md` for merge reasons; here it also gives
|
||
"the most recent key at or before `f`" as a `rsubseq` on a sorted map instead of
|
||
a scan. **Store a plain map** in the document — transit and JSON both lose
|
||
sortedness — and build the sorted index in the resolver.
|
||
|
||
### The flag lives on the channel, not the node
|
||
|
||
```clojure
|
||
:generated {:by :roto/lips-outer
|
||
:analysis "sha256:…" ; which analysis artifact
|
||
:params {:verts 8 :contour-avg 1 :aperture-cut 0.004}}
|
||
```
|
||
|
||
Present means the UI offers a parameter panel and a re-freeze button. Absent
|
||
means the UI offers the keys directly. **The renderer never reads it.**
|
||
|
||
It belongs on the channel rather than the node because a node routinely wants
|
||
both at once: a mouth whose `[:geom :pts]` is rotoscoped and whose `[:xform :pos]`
|
||
is hand-animated to sit on a plate. Putting the flag on the node would forbid the
|
||
most useful thing in the model.
|
||
|
||
### Channels are layered
|
||
|
||
A channel is a base plus optional override layers, and a layer declares how it
|
||
combines:
|
||
|
||
```clojure
|
||
{:animated? true :interp :hold
|
||
:dense {...} :generated {...}
|
||
:over [{:id :nudge :support [88 98] :op :offset
|
||
:values {:animated? true :interp :linear :keys {88 [2 0], 96 [0 0]}}}
|
||
{:id :redraw :support [104 105] :op :replace
|
||
:values {:animated? false :value [[3 7] [4 7] …]}}]}
|
||
```
|
||
|
||
A LAYER'S VALUES ARE A CHANNEL, which is what keeps a constant adjustment, a
|
||
ramp and a return motion from being three mechanisms: a framed one says the same
|
||
thing on every frame it covers, a keyed one moves. They read through `value-at`
|
||
and `cursor` like any channel, one reading head each, so the specification and
|
||
the playback path share their blending and differ only in how they read — and a
|
||
layer's values may not carry layers of their own, which the stack already
|
||
orders.
|
||
|
||
`:support` is half-open and explicit, `[in out)`. Outside it a layer is inactive
|
||
and the base evaluates exactly as it did before, which is the difference between
|
||
a bounded correction and inserting boundary keys — the latter alters the
|
||
neighbouring segments. And a layer has NO TIME SPACE of its own: its support and
|
||
its values' keys are in the frames the base channel's keys are in, the node's
|
||
own. A correction on a lane is therefore in lane frames and reaches across the
|
||
drawings exposed under it; one on a single occurrence is in that occurrence's
|
||
frames and travels with it when the exposure moves. Ownership had already
|
||
answered the question, so there is no field to disagree with.
|
||
|
||
- **`:offset`** adds a delta to the base. "Nudge the mouth two pixels right for
|
||
ten frames" survives a re-freeze at different parameters, because it was never
|
||
a position — it was a correction.
|
||
- **`:replace`** wins outright. For the frame where detection simply failed.
|
||
|
||
This is what `docs/design.md` means by an override layer, and it is why
|
||
re-freezing is safe: the base is regenerated, the layers are untouched. It is
|
||
Blender's NLA blending and AE's effect stack at one property.
|
||
|
||
WHEN THE BASE OUTGROWS A CORRECTION it is a CONFLICT, which is neither a dropped
|
||
layer nor an applied one. Turning `:verts` gives the mouth a different number of
|
||
points, and an `:offset` is a row of components that has to match: so the
|
||
regeneration records `:conflict` on the layer, the layer stays in the document,
|
||
the picture is the base meanwhile, and `clip/conflicts` is the list a view
|
||
offers to resolve. Deliberately not `problems` — the document loads and saves
|
||
fine, it just contains a decision nobody has made yet. A later regeneration
|
||
that restores the shape clears the mark. Only `:offset` can conflict; `:replace`
|
||
states a whole value and has nothing to agree with.
|
||
|
||
A correction is NOT a hand placement. `regenerate-head` leaves the head's
|
||
authored channels alone once somebody has placed it by hand, and it compares the
|
||
channels WITHOUT their layers to decide: otherwise the first correction anyone
|
||
made would stop the head following re-measurement forever, which is the opposite
|
||
of what a layer is for.
|
||
|
||
Layers are what "set it by hand" means for anything measured, and the measured
|
||
channel does not need to know. A hand-set gaze is an `:over` on
|
||
`[:xform :pos]` of the iris; a hand-set mouth shape is an `:over` on
|
||
`[:geom :pts]`. Turning the gaze-step or gaze-dwell knob regenerates the base and
|
||
leaves the correction alone, which is the entire reason a correction is stored as
|
||
a layer rather than written into the track.
|
||
|
||
**A `:replace` layer overrides absence, an `:offset` layer does not.** Sampling a
|
||
channel is: read the base, then apply the layers — and the base coming back
|
||
`absent` does not short-circuit that. `:replace` is explicitly for the frame
|
||
where detection failed, so it has to be able to supply a value where there is
|
||
none; `:offset` is a delta, and there is nothing to nudge, so an offset over an
|
||
absent base stays absent. Implemented the obvious way — bail out on absence
|
||
before reaching the layers — the one case the feature exists for is the one case
|
||
it would not cover.
|
||
|
||
### One signal, two nodes
|
||
|
||
Gaze is deliberately **one measurement shared by both eyes**: at this size the
|
||
per-eye difference is noise, and independent noise reads as wall-eyed
|
||
immediately, which is the most expensive artefact on a face. But it is stored as
|
||
`[:xform :pos]` on `:iris-r` and on `:iris-l`, which are two channels on two
|
||
nodes with two different parents — so the invariant lives in `measure` and
|
||
nothing in the document enforces it.
|
||
|
||
That matters as soon as either one can be overridden by hand, because an `:over`
|
||
on one iris alone reproduces exactly the artefact the shared measurement exists
|
||
to prevent. Until drivers exist, **the override is on both or on neither**, and
|
||
that is a rule the UI has to keep rather than one the data can.
|
||
|
||
This is the case that will eventually justify **drivers** — one value, evaluated
|
||
once, feeding several channels — which is why gaze is named in Deferred as the
|
||
obvious first one. Nothing here forecloses it: a driver needs a place in the
|
||
document and a `:driven-by` on a channel, both of which are additive, and an
|
||
absent key means "not driven". So it stays deferred, and the shape does not have
|
||
to change to allow it.
|
||
|
||
## Transform: decomposed, never a matrix
|
||
|
||
```clojure
|
||
{:pos [x y] :rot θ :scale [sx sy] :skew [kx ky] :anchor [ax ay]}
|
||
```
|
||
|
||
Stored decomposed for two reasons. Each component has to be independently
|
||
keyframable, which is the entire point of channels. And interpolating matrix
|
||
entries is meaningless — a rotation tweened through its matrix shears on the way.
|
||
|
||
Composition, per node:
|
||
|
||
```
|
||
local = T(pos) · T(anchor) · R(rot) · K(skew) · S(scale) · T(-anchor)
|
||
world = world(parent) · pinv · local
|
||
```
|
||
|
||
`:anchor` is Flash's registration point and Blender's origin: rotation and scale
|
||
happen about it, and getting it wrong is why hand-placed parts swing rather than
|
||
turn.
|
||
|
||
`:pinv` is Blender's `parent_inverse`, captured at the moment of parenting so the
|
||
child does not jump when it acquires a parent. Small, and its absence is the kind
|
||
of thing that makes a parenting feature feel broken.
|
||
|
||
**The similarity fit already produces a decomposition.** `fitSimilarity` returns
|
||
`{s θ tx ty}`, which drops straight into `[:xform :scale]`, `[:xform :rot]` and
|
||
`[:xform :pos]` with no conversion. The analysis output and the animation model
|
||
meet without an adapter, which is a sign the decomposition is the right one.
|
||
|
||
## What space geometry is in
|
||
|
||
**`[:geom :pts]` is always in the node's own local space, and the transform
|
||
chain says what that means.** There is no global geometry space and no decision
|
||
to make about one.
|
||
|
||
| Node | Its local space | Why that one |
|
||
| --- | --- | --- |
|
||
| a rotoscoped feature | head-local, isotropic, unit = one image height | what the anchor fit already produces; the `xform` to raster is not applied and not stored |
|
||
| a painted cel | the stage, in pixels, grid-snapped | the artist is placing pixels, so the pixel grid is the thing being authored |
|
||
| a primitive under a feature | its parent's | the iris is positioned on the lid ring, not on the stage |
|
||
|
||
This looks like a small clarification and it removes a whole class of argument.
|
||
The prototype bakes the framing into the numbers: `toRasterRing` applies
|
||
`makeXform`, which centres on the face oval's bounding box and zooms until the
|
||
face is 80% of the raster height, so **every stored vertex carries a cropping
|
||
decision** that was made once, at analysis time, from one frame's landmarks.
|
||
Dropping that step is a deletion, not a feature, and after it the framing is
|
||
simply a transform on a node.
|
||
|
||
Grid snapping belongs to the cel and not to the roto, for the same reason: a cel
|
||
is authored on the grid and a traced contour is not. So it is a property of a
|
||
node's space rather than a rule about all geometry, and the tension between
|
||
"integer polygons" and "arbitrary placement" was never real.
|
||
|
||
Each dense block therefore carries its own **fixed-point scale** in its header,
|
||
because a block in image-height units and a block in stage pixels need different
|
||
ones to fill an `Int16` usefully.
|
||
|
||
### There is no camera node
|
||
|
||
A camera is a global transform over everything, and nothing here wants one.
|
||
Placing the face on the stage is a transform on a node, which already exists;
|
||
what is not on the stage hangs off the edges and the canvas clips it. Every fill
|
||
in `domain/raster` already clamps rather than assuming it is inside, so drawing
|
||
past the edge is not a feature to add.
|
||
|
||
Project dimensions are therefore **independent of the footage**. A 1440x1920
|
||
portrait clip composited onto a 320x200 stage is not a problem to solve — the
|
||
head is placed and scaled where it belongs and the rest of the frame is simply
|
||
not on stage. The full frame stays *available* for tracing without being
|
||
*visible*, and those are different requirements.
|
||
|
||
## Head motion: free or anchored to measured frames
|
||
|
||
`stabilize` produces `{s, θ, tx, ty}` per source frame. Its inverse is stored
|
||
densely on `:head`'s position, rotation and scale channels. The same measured
|
||
track serves every placement choice:
|
||
|
||
```clojure
|
||
;; no :anchors — free: read the measured transform at the current frame
|
||
;; one key — lock to a chosen measured frame throughout
|
||
:anchors {0 12}
|
||
;; several keys — cut to another measured head transform at frame 40
|
||
:anchors {0 12, 40 42}
|
||
```
|
||
|
||
The map is `local change frame -> measured source frame`. A single lock is a
|
||
one-key map. Position, rotation and scale read the same held source frame. The
|
||
frame set belongs to head placement, independently of plate drawings and stage
|
||
pose cuts. No measured block is copied into authored transform keys.
|
||
|
||
**Always measure, always store factored.** The fit is computed and the geometry
|
||
is stored head-local in every mode. Only the frame address used to read the
|
||
head's measured transform changes. Two things downstream require that split:
|
||
|
||
- *Smoothing.* "Smooth the transform, never the contour" only means anything
|
||
while the two are separate.
|
||
- *Key selection.* A velocity minimum is "articulation paused" in head-local
|
||
space and "the head happened to be still" in image space.
|
||
|
||
This is a document edit, not a reason to re-analyse. A registered tracing photo
|
||
will use its own source frame's stabilising transform followed by the same
|
||
selected head placement, so it aligns with the vectors drawn over it.
|
||
|
||
### Two nodes, because two different things want that transform
|
||
|
||
```
|
||
:face group — AUTHORED. where the face sits on the stage, and how big.
|
||
:head group — MEASURED. dense head motion read at the selected frame.
|
||
:mouth :mouth-in :teeth :lid-r :lid-l :brow-r :brow-l …
|
||
```
|
||
|
||
Changing anchor keys edits `:head` and never touches `:place`, so it cannot move
|
||
something that was placed by hand. A group node is free, and keeping the authored
|
||
and the measured transform apart is the whole reason the transform is decomposed
|
||
in the first place.
|
||
|
||
## Time maps — exposure, lead and symbol timing are one thing
|
||
|
||
Every node may map the frame it is evaluated at:
|
||
|
||
```clojure
|
||
:time {:mode :inherit} ; the default, and almost always right
|
||
:time {:mode :map :expose 2 :offset -1 :rate 1.0 :loop? false}
|
||
:time {:mode :map :source-fps 30 :sample-fps 12} ; root: lower picture cadence
|
||
```
|
||
|
||
Three features that look unrelated are this one mechanism:
|
||
|
||
- **exposure** is `⌊f/n⌋·n`,
|
||
- **picture fps** quantises source time to a chosen picture grid, then reads the
|
||
latest source pose at or before that time; source analysis and audio keep their
|
||
original cadence,
|
||
- **mouth lead** is `f + k`,
|
||
- **a symbol instance's timing** is `(f - at)·rate + in`, with optional looping.
|
||
|
||
Composed along the nesting chain, outermost first. Two rules follow, and they are
|
||
different rules:
|
||
|
||
- **Exposure inherits strictly.** `docs/design.md` is emphatic that everything
|
||
rides one grid, because a head cutting on odd frames against a mouth cutting on
|
||
even ones reads as two performances. The model permits a per-node grid; the
|
||
default must be `:inherit`, and setting it lower is a deliberate act the UI
|
||
should make feel like one.
|
||
- **Offset is per-node by design.** Mouth lead applies to performance nodes and
|
||
*not* to the plate, which is the whole point of it — so the offset genuinely
|
||
belongs at the node, not the clip.
|
||
|
||
## Symbols, and why a scene is one
|
||
|
||
A **symbol** is an ordered bag of nodes in its own frame space. (Earlier drafts
|
||
and code called this a *timeline*; that word now means only the UI pane that
|
||
shows one.)
|
||
|
||
```clojure
|
||
{:frames 91
|
||
:palette {...} ; see Palettes
|
||
:nodes {id -> node}}
|
||
```
|
||
|
||
That is the whole type, and **everything that holds nodes is one of these**:
|
||
|
||
- what a document opens on is a symbol, and **no symbol is reserved** — a new
|
||
document's is called `main` only because it has to be called something,
|
||
- anything placed inside another symbol is a symbol,
|
||
- a node with `:kind :instance` is an **instance** of one.
|
||
|
||
An earlier draft of this document had a scene and a `:kind :timeline` symbol as
|
||
two structures with the same fields and never said they were the same thing.
|
||
They are. Flash's `_root` is a MovieClip; After Effects' "a pre-comp is just a
|
||
layer" is already in the prior-art table above. Collapsing them is what makes
|
||
nesting arbitrary and free, rather than a feature to be added.
|
||
|
||
### Two axes of nesting, and they are different
|
||
|
||
This is the distinction the flat-storage rule is about, and conflating the two is
|
||
why "nested" and "flat with parent pointers" sound contradictory when they are
|
||
not:
|
||
|
||
| Axis | What nests | How it is stored |
|
||
| --- | --- | --- |
|
||
| **parent / child** | transform composition within one timeline | **flat, with parent pointers** — never nested maps |
|
||
| **instance** | a timeline inside another timeline | by reference into the library |
|
||
|
||
Each timeline is flat. Timelines nest. Every argument for flat storage —
|
||
addressability, one-field reparenting, structural sharing, per-node sync leaves —
|
||
is about the first axis and is untouched by the second.
|
||
|
||
The instance boundary is also **the only place the frame space changes.** Within
|
||
a timeline, `:time` is exposure and lead: a shift inside one space. At an
|
||
instance it is `(f - at)·rate + in`, into a different one. That is why `:rate` is
|
||
meaningless on an ordinary node and why sampling one must fail loudly rather than
|
||
be ignored.
|
||
|
||
### What is scoped to a timeline
|
||
|
||
Three fields on a node only have meaning relative to a timeline, and the answer
|
||
for all three is the same — **their own**:
|
||
|
||
- **`:z`** orders among siblings; a node cannot interleave with nodes inside a
|
||
nested instance. The instance occupies one position in its parent's order and
|
||
its contents sort beneath it, which the z path gives for free by being a
|
||
vector.
|
||
- **`:stencil`** names a node in the same timeline. A colour key does not
|
||
naturally respect a boundary — it is just pixels — so this is a rule rather
|
||
than a consequence, and it is Flash's rule for masks.
|
||
- **`:span`** is in the parent node's frame space.
|
||
|
||
### Instances
|
||
|
||
A node with `:kind :instance` and `:source {:symbol :sym/blink}` places one, and
|
||
its `:playback` says how time runs inside it — which drawing is used and how it
|
||
is played are separate facts, per [the lane model](lane-model.md). Its own channels
|
||
compose *over* the symbol's, so one definition is placed many times and tinted,
|
||
offset or retimed at each placement — that is how a three-frame blink is reused
|
||
at frames 40, 88 and 200 without copying it.
|
||
|
||
This is also where `docs/design.md`'s "closed vocabulary is right for the head"
|
||
lands: a plate library is a set of `:sym/head-*` timelines, and the strip chooses
|
||
which is instanced on which frame.
|
||
|
||
**Cursors and point buffers are per-instance, not per-node.** Two instances of
|
||
one symbol sit at different frames in their own space, so they cannot share a
|
||
reading head over the same channel. The resolver keys its caches by the instance
|
||
path, not by node id — which is a detail of `Making it fast` below, and the one
|
||
place symbol nesting is not free.
|
||
|
||
### Audio placements and controls
|
||
|
||
Sound is placed on a timeline as a separate `:audio` node. It uses the same
|
||
`:span`, `:time`, and channel representation as a drawn node. A `:linked-to` id
|
||
records which picture instance it was placed with; it does not force the two
|
||
spans or source in-points to match.
|
||
|
||
```clojure
|
||
{:id :voice-right :kind :audio :parent :root :z "a4"
|
||
:linked-to :right
|
||
:source {:footage "f8cace9e-..."}
|
||
:span [48 260]
|
||
:time {:mode :map :at 48 :in 0 :rate 1}
|
||
:channels {[:audio :gain]
|
||
{:animated? true :interp :linear
|
||
:keys {48 0.0, 60 1.0, 245 1.0, 259 0.0} :over []}}}
|
||
```
|
||
|
||
`[:audio :gain]`, `[:audio :pan]`, and `[:audio :rate]` are ordinary scalar
|
||
channels. They may be framed, keyed, or dense; numeric keyed channels can ramp
|
||
linearly. The time map sets the placement's base source rate, and
|
||
`[:audio :rate]` multiplies it. Audio is mixed from the referenced immutable
|
||
footage when the clip opens. The mix is derived output; the saved document holds
|
||
the nodes and channel keys, not another audio file. One audio element plays that
|
||
mix and remains the clock for both sound and picture.
|
||
|
||
This is also the boundary for a future control surface. A control has a stable
|
||
target, such as a feature's `:verts` setting or an audio node's
|
||
`[:audio :gain]` channel. The UI and a MIDI binding can address both through the
|
||
same control interface. Their update costs differ: gain can be keyed over time;
|
||
changing the number of lip vertices changes topology and must regenerate its
|
||
dense geometry. A topology setting cannot be treated as a per-frame gain curve.
|
||
|
||
## Evaluating a frame
|
||
|
||
```clojure
|
||
(defn eval-frame
|
||
"Scene at clip frame f -> draw ops in z order. Pure."
|
||
[scene f] ...)
|
||
```
|
||
|
||
1. Walk nodes in **topological order** by parent depth (cached; recompute only
|
||
when parentage changes).
|
||
2. Skip nodes outside `:span`.
|
||
3. Apply the node's time map to get its own local frame `fn`.
|
||
4. **Sample** each channel at `fn`: a map lookup for framed, a sorted-index
|
||
lookup for keyed, an array read for dense. Then apply `:over` layers.
|
||
5. Compose `world` from the parent's.
|
||
6. Transform geometry into raster space, writing into a **preallocated buffer**
|
||
owned by the node.
|
||
7. Emit `{:kind :poly :pts buf :n 20 :color idx :stencil id}`.
|
||
8. Sort by resolved `z`.
|
||
|
||
The op list is the boundary with stage 7 in `docs/architecture.md`: the
|
||
rasteriser takes ops and knows nothing about nodes, channels or time.
|
||
|
||
**A photographic underlay is not an op.** The registered source frame that an
|
||
animator traces over is a reference, not output, and it may not enter the indexed
|
||
buffer — the same rule `docs/architecture.md` already sets for handles and
|
||
vertex boxes. It is a `drawImage` at an affine on a separate canvas, which clips
|
||
at the canvas edge for free, and the only thing it needs from the model is the
|
||
world transform of the node it rides:
|
||
|
||
```clojure
|
||
(world-of resolver :head) ;; -> Float64Array[6]
|
||
```
|
||
|
||
Composed with image-pixels-to-local — **both axes divided by `imgH`**, never by
|
||
their own dimension — the photo is registered with the shapes by construction,
|
||
and an unregistered underlay is merely decorative. The tracing editor chooses
|
||
which source frame to show under a cel. That reference choice is independent of
|
||
the finished picture fps and does not change the dense analysis track. A cel can
|
||
therefore use any useful source frame as its drawing reference, even when that
|
||
frame is not one of the displayed picture poses.
|
||
|
||
A photo that has to sit *between* two drawn layers is the case that would make it
|
||
a `:bitmap` node with an op of its own. Nothing wants that yet: a reference is
|
||
either under everything or over everything at low alpha.
|
||
|
||
### Making it fast in CLJS
|
||
|
||
Three things, and only these three matter:
|
||
|
||
- **Decomposed and persistent for storage; flat and mutable for evaluation.**
|
||
Composed transforms are 6-element `Float64Array`s, not maps. Every renderer
|
||
does this; the storage form and the evaluation form are allowed to differ.
|
||
- **A cursor per channel.** Playback is sequential, so "most recent key at or
|
||
before `f`" is an advance of a saved index, O(1) amortised. Binary search only
|
||
on a seek. This is the difference between a `rsubseq` allocation per channel per
|
||
frame and none.
|
||
- **Preallocated point buffers per node.** Fixed topology means the size is known
|
||
at freeze time, so the vertices — the overwhelming majority of the per-frame
|
||
bytes — are written into a buffer the node already owns. A frame still
|
||
allocates its op maps and the sorted op vector; that is a dozen small objects
|
||
against hundreds of points, and pooling them would buy nothing and cost the
|
||
ability to pass an op list around as plain data. At 30fps, per-vertex
|
||
allocation is the thing that will make this stutter.
|
||
|
||
Because the buffers are reused, **ops must be consumed before the next frame is
|
||
asked for.** That is the contract the rAF loop wants anyway: it reads, blits,
|
||
and dispatches nothing.
|
||
|
||
### What is in app-db, and what is not
|
||
|
||
| In app-db (tier 1) | In tier 2, behind a handle |
|
||
| --- | --- |
|
||
| nodes, parentage, z, spans, stencils | dense channel blocks |
|
||
| channel definitions, `:interp`, `:generated` | analysis artifacts |
|
||
| **framed** values, **keyed** keys, `:over` layers | preallocated eval buffers |
|
||
| library / symbol definitions | composed transform scratch |
|
||
|
||
The rule: **anything a human placed is in the document; anything a generator
|
||
produced is a handle.** Which is the same line `docs/architecture.md` draws for
|
||
sync and baking, arrived at again from the renderer's side.
|
||
|
||
## The current parts, in this model
|
||
|
||
Proof that it covers what exists, not just what is wanted:
|
||
|
||
| Now | Becomes |
|
||
| --- | --- |
|
||
| `mouth` outer ring, every frame | node `:mouth`, `[:geom :pts]` dense, `:generated {:by :roto/lips-outer}` |
|
||
| `mouth_in`, hidden below aperture | node `:mouth-in`, parent `:mouth`, `[:geom :pts]` dense + `[:vis]` dense |
|
||
| `teeth` from image content | node `:teeth`, stencil `:mouth-in`, `[:geom :pts]` dense, `:generated {:by :interior/teeth}` |
|
||
| lid rings | nodes `:lid-r/-l`, `[:geom :pts]` dense |
|
||
| lash line (`offsetRing`) | not data — a stage-6 parameter on the node, `{:grow px}` |
|
||
| iris disc | node `:iris-r`, `:kind :disc`, parent `:lid-r`, stencil `:sclera-r`, `[:xform :pos]` dense (quantised at freeze), radius framed |
|
||
| square pupil | node `:pupil-r`, `:kind :rect`, parent `:iris-r`, stencil `:iris-r` |
|
||
| brow ring + quantised raise | node `:brow-r`, `[:geom :pts]` dense (the traced ring with height removed), `[:xform :pos]` dense (the quantised raise). **The decomposition design.md insists on is two channels.** |
|
||
| head plate, kept frames | node `:head`, `:symbol` per instance, keys on `[:symbol]` at kept frames |
|
||
| `makeXform` face-oval crop | **gone.** Placement is `[:xform :*]` on the face's own `:place`; the stage clips |
|
||
| `stabilize` transforms | dense `[:xform :*]` on `:head`, read through its optional `:anchors` map |
|
||
| registered underlay | not data — a UI layer riding `(world-of resolver :head)` |
|
||
| painted background cel | node per layer, `[:geom :pts]` **framed**, `[:style :color]` framed |
|
||
| `mouth lead` | `:time {:offset k}` on performance nodes only |
|
||
| `exposure` | `:time {:expose n}` on the clip root, inherited |
|
||
| picture fps | resolver samples marked generated channels at the picture rate; authored keys keep their own time |
|
||
| hand correction | an `:over` layer, `:offset` or `:replace` |
|
||
|
||
The brow row is the one worth looking at twice. `docs/design.md` argues at length
|
||
that the traced ring already contains the height, so the quantised raise must be
|
||
measured *out* and put *back* or the brow moves twice. In this model that is not
|
||
an argument to remember — it is two channels on one node, and getting it wrong
|
||
would mean writing the height into both.
|
||
|
||
## Palettes
|
||
|
||
Three levels, and keeping them apart is what makes a palette swap a
|
||
**reinterpretation** rather than an edit:
|
||
|
||
| Level | Holds | Lives on |
|
||
| --- | --- | --- |
|
||
| **tone** | which mark this is — `:skin-dark` | `[:style :color]`, a channel on the node |
|
||
| **ramp** | what that tone looks like *here* | `:palette`, a channel on the timeline |
|
||
| **the ramps** | every named palette | the project |
|
||
|
||
A node names a **tone**, never a colour and never a ramp. Which ramp the tone is
|
||
read in is decided by the timeline the node is in. So the same drawing reads day
|
||
or night without one stored value changing — which is the entire payoff of
|
||
indexed colour, and is why `docs/design.md` forbids sampled RGB: once a shape
|
||
holds a measured colour there is nothing left to reinterpret.
|
||
|
||
Named palettes are **variants over one tone vocabulary**, not arbitrary colour
|
||
lists. `:day` and `:night` both define `:skin-dark`; that is what keeps a swap
|
||
total and keeps `docs/design.md`'s closed vocabulary closed. A tone the ramp in
|
||
scope does not define resolves to the loud magenta, like any other missing index.
|
||
|
||
### The scope rule
|
||
|
||
`:palette` on a timeline is a channel like any other:
|
||
|
||
```clojure
|
||
{:frames 91
|
||
:palette {:animated? true :interp :hold :keys {0 :day, 48 :dusk, 72 :night}}
|
||
:nodes {...}}
|
||
```
|
||
|
||
**Absent means inherit** from the instancing context. **Present means this
|
||
timeline's content is read in that ramp, and it travels with the timeline** — a
|
||
symbol authored against `:night` stays night wherever it is placed. That is
|
||
lexical scope, and deliberately: a character with their own palette is a
|
||
character, not a decoration of whichever scene they were dropped into.
|
||
|
||
Composition is the same walk as `:time` — down the instance chain, **innermost
|
||
set palette wins**. An enclosing timeline's palette therefore applies to
|
||
everything inside it that does not set its own, which is adjustment-layer
|
||
behaviour with no adjustment layer in it. It is just scope.
|
||
|
||
And because it is an ordinary channel, a project switches palette over time with
|
||
keys on the root timeline, a child timeline switches on its own, and neither
|
||
knows about the other.
|
||
|
||
### One index space, partitioned by palette
|
||
|
||
A raster is one `Uint8Array` and an index means one colour in it, so two ramps in
|
||
one frame cannot both own index 2. The resolution: **the output index space is
|
||
the concatenation of the named palettes**, and a tone resolves to
|
||
`palette-base + tone-index`.
|
||
|
||
Everything downstream is then unchanged — one buffer, one flat table for
|
||
`->rgba`, no per-frame palette construction, and an index does not change meaning
|
||
between frames, so bakes and thumbnails stay valid.
|
||
|
||
Two consequences worth stating rather than discovering:
|
||
|
||
- **The limit is real and reachable.** 256 indices over a nine-tone vocabulary is
|
||
twenty-eight palettes. Detect it and say so; do not let it arrive as wrapped
|
||
colour.
|
||
- **It makes the stencil sharper.** A stencil is a colour key, so two nodes
|
||
sharing a tone share a stencil — a genuine weakness of the technique.
|
||
Partitioning the index space by palette means two nodes in *different* palettes
|
||
no longer collide at all, and the resolved stencil picks up whichever index the
|
||
stencil node actually drew in.
|
||
|
||
### Where it is resolved
|
||
|
||
At the op boundary, and nowhere else. `[:style :color]` holds a keyword all the
|
||
way through evaluation; the walk carries the palette in scope the same way it
|
||
carries the parent transform and the local frame; the op carries a resolved
|
||
index. The rasteriser never sees a tone name and the node never sees an index.
|
||
|
||
This also means the palette is a **parameter of evaluation**, not a global. The
|
||
resolver takes it alongside the store.
|
||
|
||
## Format on disk and on the wire
|
||
|
||
Tier 1 is EDN/transit: the node tree, channel definitions, framed values, keys,
|
||
layers, library. Kilobytes, human-readable, diffable, and leaf-addressable for
|
||
sync.
|
||
|
||
Dense blocks are separate content-addressed binaries — `Int16Array` for
|
||
geometry, `Float32Array` for transforms — with a small header naming the channel
|
||
path, frame count, stride, and the **fixed-point scale** of the node-local space
|
||
the block is in. Geometry is stored in the node's own space, not in raster space;
|
||
see "What space geometry is in".
|
||
|
||
**Not Lottie internally**, despite the property shape being borrowed from it.
|
||
Lottie has no palette-indexed colour, its shapes are bezier with in/out tangents
|
||
where these are integer polygons, and its interpolation defaults are the opposite
|
||
of what is wanted. It is a fine thing to write out one day and a bad thing to
|
||
store.
|
||
|
||
Output is deliberately not specified here. The target is encoding video in the
|
||
browser, which touches the op list and nothing above it — a writer consumes
|
||
frames, and frames are what stage 7 already produces.
|
||
|
||
## Deferred
|
||
|
||
- **Per-key easing.** The structure allows it; nothing should use it until a
|
||
parented transform on a painted cel asks for it.
|
||
- **More than two channel layers.** The `:over` vector is already a list; a real
|
||
blend stack with weights is the NLA, and it is not needed to fix a bad frame.
|
||
- **Skew beyond the field.** `[:xform :skew]` is in the transform and in the
|
||
composition order from the start, because adding a component to a decomposition
|
||
later means migrating every stored transform.
|
||
- **Instance channel overrides on symbols.** Compose-over is specified; only
|
||
colour and transform need it at first.
|
||
- **Constraints and drivers.** Blender's other half. A gaze that aims at a null
|
||
object is the obvious first one, and it is a long way off. Until then the one
|
||
gaze shared by two iris nodes is a UI rule, not a stored relationship — see
|
||
"One signal, two nodes".
|