758 lines
37 KiB
Markdown
758 lines
37 KiB
Markdown
# arthur — the animation model
|
|
|
|
The data that describes a moving picture: what the primitives are, how they
|
|
nest, how they change over time, and how rotoscoped and hand-authored work end
|
|
up being the same thing with one flag between them.
|
|
|
|
`docs/design.md` is the aesthetic argument. `docs/architecture.md` is where the
|
|
code goes. This is the type that both of them are about.
|
|
|
|
## What this replaces
|
|
|
|
`docs/design.md` has a table of five kinds of part — plate, feature, interior,
|
|
primitive, scalar — each with its own source, vocabulary and interpolation. That
|
|
table is a good description of **where data comes from** and a bad description of
|
|
**what data is**, and the current code follows it too literally: eyes, brows,
|
|
teeth and mouth each get their own build function, their own key shape and their
|
|
own path through the prototype's writer.
|
|
|
|
They are all one thing. A part is a **node** with **channels**, and the five
|
|
kinds collapse into differences of which channels exist and who filled them in.
|
|
|
|
## Prior art, and what each one gets right
|
|
|
|
| System | The idea worth taking |
|
|
| --- | --- |
|
|
| **Flash / SWF** | A **library of symbols** and a timeline of **instances** at depths. "Framed" content that simply exists on a frame, versus tweened content. `DefineMorphShape` requires matching vertex counts — the fixed-topology rule, arrived at from the other direction. |
|
|
| **Blender** | Animation is **addressed by path into the data** (`location[0]`), not stored as fields on the object. An Action is a bag of F-Curves. That decoupling is what makes the dope sheet, the graph editor and the NLA three views of one dataset. Also: parenting captures a `parent_inverse` so the child does not jump. |
|
|
| **After Effects** | Every leaf property is animatable, uniformly. Property groups form a tree. Pre-comps nest arbitrarily and a pre-comp is just a layer. |
|
|
| **Lottie** | The uniform property shape: `{a: 0, k: <value>}` or `{a: 1, k: [<keys>]}`. One representation for static and animated, which is exactly "framed or keyframed". |
|
|
| **Grease Pencil** | A 2D layer holds frames at frame numbers, and a frame **holds until the next one**. Hold is the default, not a special case. |
|
|
|
|
What none of them get right for this project: colour. All four store RGB on the
|
|
shape. `docs/design.md` forbids that, so colour is a palette index here and it is
|
|
a channel like any other.
|
|
|
|
## The one idea
|
|
|
|
**Analysis is a channel generator.** It does not produce a different kind of
|
|
data; it produces keys, densely, on the same channels a hand would fill in
|
|
sparsely. So:
|
|
|
|
```
|
|
footage ──▶ analysis ──▶ FREEZE ──▶ channels on nodes ──▶ evaluate ──▶ raster
|
|
▲
|
|
hand authoring ──┘
|
|
```
|
|
|
|
Freezing is not a conversion into a second format. There is one format, and
|
|
freezing fills it in. That is what makes "the only difference is a special flag"
|
|
literally true: the flag is provenance on a channel, and nothing in the renderer
|
|
reads it.
|
|
|
|
## Node
|
|
|
|
A node is an instance in the scene. The tree is stored **flat, with parent
|
|
pointers** — never as nested maps.
|
|
|
|
```clojure
|
|
{:id :mouth
|
|
:name "mouth"
|
|
:kind :poly ; :poly :disc :rect :group :bitmap :symbol
|
|
:parent :head ; nil at the root
|
|
:z "a3" ; fractional index, ordered among all siblings
|
|
:symbol nil ; or :sym/blink — see Symbols
|
|
:stencil :mouth-in ; colour-key clip; structural, not a channel
|
|
:span [0 240] ; in/out in the parent's frame space
|
|
:pinv [1 0 0 1 0 0] ; parent-inverse, captured when parented
|
|
:channels {...}}
|
|
```
|
|
|
|
Flat with pointers, for four reasons that all point the same way: any node is
|
|
addressable without a walk; reparenting is a one-field write rather than a
|
|
subtree move; an edit to a leaf does not change the identity of its ancestors, so
|
|
re-frame's structural sharing keeps ancestor subs from invalidating; and it is
|
|
what lets every node be its own sync leaf. Flash, Blender and AE all store it
|
|
this way.
|
|
|
|
`:span` is Lottie's `ip`/`op` and Flash's `PlaceObject`/`RemoveObject`: the range
|
|
over which the node exists at all. Distinct from a `[:vis]` channel, which
|
|
blinks an existing node on and off.
|
|
|
|
### Subjects and tracked features
|
|
|
|
Scene nodes describe drawings, not tracking identity. A scene may also carry a
|
|
flat `:features` map. A feature ID stays stable for the whole clip, including
|
|
frames where that feature is occluded and later reappears:
|
|
|
|
```clojure
|
|
:subjects {:face-1 {:id :face-1}}
|
|
:features
|
|
{:eye-r {:id :eye-r :subject :face-1 :area :eye
|
|
:nodes [:eye-r :eye-r-in :iris-r :pupil-r] :params {}}
|
|
:eye-l {:id :eye-l :subject :face-1 :area :eye
|
|
:nodes [:eye-l :eye-l-in :iris-l :pupil-l] :params {}}
|
|
:mouth {:id :mouth :subject :face-1 :area :mouth
|
|
:nodes [:mouth :mouth-in] :params {}}
|
|
;; The teeth are their OWN feature and not three nodes of the mouth. A feature
|
|
;; carries the params of exactly one area, and the teeth have an `:area :teeth`
|
|
;; of their own — the otsu threshold, the tongue rejection, the radial contour's
|
|
;; vertex budget — which could not be reached if they were part of `:mouth`.
|
|
;; The coupling that made them look like the mouth's is real and is enforced
|
|
;; elsewhere: `:teeth` is STENCILLED by `:mouth-in`, and a node whose stencil drew
|
|
;; nothing is dropped, so an absent mouth takes the teeth with it without either
|
|
;; of them sharing an absence mask. An earlier draft of this block listed them
|
|
;; together; the code is right and this document was wrong.
|
|
:teeth {:id :teeth :subject :face-1 :area :teeth
|
|
:nodes [:teeth] :params {}}}
|
|
:groups
|
|
{:eyes-1 {:id :eyes-1 :kind :eye-pair :subject :face-1
|
|
:members [:eye-r :eye-l] :params {}}}
|
|
```
|
|
|
|
An eye pair is an explicit relationship between one or two eyes of the **same
|
|
subject**. It may have one member when only one eye has been identified; it does
|
|
not invent a second eye. Five subjects with nine identified eyes can have four
|
|
two-member pairs and one one-member pair. Each eye still has its own feature ID
|
|
and presence track. A group is a settings association, not a scene parent or a
|
|
tracking ID. Membership lives only on the group, avoiding a second pointer on
|
|
the feature that could disagree with it.
|
|
|
|
Each feature resolves settings from its area's definitions, then its group,
|
|
then its own `:params`. An eye can therefore inherit a pair setting or override
|
|
it without changing its partner. Removing it from a pair copies its effective
|
|
values into the feature first, so the result does not jump. Feature identity
|
|
and pair membership are clip-wide; a future parameter track can vary values
|
|
over time without splitting a feature at an observation gap.
|
|
|
|
Parameter definitions live in one registry: key, default, applicable area,
|
|
value constraints and affected areas. The registry supplies the take's defaults
|
|
today. The parameter UI and regeneration from edited values are later work.
|
|
|
|
Dense channel state records whether a measurement exists **for that feature on
|
|
that frame**. Occlusion means absent data on that frame, not a false `[:vis]`
|
|
value and not the end of the feature's identity. A full-face detection failure
|
|
makes all its features absent. A single occluded eye need only make that eye's
|
|
channels absent. Footage can carry explicit feature absence intervals in its
|
|
manifest, with one-based inclusive source frame numbers, for example
|
|
`"feature-absence": {"eye-r": [[10, 14]]}`. The loader expands these into
|
|
per-frame observation tracks before measurement. Unobserved landmarks may fill
|
|
rectangular numeric buffers, but they cannot contribute to an eye's contour,
|
|
blink or shared gaze. When one eye is absent, gaze uses the observed eye.
|
|
Until a detector supplies feature-level confidence, footage without annotations
|
|
uses the full-face detection mask as the fallback; it must not claim to detect
|
|
individual occlusions that it cannot see.
|
|
|
|
## Channel
|
|
|
|
Every animatable property is a channel, and channels are addressed **by path**:
|
|
|
|
```clojure
|
|
:channels
|
|
{[:xform :pos] {:animated? false :value [0.0 0.0]}
|
|
[:xform :rot] {:animated? false :value 0.0}
|
|
[:xform :scale] {:animated? false :value [1.0 1.0]}
|
|
[:xform :skew] {:animated? false :value [0.0 0.0]}
|
|
[:xform :anchor]{:animated? false :value [0.0 0.0]}
|
|
[:geom :pts] {:animated? true :interp :hold :dense {...} :generated {...}}
|
|
[:style :color] {:animated? false :value :skin-dark}
|
|
[:vis] {:animated? true :interp :hold :keys {0 true, 37 false}}}
|
|
```
|
|
|
|
A path is a **vector**, not a string — CLJS maps take vectors as keys natively,
|
|
so Blender's `data_path` idea arrives with no parsing. The set of valid paths for
|
|
a node follows from its `:kind`, and that is a spec, not a schema migration.
|
|
|
|
Three channel shapes, and the uniformity across them is the point:
|
|
|
|
```clojure
|
|
;; FRAMED — one static thing. No animation, no vertex correspondence to worry
|
|
;; about. A painted background cel is this.
|
|
{:animated? false :value v}
|
|
|
|
;; KEYED — sparse, authored, in the document. Undoable and syncable.
|
|
{:animated? true :interp :hold :keys {0 v, 4 v, 12 v}}
|
|
|
|
;; DENSE — generated, one value per frame, held in tier 2 as a typed array.
|
|
{:animated? true :interp :hold
|
|
:dense {:store "sha256:…" :offset 0 :stride 40 :frames 600}
|
|
:generated {...}}
|
|
```
|
|
|
|
`:interp` defaults to `:hold`, which `docs/design.md` requires of every cut part.
|
|
A key may carry its own `:interp` to override the channel's, which is how Lottie
|
|
and Blender both do per-key easing; nothing uses it yet and the door is cheap to
|
|
leave open.
|
|
|
|
### Keys are a map by frame, not a list
|
|
|
|
Already argued in `docs/architecture.md` for merge reasons; here it also gives
|
|
"the most recent key at or before `f`" as a `rsubseq` on a sorted map instead of
|
|
a scan. **Store a plain map** in the document — transit and JSON both lose
|
|
sortedness — and build the sorted index in the resolver.
|
|
|
|
### The flag lives on the channel, not the node
|
|
|
|
```clojure
|
|
:generated {:by :roto/lips-outer
|
|
:analysis "sha256:…" ; which analysis artifact
|
|
:params {:verts 8 :contour-avg 1 :aperture-cut 0.004}}
|
|
```
|
|
|
|
Present means the UI offers a parameter panel and a re-freeze button. Absent
|
|
means the UI offers the keys directly. **The renderer never reads it.**
|
|
|
|
It belongs on the channel rather than the node because a node routinely wants
|
|
both at once: a mouth whose `[:geom :pts]` is rotoscoped and whose `[:xform :pos]`
|
|
is hand-animated to sit on a plate. Putting the flag on the node would forbid the
|
|
most useful thing in the model.
|
|
|
|
### Channels are layered
|
|
|
|
A channel is a base plus optional override layers, and a layer declares how it
|
|
combines:
|
|
|
|
```clojure
|
|
{:animated? true :interp :hold
|
|
:dense {...} :generated {...}
|
|
:over [{:blend :offset :keys {88 [2 0], 96 [0 0]}}
|
|
{:blend :replace :keys {104 [[3 7] [4 7] …]}}]}
|
|
```
|
|
|
|
- **`:offset`** adds a delta to the base. "Nudge the mouth two pixels right for
|
|
ten frames" survives a re-freeze at different parameters, because it was never
|
|
a position — it was a correction.
|
|
- **`:replace`** wins outright. For the frame where detection simply failed.
|
|
|
|
This is what `docs/design.md` means by an override layer, and it is why
|
|
re-freezing is safe: the base is regenerated, the layers are untouched. It is
|
|
Blender's NLA blending and AE's effect stack at one property.
|
|
|
|
Layers are what "set it by hand" means for anything measured, and the measured
|
|
channel does not need to know. A hand-set gaze is an `:over` on
|
|
`[:xform :pos]` of the iris; a hand-set mouth shape is an `:over` on
|
|
`[:geom :pts]`. Turning the gaze-step or gaze-dwell knob regenerates the base and
|
|
leaves the correction alone, which is the entire reason a correction is stored as
|
|
a layer rather than written into the track.
|
|
|
|
**A `:replace` layer overrides absence, an `:offset` layer does not.** Sampling a
|
|
channel is: read the base, then apply the layers — and the base coming back
|
|
`absent` does not short-circuit that. `:replace` is explicitly for the frame
|
|
where detection failed, so it has to be able to supply a value where there is
|
|
none; `:offset` is a delta, and there is nothing to nudge, so an offset over an
|
|
absent base stays absent. Implemented the obvious way — bail out on absence
|
|
before reaching the layers — the one case the feature exists for is the one case
|
|
it would not cover.
|
|
|
|
### One signal, two nodes
|
|
|
|
Gaze is deliberately **one measurement shared by both eyes**: at this size the
|
|
per-eye difference is noise, and independent noise reads as wall-eyed
|
|
immediately, which is the most expensive artefact on a face. But it is stored as
|
|
`[:xform :pos]` on `:iris-r` and on `:iris-l`, which are two channels on two
|
|
nodes with two different parents — so the invariant lives in `measure` and
|
|
nothing in the document enforces it.
|
|
|
|
That matters as soon as either one can be overridden by hand, because an `:over`
|
|
on one iris alone reproduces exactly the artefact the shared measurement exists
|
|
to prevent. Until drivers exist, **the override is on both or on neither**, and
|
|
that is a rule the UI has to keep rather than one the data can.
|
|
|
|
This is the case that will eventually justify **drivers** — one value, evaluated
|
|
once, feeding several channels — which is why gaze is named in Deferred as the
|
|
obvious first one. Nothing here forecloses it: a driver needs a place in the
|
|
document and a `:driven-by` on a channel, both of which are additive, and an
|
|
absent key means "not driven". So it stays deferred, and the shape does not have
|
|
to change to allow it.
|
|
|
|
## Transform: decomposed, never a matrix
|
|
|
|
```clojure
|
|
{:pos [x y] :rot θ :scale [sx sy] :skew [kx ky] :anchor [ax ay]}
|
|
```
|
|
|
|
Stored decomposed for two reasons. Each component has to be independently
|
|
keyframable, which is the entire point of channels. And interpolating matrix
|
|
entries is meaningless — a rotation tweened through its matrix shears on the way.
|
|
|
|
Composition, per node:
|
|
|
|
```
|
|
local = T(pos) · T(anchor) · R(rot) · K(skew) · S(scale) · T(-anchor)
|
|
world = world(parent) · pinv · local
|
|
```
|
|
|
|
`:anchor` is Flash's registration point and Blender's origin: rotation and scale
|
|
happen about it, and getting it wrong is why hand-placed parts swing rather than
|
|
turn.
|
|
|
|
`:pinv` is Blender's `parent_inverse`, captured at the moment of parenting so the
|
|
child does not jump when it acquires a parent. Small, and its absence is the kind
|
|
of thing that makes a parenting feature feel broken.
|
|
|
|
**The similarity fit already produces a decomposition.** `fitSimilarity` returns
|
|
`{s θ tx ty}`, which drops straight into `[:xform :scale]`, `[:xform :rot]` and
|
|
`[:xform :pos]` with no conversion. The analysis output and the animation model
|
|
meet without an adapter, which is a sign the decomposition is the right one.
|
|
|
|
## What space geometry is in
|
|
|
|
**`[:geom :pts]` is always in the node's own local space, and the transform
|
|
chain says what that means.** There is no global geometry space and no decision
|
|
to make about one.
|
|
|
|
| Node | Its local space | Why that one |
|
|
| --- | --- | --- |
|
|
| a rotoscoped feature | head-local, isotropic, unit = one image height | what the anchor fit already produces; the `xform` to raster is not applied and not stored |
|
|
| a painted cel | the stage, in pixels, grid-snapped | the artist is placing pixels, so the pixel grid is the thing being authored |
|
|
| a primitive under a feature | its parent's | the iris is positioned on the lid ring, not on the stage |
|
|
|
|
This looks like a small clarification and it removes a whole class of argument.
|
|
The prototype bakes the framing into the numbers: `toRasterRing` applies
|
|
`makeXform`, which centres on the face oval's bounding box and zooms until the
|
|
face is 80% of the raster height, so **every stored vertex carries a cropping
|
|
decision** that was made once, at analysis time, from one frame's landmarks.
|
|
Dropping that step is a deletion, not a feature, and after it the framing is
|
|
simply a transform on a node.
|
|
|
|
Grid snapping belongs to the cel and not to the roto, for the same reason: a cel
|
|
is authored on the grid and a traced contour is not. So it is a property of a
|
|
node's space rather than a rule about all geometry, and the tension between
|
|
"integer polygons" and "arbitrary placement" was never real.
|
|
|
|
Each dense block therefore carries its own **fixed-point scale** in its header,
|
|
because a block in image-height units and a block in stage pixels need different
|
|
ones to fill an `Int16` usefully.
|
|
|
|
### There is no camera node
|
|
|
|
A camera is a global transform over everything, and nothing here wants one.
|
|
Placing the face on the stage is a transform on a node, which already exists;
|
|
what is not on the stage hangs off the edges and the canvas clips it. Every fill
|
|
in `domain/raster` already clamps rather than assuming it is inside, so drawing
|
|
past the edge is not a feature to add.
|
|
|
|
Project dimensions are therefore **independent of the footage**. A 1440x1920
|
|
portrait clip composited onto a 320x200 stage is not a problem to solve — the
|
|
head is placed and scaled where it belongs and the rest of the frame is simply
|
|
not on stage. The full frame stays *available* for tracing without being
|
|
*visible*, and those are different requirements.
|
|
|
|
## The anchor: stabilisation is a channel, not a mode
|
|
|
|
`stabilize` produces `{s, θ, tx, ty}` per frame, which is exactly
|
|
`[:xform :scale]`, `[:xform :rot]` and `[:xform :pos]`. So removing the head's
|
|
motion is not a pipeline setting — it is a question of **which node holds that
|
|
motion**, and the answer is one channel definition:
|
|
|
|
```clojure
|
|
;; locked: the head sits still, for tracing and for judging articulation
|
|
[:xform :pos] {:animated? false :value [0.0 0.0]}
|
|
|
|
;; as filmed: the head moves around the stage
|
|
[:xform :pos] {:animated? true :interp :hold
|
|
:dense {:store "sha256:…" :stride 2 :frames 600}
|
|
:generated {:by :anchor/similarity}}
|
|
|
|
;; per plate: the head snaps at each selected frame and holds
|
|
[:xform :pos] {:animated? true :interp :hold :keys {0 […], 12 […], 23 […]}}
|
|
```
|
|
|
|
The three modes are the three channel shapes, on one channel, on one node. The
|
|
third is the one a plate strip wants — the head pose is stable for exactly as
|
|
long as a drawing is on screen — and it costs nothing because `:keys` already
|
|
exists. Its frame set is the kept-frame set, which is `suggestPlateFrames` in the
|
|
prototype and belongs to painting rather than to measurement.
|
|
|
|
**Always measure, always store factored, toggle the parent.** The fit is computed
|
|
and the geometry is stored head-local in every mode, and only the parent's
|
|
channel changes. Two things downstream require it, and both would be lost by
|
|
making this an analysis-time switch:
|
|
|
|
- *Smoothing.* "Smooth the transform, never the contour" only means anything
|
|
while the two are separate.
|
|
- *Key selection.* A velocity minimum is "articulation paused" in head-local
|
|
space and "the head happened to be still" in image space.
|
|
|
|
It also makes the toggle an edit to the document rather than a reason to
|
|
re-analyse: tier 1, undoable, syncable, and instant.
|
|
|
|
### Two nodes, because two different things want that transform
|
|
|
|
```
|
|
:face group — AUTHORED. where the face sits on the stage, and how big.
|
|
:head group — MEASURED. the head's motion, or identity.
|
|
:mouth :mouth-in :teeth :lid-r :lid-l :brow-r :brow-l …
|
|
```
|
|
|
|
Switching modes rewrites `:head` and never touches `:face`, so it cannot move
|
|
something that was placed by hand. A group node is free, and keeping the authored
|
|
and the measured transform apart is the whole reason the transform is decomposed
|
|
in the first place.
|
|
|
|
## Time maps — exposure, lead and symbol timing are one thing
|
|
|
|
Every node may map the frame it is evaluated at:
|
|
|
|
```clojure
|
|
:time {:mode :inherit} ; the default, and almost always right
|
|
:time {:mode :map :expose 2 :offset -1 :rate 1.0 :loop? false}
|
|
:time {:mode :map :source-fps 30 :sample-fps 12} ; root: lower picture cadence
|
|
```
|
|
|
|
Three features that look unrelated are this one mechanism:
|
|
|
|
- **exposure** is `⌊f/n⌋·n`,
|
|
- **picture fps** quantises source time to a chosen picture grid, then reads the
|
|
latest source pose at or before that time; source analysis and audio keep their
|
|
original cadence,
|
|
- **mouth lead** is `f + k`,
|
|
- **a symbol instance's timing** is `(f - at)·rate + in`, with optional looping.
|
|
|
|
Composed along the nesting chain, outermost first. Two rules follow, and they are
|
|
different rules:
|
|
|
|
- **Exposure inherits strictly.** `docs/design.md` is emphatic that everything
|
|
rides one grid, because a head cutting on odd frames against a mouth cutting on
|
|
even ones reads as two performances. The model permits a per-node grid; the
|
|
default must be `:inherit`, and setting it lower is a deliberate act the UI
|
|
should make feel like one.
|
|
- **Offset is per-node by design.** Mouth lead applies to performance nodes and
|
|
*not* to the plate, which is the whole point of it — so the offset genuinely
|
|
belongs at the node, not the clip.
|
|
|
|
## Timelines, and why a scene is one
|
|
|
|
A **timeline** is an ordered bag of nodes in its own frame space:
|
|
|
|
```clojure
|
|
{:frames 91
|
|
:palette {...} ; see Palettes
|
|
:nodes {id -> node}}
|
|
```
|
|
|
|
That is the whole type, and **everything that holds nodes is one of these**:
|
|
|
|
- a clip's **scene** is its root timeline,
|
|
- a **symbol** in the library is a timeline,
|
|
- a node with `:kind :symbol` is an **instance** of one.
|
|
|
|
An earlier draft of this document had a scene and a `:kind :timeline` symbol as
|
|
two structures with the same fields and never said they were the same thing.
|
|
They are. Flash's `_root` is a MovieClip; After Effects' "a pre-comp is just a
|
|
layer" is already in the prior-art table above. Collapsing them is what makes
|
|
nesting arbitrary and free, rather than a feature to be added.
|
|
|
|
### Two axes of nesting, and they are different
|
|
|
|
This is the distinction the flat-storage rule is about, and conflating the two is
|
|
why "nested" and "flat with parent pointers" sound contradictory when they are
|
|
not:
|
|
|
|
| Axis | What nests | How it is stored |
|
|
| --- | --- | --- |
|
|
| **parent / child** | transform composition within one timeline | **flat, with parent pointers** — never nested maps |
|
|
| **instance** | a timeline inside another timeline | by reference into the library |
|
|
|
|
Each timeline is flat. Timelines nest. Every argument for flat storage —
|
|
addressability, one-field reparenting, structural sharing, per-node sync leaves —
|
|
is about the first axis and is untouched by the second.
|
|
|
|
The instance boundary is also **the only place the frame space changes.** Within
|
|
a timeline, `:time` is exposure and lead: a shift inside one space. At an
|
|
instance it is `(f - at)·rate + in`, into a different one. That is why `:rate` is
|
|
meaningless on an ordinary node and why sampling one must fail loudly rather than
|
|
be ignored.
|
|
|
|
### What is scoped to a timeline
|
|
|
|
Three fields on a node only have meaning relative to a timeline, and the answer
|
|
for all three is the same — **their own**:
|
|
|
|
- **`:z`** orders among siblings; a node cannot interleave with nodes inside a
|
|
nested instance. The instance occupies one position in its parent's order and
|
|
its contents sort beneath it, which the z path gives for free by being a
|
|
vector.
|
|
- **`:stencil`** names a node in the same timeline. A colour key does not
|
|
naturally respect a boundary — it is just pixels — so this is a rule rather
|
|
than a consequence, and it is Flash's rule for masks.
|
|
- **`:span`** is in the parent node's frame space.
|
|
|
|
### Instances
|
|
|
|
A node with `:kind :symbol` and `:of :sym/blink` places one. Its own channels
|
|
compose *over* the symbol's, so one definition is placed many times and tinted,
|
|
offset or retimed at each placement — that is how a three-frame blink is reused
|
|
at frames 40, 88 and 200 without copying it.
|
|
|
|
This is also where `docs/design.md`'s "closed vocabulary is right for the head"
|
|
lands: a plate library is a set of `:sym/head-*` timelines, and the strip chooses
|
|
which is instanced on which frame.
|
|
|
|
**Cursors and point buffers are per-instance, not per-node.** Two instances of
|
|
one symbol sit at different frames in their own space, so they cannot share a
|
|
reading head over the same channel. The resolver keys its caches by the instance
|
|
path, not by node id — which is a detail of `Making it fast` below, and the one
|
|
place symbol nesting is not free.
|
|
|
|
### Audio placements and controls
|
|
|
|
Sound is placed on a timeline as a separate `:audio` node. It uses the same
|
|
`:span`, `:time`, and channel representation as a drawn node. A `:linked-to` id
|
|
records which picture instance it was placed with; it does not force the two
|
|
spans or source in-points to match.
|
|
|
|
```clojure
|
|
{:id :voice-right :kind :audio :parent :root :z "a4"
|
|
:linked-to :right
|
|
:source {:footage "f8cace9e-..."}
|
|
:span [48 260]
|
|
:time {:mode :map :at 48 :in 0 :rate 1}
|
|
:channels {[:audio :gain]
|
|
{:animated? true :interp :linear
|
|
:keys {48 0.0, 60 1.0, 245 1.0, 259 0.0} :over []}}}
|
|
```
|
|
|
|
`[:audio :gain]`, `[:audio :pan]`, and `[:audio :rate]` are ordinary scalar
|
|
channels. They may be framed, keyed, or dense; numeric keyed channels can ramp
|
|
linearly. The time map sets the placement's base source rate, and
|
|
`[:audio :rate]` multiplies it. Audio is mixed from the referenced immutable
|
|
footage when the clip opens. The mix is derived output; the saved document holds
|
|
the nodes and channel keys, not another audio file. One audio element plays that
|
|
mix and remains the clock for both sound and picture.
|
|
|
|
This is also the boundary for a future control surface. A control has a stable
|
|
target, such as a feature's `:verts` setting or an audio node's
|
|
`[:audio :gain]` channel. The UI and a MIDI binding can address both through the
|
|
same control interface. Their update costs differ: gain can be keyed over time;
|
|
changing the number of lip vertices changes topology and must regenerate its
|
|
dense geometry. A topology setting cannot be treated as a per-frame gain curve.
|
|
|
|
## Evaluating a frame
|
|
|
|
```clojure
|
|
(defn eval-frame
|
|
"Scene at clip frame f -> draw ops in z order. Pure."
|
|
[scene f] ...)
|
|
```
|
|
|
|
1. Walk nodes in **topological order** by parent depth (cached; recompute only
|
|
when parentage changes).
|
|
2. Skip nodes outside `:span`.
|
|
3. Apply the node's time map to get its own local frame `fn`.
|
|
4. **Sample** each channel at `fn`: a map lookup for framed, a sorted-index
|
|
lookup for keyed, an array read for dense. Then apply `:over` layers.
|
|
5. Compose `world` from the parent's.
|
|
6. Transform geometry into raster space, writing into a **preallocated buffer**
|
|
owned by the node.
|
|
7. Emit `{:kind :poly :pts buf :n 20 :color idx :stencil id}`.
|
|
8. Sort by resolved `z`.
|
|
|
|
The op list is the boundary with stage 7 in `docs/architecture.md`: the
|
|
rasteriser takes ops and knows nothing about nodes, channels or time.
|
|
|
|
**A photographic underlay is not an op.** The registered source frame that an
|
|
animator traces over is a reference, not output, and it may not enter the indexed
|
|
buffer — the same rule `docs/architecture.md` already sets for handles and
|
|
vertex boxes. It is a `drawImage` at an affine on a separate canvas, which clips
|
|
at the canvas edge for free, and the only thing it needs from the model is the
|
|
world transform of the node it rides:
|
|
|
|
```clojure
|
|
(world-of resolver :head) ;; -> Float64Array[6]
|
|
```
|
|
|
|
Composed with image-pixels-to-local — **both axes divided by `imgH`**, never by
|
|
their own dimension — the photo is registered with the shapes by construction,
|
|
and an unregistered underlay is merely decorative. The tracing editor chooses
|
|
which source frame to show under a cel. That reference choice is independent of
|
|
the finished picture fps and does not change the dense analysis track. A cel can
|
|
therefore use any useful source frame as its drawing reference, even when that
|
|
frame is not one of the displayed picture poses.
|
|
|
|
A photo that has to sit *between* two drawn layers is the case that would make it
|
|
a `:bitmap` node with an op of its own. Nothing wants that yet: a reference is
|
|
either under everything or over everything at low alpha.
|
|
|
|
### Making it fast in CLJS
|
|
|
|
Three things, and only these three matter:
|
|
|
|
- **Decomposed and persistent for storage; flat and mutable for evaluation.**
|
|
Composed transforms are 6-element `Float64Array`s, not maps. Every renderer
|
|
does this; the storage form and the evaluation form are allowed to differ.
|
|
- **A cursor per channel.** Playback is sequential, so "most recent key at or
|
|
before `f`" is an advance of a saved index, O(1) amortised. Binary search only
|
|
on a seek. This is the difference between a `rsubseq` allocation per channel per
|
|
frame and none.
|
|
- **Preallocated point buffers per node.** Fixed topology means the size is known
|
|
at freeze time, so the vertices — the overwhelming majority of the per-frame
|
|
bytes — are written into a buffer the node already owns. A frame still
|
|
allocates its op maps and the sorted op vector; that is a dozen small objects
|
|
against hundreds of points, and pooling them would buy nothing and cost the
|
|
ability to pass an op list around as plain data. At 30fps, per-vertex
|
|
allocation is the thing that will make this stutter.
|
|
|
|
Because the buffers are reused, **ops must be consumed before the next frame is
|
|
asked for.** That is the contract the rAF loop wants anyway: it reads, blits,
|
|
and dispatches nothing.
|
|
|
|
### What is in app-db, and what is not
|
|
|
|
| In app-db (tier 1) | In tier 2, behind a handle |
|
|
| --- | --- |
|
|
| nodes, parentage, z, spans, stencils | dense channel blocks |
|
|
| channel definitions, `:interp`, `:generated` | analysis artifacts |
|
|
| **framed** values, **keyed** keys, `:over` layers | preallocated eval buffers |
|
|
| library / symbol definitions | composed transform scratch |
|
|
|
|
The rule: **anything a human placed is in the document; anything a generator
|
|
produced is a handle.** Which is the same line `docs/architecture.md` draws for
|
|
sync and baking, arrived at again from the renderer's side.
|
|
|
|
## The current parts, in this model
|
|
|
|
Proof that it covers what exists, not just what is wanted:
|
|
|
|
| Now | Becomes |
|
|
| --- | --- |
|
|
| `mouth` outer ring, every frame | node `:mouth`, `[:geom :pts]` dense, `:generated {:by :roto/lips-outer}` |
|
|
| `mouth_in`, hidden below aperture | node `:mouth-in`, parent `:mouth`, `[:geom :pts]` dense + `[:vis]` dense |
|
|
| `teeth` from image content | node `:teeth`, stencil `:mouth-in`, `[:geom :pts]` dense, `:generated {:by :interior/teeth}` |
|
|
| lid rings | nodes `:lid-r/-l`, `[:geom :pts]` dense |
|
|
| lash line (`offsetRing`) | not data — a stage-6 parameter on the node, `{:grow px}` |
|
|
| iris disc | node `:iris-r`, `:kind :disc`, parent `:lid-r`, stencil `:sclera-r`, `[:xform :pos]` dense (quantised at freeze), radius framed |
|
|
| square pupil | node `:pupil-r`, `:kind :rect`, parent `:iris-r`, stencil `:iris-r` |
|
|
| brow ring + quantised raise | node `:brow-r`, `[:geom :pts]` dense (the traced ring with height removed), `[:xform :pos]` dense (the quantised raise). **The decomposition design.md insists on is two channels.** |
|
|
| head plate, kept frames | node `:head`, `:symbol` per instance, keys on `[:symbol]` at kept frames |
|
|
| `makeXform` face-oval crop | **gone.** Placement is `[:xform :*]` on `:face`; the stage clips |
|
|
| `stabilize` transforms | `[:xform :*]` on `:head` — framed identity, dense, or keyed at kept frames |
|
|
| registered underlay | not data — a UI layer riding `(world-of resolver :head)` |
|
|
| painted background cel | node per layer, `[:geom :pts]` **framed**, `[:style :color]` framed |
|
|
| `mouth lead` | `:time {:offset k}` on performance nodes only |
|
|
| `exposure` | `:time {:expose n}` on the clip root, inherited |
|
|
| picture fps | `:time {:source-fps s :sample-fps p}` on the clip root, applied after analysis |
|
|
| hand correction | an `:over` layer, `:offset` or `:replace` |
|
|
|
|
The brow row is the one worth looking at twice. `docs/design.md` argues at length
|
|
that the traced ring already contains the height, so the quantised raise must be
|
|
measured *out* and put *back* or the brow moves twice. In this model that is not
|
|
an argument to remember — it is two channels on one node, and getting it wrong
|
|
would mean writing the height into both.
|
|
|
|
## Palettes
|
|
|
|
Three levels, and keeping them apart is what makes a palette swap a
|
|
**reinterpretation** rather than an edit:
|
|
|
|
| Level | Holds | Lives on |
|
|
| --- | --- | --- |
|
|
| **tone** | which mark this is — `:skin-dark` | `[:style :color]`, a channel on the node |
|
|
| **ramp** | what that tone looks like *here* | `:palette`, a channel on the timeline |
|
|
| **the ramps** | every named palette | the project |
|
|
|
|
A node names a **tone**, never a colour and never a ramp. Which ramp the tone is
|
|
read in is decided by the timeline the node is in. So the same drawing reads day
|
|
or night without one stored value changing — which is the entire payoff of
|
|
indexed colour, and is why `docs/design.md` forbids sampled RGB: once a shape
|
|
holds a measured colour there is nothing left to reinterpret.
|
|
|
|
Named palettes are **variants over one tone vocabulary**, not arbitrary colour
|
|
lists. `:day` and `:night` both define `:skin-dark`; that is what keeps a swap
|
|
total and keeps `docs/design.md`'s closed vocabulary closed. A tone the ramp in
|
|
scope does not define resolves to the loud magenta, like any other missing index.
|
|
|
|
### The scope rule
|
|
|
|
`:palette` on a timeline is a channel like any other:
|
|
|
|
```clojure
|
|
{:frames 91
|
|
:palette {:animated? true :interp :hold :keys {0 :day, 48 :dusk, 72 :night}}
|
|
:nodes {...}}
|
|
```
|
|
|
|
**Absent means inherit** from the instancing context. **Present means this
|
|
timeline's content is read in that ramp, and it travels with the timeline** — a
|
|
symbol authored against `:night` stays night wherever it is placed. That is
|
|
lexical scope, and deliberately: a character with their own palette is a
|
|
character, not a decoration of whichever scene they were dropped into.
|
|
|
|
Composition is the same walk as `:time` — down the instance chain, **innermost
|
|
set palette wins**. An enclosing timeline's palette therefore applies to
|
|
everything inside it that does not set its own, which is adjustment-layer
|
|
behaviour with no adjustment layer in it. It is just scope.
|
|
|
|
And because it is an ordinary channel, a project switches palette over time with
|
|
keys on the root timeline, a child timeline switches on its own, and neither
|
|
knows about the other.
|
|
|
|
### One index space, partitioned by palette
|
|
|
|
A raster is one `Uint8Array` and an index means one colour in it, so two ramps in
|
|
one frame cannot both own index 2. The resolution: **the output index space is
|
|
the concatenation of the named palettes**, and a tone resolves to
|
|
`palette-base + tone-index`.
|
|
|
|
Everything downstream is then unchanged — one buffer, one flat table for
|
|
`->rgba`, no per-frame palette construction, and an index does not change meaning
|
|
between frames, so bakes and thumbnails stay valid.
|
|
|
|
Two consequences worth stating rather than discovering:
|
|
|
|
- **The limit is real and reachable.** 256 indices over a nine-tone vocabulary is
|
|
twenty-eight palettes. Detect it and say so; do not let it arrive as wrapped
|
|
colour.
|
|
- **It makes the stencil sharper.** A stencil is a colour key, so two nodes
|
|
sharing a tone share a stencil — a genuine weakness of the technique.
|
|
Partitioning the index space by palette means two nodes in *different* palettes
|
|
no longer collide at all, and the resolved stencil picks up whichever index the
|
|
stencil node actually drew in.
|
|
|
|
### Where it is resolved
|
|
|
|
At the op boundary, and nowhere else. `[:style :color]` holds a keyword all the
|
|
way through evaluation; the walk carries the palette in scope the same way it
|
|
carries the parent transform and the local frame; the op carries a resolved
|
|
index. The rasteriser never sees a tone name and the node never sees an index.
|
|
|
|
This also means the palette is a **parameter of evaluation**, not a global. The
|
|
resolver takes it alongside the store.
|
|
|
|
## Format on disk and on the wire
|
|
|
|
Tier 1 is EDN/transit: the node tree, channel definitions, framed values, keys,
|
|
layers, library. Kilobytes, human-readable, diffable, and leaf-addressable for
|
|
sync.
|
|
|
|
Dense blocks are separate content-addressed binaries — `Int16Array` for
|
|
geometry, `Float32Array` for transforms — with a small header naming the channel
|
|
path, frame count, stride, and the **fixed-point scale** of the node-local space
|
|
the block is in. Geometry is stored in the node's own space, not in raster space;
|
|
see "What space geometry is in".
|
|
|
|
**Not Lottie internally**, despite the property shape being borrowed from it.
|
|
Lottie has no palette-indexed colour, its shapes are bezier with in/out tangents
|
|
where these are integer polygons, and its interpolation defaults are the opposite
|
|
of what is wanted. It is a fine thing to write out one day and a bad thing to
|
|
store.
|
|
|
|
Output is deliberately not specified here. The target is encoding video in the
|
|
browser, which touches the op list and nothing above it — a writer consumes
|
|
frames, and frames are what stage 7 already produces.
|
|
|
|
## Deferred
|
|
|
|
- **Per-key easing.** The structure allows it; nothing should use it until a
|
|
parented transform on a painted cel asks for it.
|
|
- **More than two channel layers.** The `:over` vector is already a list; a real
|
|
blend stack with weights is the NLA, and it is not needed to fix a bad frame.
|
|
- **Skew beyond the field.** `[:xform :skew]` is in the transform and in the
|
|
composition order from the start, because adding a component to a decomposition
|
|
later means migrating every stored transform.
|
|
- **Instance channel overrides on symbols.** Compose-over is specified; only
|
|
colour and transform need it at first.
|
|
- **Constraints and drivers.** Blender's other half. A gaze that aims at a null
|
|
object is the obvious first one, and it is a long way off. Until then the one
|
|
gaze shared by two iris nodes is a UI rule, not a stored relationship — see
|
|
"One signal, two nodes".
|