Model head anchors and independent pose timing

This commit is contained in:
Olive Vaughn 2026-09-29 00:46:08 -04:00
parent a45e89f4e4
commit 49ece8dee6
17 changed files with 521 additions and 215 deletions

View file

@ -339,54 +339,47 @@ head is placed and scaled where it belongs and the rest of the frame is simply
not on stage. The full frame stays *available* for tracing without being
*visible*, and those are different requirements.
## The anchor: stabilisation is a channel, not a mode
## Head motion: free or anchored to measured frames
`stabilize` produces `{s, θ, tx, ty}` per frame, which is exactly
`[:xform :scale]`, `[:xform :rot]` and `[:xform :pos]`. So removing the head's
motion is not a pipeline setting — it is a question of **which node holds that
motion**, and the answer is one channel definition:
`stabilize` produces `{s, θ, tx, ty}` per source frame. Its inverse is stored
densely on `:head`'s position, rotation and scale channels. The same measured
track serves every placement choice:
```clojure
;; locked: the head sits still, for tracing and for judging articulation
[:xform :pos] {:animated? false :value [0.0 0.0]}
;; as filmed: the head moves around the stage
[:xform :pos] {:animated? true :interp :hold
:dense {:store "sha256:…" :stride 2 :frames 600}
:generated {:by :anchor/similarity}}
;; per plate: the head snaps at each selected frame and holds
[:xform :pos] {:animated? true :interp :hold :keys {0 […], 12 […], 23 […]}}
;; no :anchors — free: read the measured transform at the current frame
;; one key — lock to a chosen measured frame throughout
:anchors {0 12}
;; several keys — cut to another measured head transform at frame 40
:anchors {0 12, 40 42}
```
The three modes are the three channel shapes, on one channel, on one node. The
third is the one a plate strip wants — the head pose is stable for exactly as
long as a drawing is on screen — and it costs nothing because `:keys` already
exists. Its frame set is the kept-frame set, which is `suggestPlateFrames` in the
prototype and belongs to painting rather than to measurement.
The map is `local change frame -> measured source frame`. A single lock is a
one-key map. Position, rotation and scale read the same held source frame. The
frame set belongs to head placement, independently of plate drawings and stage
pose cuts. No measured block is copied into authored transform keys.
**Always measure, always store factored, toggle the parent.** The fit is computed
and the geometry is stored head-local in every mode, and only the parent's
channel changes. Two things downstream require it, and both would be lost by
making this an analysis-time switch:
**Always measure, always store factored.** The fit is computed and the geometry
is stored head-local in every mode. Only the frame address used to read the
head's measured transform changes. Two things downstream require that split:
- *Smoothing.* "Smooth the transform, never the contour" only means anything
while the two are separate.
- *Key selection.* A velocity minimum is "articulation paused" in head-local
space and "the head happened to be still" in image space.
It also makes the toggle an edit to the document rather than a reason to
re-analyse: tier 1, undoable, syncable, and instant.
This is a document edit, not a reason to re-analyse. A registered tracing photo
will use its own source frame's stabilising transform followed by the same
selected head placement, so it aligns with the vectors drawn over it.
### Two nodes, because two different things want that transform
```
:face group — AUTHORED. where the face sits on the stage, and how big.
:head group — MEASURED. the head's motion, or identity.
:head group — MEASURED. dense head motion read at the selected frame.
:mouth :mouth-in :teeth :lid-r :lid-l :brow-r :brow-l …
```
Switching modes rewrites `:head` and never touches `:face`, so it cannot move
Changing anchor keys edits `:head` and never touches `:face`, so it cannot move
something that was placed by hand. A group node is free, and keeping the authored
and the measured transform apart is the whole reason the transform is decomposed
in the first place.
@ -627,12 +620,12 @@ Proof that it covers what exists, not just what is wanted:
| brow ring + quantised raise | node `:brow-r`, `[:geom :pts]` dense (the traced ring with height removed), `[:xform :pos]` dense (the quantised raise). **The decomposition design.md insists on is two channels.** |
| head plate, kept frames | node `:head`, `:symbol` per instance, keys on `[:symbol]` at kept frames |
| `makeXform` face-oval crop | **gone.** Placement is `[:xform :*]` on `:face`; the stage clips |
| `stabilize` transforms | `[:xform :*]` on `:head` — framed identity, dense, or keyed at kept frames |
| `stabilize` transforms | dense `[:xform :*]` on `:head`, read through its optional `:anchors` map |
| registered underlay | not data — a UI layer riding `(world-of resolver :head)` |
| painted background cel | node per layer, `[:geom :pts]` **framed**, `[:style :color]` framed |
| `mouth lead` | `:time {:offset k}` on performance nodes only |
| `exposure` | `:time {:expose n}` on the clip root, inherited |
| picture fps | `:time {:source-fps s :sample-fps p}` on the clip root, applied after analysis |
| picture fps | resolver samples marked generated channels at the picture rate; authored keys keep their own time |
| hand correction | an `:over` layer, `:offset` or `:replace` |
The brow row is the one worth looking at twice. `docs/design.md` argues at length

86
docs/timing-model.md Normal file
View file

@ -0,0 +1,86 @@
# Timing model
The source footage, authored drawings, generated face motion, and stage placement
have different frame decisions. They share a clock but do not share one kept-frame
list. `timing-handoff.md` records earlier implementation notes.
## Frame spaces
- A source frame addresses a decoded image and its measured face data. Keep the
source cadence and, for variable-rate video, its presentation timestamp.
- A timeline frame addresses authored keys in the clip or symbol's local space.
- A stage frame is mapped through the symbol instance's offset and rate before
local frame decisions are read. Moving a placement does not rewrite its keys.
The analyzed source poses remain dense. A lower picture rate or a skipped pose
never removes source data or shortens audio.
## Head placement
Analysis fits each source frame's rigid landmarks into one common head-local
space. Its inverse is the measured head transform, stored densely on `:head`.
The head node has one optional anchor map:
```clojure
;; no :anchors free movement: read measured frame f at f
:anchors {0 12} ; one lock: use frame 12's transform throughout
:anchors {0 12, 40 42} ; keyed locks: switch to frame 42 at local frame 40
```
A key is `(local change frame -> measured source frame)`. Its value holds to the
next key. The map chooses position, rotation and scale together. Frame zero must
have a key when the map exists. The dense transform blocks remain intact, so
editing anchors is a small document change and re-freezing can replace the
measurements without losing the anchor choices.
A source image used for tracing should be registered with that image's measured
stabilizing transform, then the selected head transform, then the authored
`:face` placement. This makes the photo and head-local vectors share the same
orientation and position. Tracing-photo selection is a separate editor address;
it does not choose the head anchor.
The prototype stabilizes into the shot's mean rigid pose and uses an early
closed-mouth frame for raster framing. Those are internal analysis and framing
choices. The authored head-anchor map above controls which measured head pose is
shown over each range. It is independent of plate drawing starts.
## Performance poses
Generated mouth, eye, and brow channels can be sampled at a lower picture rate
without retiming authored keys. The normal rule picks the latest available pose
at or before a picture-grid time. A future performance policy may add important
closed-mouth poses and store manual keeps/drops separately from the rate's
proposal. Related parts should share a selected pose by default: a mouth outline,
interior, teeth and generated visibility must not disagree about its frame.
## Stage placement
A symbol placement has optional pose-cut tracks, separate from head anchors:
```clojure
:playback {:tracks {:mouth {0 12, 8 27}
:eye-r {0 0, 4 6}}}
```
These maps are also `(local change frame -> source pose frame)`. They select which
baked/generated shape pose appears on that placement. Before the first explicit
cut, normal generated motion continues. Cuts hold, without interpolation, until
the next cut. A `[:node id]` track can override one shape in a shared group.
Authored cels, transforms, and audio remain on their normal local time.
The current implementation reads retained frozen channels. A separate resolved
geometry bake is not implemented; when added, it must preserve addressable
candidate poses so stage cuts can still select any of them.
## Ownership
| Choice | Owner | Current state |
| --- | --- | --- |
| Source frames and timestamps | Footage/analysis | Constant-rate frame indexing exists; variable timestamps remain future work |
| Head anchor map | `:head` node | Implemented, stored with the node |
| Tracing cel starts and photo address | Authored cel | Separate future work |
| Generated picture-rate proposal and closure protection | Roto clip/symbol | Generated-only picture sampling exists; closure protection remains future work |
| Stage pose cuts | Symbol instance | Implemented, stored with the instance |
Preview and export use the same resolver for generated picture sampling and stage
cuts. Export still emits every timeline frame at the clip's audio rate.