arthur/docs/multi-face-representation.md

104 lines
5.2 KiB
Markdown
Raw Normal View History

2026-09-29 02:34:53 -04:00
# Multi-face representation
Status (2026-09-29): representation work complete. Reopen it for a concrete
requirement, rather than another round of abstract alternatives.
Implemented: each tracked face has a drawing timeline, placed by an ordinary
symbol instance. Timelines already provide local node names, independent playback,
and persistence. No new kind of scene container is needed.
```clojure
:symbols
2026-09-29 02:34:53 -04:00
{:main {:nodes {:root {:time {:mode :map :expose 2}}
:face {:parent :root :channels <source-to-stage placement>}
An occurrence is a node, with a clock of its own A lane's drawings were going to be one instance whose source was a KEYED channel: frame 0 says `:drawing-a`, frame 4 says `:drawing-b`, and the cels of a row are that channel's keys. Two things followed from it, and both were wrong. The first is that playback meant whichever shape the channel happened to have. A framed source played its symbol; a keyed source froze the selected frame. So `node/placed-at` read animation out of storage, and adding an ordinary key to a still turned it into an animation — the last-key bug, which was not a bug in the code so much as the rule working as written. But WHICH drawing is used and HOW time runs inside it are independent questions, and all four combinations are ordinary: hold one drawing, play one animation, cut between held drawings, cut between playing ones. So an occurrence names one symbol in `:source {:symbol ...}` and says how its source time advances in `:playback {:in :speed :end}` — `source = in + speed * f`, a hold being speed 0, with `:stop`, `:hold` or `:loop` at the end named rather than guessed. `node/placed-frame` samples it forwards, which works for holds too, and `node/source-time` is the separate, invertible edit map, nil where inversion is meaningless. The two were one function before, and a hold had to lie about one of them. The second is that a keyed source only looked necessary because an occurrence was assumed to need a ROW. It does not. A lane is a group with `:layout :sequence`, its occurrences are ordinary instances in the same flat node map, and `timeline/rows` draws them as cel blocks on the lane's own row: twelve exposures, one row, each cel still separately selectable and addressable. The vertical growth that justified the keyed source is a presentation question, and it is answered in the view. `arthur.domain.sequence` holds the first commands over that shape — add lane, append drawing, extend hold — each one history step, each refusing rather than half-applying. Extending a hold leaves the lane's keys at their authored times, because you are adjusting drawings underneath timed motion; a correction owned by an occurrence travels with it. Ownership does that work, so no key needs a flag saying what it follows. Ripple past the symbol's end is refused with the frame count it would need, and `:extent :grow-symbol` is the caller saying yes. `clip/blank` no longer carries `:subjects {} :features {} :groups {}`. Empty maps write no leaf, so a blank document could not survive its own round trip — `leaf/leaves` promises exactness and was the only honest side of that. Documents are schema 3. A version 2 document is not read; nothing here converts one. `docs/lane-model.md` is the design, and says which of its parts are built. 392 tests, 5,525 assertions, and `test/browser/sequence.mjs` drives the editor through create, hold, explicit overflow and undo. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-30 15:20:23 -04:00
:face-1 {:kind :instance :source {:symbol :face-1}
:parent :face :z "a0"}
:face-2 {:kind :instance :source {:symbol :face-2}
:parent :face :z "a1"}}}
2026-09-29 02:34:53 -04:00
:face-1 {:nodes {:head {...} :mouth {:parent :head ...} ...}}
:face-2 {:nodes {:head {...} :mouth {:parent :head ...} ...}}}
:features
{:face-1/mouth {:subject :face-1 :timeline :face-1 :area :mouth
:nodes [:mouth :mouth-in]}
:face-2/mouth {:subject :face-2 :timeline :face-2 :area :mouth
:nodes [:mouth :mouth-in]}}
```
The example omits ordinary ids, frame counts and channel details.
## What belongs where
- A **subject** identifies a source track and supplies shared measurement settings.
Its timeline has the same id and contains its measured `:head`.
- A **feature** owns nodes in an explicitly named timeline. Its clip-level id is
qualified when generated; ownership is read from fields, never parsed from ids.
- A **node** has a local name. Parents, stencils and pose groups use local names too.
- An **instance** places and retimes a drawing. Its pose tracks can hold one face's
mouth while the other face continues moving.
Subject metadata and drawing timelines remain separate facts. Hand-drawn timelines
need no subject. Features retain explicit timeline references, so their locations
are not inferred from their labels.
Both filmed faces share one source-to-stage transform. Fitting them independently
would stack them at the center. Additional placement uses each instance's ordinary
channels. Cross-face draw order is the instances' `:z` order.
## Consequences
Regeneration updates the addressed timeline directly. There is no temporary swap
into `:main`, no special stage regeneration path, and no renaming of parents or
stencils. Composing a stage moves the take's root into the library and preserves
its child timelines. Settings appear once per tracked object, since all placements
read that same drawing.
Presence masks inside a subject use local feature names, matching measurement.
Freeze qualifies them when building block descriptors. Head and retained-source
blocks explicitly name their subject; otherwise two faces with identical detection
masks could produce different bytes under the same key. Analysis addresses also
include detection capacity and assignment settings, so old single-face detection
results cannot satisfy a new multi-face request.
Validation counts node ownership by `[timeline node]`. The server requires one
complete set of retained source roles **per subject**, rather than exactly three
blocks for the entire analysis.
Nested rectangles retain fractional sizes until rasterization. Rounding inside a
face timeline discarded small head-local pupils before the source-to-stage scale
was applied. This was a real rendering error missed by the earlier proposal's
coordinate-only benchmark; the regression now compares all mark extents as well.
## Verification and limits
`frontend/test/arthur/flow/multi_face_test.cljs` exercises two distinct subjects,
source block separation, detection and feature gaps, independent pose cuts and
regeneration, nested stage save/load, and equivalence to a flat single-face scene.
Existing geometry, raster, source, and regeneration tests cover the same paths.
`clips/tests/test_api.py` checks complete source roles per subject and immutability.
The browser suite checks rendering, pupils, playback, save/open, drawing and upload.
Assignment remains a nearest-centroid heuristic, with a version and distance gate
recorded in the analysis. Reordered detections, late arrivals and gaps are tested;
identity through crossings or long disappearances is not guaranteed. Assignment
happens before measurement, so correcting it requires measuring again.
This changes the freeze and retained-source contracts. It does not migrate older
flat captures; reanalyze their footage to use the new regeneration path. Existing
rendering and leaf codecs still understand their node/channel representation.
## Next steps
1. Commit the verified checkpoint: 319 frontend tests, 44 API tests, browser
checks, and app/test builds passed. Builds reported no warnings.
2. Exercise real two-person footage, especially crossings, late arrivals and
disappearances. Check assignment before treating the resulting geometry as
evidence about the representation.
3. Build performance-pose selection and instance-scoped picture rates, reusing
the existing held-frame lookup and pose groups.
4. Then build plate-drawing selection and independent tracing references.
The [timing handoff](timing-handoff.md) owns the detailed next implementation
sequence. Older flat captures need reanalysis unless a migration is separately
undertaken to preserve their authored edits.