multi fce stuff

This commit is contained in:
Olive Vaughn 2026-09-29 02:34:53 -04:00
parent 49ece8dee6
commit ddabfbeaa8
32 changed files with 1309 additions and 700 deletions

View file

@ -0,0 +1,101 @@
# Multi-face representation
Status (2026-09-29): representation work complete. Reopen it for a concrete
requirement, rather than another round of abstract alternatives.
Implemented: each tracked face has a drawing timeline, placed by an ordinary
symbol instance. Timelines already provide local node names, independent playback,
and persistence. No new kind of scene container is needed.
```clojure
:timelines
{:main {:nodes {:root {:time {:mode :map :expose 2}}
:face {:parent :root :channels <source-to-stage placement>}
:face-1 {:kind :symbol :of :face-1 :parent :face :z "a0"}
:face-2 {:kind :symbol :of :face-2 :parent :face :z "a1"}}}
:face-1 {:nodes {:head {...} :mouth {:parent :head ...} ...}}
:face-2 {:nodes {:head {...} :mouth {:parent :head ...} ...}}}
:features
{:face-1/mouth {:subject :face-1 :timeline :face-1 :area :mouth
:nodes [:mouth :mouth-in]}
:face-2/mouth {:subject :face-2 :timeline :face-2 :area :mouth
:nodes [:mouth :mouth-in]}}
```
The example omits ordinary ids, frame counts and channel details.
## What belongs where
- A **subject** identifies a source track and supplies shared measurement settings.
Its timeline has the same id and contains its measured `:head`.
- A **feature** owns nodes in an explicitly named timeline. Its clip-level id is
qualified when generated; ownership is read from fields, never parsed from ids.
- A **node** has a local name. Parents, stencils and pose groups use local names too.
- An **instance** places and retimes a drawing. Its pose tracks can hold one face's
mouth while the other face continues moving.
Subject metadata and drawing timelines remain separate facts. Hand-drawn timelines
need no subject. Features retain explicit timeline references, so their locations
are not inferred from their labels.
Both filmed faces share one source-to-stage transform. Fitting them independently
would stack them at the center. Additional placement uses each instance's ordinary
channels. Cross-face draw order is the instances' `:z` order.
## Consequences
Regeneration updates the addressed timeline directly. There is no temporary swap
into `:main`, no special stage regeneration path, and no renaming of parents or
stencils. Composing a stage moves the take's root into the library and preserves
its child timelines. Settings appear once per tracked object, since all placements
read that same drawing.
Presence masks inside a subject use local feature names, matching measurement.
Freeze qualifies them when building block descriptors. Head and retained-source
blocks explicitly name their subject; otherwise two faces with identical detection
masks could produce different bytes under the same key. Analysis addresses also
include detection capacity and assignment settings, so old single-face detection
results cannot satisfy a new multi-face request.
Validation counts node ownership by `[timeline node]`. The server requires one
complete set of retained source roles **per subject**, rather than exactly three
blocks for the entire analysis.
Nested rectangles retain fractional sizes until rasterization. Rounding inside a
face timeline discarded small head-local pupils before the source-to-stage scale
was applied. This was a real rendering error missed by the earlier proposal's
coordinate-only benchmark; the regression now compares all mark extents as well.
## Verification and limits
`frontend/test/arthur/flow/multi_face_test.cljs` exercises two distinct subjects,
source block separation, detection and feature gaps, independent pose cuts and
regeneration, nested stage save/load, and equivalence to a flat single-face scene.
Existing geometry, raster, source, and regeneration tests cover the same paths.
`clips/tests/test_api.py` checks complete source roles per subject and immutability.
The browser suite checks rendering, pupils, playback, save/open, drawing and upload.
Assignment remains a nearest-centroid heuristic, with a version and distance gate
recorded in the analysis. Reordered detections, late arrivals and gaps are tested;
identity through crossings or long disappearances is not guaranteed. Assignment
happens before measurement, so correcting it requires measuring again.
This changes the freeze and retained-source contracts. It does not migrate older
flat captures; reanalyze their footage to use the new regeneration path. Existing
rendering and leaf codecs still understand their node/channel representation.
## Next steps
1. Commit the verified checkpoint: 319 frontend tests, 44 API tests, browser
checks, and app/test builds passed. Builds reported no warnings.
2. Exercise real two-person footage, especially crossings, late arrivals and
disappearances. Check assignment before treating the resulting geometry as
evidence about the representation.
3. Build performance-pose selection and instance-scoped picture rates, reusing
the existing held-frame lookup and pose groups.
4. Then build plate-drawing selection and independent tracing references.
The [timing handoff](timing-handoff.md) owns the detailed next implementation
sequence. Older flat captures need reanalysis unless a migration is separately
undertaken to preserve their authored edits.

View file

@ -2,7 +2,7 @@
Self-contained. You should not need any prior conversation to execute this.
**Implementation status (2026-09-28):** steps 0–9 are in. Step 6 reads extracted
**Implementation status (2026-09-29):** steps 0–9 are in. Step 6 reads extracted
footage, detects landmarks with local MediaPipe assets at full source cadence, and
runs the same freeze path as the synthetic take. The scene time map can sample the
frozen roto at a lower picture fps without changing source analysis, duration or
@ -13,9 +13,24 @@ absence intervals through measurement and freeze. Step 9 adds the Django backend
the three-tier split, content-addressed tier 2 with the detector version inside
every key, leaf addressing for tier 1, and project load/save that round-trips.
**Still open.** Step 8's parameter UI and scoped regeneration, and automatic
per-feature detection. Everything under "Out, and do not build it" below, which
step 9 did not touch.
Step 8 now has parameter controls and scoped regeneration from retained source.
Multi-face representation is complete: each tracked subject has a drawing
timeline, placed by an ordinary symbol instance. See
[multi-face representation](multi-face-representation.md) for the implemented
model, verification and compatibility limits.
**Next, in order:** commit the verified checkpoint; exercise real two-person
footage, including crossings and disappearances; build performance-pose
Suggest/Keep/Drop and instance-scoped picture rates; then add plate-drawing
selection and independent tracing references. The
[timing handoff](timing-handoff.md) records current code and implementation order.
Reopen the representation only for a concrete requirement it cannot express.
**Still open:** real-footage identity validation, automatic per-feature detection,
the timing and tracing work above, and time-varying parameter settings. Older flat
captures need reanalysis for the new regeneration path; no migration is included.
The step descriptions below retain the original port scope; this status and the
linked handoffs describe subsequent work.
## What arthur is
@ -355,7 +370,7 @@ stencilled by the sclera.
minus paint. The fixed pixel thresholds remain provisional; step 8 exposes their
parameters for tuning without changing the source track or picture timing.
### 8 — knobs
### 8 — knobs — DONE for static settings and scoped regeneration
Build the parameter model before its UI. Define each parameter once with its
default, validation, applicable area and regeneration dependencies. Store values
by stable subject and feature ID. Represent an eye pair as one group with one or
@ -376,8 +391,9 @@ visible eye. This is an input format, not a control UI or an automatic detector.
Use leaf-addressable settings under the clip, subject, feature and optional
group. Retain source measurements so a setting change can regenerate affected
channels without re-detecting footage. Time-varying parameter values and all
parameter controls are deferred to the UI pass.
channels without re-detecting footage. Static parameter controls and scoped
regeneration are implemented for takes and composed stages. Time-varying
parameter values remain deferred.
### 9 — backend — DONE
Django project, the `clips` app, models for
@ -431,7 +447,7 @@ nobody should pre-empt by porting the old one.
## Two things to not foreclose
Feature controls will later handle more than one face and editing presence.
Feature controls now handle more than one face; editing presence remains future work.
The underlying identity, occlusion and group association model begins in step 8:
- **Presence is not visibility.** An occluded feature has *no value* on a frame,
@ -440,5 +456,6 @@ The underlying identity, occlusion and group association model begins in step 8:
- **Params carry stable identity.** A subject and its features keep their IDs
across observation gaps. A run of visible frames is not a new identity.
The identity tracker, when it comes, should use the same pattern the iris and brow
correspondences already use: vote across every frame rather than trusting one.
The current identity tracker uses nearest-centroid assignment. Validate it on
real crossings and disappearances before choosing a more elaborate policy; the
iris and brow correspondence code offers whole-take voting as one option.

128
docs/timing-handoff.md Normal file
View file

@ -0,0 +1,128 @@
# Timing and frame-selection handoff
Status (2026-09-29): the multi-face representation is complete. Each face has a
local drawing timeline and an ordinary symbol instance. Keep that model; the next
feature is performance-pose selection, followed by plate drawings and tracing.
See [multi-face representation](multi-face-representation.md) for verification
and compatibility limits.
## Next steps, in order
1. **Commit the verified checkpoint.** Representation, scoped regeneration,
nested stage composition and source persistence are implemented and tested.
2. **Exercise real two-person footage.** Include crossings, late arrivals and
disappearances. Assignment is still a nearest-centroid heuristic; inspect
whether identities, landmarks and mouth crops stay together. Correcting an
assignment requires measuring again. Do not redesign the representation to
compensate for an assignment failure.
3. **Build performance-pose selection.** Propose frames from a target picture
rate, allow explicit Keep/Drop edits, and apply requests per instance. Reuse
the existing held-frame lookup and generated pose groups. Keep authored keys
and audio timing intact.
4. **Then build plate drawings and tracing.** Suggest drawing frames from head
displacement, allow manual choices, and give each cel an independently
selectable tracing reference.
Older flat captures need reanalysis for the new regeneration path. Migrating their
existing authored edits is separate work; it is not implemented by this checkpoint.
## Timing decisions
Keep the dense analyzed frames. Generated motion holds the most recent selected
source pose; removing a selected pose never deletes source data or shortens the
clip. Store edits in the animation's local frame space, so moving an instance
does not move its edits. Authored keys follow intentional instance retiming but
must not be quantized by a picture-rate request. Clip FPS and audio duration stay
fixed.
There are two selections with different owners, sharing held-frame lookup:
- **Performance poses:** propose a kept-frame list from the target picture rate,
then apply explicit keep/drop edits. A parent instance may request a lower
rate. Mouth outline, interior, teeth and visibility read the same selected
source frame; likewise each eye's coupled parts. Use group overrides when
needed, rather than a setting on every channel. Head motion currently has its
own anchor selection; do not silently put it under mouth timing.
- **Plate drawings:** start with frame 0, walk measured rigid head poses, and
suggest a frame when maximum landmark displacement from the last kept pose
exceeds tolerance. Let the artist add/remove frames. A removed drawing stays
stored so it can reappear if restored. This selection does not thin the mouth.
A target rate is approximate. Pin a useful closed-mouth pose at its actual frame,
even if that produces more changes than the target. Do not show a future pose
early to fit a grid. Manual drop wins over an automatic suggestion; make removal
of the only closed pose in a beat visible in the UI. Skip missing detections when
suggesting a replacement. Keep a frame-zero selection and hold the last selection
through the end. A skipped pose (hold), `[:vis] false` (hidden), and an absent
measurement remain different facts.
Store manual edits separately from generated proposals so changing the rate or
tolerance retains hand decisions. Selection edits change the document, not dense
blocks or analysis addresses. Verify save/open for every new field; extend leaf
handling and the relevant key whitelist if its storage location requires it.
## Current code: reuse these mechanisms
- `domain/pose.cljs` already has `prepare`, `held-frame` and `source-frame`.
Instance `:playback :tracks` map local change frames to held source frames,
keyed by pose group. Reuse this lookup; frame suggestion and Keep/Drop policy
are the missing layer. An explicit cut is not itself a complete selection UI.
- `freeze/performance-nodes` marks generated animated channels with
`:pose-sampled?` and local `:pose-group` names. This includes keyed visibility
as well as dense geometry. `:generated` remains provenance for regeneration.
- `timeline/channel-frame` already applies explicit pose choices and default
picture sampling to marked channels. Playback and export both use
`clip/resolver` with `:picture-fps`; there is no need for a second sampling
implementation. Export's pose count is still a rate-based estimate.
- The picture-rate option is currently passed through the resolver tree
unchanged. Instance-specific parent requests are still to be implemented.
Instance offset/rate must apply before selecting the local source pose.
- `pose/put-cut` and `remove-cut` currently address instances in `:main`.
A take's face instances are there, but a composed stage nests them inside a
shared source timeline. Make the editing scope explicit when adding nested
controls. A request on one outer placement must not rewrite the shared
drawing's playback settings for every placement.
- Generic root `:time :expose` still retimes descendants, and frozen takes still
store it. Paint nodes are rootless to escape it. When the selection path
replaces take picture cadence, remove that redundant quantization from the
take default; preserve intentional generic time maps. Moving exposure to
`:head` would still retime authored children.
- `freeze/head-mode` supports `:free` and `:anchored`. It keeps measured channels
dense and writes optional per-subject `:anchors` maps; it does **not** implement
`:per-plate` mode or materialize transform keys from `:kept`. Plate selection
should reuse held measured-frame addresses where appropriate, without
rerunning analysis or copying the measurements.
- `:over` hand corrections are currently refused by the channel reader. Their
future application belongs after generated pose selection.
## Performance-pose implementation sequence
1. Add pure proposal and Keep/Drop policy around the existing held-frame lookup.
Cover frame zero, nondivisible rates, manual precedence, missing poses and a
protected mouth closure. Preserve all source frames.
2. Feed instance requests and group selections into the existing channel read
path. Cover two faces, two differently timed placements of one source, nested
instances, coupled visibility/geometry, and authored keys at their normal time.
Use this same path for preview and export; keep audio duration unchanged.
3. Wire the performance strip's Suggest/Keep/Drop controls and persistence.
Replace the export pose estimate with the actual selection count. Retire the
take's redundant root exposure only when this path replaces its behavior.
## Plate drawings and tracing, afterward
The old suggestion algorithm is `js/pipeline.js:suggestPlateFrames`; the strip,
worksheet and tracing photo are in `js/app.js`. Port the useful policy over the
measured head poses and reuse held-frame lookup for the resulting drawing set.
Give a cel an editor-only source-frame reference, defaulting to its plate frame
but independently changeable. It may point to a frame omitted from either rendered
selection. Register the photo using that source frame's measured transform. The
old prototype coupled photo and cel addresses; independent tracing is new work.
The old iris socket lock, gaze origin, CLJS head anchors and registration pivot
are separate settings. Clarify what an "origin-lock" request means before adding
that control.
Keep the UI to two scopes: **performance poses** and **plate drawings**, each with
Suggest/Keep/Drop. Tracing reference and lock controls live with the cel or feature
they affect. No general keyframe framework is needed for this work.