Give features stable identity, eye pairs and per-feature presence
Step 8's data model, ahead of its controls. Nothing here is a UI. domain/params holds every knob's definition once — default, applicable area, value constraints and the areas a change would force to regenerate. flow/take's literal knob map becomes a view of it, so the take's defaults and the future parameter panel cannot drift apart. domain/feature adds subjects, features and groups as document data the renderer never reads. A feature ID is stable for the whole clip, across occlusion: a run of visible frames is not a new identity. An eye pair is an explicit group of one or two eyes of the same subject, so a profile view with one identified eye needs no invented partner. Settings resolve area -> subject -> group -> feature, and dropping an eye from a pair materialises its effective values first so playback does not jump. scene/problems now validates all of it. Presence becomes per-feature rather than per-subject. freeze's :absent predicate takes a track as well as a frame, so one occluded eye can be absent while its partner still has a value; a full-face miss still marks everything absent. A manifest may annotate known gaps as one-based inclusive intervals, which ingest expands into observation tracks before measurement. An unobserved eye then gets no vote in the iris pairing and cannot steer the shared gaze — gaze falls back to whichever eye is visible. Temporal filters still see a sample on every frame, held from the last observed one, because the numbers are a rectangular buffer; the state mask, not the buffer, is what says the frame has no value. js/app.js gets the same occlusion lesson: leading nulls from a face that starts occluded used to throw away the whole take, and the neutral frame could be chosen from a held duplicate pose. Parameter editing, scoped regeneration and a feature-level detector remain. Until one exists, footage without annotations falls back to the full-face mask rather than claiming occlusions it cannot see. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01B87NVmiU36qQmN9gmFYnJ9
This commit is contained in:
parent
ccca93e233
commit
35ef150b48
18 changed files with 631 additions and 86 deletions
|
|
@ -79,6 +79,59 @@ this way.
|
|||
over which the node exists at all. Distinct from a `[:vis]` channel, which
|
||||
blinks an existing node on and off.
|
||||
|
||||
### Subjects and tracked features
|
||||
|
||||
Scene nodes describe drawings, not tracking identity. A scene may also carry a
|
||||
flat `:features` map. A feature ID stays stable for the whole clip, including
|
||||
frames where that feature is occluded and later reappears:
|
||||
|
||||
```clojure
|
||||
:subjects {:face-1 {:id :face-1}}
|
||||
:features
|
||||
{:eye-r {:id :eye-r :subject :face-1 :area :eye
|
||||
:nodes [:eye-r :eye-r-in :iris-r :pupil-r] :params {}}
|
||||
:eye-l {:id :eye-l :subject :face-1 :area :eye
|
||||
:nodes [:eye-l :eye-l-in :iris-l :pupil-l] :params {}}
|
||||
:mouth {:id :mouth :subject :face-1 :area :mouth
|
||||
:nodes [:mouth :mouth-in :teeth] :params {}}}
|
||||
:groups
|
||||
{:eyes-1 {:id :eyes-1 :kind :eye-pair :subject :face-1
|
||||
:members [:eye-r :eye-l] :params {}}}
|
||||
```
|
||||
|
||||
An eye pair is an explicit relationship between one or two eyes of the **same
|
||||
subject**. It may have one member when only one eye has been identified; it does
|
||||
not invent a second eye. Five subjects with nine identified eyes can have four
|
||||
two-member pairs and one one-member pair. Each eye still has its own feature ID
|
||||
and presence track. A group is a settings association, not a scene parent or a
|
||||
tracking ID. Membership lives only on the group, avoiding a second pointer on
|
||||
the feature that could disagree with it.
|
||||
|
||||
Each feature resolves settings from its area's definitions, then its group,
|
||||
then its own `:params`. An eye can therefore inherit a pair setting or override
|
||||
it without changing its partner. Removing it from a pair copies its effective
|
||||
values into the feature first, so the result does not jump. Feature identity
|
||||
and pair membership are clip-wide; a future parameter track can vary values
|
||||
over time without splitting a feature at an observation gap.
|
||||
|
||||
Parameter definitions live in one registry: key, default, applicable area,
|
||||
value constraints and affected areas. The registry supplies the take's defaults
|
||||
today. The parameter UI and regeneration from edited values are later work.
|
||||
|
||||
Dense channel state records whether a measurement exists **for that feature on
|
||||
that frame**. Occlusion means absent data on that frame, not a false `[:vis]`
|
||||
value and not the end of the feature's identity. A full-face detection failure
|
||||
makes all its features absent. A single occluded eye need only make that eye's
|
||||
channels absent. Footage can carry explicit feature absence intervals in its
|
||||
manifest, with one-based inclusive source frame numbers, for example
|
||||
`"feature-absence": {"eye-r": [[10, 14]]}`. The loader expands these into
|
||||
per-frame observation tracks before measurement. Unobserved landmarks may fill
|
||||
rectangular numeric buffers, but they cannot contribute to an eye's contour,
|
||||
blink or shared gaze. When one eye is absent, gaze uses the observed eye.
|
||||
Until a detector supplies feature-level confidence, footage without annotations
|
||||
uses the full-face detection mask as the fallback; it must not claim to detect
|
||||
individual occlusions that it cannot see.
|
||||
|
||||
## Channel
|
||||
|
||||
Every animatable property is a channel, and channels are addressed **by path**:
|
||||
|
|
|
|||
|
|
@ -7,8 +7,11 @@ Step 6 reads extracted footage from the manifest, detects landmarks with local
|
|||
MediaPipe assets at full source cadence, and runs the same freeze path as the
|
||||
synthetic take. The scene time map can sample the frozen roto at a lower picture
|
||||
fps without changing source analysis, duration or audio. Step 7 adds dense
|
||||
eyelids, shared gaze, brows and pixel-derived teeth. Step 8 is next: authored
|
||||
parameter controls and their scoped recomputation.
|
||||
eyelids, shared gaze, brows and pixel-derived teeth. Step 8's data model now
|
||||
has stable feature identity, feature-level presence, explicit eye pairs and
|
||||
shared parameter definitions. A manifest can now supply known feature absence
|
||||
intervals through measurement and freeze. Automatic per-feature detection,
|
||||
parameter editing and scoped regeneration remain.
|
||||
|
||||
## What arthur is
|
||||
|
||||
|
|
@ -326,9 +329,28 @@ minus paint. The fixed pixel thresholds remain provisional; step 8 exposes their
|
|||
parameters for tuning without changing the source track or picture timing.
|
||||
|
||||
### 8 — knobs
|
||||
The parameter UI, as leaf-addressed params in app-db
|
||||
(`clip/:cid/params/:subject/:area` — see below), so the sync layer added later has
|
||||
nothing to retrofit.
|
||||
Build the parameter model before its UI. Define each parameter once with its
|
||||
default, validation, applicable area and regeneration dependencies. Store values
|
||||
by stable subject and feature ID. Represent an eye pair as one group with one or
|
||||
two eye member IDs from the same subject; a profile view with one identified eye
|
||||
needs no invented partner. Each eye may override a pair value. Removing an eye
|
||||
from a pair materialises its effective values so playback does not change. Keep the
|
||||
existing frozen channels as renderer input; settings and provenance do not enter
|
||||
the render path.
|
||||
|
||||
Carry feature-level presence through freeze and dense channel state. The same
|
||||
feature ID covers every observed run across occlusion; a missing measurement
|
||||
has no channel value on that frame. Full-face detection is the fallback mask
|
||||
until there is a feature-level detector or authored presence data. A shared gaze
|
||||
measurement may still feed two independently identified eyes. A small manifest
|
||||
annotation can supply feature absence intervals now: the loader expands them
|
||||
before measurement, so invalid eye landmarks are ignored and gaze uses the
|
||||
visible eye. This is an input format, not a control UI or an automatic detector.
|
||||
|
||||
Use leaf-addressable settings under the clip, subject, feature and optional
|
||||
group. Retain source measurements so a setting change can regenerate affected
|
||||
channels without re-detecting footage. Time-varying parameter values and all
|
||||
parameter controls are deferred to the UI pass.
|
||||
|
||||
### 9 — backend
|
||||
Django project, the `clips` app, models for Project/Clip/Footage/Analysis/Leaf,
|
||||
|
|
@ -342,16 +364,14 @@ nobody should pre-empt by porting the old one.
|
|||
|
||||
## Two things to not foreclose
|
||||
|
||||
The feature controls will later be rethought to handle more than one face,
|
||||
periodic occlusion, stable identity across frames, and feature groups with their
|
||||
own parameters. That design can wait; two decisions here are free now and
|
||||
annoying to reverse:
|
||||
Feature controls will later handle more than one face and editing presence.
|
||||
The underlying identity, occlusion and group association model begins in step 8:
|
||||
|
||||
- **Presence is not visibility.** An occluded subject has *no value* on a frame,
|
||||
which is different from a part being hidden. Give every dense block a
|
||||
`Uint8Array` state mask per frame and let it mean *absent* as well as hidden.
|
||||
- **Params carry a subject segment.** `clip/:cid/params/:subject/:area`, with one
|
||||
subject today. Adding a path segment later touches every read and write.
|
||||
- **Presence is not visibility.** An occluded feature has *no value* on a frame,
|
||||
which is different from a part being hidden. Dense blocks carry a per-track,
|
||||
per-frame absence mask; `[:vis]` remains the sole hiding mechanism.
|
||||
- **Params carry stable identity.** A subject and its features keep their IDs
|
||||
across observation gaps. A run of visible frames is not a new identity.
|
||||
|
||||
The identity tracker, when it comes, should use the same pattern the iris and brow
|
||||
correspondences already use: vote across every frame rather than trusting one.
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue