arthur/docs/frame-selection.md

27 KiB
Raw Blame History

Frame selection

Since docs/tracing-symbol-plan.md: domain/trace is gone. Trace keys are the face's :plate placement's :time :holds, the origin is the head's :reads, and the photo's registration is the plate's own measured channels. Where this document names trace/measured-local or trace/prepare, read the plate's or the head's measured channels and node/hold.

Two mechanisms. One vocabulary. An earlier draft of this document claimed they were one component used twice — because suggestPlateFrames in the old js/pipeline.js and the never-built "performance poses" of timing-handoff looked like the same function — and that claim is wrong. They share how a selection is read and how the hand overrides one. They do not share how frames get chosen, because the two are answering questions of different shapes.

Time selection is the floor both stand on: an output frame reads the latest native frame at or before its time, and nothing rewrites the dense measurements. That is a cadence: an answer with no opinion about content. It cannot know that the one frame where the eye is fully closed is worth more than its neighbours, so at 12fps out of 30 it drops that frame two times in three. This document is how the picture gets an opinion.

The one idea

A selection is a set of frames chosen out of a dense measurement, and read by holding the latest one at or before now.

The holding half already exists and is already shared: pose/held-frame is called by node/hold (a placement's :time :holds) and by pose/source-frame, which is the two sites agreeing about reading. The hand half is shared too — see Three layers below. Choosing is what differs.

Plate drawings (tracing) Performance poses
the question which frames does an artist have to draw a head on? which frames does the picture change a shape on?
the cost being managed a person drawing a pose looking wrong
signal the measured head's motion — none; a stored cut
the baseline it improves on drawing on 2s the cadence, or the exposure grid
the shape of the answer a non-uniform set out of dense the same grid, nudged
lives on the face's :plate :time :holds the instance's :playback :tracks
hand edit today ::project/toggle-hold pose/put-cut / pose/remove-cut
UI today params/layer-section none
proposes today nothing nothing

Why they are not one function

A plate selection has to be non-uniform, and that is the whole reason it exists. A head still for sixty frames and then whipping across in ten wants two drawings for the first stretch and eight for the second. Drawing on 2s gives thirty-five drawings, most of them identical, and no amount of nudging a uniform grid will produce the distribution that is wanted — the spacing itself is the answer. That is what the prototype's walk was for, and it is why a cost knob (:tolerance) belongs on this side: the artist is buying drawings.

A performance selection is not choosing sparseness at all. The output rate or the exposure setting has already chosen it. The question left over is only which native frame each already-decided slot reads, and the failure it fixes is narrow: a slot landing one or two frames off the closure. Nudging the grid is the right size of answer, and there is nothing for a tolerance to mean.

There is a second, harder reason, and it is the one that settles it:

A selection cannot put a frame on screen that the output grid never samples. At 12fps out of 30, output frame 5 reads native 12 and output frame 6 reads native 15. A closure at native 13 is between them. Protecting frame 13 in a set of kept frames makes it available and makes it the frame held across 13 and 14 in native space — and at a 12fps output it still never appears, exactly as time.md says: an event between output frames cannot create an extra frame in a 12fps output. Only moving what output frame 6 reads can show it. So the performance side has to act on the grid, not on a set beside it.

Three layers, and the middle one is derived

The trap this is designed around is stated in timing-handoff and is worth repeating because it is the only hard rule here:

Store manual edits separately from generated proposals so changing the rate or tolerance retains hand decisions.

So a selection is:

{:policy {:tolerance 0.02}   ; what the proposer was asked for
 :keep   #{47}               ; frames the hand insists on
 :drop   #{30}}              ; frames the hand refuses

and the effective set is (proposed ∪ keep) \ drop, always containing frame 0.

:keep and :drop are the document. The proposal is not: it is recomputed from :policy and the dense signal whenever either changes. Re-suggesting at a new tolerance must never cost somebody their pinned blink, and that is the entire reason the hand decisions are stored as their own two sets rather than as the resulting frame list.

This layering is the part that really is shared. On the performance side there is no :policy worth storing — the grid is the policy — but :keep and :drop mean exactly what they mean on the plate side, and select/effective is the one implementation for both. A preserve mark is a keep.

Materialise the result, do not derive it on the render path

The effective set is written back to where each site already reads it — the plate's :time :holds, or the pose track — so that every existing reader is untouched and nothing on the per-frame path has to open a dense block. Proposing is a command, not a subscription. ch/value-at allocates per call and says so; that is fine for a button press over a few hundred frames and would not be fine at 30fps.

This means the stored frame list is redundant with policy + keep + drop. That is deliberate and it is the cheap direction of the trade: a stale list is recoverable by pressing Suggest again, and a dense read per node per frame is not recoverable at all.

The plate selection: a non-uniform chooser

arthur.domain.select, built — see What is revertible for the one part of it that is not yet wanted. Pure, no store access, no clip access: the site hands it a signal it has already read.

(defn propose
  "Frames worth keeping out of `n`, given `signal`."
  [n signal {:keys [tolerance protect]}])

(defn effective
  "`proposed` with the hand's decisions applied. Always contains 0."
  [proposed keep drop])

propose is the prototype's walk, generalised off landmarks:

  1. keep frame 0, make it the anchor;
  2. settle the protected frames from :protect before walking;
  3. for each later frame, keep it when it is protected, or when distance(anchor, f) > tolerance. Either way it becomes the anchor.

distance is the max absolute difference over components, so a signal of landmark pairs and a signal of one number both work without the caller saying which it handed over. Every sample in a signal is the same width, and a signal of two widths is refused: comparing the prefix two samples happen to share would let a reader that drops a component read as no movement at all.

A protected frame anchors the walk like any other kept frame, which is why protection is settled first and is not unioned onto the walk's result. The anchor is what is on screen; once a protected frame is kept the viewer is looking at it, so measuring the next frame's drift from a frame no longer displayed is wrong.

An absent measurement is nil, and converting to that is the reader's job. A frame where the face was not found says nothing about the signal: it cannot move the anchor and it is not a frame worth keeping. trace/measured-local already returns nil there. A protected frame with no measurement is still kept — a plate frame is a frame somebody draws on whether or not the detector found a face.

A non-finite tolerance falls back to nought. A cleared slider reads as NaN and every comparison against NaN is false, which taken literally proposes frame 0 alone and collapses the whole take to one drawing. Nought proposes every frame that changes, which is merely the baseline back again: wrong in a way somebody can see and undo.

The signal

The head's measured transform, applied to a fixed reference quad, giving displacement in stage units — so a tolerance means "the head has moved this far" and is a number a person can reason about. trace/measured-local already builds that matrix per frame and is private; make it public rather than writing a second one. Map it over the frames and transform four corners through node/apply-pt!.

The quad's size is a real parameter hiding in the word "fixed": it sets how much rotation registers against translation. Give it a name and a comment rather than an inline literal.

Do not reach for raw landmarks. The prototype used them because it had them lying around; the transform is what the drawing actually follows, it is already on the node, and it is three channels instead of a block.

The upgrade path, and do not start here

The greedy walk is order-dependent and slightly suboptimal. The optimal version is a dynamic program — choose k frames minimising held-reconstruction error, which is textbook segmented least squares and is O(n²k), nothing at n≈300 — and it keeps extrema for free, because an extremum is exactly where a zero-order hold is most wrong.

Build the greedy one first anyway. It is proven, it shipped in the prototype, and having two implementations to compare is how the DP gets tested. Swap it behind propose afterwards, where the signature already permits it.

Most of select_test pins the greedy walk's exact output, deliberately, for that comparison — so expect to rewrite those expectations when the DP lands, and keep them as greedy-specific tests rather than deleting them. The assertion that is a spec rather than a pinned vector, and should be written on this side before the swap, is:

reading the signal through the selection, held, never differs from the dense measurement by more than tolerance

That is what makes the tolerance number mean something to a person. It is true of the greedy walk by construction and it is what the DP optimises, so it survives the swap untouched.

The performance selection: a preserve-snap on the grid

Not built. The rule is ten lines; getting the two halves of it into the same place is the work. An earlier draft of this section said "about thirty lines" and that was understated — see Where it goes below.

A grid slot already picks a native frame — cadence/frame for the output rate, or the exposure fold for a deliberate hold at full rate. Write d(k) for the native frame slot k defaults to. Slot k is the first slot to cover everything in (d(k-1), d(k)], and the frames strictly inside that interval are the ones the grid shows to nobody. So:

Slot k reads the latest preserved frame in (d(k-1), d(k)], and d(k) when there is none.

That is the whole mechanism. It recovers a dropped frame out of the slot's own gap, it can never read a frame another slot already showed, and it cannot reach past d(k).

Snap backward only, never forward, and note which direction that actually is, because it is easy to get backwards. The closure at native 13 in a 12-from-30 output is recovered by slot 6 — whose default is 15 — reading 13. It is not recovered by slot 5, whose default is 12, reaching forward to 13: slot 5's instant is 5/12s = 0.4167s and native 13's is 13/30s = 0.4333s, so that would show the closure 17ms before the mouth shut. cadence/frame's contract is the latest native frame at or before the slot's time and cadence_test asserts selected <= f*native/grid over every grid and native pair, so reaching forward breaks a tested invariant as well as the no-lead rule. Reading 13 at slot 6 shows the closure two native frames late, which is the same lateness every hold already has.

The marks come from a cut that is already stored. flow/freeze computes both closures and keys them as [:vis], with the thresholding and hysteresis already decided:

  • the mouth, from the aperture relative to the take's peak — [:vis] on :mouth-in, provenance :roto/mouth-aperture (freeze.cljs:473);
  • the eyes, from condition/resolve-blink with its cut, dwell and hold — [:vis] keyed per eye part, provenance :roto/blink (freeze.cljs:533).

So "preserve the frames the cut says shut" reads what the document already holds. No new signal, no new dense track, no threshold decided twice. Brows have no closure and get no marks, which is correct: there is no extreme brow position worth protecting.

Precompute the mark set when the resolver is built, the way pose/prepare and trace/prepare already do (symbol.cljs:420, :470, :536). A [:vis] channel is keys, not dense, so reading it per node per frame would be cheap — but the snap also needs the marks sorted for a backward lookup, and building that per frame is the one thing animation-model.md and the render-path note above both forbid.

Where it goes, and why it is not a one-liner

The rule needs two things that currently live at opposite ends of the resolver:

  • the slot interval (d(k-1), d(k)], which needs the output slot index and the fps ratio. Both exist at clip.cljs:261, the single place the grid becomes a native frame: (cadence/frame f (or (:grid-fps opts) (:fps clip)) (fps clip sid)).
  • the marks, which are per pose group, and so belong where group identity exists: symbol/base-channel-frame (symbol.cljs:399), whose :pose-sampled? branch already calls pose/source-frame with (js/Math.floor lf) as the default pose. That default is the snap's seat. Replacing it leaves an explicit hand cut winning over a snap, which is correct — manual precedence is absolute — and costs no new plumbing on the pose side, because per-group choices are already threaded and already prepared.

By the time control reaches base-channel-frame there is only lf, a native local frame that placement and retime have already been through, so the slot interval cannot be recovered there. It has to be threaded down from clip.cljs:261 alongside the frame. That is the actual work of this step: two namespaces' internal signatures, not a drop-in.

Do not take the shortcut of snapping at clip.cljs:261 itself. It is right there, it needs no threading, and it is wrong: one native frame per output frame means the whole picture reads 13 instead of 15, so the head goes two frames stale for one output frame to fix the mouth. At 12fps that is a 67ms hitch on a moving head, and it fights the trace selection, which has its own opinion about which head frame to show. The snap is per group because the thing being recovered is one group's closure.

Why not an aperture signal

An earlier draft had this side read the group's aperture as a single-component signal and hand it to propose with :protect :extrema. It cannot. The aperture is the separation of landmarks 13 and 14, at positions 5 and 15 of the 20-slot LIPS-INNER ring, and freeze/rings->flat subsamples the ring to the verts budget: those two positions survive only when verts is a multiple of four. At verts 6, 10, 14 and 18 — all legal, all even — they are not in the stored data at all. ring/subsample-slots used to claim otherwise and has been corrected.

flow/measure/mouth does compute the exact scalar and freeze does throw it away after thresholding, so storing it was an option. Reading the cut is strictly less work and decides nothing twice.

How the two interact

They are keyed in the same frame space: clip/resolver hands an instance's :playback :tracks down into the child it places, and symbol/base-channel-frame reads both the trace and the pose choices at lf, the node's local frame inside that face. A trace frame and a pose cut are the same kind of number.

What differs is the owner, and that asymmetry is load-bearing:

  • the trace is the face's, on its :head — every instance of that face shares it, because it says how the drawings were made;
  • the pose tracks are the instance's — two placements of one face can be timed differently.

The hazard, which is already written down

animation-model.md states it for exposure and it is the same hazard here:

a head cutting on odd frames against a mouth cutting on even ones reads as two performances

It is not a correctness problem. The mouth is a child of :head and its geometry is stored head-local, so a mouth from frame 17 composed onto a head held at frame 12 is exactly lip-sync on a held drawing — the decomposition already decoupled them and nothing is geometrically wrong. The problem is perceptual, and perceptual problems want a constraint rather than a repair.

Nest them, do not couple them

The two are not peers. One is coarse and expensive — the prototype's own comment says it: "The cost being managed is an artist drawing a head, which is why the signal is head pose and not the mouth — the mouth is traced and free." The other is fine and cheap.

So the rule is a subset, in one direction only:

Every kept plate frame is a preserved frame of the performance selection.

When the drawing changes, the performance changes with it, so the two can never cut against each other on neighbouring frames. The mouth stays free to change on frames where the head does not, which is what shooting a held drawing with a live mouth is.

This is why the preserve set is a set and not a closure predicate: plate frames and shut-mouth frames go into the same pile, and the snap does not care which is which. It needs no new mechanism on either side.

The reverse is a suggestion and never automatic. A mouth closure is a reasonable place to want a new drawing, but proposing one spends somebody's afternoon. Offer it; do not take it.

It also composes with :origin for free. A head on :continuous has opted out of its own selection, so there are no plate frames to preserve and the constraint is vacuous — which is correct, because a continuously moving head cannot cut against anything.

Between the kept frames is a third shared field

The head's :reads (once :trace :origin) is not a tracing setting. It is the answer to what happens between kept frames, and the plate selection has to answer it:

  • :continuous — ignore the selection for this purpose and read the frame you are on,
  • :keys — jump to each kept frame and hold it, a hold and not a tween,
  • :start — hold the first forever.

The performance side does not need the field: a snap picks which frame a slot reads and the grid does the holding, so :keys is the only behaviour there is. Do not rename :origin to match anything: for a head it genuinely means where the face's origin goes, and a saved field in a shipped UI is not worth churning for a vocabulary tidy.

The modes

One setting, two states, and "manual" is not a third state.

  • off — the cadence alone, which is what ships today. The output frame reads the latest native frame at or before it, and nothing has an opinion.
  • smart frame picking — on the plate side the proposal is live, and :policy holds the tolerance; on the performance side the snap is active.

Hand keeps and drops apply in both states, which is why they are not a mode: turning smart picking off must not throw away the frames somebody pinned, and pinning a frame with smart picking off is a perfectly reasonable thing to want. Manual precedence is absolute — a drop beats a proposal, always.

The name on the toggle should be the same word in both sections. "Smart frame picking" is fine. What it must not be is two different names for the one idea, which is how these became two features the first time. That the two sections are now backed by different code is an implementation fact and must not reach the UI.

The UI

One component rendered twice, in ui/params.cljs:

smart frame picking   [ off | on ]
tolerance             [ ----•------- ]   0.02      (plate section only)
                      [ Suggest ]
frames   0  12  30  47*  61        (* = kept by hand, strikethrough = dropped)

The frame strip already exists in miniature — params/trace-keys draws the trace keys as seek buttons (params.cljs:386). Lift it into a shared component and give it three affordances: click to seek, a modifier to pin, a modifier to drop. A pinned frame and a proposed frame must be visually distinct, because "will this survive me moving the slider?" is the question the strip exists to answer.

Render it in two sections:

  • tracing · <face> (params.cljs:426), beside the existing origin row, with the tolerance slider and Suggest.
  • performance · <group> — new, on an instance's inspector, one per pose group the placed symbol has. No tolerance: the strip shows the preserved frames and the grid slots that snapped to them.

Order to build it

  1. domain/select with the greedy walk, and its tests. Done, in commit 02069e8. Pure, no store, no clip. Note that the extrema part of it is not wanted by anything below — see What is revertible.
  2. The preserve-snap, pulled forward ahead of the plate side because it is the half with no UI and nothing proposing today, and because it needs nothing from the plate side except a fold that can land last. Three pieces: the mark set from the stored [:vis] cuts, built in a prepare step beside pose/prepare; the slot interval threaded from clip.cljs:261; the rule itself, seated in base-channel-frame's default pose. The test that matters is the same synthetic case select_test uses — a one-frame closure at native 13 that a 12-from-30 grid drops — asserted end to end this time: the resolver shows a shut mouth on exactly one output frame, and the head's frame does not move while it happens. Write it first.
  3. The plate signal reader. trace/head-signal: make trace/measured-local public, map it over the frames, four corners through node/apply-pt!. Test that it returns a vector of the right length, that a motionless take proposes [0], and that a stretch where the face was not found becomes nil rather than a pose, a zero or a gap in the vector. Add the held-reconstruction invariant from The upgrade path here, since this is the first place a real signal exists to assert it over.
  4. Storage for the plate selection. :policy/:keep/:drop beside the plate's :time :holds. Extend leaf/leaves and the key whitelists in the same commit — a field without a leaf saves silently and comes back missing, which is the one bug persistence must not be able to have. Round-trip test.
  5. Re-suggest preserves hand decisions. Propose at one tolerance, pin a frame, drop a frame, propose at another, assert both survive. Settle here whether the hand's :keep is also fed to propose as :protect: under the layering as written it is not, so a pinned frame does not re-anchor the walk even though it is on screen, which contradicts the anchor rule above. It is the one live caller for propose's :protect frames.
  6. Fold the plate frames into the mark set, which is the nesting rule and is one line once both sides exist.
  7. The shared UI component, then its two mountings.

What is revertible, and where it is

Commit 02069e8 contains one part that nothing below asks for: extrema detection. Specifically segments, turns, the :extrema branch of propose, the multi-component refusal that exists only to guard it, and three tests — a-one-frame-closure-survives-only-because-it-is-protected, extrema-are-turns-worth-more-than-the-tolerance and extrema-are-refused-on-a-multi-component-signal.

It has no caller because plate selections take no extrema by design — displacement is the whole story for a head — and the performance side reads a stored cut instead of finding extrema in a signal. It is correct, tested and speculative.

Revert it if the shape above holds. Keep it if either of these turns out to be wanted: a group whose extreme is not a closure and therefore has no [:vis] cut to read (a mouth at its widest, a head at the top of a nod), or a take whose mouth never shuts far enough to cross aperture-cut, where a peak-relative threshold marks nothing and an extremum would still find the most closed frame. Neither is asked for today. The DP in The upgrade path gets extrema for free regardless, so reverting costs nothing that cannot be had again more cheaply.

Traps

  • Do not let Suggest write :keep. The proposal and the hand are different layers; collapsing them is the bug this whole shape exists to avoid, and it will look like it works right up until somebody moves the tolerance slider.
  • Do not snap a grid slot forward. Every hold and every pick is "at or before". Showing a closure before the mouth shut is a lead, which is a different control for a different reason.
  • Do not quantise a plate selection onto the output grid. A kept frame is a native frame and lands where it lands. time.md is explicit that an event between output frames appears on the next one; forcing kept frames onto the grid would re-create the problem the selection exists to solve. The snap is the opposite operation and is on the other side of the fence: it moves the grid's pick, never the kept frame.
  • Do not read the aperture pair off a subsampled ring. Positions 5 and 15 of LIPS-INNER survive only at verts divisible by four.
  • Do not thin the plate selection with the mouth's. Two scopes, two owners, and the nesting runs one way only: plate frames preserve performance frames, never the reverse.
  • Do not give the performance side a tolerance. The grid has already chosen the sparseness. A second knob there would be a control with nothing to control.
  • A skipped frame, a hidden feature and an absent measurement remain three different facts. A selection says nothing about visibility and nothing about whether a face was found. Note that the preserve marks are derived from a visibility cut, which makes this easy to blur: the cut says the interior is hidden, the mark says the frame is worth landing on, and one is not the other.

Not in scope

  • Automatic grouping of related parts beyond the existing :pose-group. Related parts must share one selection — a mouth outline, its interior, the teeth and the generated visibility reading different frames is the bug that grouping prevents — and :pose-group is where the grouping already lives.
  • A per-instance request for a different tolerance than the symbol's. The resolver threads no such option today and should not grow one until something needs it.
  • Variable frame rate. time.md assumes constant fps and so does this; a VFR source needs presentation timestamps before any of this means anything.