Eyes: lids, blinking, line of sight
Three parts per eye, stacked the way the mouth is - dark lash ring, sclera inside it, iris inside that, square pupil in the iris. A blink then costs nothing: when the lid shuts the traced ring goes flat and the lash line collapses to a lens, which is a closed eye, drawn correctly, for free. Lids are a FEATURE, rotoscoped like the mouth: head-local, a key on every frame, the same contour avg knob. The iris is a PRIMITIVE - a disc at a quantised position - and that is where the stylisation lives. Line of sight. Gaze is the iris centre relative to the midpoint of the eye's two corners, in units of corner distance. Both corners are in RIGID, so the origin and the scale are immune to the performance being measured; against the lid ring's centroid instead, every blink would drag the origin down and fake a glance at the floor on exactly the frames where the eye is most visible. Both eyes share one gaze - at this size the difference between the two measurements is noise, not vergence, and independent per-eye noise reads as wall-eyed immediately. Openness stays per-eye so a wink survives. Gaze is then quantised to a pixel grid with a dwell, which is not a stylisation imposed on the truth: real eyes move in saccades, and the smooth drift left in the measurement is tracker noise plus head-compensation error. Snapping to a grid removes the noise and recovers the saccade in one operation. The iris is placed in the frame of the already-smoothed, already-subsampled lid ring - slots 0 and 8 of a 16-slot ring are the corners, and subsampling to any even budget keeps them at 0 and n/2 - so it cannot drift relative to its own eye. Size is authored from the take mean, never remeasured per frame: a radius that breathes by a fraction of a pixel flickers a pixel on and off around the whole silhouette. iris anchor toggles steady/free/locked, because how much the eye wanders turns out to be an aesthetic choice and not only a correctness one. Blinking gets hysteresis and a dwell like the teeth, plus one knob they do not have: blink hold. A blink is one frame at 12fps and a single frame of closed eye reads as a dropped frame, so once the eye shuts it stays shut long enough to be legible. Detection accuracy is not the problem; legibility is. The pupil is a square because at three pixels a circle is a plus sign with the corners gnawed off, and it changes shape as it moves. Drawn from a rounded centre shared with the iris so it is exactly its nominal size on every frame. Iris/pupil clip by colour key against the indexed buffer, the way Animator Pro would: the lid crops the iris at extreme gaze for free, so nothing has to clamp the gaze, which would flatten the performance at the extremes that carry it. Which iris block belongs to which eye is RESOLVED from geometry, not declared. A swap looks almost right - each eye still has a disc roughly where it belongs - so it survives an eyeball and then reads as a subtly wall-eyed character forever. Voted across every frame; the test feeds a deliberately swapped track. Also: exposure. Aesthetic sparseness was set by the extraction rate, which made the timing a property of a directory of PNGs - auditioning 12 against 24 meant re-ripping and re-detecting the whole clip. It is now a render-time grid, on 1s/2s/3s/4s, so the dense track keeps everything and the audio clock is untouched. The take format already carried an exposure field; it was never driven. Everything rides the same grid, because a head cutting on the odd frames while the mouth cuts on the even ones reads as two performances laid over each other. 41 -> 91 assertions. The load-bearing new ones: the iris pairing follows a swapped track, a blink does not fake a change of gaze, a stencilled disc cannot spill past its clip, a 3px pupil is 3x3 at every sub-pixel centre, and exposure never reads a pose from the future. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
parent
c2603102da
commit
e497e9ac4d
11 changed files with 1342 additions and 42 deletions
|
|
@ -70,9 +70,19 @@ avoidance.
|
|||
|
||||
## Two kinds of sparseness
|
||||
|
||||
Conflating these was the original design error. **Aesthetic** sparseness is set
|
||||
by the extraction rate: pick 12fps and the timing is already chosen. **Labour**
|
||||
sparseness is a human drawing each one, and it binds only on the plate.
|
||||
Conflating these was the original design error. **Aesthetic** sparseness is the
|
||||
rate the picture changes at. **Labour** sparseness is a human drawing each one,
|
||||
and it binds only on the plate.
|
||||
|
||||
Aesthetic sparseness used to be set by the extraction rate — rip at 12 and the
|
||||
timing is chosen. That was wrong in a small way: it makes the timing a property
|
||||
of a directory of PNGs, so auditioning 12 against 24 means re-ripping the clip
|
||||
and re-running detection over the whole of it, and the decision you most want to
|
||||
play with is the one that costs the most to change. Rip at the camera's rate and
|
||||
quantise at render time instead — an **exposure** grid, on 1s, 2s, 3s — so the
|
||||
dense track keeps everything, the audio clock is untouched, and the timing is a
|
||||
dropdown rather than a re-rip. The take format already carried an `exposure`
|
||||
field for this; it was simply never driven.
|
||||
|
||||
So the mouth keeps **every** frame — it is traced, and therefore free. In limited
|
||||
animation lip sync is routinely the densest element, on 1s, while heads hold on
|
||||
|
|
@ -136,6 +146,83 @@ Two escapes, both used:
|
|||
temporal smoothing is well defined, and the star-shaped result suits flat
|
||||
colour.
|
||||
|
||||
## Every part is measured in the frame of the thing it is attached to
|
||||
|
||||
The mouth is expressed against the head. The iris is expressed against its own
|
||||
eye — specifically against the midpoint of that eye's two corners, in units of
|
||||
corner distance. Both corners are rigid landmarks, so the origin and the scale
|
||||
of the measurement are immune to the performance being measured. Against the lid
|
||||
ring's centroid instead, every blink would drag the origin down and fake a glance
|
||||
at the floor on exactly the frames where the eye is most visible.
|
||||
|
||||
The rule generalises: **measure a feature in a frame built only from landmarks
|
||||
that do not move with it.** It is the same argument as "rigid landmarks only" for
|
||||
the anchor fit, one level down.
|
||||
|
||||
There is a tempting over-application. An eye can be pinned into a fixed socket
|
||||
fitted to its corners' mean over the shot, which removes the residual wobble a
|
||||
2D similarity cannot — and it is wrong. That residual is real motion of the eye
|
||||
relative to the head, it is still there in the footage, and removing it leaves
|
||||
the drawn eyes hanging still over a registered photo whose eyes are moving. A
|
||||
part must track the face in the same space the underlay is drawn in. The wobble
|
||||
is a job for the bounded contour average below, not for a second anchor.
|
||||
|
||||
Placement follows from the same idea. The iris is drawn in the frame of the
|
||||
already-smoothed, already-subsampled lid ring, read off the ring's own corner
|
||||
vertices, so it cannot drift relative to the eye it sits inside and it inherits
|
||||
the contour average for free. Size, by contrast, is authored from the take's
|
||||
mean, never remeasured per frame: a radius that breathes by a fraction of a pixel
|
||||
flickers a pixel on and off around the whole silhouette.
|
||||
|
||||
## The indexed buffer is its own stencil
|
||||
|
||||
Parts that nest — iris inside sclera, pupil inside iris — clip by colour key:
|
||||
paint only where the buffer already holds the parent's index. This is how
|
||||
Animator Pro would do it, it costs one comparison per pixel, and it composes
|
||||
transitively, so a blink takes the right bite out of the pupil without anything
|
||||
computing where.
|
||||
|
||||
It also removes a temptation. Without a stencil the gaze has to be clamped to
|
||||
keep the iris inside the lid, and a clamp flattens the performance at exactly the
|
||||
extremes that carry it.
|
||||
|
||||
## Quantisation can be the truthful choice
|
||||
|
||||
Gaze snapped to a pixel grid with a dwell is the "Primitive — quantised" row of
|
||||
the part table, and it looks like a stylisation imposed on a continuous
|
||||
measurement. It is not. Real eyes move in saccades: hold a fixation, jump, hold.
|
||||
The smooth drift left in the measured signal is tracker noise plus
|
||||
head-compensation error. Snapping to a grid and requiring a dwell removes the
|
||||
noise and recovers the saccade in the same operation — the rare case where the
|
||||
aesthetic rule and the physiology agree.
|
||||
|
||||
The count of distinct cells the iris ever occupies is the number the knobs exist
|
||||
to control. Three or four is a character who looks at things; forty is an
|
||||
unquantised iris sliding around.
|
||||
|
||||
## Some thresholds need a minimum duration, not just a dwell
|
||||
|
||||
A dwell delays a change until it has persisted, which is the right guard against
|
||||
chatter and is what the teeth use. A blink needs the opposite guard as well. It
|
||||
lasts 100–150ms — one frame at 12fps — and a single frame of closed eye reads as
|
||||
a dropped frame rather than as a blink. Animators draw a blink over two or three
|
||||
drawings for that reason, so once the eye shuts it must stay shut for a minimum
|
||||
number of frames. Detection accuracy is not the problem; legibility is.
|
||||
|
||||
## Resolve correspondences from data when a wrong guess is survivable
|
||||
|
||||
The refined mesh appends ten iris points, five per eye, and the upstream
|
||||
left/right naming is viewer-relative in some documentation and subject-relative
|
||||
in others. Swapping them looks *almost* right — each eye still has a disc roughly
|
||||
where it belongs — so the error survives inspection and then reads as a subtly
|
||||
wall-eyed character for the life of the project.
|
||||
|
||||
A hardcoded table is the wrong shape for a fact like that. Voting each block's
|
||||
distance to each eye's corner midpoint across every frame settles it from the
|
||||
geometry, cannot be got wrong, and keeps working if the model is renumbered. The
|
||||
test feeds it a track built the other way round, because a resolver checked only
|
||||
against the convention it was written for is checking nothing.
|
||||
|
||||
## The bounded smoothing exception
|
||||
|
||||
*Smooth the transform, never the contour* held while keys were sparse: sampling
|
||||
|
|
@ -188,6 +275,7 @@ Current modules:
|
|||
| Module | Role |
|
||||
| --- | --- |
|
||||
| `landmarks.js` | Index tables. Ring arrays are ordered traversals: slot position *is* vertex identity. |
|
||||
| `pipeline.js` | …also eye openness, gaze, blink resolution and the iris pairing vote. |
|
||||
| `mathutil.js` | Similarity fit, Procrustes mean, temporal smoothing. |
|
||||
| `pipeline.js` | Stabilise → subsample → key-select → frame-removal. |
|
||||
| `interior.js` | Teeth from image content: Otsu, morphology, components, radial contour. |
|
||||
|
|
@ -202,7 +290,7 @@ Current modules:
|
|||
- **A paint surface.** The plates have nowhere to be drawn. This is the largest
|
||||
gap between "tool" and "suite": a pixel paint canvas with onion skin, palette
|
||||
constraint, and the registered underlay behind it.
|
||||
- Eyes, irises, brows as parts.
|
||||
- Brows as parts, and a tongue.
|
||||
- Plate libraries with per-plate mouth slots.
|
||||
- Real performer→character calibration (currently identity).
|
||||
- The override layer.
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue