Eyes: lids, blinking, line of sight
Three parts per eye, stacked the way the mouth is - dark lash ring, sclera inside it, iris inside that, square pupil in the iris. A blink then costs nothing: when the lid shuts the traced ring goes flat and the lash line collapses to a lens, which is a closed eye, drawn correctly, for free. Lids are a FEATURE, rotoscoped like the mouth: head-local, a key on every frame, the same contour avg knob. The iris is a PRIMITIVE - a disc at a quantised position - and that is where the stylisation lives. Line of sight. Gaze is the iris centre relative to the midpoint of the eye's two corners, in units of corner distance. Both corners are in RIGID, so the origin and the scale are immune to the performance being measured; against the lid ring's centroid instead, every blink would drag the origin down and fake a glance at the floor on exactly the frames where the eye is most visible. Both eyes share one gaze - at this size the difference between the two measurements is noise, not vergence, and independent per-eye noise reads as wall-eyed immediately. Openness stays per-eye so a wink survives. Gaze is then quantised to a pixel grid with a dwell, which is not a stylisation imposed on the truth: real eyes move in saccades, and the smooth drift left in the measurement is tracker noise plus head-compensation error. Snapping to a grid removes the noise and recovers the saccade in one operation. The iris is placed in the frame of the already-smoothed, already-subsampled lid ring - slots 0 and 8 of a 16-slot ring are the corners, and subsampling to any even budget keeps them at 0 and n/2 - so it cannot drift relative to its own eye. Size is authored from the take mean, never remeasured per frame: a radius that breathes by a fraction of a pixel flickers a pixel on and off around the whole silhouette. iris anchor toggles steady/free/locked, because how much the eye wanders turns out to be an aesthetic choice and not only a correctness one. Blinking gets hysteresis and a dwell like the teeth, plus one knob they do not have: blink hold. A blink is one frame at 12fps and a single frame of closed eye reads as a dropped frame, so once the eye shuts it stays shut long enough to be legible. Detection accuracy is not the problem; legibility is. The pupil is a square because at three pixels a circle is a plus sign with the corners gnawed off, and it changes shape as it moves. Drawn from a rounded centre shared with the iris so it is exactly its nominal size on every frame. Iris/pupil clip by colour key against the indexed buffer, the way Animator Pro would: the lid crops the iris at extreme gaze for free, so nothing has to clamp the gaze, which would flatten the performance at the extremes that carry it. Which iris block belongs to which eye is RESOLVED from geometry, not declared. A swap looks almost right - each eye still has a disc roughly where it belongs - so it survives an eyeball and then reads as a subtly wall-eyed character forever. Voted across every frame; the test feeds a deliberately swapped track. Also: exposure. Aesthetic sparseness was set by the extraction rate, which made the timing a property of a directory of PNGs - auditioning 12 against 24 meant re-ripping and re-detecting the whole clip. It is now a render-time grid, on 1s/2s/3s/4s, so the dense track keeps everything and the audio clock is untouched. The take format already carried an exposure field; it was never driven. Everything rides the same grid, because a head cutting on the odd frames while the mouth cuts on the even ones reads as two performances laid over each other. 41 -> 91 assertions. The load-bearing new ones: the iris pairing follows a swapped track, a blink does not fake a change of gaze, a stencilled disc cannot spill past its clip, a 3px pupil is 3x3 at every sub-pixel centre, and exposure never reads a pose from the future. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
parent
c2603102da
commit
e497e9ac4d
11 changed files with 1342 additions and 42 deletions
110
README.md
110
README.md
|
|
@ -43,6 +43,13 @@ playback speed — neither gives a deterministic per-frame pass.
|
|||
assuming, because a guessed fps desynchronises audio from picture — and sync is
|
||||
the one thing this view exists to show.
|
||||
|
||||
**Exposure** decides how often the picture gets a new drawing: rip at 24 and
|
||||
render `on 2s` for 12, `on 3s` for 8. The dense track and the audio are
|
||||
untouched, so it is a dropdown rather than a re-rip, and the export emits keys
|
||||
only on the grid instead of the same pose twice. Everything rides the same grid
|
||||
— mouth, eyes, teeth, plate — because a head cutting on the odd frames while the
|
||||
mouth cuts on the even ones reads as two performances laid over each other.
|
||||
|
||||
**Audio is the playback clock**: `frame = floor(audio.currentTime * fps)`. A slow
|
||||
render loop therefore drops frames instead of drifting, and ½x / ¼x work by
|
||||
setting `playbackRate` with the picture following for free.
|
||||
|
|
@ -62,6 +69,7 @@ is hand-drawn head plates, which this tool does not yet do.
|
|||
| Knob | What it does |
|
||||
| --- | --- |
|
||||
| vertices | Lip vertex budget. The reduction past what the footage supports *is* the style. |
|
||||
| exposure | How often the picture changes: on 1s, 2s, 3s, 4s. Rip dense, choose timing here. |
|
||||
| mouth lead ±f | Shifts the performance tracks earlier against the audio and the head. `[` `]`. |
|
||||
| contour avg ±f | Radius in frames. 0 off, 1 = ±1. Removes per-frame landmark jitter. |
|
||||
| anchor avg ±f | Radius on the four similarity parameters. Smooths the *transform*. |
|
||||
|
|
@ -99,6 +107,100 @@ kept pixels green, extracted contour amber. Tune against that, not the numbers.
|
|||
brightness and biased low rather than high. Not implemented: it is not visible in
|
||||
the test footage, which reads as a dark cavity with a bright upper-teeth band.
|
||||
|
||||
## Eyes
|
||||
|
||||
Three parts per eye, stacked the way the mouth is: a dark **lash ring**, the
|
||||
**sclera** inside it, and the **iris** inside that, with a square **pupil** in
|
||||
the iris. The dark ring outside a pale interior is what makes a flat shape read
|
||||
as an opening rather than a blob, and it is why a blink costs nothing — when the
|
||||
lid shuts, the traced ring goes flat and the lash line collapses to a lens,
|
||||
which is a closed eye, drawn correctly, for free.
|
||||
|
||||
The **lids are a feature**, rotoscoped like the mouth: head-local, a key on every
|
||||
frame, the same `contour avg` knob. They track the face, because the face is what
|
||||
they are attached to.
|
||||
|
||||
The **iris is a primitive** — a disc at a quantised position — and that is where
|
||||
the stylisation is.
|
||||
|
||||
### Line of sight
|
||||
|
||||
Gaze is the iris centre relative to the **midpoint of the eye's two corners**,
|
||||
in units of corner distance. Both corners are in `RIGID`, which is the point:
|
||||
the origin and the scale are built only from landmarks that do not move under
|
||||
performance. Measure against the lid ring's centroid instead and every blink
|
||||
drags that centroid down and fakes a glance at the floor, on exactly the frames
|
||||
where the eye is most conspicuous.
|
||||
|
||||
**Both eyes share one gaze.** At 320×200 an iris is a handful of pixels and its
|
||||
centre comes from five landmarks on an eye twenty pixels wide, so the difference
|
||||
between the two measurements is noise, not vergence — and independent per-eye
|
||||
noise reads as wall-eyed immediately, which is the most expensive artefact on a
|
||||
face. Openness stays per-eye, so a wink survives.
|
||||
|
||||
Then the gaze is **quantised to a pixel grid with a dwell**, which is not a
|
||||
stylisation imposed on the truth: real eyes move in saccades, holding a fixation
|
||||
and then jumping. The smooth drift left in the measurement is tracker noise plus
|
||||
head-compensation error, so snapping to a grid and requiring a dwell removes the
|
||||
noise and recovers the saccade in one operation. The readout reports how many
|
||||
distinct cells the iris ever occupies — three or four is a character who looks at
|
||||
things, forty is an unquantised iris sliding around.
|
||||
|
||||
The iris is drawn at the socket read back off the **already-smoothed, already-
|
||||
subsampled lid ring** — slots 0 and 8 of a 16-slot ring are the corners, and
|
||||
subsampling to any even budget keeps them at output indices 0 and `n/2`. So the
|
||||
iris is placed in the frame of the exact polygon it sits inside and cannot drift
|
||||
relative to its own eye. Size is authored from the take's mean eye width, not
|
||||
remeasured per frame: a radius that breathes by a fraction of a pixel flickers a
|
||||
pixel on and off around the whole silhouette.
|
||||
|
||||
The iris is **stencilled to the sclera** and the pupil to the iris — the indexed
|
||||
buffer is its own clip mask, the way Animator Pro would do it. So the lid crops
|
||||
the iris at extreme gaze automatically, and nothing needs to clamp the gaze,
|
||||
which would flatten the performance at exactly the extremes that carry it.
|
||||
|
||||
### Blinking
|
||||
|
||||
Openness is the lid gap over the corner distance — normalised, so one threshold
|
||||
carries across takes and faces. It gets hysteresis and a dwell like the teeth,
|
||||
plus one knob the teeth do not have: **blink hold**. A blink is 100–150ms, which
|
||||
is one frame at 12fps, and a single frame of closed eye reads as a dropped frame
|
||||
rather than as a blink. Animators draw a blink over two or three drawings for
|
||||
that reason, so once the eye shuts it stays shut for `hold` frames.
|
||||
|
||||
### The pupil is a square
|
||||
|
||||
At this size a pupil is three pixels across, and a circle of radius 1.5 is not a
|
||||
circle — it is a plus sign with the corners gnawed off, and it changes shape as
|
||||
it moves. A square that size is a deliberate mark that stays the same mark
|
||||
wherever it lands. It is drawn from a rounded centre shared with the iris, so it
|
||||
is exactly its nominal size on every frame instead of spilling to the next pixel
|
||||
on some and not others.
|
||||
|
||||
### Which iris is which
|
||||
|
||||
The refined mesh appends ten iris points, five per eye, and MediaPipe's own
|
||||
left/right naming is viewer-relative in some places and subject-relative in
|
||||
others. Getting it backwards swaps the irises, which looks *almost* right — each
|
||||
eye still has a disc roughly where it belongs — so it survives an eyeball and
|
||||
then reads as a subtly wall-eyed character forever. The pairing is therefore
|
||||
**resolved from the geometry**, by voting each block's distance to each eye's
|
||||
corner midpoint across every frame, and the selftest feeds it a track built the
|
||||
other way round to prove it actually looks.
|
||||
|
||||
| Knob | What it does |
|
||||
| --- | --- |
|
||||
| eye vertices | Lid ring vertex budget, off a 16-slot ring. |
|
||||
| lash line | How far the dark ring sits outside the lid, in pixels. |
|
||||
| blink cut | Openness below which the eye is shut. Normalised by corner distance. |
|
||||
| blink hold | Minimum frames a blink stays on screen. A one-frame blink is a dropout. |
|
||||
| blink dwell | Frames a change must persist. Usually 0 — unlike the teeth, a real blink *is* one frame. |
|
||||
| gaze gain | Exaggerates or damps the throw. Measured excursion is small; a character usually wants more. |
|
||||
| gaze step | The pixel grid the iris snaps to. 0 = off, and then dwell does nothing either. |
|
||||
| gaze dwell | How long a new cell must hold before it takes. Together with step, this is what makes saccades. |
|
||||
| iris size | Diameter as a percentage of eye width. |
|
||||
| pupil | Square pupil in whole pixels. 0 = off. |
|
||||
|
||||
## The plate is reference, not art
|
||||
|
||||
The plate layer has several representations because its job changes. Cycle with
|
||||
|
|
@ -131,6 +233,10 @@ error here. **Aesthetic** sparseness is set by the extraction rate — pick 12fp
|
|||
you have already chosen your timing. **Labour** sparseness is a human drawing
|
||||
each one, and it binds only on the plate.
|
||||
|
||||
Aesthetic sparseness is the **exposure** control, not the extraction rate —
|
||||
making it a render-time grid means auditioning 12 against 24 costs a dropdown
|
||||
instead of a re-rip and a full re-detection.
|
||||
|
||||
So the mouth keeps **every** frame: it is traced, and therefore free. In limited
|
||||
animation lip sync is routinely the densest element, on 1s, while heads hold on
|
||||
2s and 3s.
|
||||
|
|
@ -162,7 +268,7 @@ chromium --headless --virtual-time-budget=8000 --dump-dom \
|
|||
http://127.0.0.1:8777/selftest.html | grep -oE '(PASS|FAIL) [0-9/]+'
|
||||
```
|
||||
|
||||
Or open `selftest.html`. 41 assertions over the stages below detection, plus a
|
||||
Or open `selftest.html`. 88 assertions over the stages below detection, plus a
|
||||
wiring cross-check: every `el('id')` in `app.js` must exist in `index.html`. A
|
||||
knob wired in one but not the other throws during wiring, which aborts the rest
|
||||
of the module and leaves a blank page — a symptom that points nowhere near its
|
||||
|
|
@ -176,7 +282,7 @@ eyeball.
|
|||
|
||||
## Not done yet
|
||||
|
||||
Eyes and irises; hand-drawn head plates and per-plate mouth slots (the strip
|
||||
Brows; hand-drawn head plates and per-plate mouth slots (the strip
|
||||
decides *which frames need one*, but you cannot yet supply the drawing); real
|
||||
performer→character calibration (currently identity, fitting the face oval to the
|
||||
canvas); the override layer; anything on the Animator Pro side. The plate is a
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue