Eyes: lids, blinking, line of sight

Three parts per eye, stacked the way the mouth is - dark lash ring, sclera
inside it, iris inside that, square pupil in the iris. A blink then costs
nothing: when the lid shuts the traced ring goes flat and the lash line
collapses to a lens, which is a closed eye, drawn correctly, for free.

Lids are a FEATURE, rotoscoped like the mouth: head-local, a key on every
frame, the same contour avg knob. The iris is a PRIMITIVE - a disc at a
quantised position - and that is where the stylisation lives.

Line of sight. Gaze is the iris centre relative to the midpoint of the eye's
two corners, in units of corner distance. Both corners are in RIGID, so the
origin and the scale are immune to the performance being measured; against the
lid ring's centroid instead, every blink would drag the origin down and fake a
glance at the floor on exactly the frames where the eye is most visible. Both
eyes share one gaze - at this size the difference between the two measurements
is noise, not vergence, and independent per-eye noise reads as wall-eyed
immediately. Openness stays per-eye so a wink survives.

Gaze is then quantised to a pixel grid with a dwell, which is not a
stylisation imposed on the truth: real eyes move in saccades, and the smooth
drift left in the measurement is tracker noise plus head-compensation error.
Snapping to a grid removes the noise and recovers the saccade in one operation.

The iris is placed in the frame of the already-smoothed, already-subsampled lid
ring - slots 0 and 8 of a 16-slot ring are the corners, and subsampling to any
even budget keeps them at 0 and n/2 - so it cannot drift relative to its own
eye. Size is authored from the take mean, never remeasured per frame: a radius
that breathes by a fraction of a pixel flickers a pixel on and off around the
whole silhouette. iris anchor toggles steady/free/locked, because how much the
eye wanders turns out to be an aesthetic choice and not only a correctness one.

Blinking gets hysteresis and a dwell like the teeth, plus one knob they do not
have: blink hold. A blink is one frame at 12fps and a single frame of closed
eye reads as a dropped frame, so once the eye shuts it stays shut long enough
to be legible. Detection accuracy is not the problem; legibility is.

The pupil is a square because at three pixels a circle is a plus sign with the
corners gnawed off, and it changes shape as it moves. Drawn from a rounded
centre shared with the iris so it is exactly its nominal size on every frame.

Iris/pupil clip by colour key against the indexed buffer, the way Animator Pro
would: the lid crops the iris at extreme gaze for free, so nothing has to clamp
the gaze, which would flatten the performance at the extremes that carry it.

Which iris block belongs to which eye is RESOLVED from geometry, not declared.
A swap looks almost right - each eye still has a disc roughly where it belongs
- so it survives an eyeball and then reads as a subtly wall-eyed character
forever. Voted across every frame; the test feeds a deliberately swapped track.

Also: exposure. Aesthetic sparseness was set by the extraction rate, which made
the timing a property of a directory of PNGs - auditioning 12 against 24 meant
re-ripping and re-detecting the whole clip. It is now a render-time grid, on
1s/2s/3s/4s, so the dense track keeps everything and the audio clock is
untouched. The take format already carried an exposure field; it was never
driven. Everything rides the same grid, because a head cutting on the odd
frames while the mouth cuts on the even ones reads as two performances laid
over each other.

41 -> 91 assertions. The load-bearing new ones: the iris pairing follows a
swapped track, a blink does not fake a change of gaze, a stencilled disc cannot
spill past its clip, a 3px pupil is 3x3 at every sub-pixel centre, and exposure
never reads a pose from the future.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
Your Name 2026-09-24 18:06:04 -04:00
parent c2603102da
commit e497e9ac4d
11 changed files with 1342 additions and 42 deletions

110
README.md
View file

@ -43,6 +43,13 @@ playback speed — neither gives a deterministic per-frame pass.
assuming, because a guessed fps desynchronises audio from picture — and sync is assuming, because a guessed fps desynchronises audio from picture — and sync is
the one thing this view exists to show. the one thing this view exists to show.
**Exposure** decides how often the picture gets a new drawing: rip at 24 and
render `on 2s` for 12, `on 3s` for 8. The dense track and the audio are
untouched, so it is a dropdown rather than a re-rip, and the export emits keys
only on the grid instead of the same pose twice. Everything rides the same grid
— mouth, eyes, teeth, plate — because a head cutting on the odd frames while the
mouth cuts on the even ones reads as two performances laid over each other.
**Audio is the playback clock**: `frame = floor(audio.currentTime * fps)`. A slow **Audio is the playback clock**: `frame = floor(audio.currentTime * fps)`. A slow
render loop therefore drops frames instead of drifting, and ½x / ¼x work by render loop therefore drops frames instead of drifting, and ½x / ¼x work by
setting `playbackRate` with the picture following for free. setting `playbackRate` with the picture following for free.
@ -62,6 +69,7 @@ is hand-drawn head plates, which this tool does not yet do.
| Knob | What it does | | Knob | What it does |
| --- | --- | | --- | --- |
| vertices | Lip vertex budget. The reduction past what the footage supports *is* the style. | | vertices | Lip vertex budget. The reduction past what the footage supports *is* the style. |
| exposure | How often the picture changes: on 1s, 2s, 3s, 4s. Rip dense, choose timing here. |
| mouth lead ±f | Shifts the performance tracks earlier against the audio and the head. `[` `]`. | | mouth lead ±f | Shifts the performance tracks earlier against the audio and the head. `[` `]`. |
| contour avg ±f | Radius in frames. 0 off, 1 = ±1. Removes per-frame landmark jitter. | | contour avg ±f | Radius in frames. 0 off, 1 = ±1. Removes per-frame landmark jitter. |
| anchor avg ±f | Radius on the four similarity parameters. Smooths the *transform*. | | anchor avg ±f | Radius on the four similarity parameters. Smooths the *transform*. |
@ -99,6 +107,100 @@ kept pixels green, extracted contour amber. Tune against that, not the numbers.
brightness and biased low rather than high. Not implemented: it is not visible in brightness and biased low rather than high. Not implemented: it is not visible in
the test footage, which reads as a dark cavity with a bright upper-teeth band. the test footage, which reads as a dark cavity with a bright upper-teeth band.
## Eyes
Three parts per eye, stacked the way the mouth is: a dark **lash ring**, the
**sclera** inside it, and the **iris** inside that, with a square **pupil** in
the iris. The dark ring outside a pale interior is what makes a flat shape read
as an opening rather than a blob, and it is why a blink costs nothing — when the
lid shuts, the traced ring goes flat and the lash line collapses to a lens,
which is a closed eye, drawn correctly, for free.
The **lids are a feature**, rotoscoped like the mouth: head-local, a key on every
frame, the same `contour avg` knob. They track the face, because the face is what
they are attached to.
The **iris is a primitive** — a disc at a quantised position — and that is where
the stylisation is.
### Line of sight
Gaze is the iris centre relative to the **midpoint of the eye's two corners**,
in units of corner distance. Both corners are in `RIGID`, which is the point:
the origin and the scale are built only from landmarks that do not move under
performance. Measure against the lid ring's centroid instead and every blink
drags that centroid down and fakes a glance at the floor, on exactly the frames
where the eye is most conspicuous.
**Both eyes share one gaze.** At 320×200 an iris is a handful of pixels and its
centre comes from five landmarks on an eye twenty pixels wide, so the difference
between the two measurements is noise, not vergence — and independent per-eye
noise reads as wall-eyed immediately, which is the most expensive artefact on a
face. Openness stays per-eye, so a wink survives.
Then the gaze is **quantised to a pixel grid with a dwell**, which is not a
stylisation imposed on the truth: real eyes move in saccades, holding a fixation
and then jumping. The smooth drift left in the measurement is tracker noise plus
head-compensation error, so snapping to a grid and requiring a dwell removes the
noise and recovers the saccade in one operation. The readout reports how many
distinct cells the iris ever occupies — three or four is a character who looks at
things, forty is an unquantised iris sliding around.
The iris is drawn at the socket read back off the **already-smoothed, already-
subsampled lid ring** — slots 0 and 8 of a 16-slot ring are the corners, and
subsampling to any even budget keeps them at output indices 0 and `n/2`. So the
iris is placed in the frame of the exact polygon it sits inside and cannot drift
relative to its own eye. Size is authored from the take's mean eye width, not
remeasured per frame: a radius that breathes by a fraction of a pixel flickers a
pixel on and off around the whole silhouette.
The iris is **stencilled to the sclera** and the pupil to the iris — the indexed
buffer is its own clip mask, the way Animator Pro would do it. So the lid crops
the iris at extreme gaze automatically, and nothing needs to clamp the gaze,
which would flatten the performance at exactly the extremes that carry it.
### Blinking
Openness is the lid gap over the corner distance — normalised, so one threshold
carries across takes and faces. It gets hysteresis and a dwell like the teeth,
plus one knob the teeth do not have: **blink hold**. A blink is 100–150ms, which
is one frame at 12fps, and a single frame of closed eye reads as a dropped frame
rather than as a blink. Animators draw a blink over two or three drawings for
that reason, so once the eye shuts it stays shut for `hold` frames.
### The pupil is a square
At this size a pupil is three pixels across, and a circle of radius 1.5 is not a
circle — it is a plus sign with the corners gnawed off, and it changes shape as
it moves. A square that size is a deliberate mark that stays the same mark
wherever it lands. It is drawn from a rounded centre shared with the iris, so it
is exactly its nominal size on every frame instead of spilling to the next pixel
on some and not others.
### Which iris is which
The refined mesh appends ten iris points, five per eye, and MediaPipe's own
left/right naming is viewer-relative in some places and subject-relative in
others. Getting it backwards swaps the irises, which looks *almost* right — each
eye still has a disc roughly where it belongs — so it survives an eyeball and
then reads as a subtly wall-eyed character forever. The pairing is therefore
**resolved from the geometry**, by voting each block's distance to each eye's
corner midpoint across every frame, and the selftest feeds it a track built the
other way round to prove it actually looks.
| Knob | What it does |
| --- | --- |
| eye vertices | Lid ring vertex budget, off a 16-slot ring. |
| lash line | How far the dark ring sits outside the lid, in pixels. |
| blink cut | Openness below which the eye is shut. Normalised by corner distance. |
| blink hold | Minimum frames a blink stays on screen. A one-frame blink is a dropout. |
| blink dwell | Frames a change must persist. Usually 0 — unlike the teeth, a real blink *is* one frame. |
| gaze gain | Exaggerates or damps the throw. Measured excursion is small; a character usually wants more. |
| gaze step | The pixel grid the iris snaps to. 0 = off, and then dwell does nothing either. |
| gaze dwell | How long a new cell must hold before it takes. Together with step, this is what makes saccades. |
| iris size | Diameter as a percentage of eye width. |
| pupil | Square pupil in whole pixels. 0 = off. |
## The plate is reference, not art ## The plate is reference, not art
The plate layer has several representations because its job changes. Cycle with The plate layer has several representations because its job changes. Cycle with
@ -131,6 +233,10 @@ error here. **Aesthetic** sparseness is set by the extraction rate — pick 12fp
you have already chosen your timing. **Labour** sparseness is a human drawing you have already chosen your timing. **Labour** sparseness is a human drawing
each one, and it binds only on the plate. each one, and it binds only on the plate.
Aesthetic sparseness is the **exposure** control, not the extraction rate —
making it a render-time grid means auditioning 12 against 24 costs a dropdown
instead of a re-rip and a full re-detection.
So the mouth keeps **every** frame: it is traced, and therefore free. In limited So the mouth keeps **every** frame: it is traced, and therefore free. In limited
animation lip sync is routinely the densest element, on 1s, while heads hold on animation lip sync is routinely the densest element, on 1s, while heads hold on
2s and 3s. 2s and 3s.
@ -162,7 +268,7 @@ chromium --headless --virtual-time-budget=8000 --dump-dom \
http://127.0.0.1:8777/selftest.html | grep -oE '(PASS|FAIL) [0-9/]+' http://127.0.0.1:8777/selftest.html | grep -oE '(PASS|FAIL) [0-9/]+'
``` ```
Or open `selftest.html`. 41 assertions over the stages below detection, plus a Or open `selftest.html`. 88 assertions over the stages below detection, plus a
wiring cross-check: every `el('id')` in `app.js` must exist in `index.html`. A wiring cross-check: every `el('id')` in `app.js` must exist in `index.html`. A
knob wired in one but not the other throws during wiring, which aborts the rest knob wired in one but not the other throws during wiring, which aborts the rest
of the module and leaves a blank page — a symptom that points nowhere near its of the module and leaves a blank page — a symptom that points nowhere near its
@ -176,7 +282,7 @@ eyeball.
## Not done yet ## Not done yet
Eyes and irises; hand-drawn head plates and per-plate mouth slots (the strip Brows; hand-drawn head plates and per-plate mouth slots (the strip
decides *which frames need one*, but you cannot yet supply the drawing); real decides *which frames need one*, but you cannot yet supply the drawing); real
performer→character calibration (currently identity, fitting the face oval to the performer→character calibration (currently identity, fitting the face oval to the
canvas); the override layer; anything on the Animator Pro side. The plate is a canvas); the override layer; anything on the Animator Pro side. The plate is a

View file

@ -70,9 +70,19 @@ avoidance.
## Two kinds of sparseness ## Two kinds of sparseness
Conflating these was the original design error. **Aesthetic** sparseness is set Conflating these was the original design error. **Aesthetic** sparseness is the
by the extraction rate: pick 12fps and the timing is already chosen. **Labour** rate the picture changes at. **Labour** sparseness is a human drawing each one,
sparseness is a human drawing each one, and it binds only on the plate. and it binds only on the plate.
Aesthetic sparseness used to be set by the extraction rate — rip at 12 and the
timing is chosen. That was wrong in a small way: it makes the timing a property
of a directory of PNGs, so auditioning 12 against 24 means re-ripping the clip
and re-running detection over the whole of it, and the decision you most want to
play with is the one that costs the most to change. Rip at the camera's rate and
quantise at render time instead — an **exposure** grid, on 1s, 2s, 3s — so the
dense track keeps everything, the audio clock is untouched, and the timing is a
dropdown rather than a re-rip. The take format already carried an `exposure`
field for this; it was simply never driven.
So the mouth keeps **every** frame — it is traced, and therefore free. In limited So the mouth keeps **every** frame — it is traced, and therefore free. In limited
animation lip sync is routinely the densest element, on 1s, while heads hold on animation lip sync is routinely the densest element, on 1s, while heads hold on
@ -136,6 +146,83 @@ Two escapes, both used:
temporal smoothing is well defined, and the star-shaped result suits flat temporal smoothing is well defined, and the star-shaped result suits flat
colour. colour.
## Every part is measured in the frame of the thing it is attached to
The mouth is expressed against the head. The iris is expressed against its own
eye — specifically against the midpoint of that eye's two corners, in units of
corner distance. Both corners are rigid landmarks, so the origin and the scale
of the measurement are immune to the performance being measured. Against the lid
ring's centroid instead, every blink would drag the origin down and fake a glance
at the floor on exactly the frames where the eye is most visible.
The rule generalises: **measure a feature in a frame built only from landmarks
that do not move with it.** It is the same argument as "rigid landmarks only" for
the anchor fit, one level down.
There is a tempting over-application. An eye can be pinned into a fixed socket
fitted to its corners' mean over the shot, which removes the residual wobble a
2D similarity cannot — and it is wrong. That residual is real motion of the eye
relative to the head, it is still there in the footage, and removing it leaves
the drawn eyes hanging still over a registered photo whose eyes are moving. A
part must track the face in the same space the underlay is drawn in. The wobble
is a job for the bounded contour average below, not for a second anchor.
Placement follows from the same idea. The iris is drawn in the frame of the
already-smoothed, already-subsampled lid ring, read off the ring's own corner
vertices, so it cannot drift relative to the eye it sits inside and it inherits
the contour average for free. Size, by contrast, is authored from the take's
mean, never remeasured per frame: a radius that breathes by a fraction of a pixel
flickers a pixel on and off around the whole silhouette.
## The indexed buffer is its own stencil
Parts that nest — iris inside sclera, pupil inside iris — clip by colour key:
paint only where the buffer already holds the parent's index. This is how
Animator Pro would do it, it costs one comparison per pixel, and it composes
transitively, so a blink takes the right bite out of the pupil without anything
computing where.
It also removes a temptation. Without a stencil the gaze has to be clamped to
keep the iris inside the lid, and a clamp flattens the performance at exactly the
extremes that carry it.
## Quantisation can be the truthful choice
Gaze snapped to a pixel grid with a dwell is the "Primitive — quantised" row of
the part table, and it looks like a stylisation imposed on a continuous
measurement. It is not. Real eyes move in saccades: hold a fixation, jump, hold.
The smooth drift left in the measured signal is tracker noise plus
head-compensation error. Snapping to a grid and requiring a dwell removes the
noise and recovers the saccade in the same operation — the rare case where the
aesthetic rule and the physiology agree.
The count of distinct cells the iris ever occupies is the number the knobs exist
to control. Three or four is a character who looks at things; forty is an
unquantised iris sliding around.
## Some thresholds need a minimum duration, not just a dwell
A dwell delays a change until it has persisted, which is the right guard against
chatter and is what the teeth use. A blink needs the opposite guard as well. It
lasts 100–150ms — one frame at 12fps — and a single frame of closed eye reads as
a dropped frame rather than as a blink. Animators draw a blink over two or three
drawings for that reason, so once the eye shuts it must stay shut for a minimum
number of frames. Detection accuracy is not the problem; legibility is.
## Resolve correspondences from data when a wrong guess is survivable
The refined mesh appends ten iris points, five per eye, and the upstream
left/right naming is viewer-relative in some documentation and subject-relative
in others. Swapping them looks *almost* right — each eye still has a disc roughly
where it belongs — so the error survives inspection and then reads as a subtly
wall-eyed character for the life of the project.
A hardcoded table is the wrong shape for a fact like that. Voting each block's
distance to each eye's corner midpoint across every frame settles it from the
geometry, cannot be got wrong, and keeps working if the model is renumbered. The
test feeds it a track built the other way round, because a resolver checked only
against the convention it was written for is checking nothing.
## The bounded smoothing exception ## The bounded smoothing exception
*Smooth the transform, never the contour* held while keys were sparse: sampling *Smooth the transform, never the contour* held while keys were sparse: sampling
@ -188,6 +275,7 @@ Current modules:
| Module | Role | | Module | Role |
| --- | --- | | --- | --- |
| `landmarks.js` | Index tables. Ring arrays are ordered traversals: slot position *is* vertex identity. | | `landmarks.js` | Index tables. Ring arrays are ordered traversals: slot position *is* vertex identity. |
| `pipeline.js` | …also eye openness, gaze, blink resolution and the iris pairing vote. |
| `mathutil.js` | Similarity fit, Procrustes mean, temporal smoothing. | | `mathutil.js` | Similarity fit, Procrustes mean, temporal smoothing. |
| `pipeline.js` | Stabilise → subsample → key-select → frame-removal. | | `pipeline.js` | Stabilise → subsample → key-select → frame-removal. |
| `interior.js` | Teeth from image content: Otsu, morphology, components, radial contour. | | `interior.js` | Teeth from image content: Otsu, morphology, components, radial contour. |
@ -202,7 +290,7 @@ Current modules:
- **A paint surface.** The plates have nowhere to be drawn. This is the largest - **A paint surface.** The plates have nowhere to be drawn. This is the largest
gap between "tool" and "suite": a pixel paint canvas with onion skin, palette gap between "tool" and "suite": a pixel paint canvas with onion skin, palette
constraint, and the registered underlay behind it. constraint, and the registered underlay behind it.
- Eyes, irises, brows as parts. - Brows as parts, and a tongue.
- Plate libraries with per-plate mouth slots. - Plate libraries with per-plate mouth slots.
- Real performer→character calibration (currently identity). - Real performer→character calibration (currently identity).
- The override layer. - The override layer.

View file

@ -63,6 +63,21 @@
<input type="text" id="framedir" value="frames" size="7" title="frame directory"> <input type="text" id="framedir" value="frames" size="7" title="frame directory">
<button id="btn-frames">Load frames</button> <button id="btn-frames">Load frames</button>
<button id="btn-play">Play</button> <button id="btn-play">Play</button>
<select id="irisAnchor" title="what the iris hangs off">
<option value="steady" selected>iris: steady</option>
<option value="free">iris: free</option>
<option value="locked">iris: locked</option>
</select>
<select id="gazeOrigin" title="what counts as looking straight ahead">
<option value="median" selected>origin: median</option>
<option value="neutral">origin: neutral f</option>
</select>
<select id="exposure" title="exposure — how often the picture gets a new drawing">
<option value="1" selected>on 1s</option>
<option value="2">on 2s</option>
<option value="3">on 3s</option>
<option value="4">on 4s</option>
</select>
<select id="speed" title="playback speed"><option value="1">1x</option><option value="0.5">½x</option><option value="0.25">¼x</option></select> <select id="speed" title="playback speed"><option value="1">1x</option><option value="0.5">½x</option><option value="0.25">¼x</option></select>
<audio id="audio" controls hidden style="height:28px;vertical-align:middle"></audio> <audio id="audio" controls hidden style="height:28px;vertical-align:middle"></audio>
<button id="btn-keepall">Keep all</button> <button id="btn-keepall">Keep all</button>
@ -86,18 +101,22 @@
<div class="panel"> <div class="panel">
<h2>source + landmarks</h2> <h2>source + landmarks</h2>
<canvas id="cv-source"></canvas> <canvas id="cv-source"></canvas>
<div class="legend">outer lip <b style="color:#4ade80">—</b> · inner lip <b style="color:#f87171">—</b></div> <div class="legend">outer lip <b style="color:#4ade80">—</b> · inner lip <b style="color:#f87171">—</b> ·
lids <b style="color:#60a5fa">—</b> · iris <b style="color:#fbbf24">—</b></div>
</div> </div>
<div class="panel"> <div class="panel">
<h2>stabilised (head-local)</h2> <h2>stabilised (head-local)</h2>
<canvas id="cv-stab"></canvas> <canvas id="cv-stab"></canvas>
<div class="legend">should sit still except the mouth · grey = held plate outline<br>dark green ghost = unshifted mouth when lead ≠ 0</div> <div class="legend">should sit still except the mouth and eyes · grey = held plate outline<br>
dark green ghost = unshifted mouth when lead ≠ 0 · red lid ring = blink</div>
</div> </div>
<div class="panel"> <div class="panel">
<h2>flat render — 320×200 indexed</h2> <h2>flat render — 320×200 indexed</h2>
<canvas id="cv-render"></canvas> <canvas id="cv-render"></canvas>
<div class="legend" id="framelabel"></div> <div class="legend" id="framelabel"></div>
<div class="legend">audio drives the clock — dropped frames, never drift<br> <div class="legend">audio drives the clock — dropped frames, never drift<br>
<b>exposure</b> holds the picture on a grid: rip at 24, render on 2s for
12. The dense track and the audio are untouched, so it is reversible<br>
plate representation: <kbd>B</kbd> cycles · photo modes are <b>registered</b> plate representation: <kbd>B</kbd> cycles · photo modes are <b>registered</b>
into raster space, so tracing them lands on the mouth</div> into raster space, so tracing them lands on the mouth</div>
</div> </div>
@ -132,6 +151,16 @@
<label class="ctl"><span>teeth vertices</span><input type="range" id="teethVerts" min="5" max="20" value="10"><output id="teethVertsv"></output></label> <label class="ctl"><span>teeth vertices</span><input type="range" id="teethVerts" min="5" max="20" value="10"><output id="teethVertsv"></output></label>
<label class="ctl"><span>teeth avg ±f</span><input type="range" id="teethSmooth" min="0" max="4" value="1"><output id="teethSmoothv"></output></label> <label class="ctl"><span>teeth avg ±f</span><input type="range" id="teethSmooth" min="0" max="4" value="1"><output id="teethSmoothv"></output></label>
<label class="ctl"><span>teeth dwell</span><input type="range" id="teethDwell" min="0" max="6" value="1"><output id="teethDwellv"></output></label> <label class="ctl"><span>teeth dwell</span><input type="range" id="teethDwell" min="0" max="6" value="1"><output id="teethDwellv"></output></label>
<label class="ctl"><span>eye vertices</span><input type="range" id="eyeVerts" min="4" max="12" step="2" value="8"><output id="eyeVertsv"></output></label>
<label class="ctl"><span>lash line</span><input type="range" id="lashPx" min="0" max="3" value="1"><output id="lashPxv"></output></label>
<label class="ctl"><span>blink cut</span><input type="range" id="blinkCut" min="20" max="300" value="130"><output id="blinkCutv"></output></label>
<label class="ctl"><span>blink hold ±f</span><input type="range" id="blinkHold" min="1" max="5" value="2"><output id="blinkHoldv"></output></label>
<label class="ctl"><span>blink dwell</span><input type="range" id="blinkDwell" min="0" max="4" value="0"><output id="blinkDwellv"></output></label>
<label class="ctl"><span>gaze gain</span><input type="range" id="gazeGain" min="50" max="400" value="100"><output id="gazeGainv"></output></label>
<label class="ctl"><span>gaze step</span><input type="range" id="gazeStep" min="0" max="6" value="2"><output id="gazeStepv"></output></label>
<label class="ctl"><span>gaze dwell</span><input type="range" id="gazeDwell" min="0" max="6" value="2"><output id="gazeDwellv"></output></label>
<label class="ctl"><span>iris size</span><input type="range" id="irisSize" min="20" max="70" value="42"><output id="irisSizev"></output></label>
<label class="ctl"><span>pupil</span><input type="range" id="pupilPx" min="0" max="7" value="3"><output id="pupilPxv"></output></label>
<label class="ctl"><span>suggest tolerance</span><input type="range" id="tol" min="2" max="60" value="14"><output id="tolv"></output></label> <label class="ctl"><span>suggest tolerance</span><input type="range" id="tol" min="2" max="60" value="14"><output id="tolv"></output></label>
<div class="legend"> <div class="legend">
<b>mouth lead</b> shifts the performance tracks earlier (positive) against <b>mouth lead</b> shifts the performance tracks earlier (positive) against
@ -150,10 +179,28 @@
brightness; <b>prefer upper</b> biases component choice toward the top of brightness; <b>prefer upper</b> biases component choice toward the top of
the cavity, where teeth are and the tongue is not. the cavity, where teeth are and the tongue is not.
<b>dwell</b> is how many frames a presence change must persist.<br> <b>dwell</b> is how many frames a presence change must persist.<br>
<b>blink cut</b> is lid gap over corner distance — normalised, so one
value carries across takes. <b>blink hold</b> is the minimum length of a
blink: a real blink is one frame at 12fps and a single frame of closed
eye reads as a dropout, so it is extended to a beat.
<b>pupil</b> is a square, in whole pixels, 0 to turn it off: at this size
a circle of radius 1.5 is a plus sign with the corners gnawed off and it
changes shape as it moves, where a square stays the mark you drew.<br>
<b>gaze step</b> is the grid the iris snaps to, in raster pixels, and
<b>gaze dwell</b> is how long a new cell must hold — together they turn
drift into saccades. <b>gaze gain</b> exaggerates or damps the throw;
measured excursion is small and a character usually wants more of it.<br>
<b>suggest tolerance</b> only affects the Suggest button: max head movement <b>suggest tolerance</b> only affects the Suggest button: max head movement
allowed before a new drawing is required. allowed before a new drawing is required.
</div> </div>
</div> </div>
<div class="panel" style="flex:0 1 170px">
<h2>gaze field</h2>
<canvas id="cv-gaze"></canvas>
<div class="legend" id="eyeinfo" style="white-space:pre-line"></div>
<div class="legend">green = every cell the iris visits in the take ·
grey = raw · amber = where it is now, quantised</div>
</div>
<div class="panel" style="flex:0 1 190px"> <div class="panel" style="flex:0 1 190px">
<h2>teeth measurement</h2> <h2>teeth measurement</h2>
<div id="cv-teeth"></div> <div id="cv-teeth"></div>

428
js/app.js
View file

@ -1,10 +1,12 @@
import { FaceLandmarker, FilesetResolver } from 'https://cdn.jsdelivr.net/npm/@mediapipe/tasks-vision@1.0.1/vision_bundle.mjs'; import { FaceLandmarker, FilesetResolver } from 'https://cdn.jsdelivr.net/npm/@mediapipe/tasks-vision@1.0.1/vision_bundle.mjs';
import { LIPS_OUTER, LIPS_INNER, FACE_OVAL } from './landmarks.js'; import { LIPS_OUTER, LIPS_INNER, FACE_OVAL,
import { stabilize, toRasterRing, smoothContours, suggestPlateFrames, heldFrame, shiftIndex } from './pipeline.js'; EYE_R_RING, EYE_L_RING, IRIS_A, IRIS_B } from './landmarks.js';
import { stabilize, toRasterRing, smoothContours, suggestPlateFrames, heldFrame, shiftIndex,
exposeIndex, eyeSignals, gazeOrigin, quantizeGaze, resolveBlink } from './pipeline.js';
import { IndexedRaster } from './raster.js'; import { IndexedRaster } from './raster.js';
import { drawRegistered, posterizeInto } from './underlay.js'; import { drawRegistered, posterizeInto } from './underlay.js';
import { extractTeeth } from './interior.js'; import { extractTeeth } from './interior.js';
import { applySim } from './mathutil.js'; import { applySim, offsetRing } from './mathutil.js';
import { writeTake } from './take.js'; import { writeTake } from './take.js';
import { synthDense } from './synth.js'; import { synthDense } from './synth.js';
@ -16,8 +18,21 @@ const PALETTE = [
{ name: 'skin_dark', hex: '#7a4f3a' }, { name: 'skin_dark', hex: '#7a4f3a' },
{ name: 'mouth_dark', hex: '#24161a' }, { name: 'mouth_dark', hex: '#24161a' },
{ name: 'teeth', hex: '#d9cfc2' }, { name: 'teeth', hex: '#d9cfc2' },
// Sclera is not white, and that is authored, not measured. A true white at
// 320x200 next to a warm skin ramp reads as a hole punched in the face; the
// eye sits in a socket, in shadow, so it is a dimmer and cooler tone than the
// teeth, which catch the light. The iris is one dark tone: at this size an
// iris is about five pixels across and a pupil inside it would be one, so the
// iris IS the pupil. Resolving it further would be drawing detail the format
// cannot hold.
{ name: 'eye_white', hex: '#c9c3b4' },
// Three tones for the eye - sclera, iris, pupil - which is the "two or three
// tones per part" budget, spent where it buys the most: an eye with no tonal
// step inside it reads as a hole.
{ name: 'iris', hex: '#4a5468' },
{ name: 'pupil', hex: '#171a22' },
]; ];
const IDX = { bg: 0, base: 1, dark: 2, mouth: 3, teeth: 4 }; const IDX = { bg: 0, base: 1, dark: 2, mouth: 3, teeth: 4, white: 5, iris: 6, pupil: 7 };
const state = { const state = {
dense: null, images: [], stab: null, xform: null, dense: null, images: [], stab: null, xform: null,
@ -27,8 +42,11 @@ const state = {
fps: 12, audio: null, // fps comes from manifest.json, never guessed fps: 12, audio: null, // fps comes from manifest.json, never guessed
aspect: 1, // imgW/imgH; converts MediaPipe's anisotropic space aspect: 1, // imgW/imgH; converts MediaPipe's anisotropic space
lead: 0, // performance-track offset in frames lead: 0, // performance-track offset in frames
exposure: 1, // 1 = on 1s, 2 = on 2s. Picture holds; audio does not.
interior: null, // per-frame teeth measurement from image content interior: null, // per-frame teeth measurement from image content
teeth: null, // resolved per-frame {show, t} after knobs teeth: null, // resolved per-frame {show, t} after knobs
eyes: null, // resolved per-frame lid rings, shut flags, iris discs
eyeSig: null, // raw eye measurement, kept for the gaze readout
}; };
const el = (id) => { const el = (id) => {
@ -53,6 +71,19 @@ const opts = () => ({
contourSmooth: +el('contourSmooth').value, contourSmooth: +el('contourSmooth').value,
apertureThresh: +el('apertureThresh').value / 1000, apertureThresh: +el('apertureThresh').value / 1000,
tol: +el('tol').value / 1000, tol: +el('tol').value / 1000,
exposure: +el('exposure').value,
irisAnchor: el('irisAnchor').value,
gazeOrigin: el('gazeOrigin').value,
eyeVerts: +el('eyeVerts').value,
lashPx: +el('lashPx').value,
irisSize: +el('irisSize').value / 100,
gazeGain: +el('gazeGain').value / 100,
gazeStep: +el('gazeStep').value, // whole raster pixels
pupilPx: +el('pupilPx').value,
gazeDwell: +el('gazeDwell').value,
blinkCut: +el('blinkCut').value / 1000,
blinkHold: +el('blinkHold').value,
blinkDwell: +el('blinkDwell').value,
}); });
function status(msg, kind = '') { function status(msg, kind = '') {
@ -167,6 +198,7 @@ function rebuild(resetKeep) {
if (!state.dense) return; if (!state.dense) return;
const o = opts(); const o = opts();
state.lead = o.lead; state.lead = o.lead;
state.exposure = o.exposure;
const N = state.dense.length; const N = state.dense.length;
state.stab = stabilize(state.dense, o.smoothWin, state.aspect); state.stab = stabilize(state.dense, o.smoothWin, state.aspect);
@ -193,6 +225,7 @@ function rebuild(resetKeep) {
state.extractKey = extractKey(o); state.extractKey = extractKey(o);
} }
state.teeth = resolveTeeth(o); state.teeth = resolveTeeth(o);
state.eyes = buildEyes(o);
// Plate outline per frame, so a kept frame shows its own head shape. // Plate outline per frame, so a kept frame shows its own head shape.
state.plates = state.stab.oval.map((r) => r.map(state.xform)); state.plates = state.stab.oval.map((r) => r.map(state.xform));
@ -220,6 +253,124 @@ function faceBoxes() {
}); });
} }
// Eyes: lid rings traced per frame, blinks resolved per eye, one gaze shared.
//
// Lids are a FEATURE in the part table - rotoscoped, open vocabulary, a key on
// every frame - so they get exactly the mouth's treatment, including the same
// bounded contour average. The iris is a PRIMITIVE: a disc whose position is
// quantised, which is where the stylisation lives.
function buildEyes(o) {
const st = state.stab, N = state.dense.length;
const sig = eyeSignals(st);
state.eyeSig = sig;
const blink = { cut: o.blinkCut, dwell: o.blinkDwell, hold: o.blinkHold };
const shutR = resolveBlink(sig.openR, blink);
const shutL = resolveBlink(sig.openL, blink);
// Head-local, subsampled, contour-averaged - the identical chain the mouth
// takes, with the identical knob. The eye tracks the face, because the face
// is what it is attached to; what gets removed is per-frame detector jitter,
// not the motion.
const lidR = smoothContours(
st.lidR.map((r) => toRasterRing(r, EYE_R_RING, o.eyeVerts, state.xform)), o.contourSmooth);
const lidL = smoothContours(
st.lidL.map((r) => toRasterRing(r, EYE_L_RING, o.eyeVerts, state.xform)), o.contourSmooth);
// Where the iris hangs. Three behaviours, because this turns out to be an
// aesthetic choice and not only a correctness one.
//
// STEADY (default) reads the socket back off the DRAWN ring. Slots 0 and 8 of
// a 16-slot lid ring are the two corners, and subsampling to any even budget n
// keeps them at output indices 0 and n/2 - so the ring that gets rendered
// carries its own corners with it. The iris is then placed in the frame of the
// exact polygon it sits inside, after smoothing, after subsampling: it cannot
// drift relative to its own eye, and it inherits the contour average for free.
//
// FREE reads the raw per-frame corners instead, jitter and all. It is what the
// eyes did before any of this, and it is not simply worse - the detector noise
// reads as liveliness, the eye never sits perfectly still, and against flat
// hand-drawn plates that restlessness can be the thing that sells it. It is
// also the honest baseline to compare the other two against.
//
// LOCKED pins the socket to the take's mean, so the eye never moves in the
// head at all. Watch it against a photo underlay and the drawn eyes hang still
// over a face whose eyes are moving - that is the registration cost, and it is
// real - but once the plate is a drawing rather than a photograph, nothing is
// being registered against and it reads as a deliberately locked-off stare.
const ringSocket = (ring) => {
const a = ring[0], b = ring[ring.length / 2];
return { cx: (a.x + b.x) / 2, cy: (a.y + b.y) / 2, w: Math.hypot(a.x - b.x, a.y - b.y) };
};
const rawSocket = (corners, f) => {
const a = state.xform(corners[f][0]), b = state.xform(corners[f][1]);
return { cx: (a.x + b.x) / 2, cy: (a.y + b.y) / 2, w: Math.hypot(a.x - b.x, a.y - b.y) };
};
const meanSocket = (rings) => {
const acc = rings.reduce((a, r) => {
const k = ringSocket(r);
return { cx: a.cx + k.cx, cy: a.cy + k.cy, w: a.w + k.w };
}, { cx: 0, cy: 0, w: 0 });
const n = rings.length;
return { cx: acc.cx / n, cy: acc.cy / n, w: acc.w / n };
};
const socketFor = (rings, corners) => {
if (o.irisAnchor === 'locked') { const k = meanSocket(rings); return () => k; }
if (o.irisAnchor === 'free') return (f) => rawSocket(corners, f);
return (f) => ringSocket(rings[f]);
};
const skR = socketFor(lidR, st.cornersR), skL = socketFor(lidL, st.cornersL);
const socket = (ring) => ringSocket(ring);
// Iris radius comes from the take's MEAN eye width, not the current frame's.
// Size is authored; only position is tracked. A radius recomputed per frame
// would breathe by a fraction of a pixel as the fit's depth-scale wanders,
// and at this resolution a fraction of a pixel is a pixel flicking on and off
// around the whole silhouette.
const meanW = (rings) => rings.reduce((a, r) => a + ringSocket(r).w, 0) / rings.length;
const wR = meanW(lidR), wL = meanW(lidL), w = (wR + wL) / 2;
// Calibrate against the neutral, apply the artist's gain, and only then
// quantise - the grid should be a grid of DRAWN positions, because that is
// what a viewer reads. Gain is an authored parameter: measured gaze excursion
// is small and a character's eye usually wants more throw than a performer's,
// which is a decision for a person and not for the detector.
const origin = gazeOrigin(sig.gazeRaw, o.gazeOrigin, state.neutral);
state.gazeOriginValue = origin;
const px = sig.gazeRaw.map((g) => ({
x: (g.x - origin.x) * o.gazeGain * w,
y: (g.y - origin.y) * o.gazeGain * w,
}));
const gaze = quantizeGaze(px, o.gazeStep, o.gazeDwell);
const eye = (sk, lids, shut, rad, f) => {
const e = sk(f);
return {
// The lash line is the lid ring pushed outward by a fixed number of
// pixels, exactly as the mouth's outer ring sits outside its inner one.
// When the eye shuts, the traced ring goes near-degenerate and this
// collapses to a lens - which is a closed eye, drawn correctly, for free.
lash: offsetRing(lids[f], o.lashPx),
lid: lids[f],
shut: shut[f],
// Rounded to whole pixels. The rasteriser quantises everything anyway, so
// this costs nothing - but it means the iris and the square pupil share
// one integer centre, so the pupil is exactly its nominal size on every
// frame instead of spilling to the next pixel on some and not others.
iris: { x: Math.round(e.cx + gaze[f].x), y: Math.round(e.cy + gaze[f].y), r: rad },
pupil: o.pupilPx,
};
};
return {
gazePx: px, gaze, shutR, shutL, hasIris: sig.hasIris,
frames: Array.from({ length: N }, (_, f) => ({
r: eye(skR, lidR, shutR, (wR * o.irisSize) / 2, f),
l: eye(skL, lidL, shutL, (wL * o.irisSize) / 2, f),
})),
};
}
// Presence gets hysteresis and a minimum dwell, the same treatment plate // Presence gets hysteresis and a minimum dwell, the same treatment plate
// selection gets: a teeth block that blinks on and off for single frames is // selection gets: a teeth block that blinks on and off for single frames is
// worse than one that is simply absent. Appearing needs a clear signal, staying // worse than one that is simply absent. Appearing needs a clear signal, staying
@ -275,7 +426,7 @@ function resolveTeeth(o) {
// stand-in until a drawing exists. // stand-in until a drawing exists.
function renderFrame(f, mode = plateMode()) { function renderFrame(f, mode = plateMode()) {
const r = new IndexedRaster(RW, RH); const r = new IndexedRaster(RW, RH);
const pf = heldFrame(keptSorted(), f); // the plate frame on screen const pf = plateIndex(f); // the plate frame on screen
if (mode === 'posterize' && state.images[pf]) { if (mode === 'posterize' && state.images[pf]) {
posterizeInto(r, state.images[pf], state.stab.transforms[pf], state.xform, posterizeInto(r, state.images[pf], state.stab.transforms[pf], state.xform,
@ -285,6 +436,12 @@ function renderFrame(f, mode = plateMode()) {
if (mode === 'oval' || mode === 'oval+photo') r.fillPoly(state.plates[pf], IDX.base); if (mode === 'oval' || mode === 'oval+photo') r.fillPoly(state.plates[pf], IDX.base);
} }
// Eyes run on the CLOCK, not on the mouth lead. The lead is a lip-sync
// device: it exists because a mouth shape anticipates the sound it makes.
// Nothing about a blink or a glance is tied to the audio, so shifting the
// eyes would only slide them off the head that carries them.
if (state.eyes) drawEyes(r, state.eyes.frames[f]);
const mf = leadIndex(f); // performance frame, possibly ahead const mf = leadIndex(f); // performance frame, possibly ahead
r.fillPoly(state.outer[mf], IDX.dark); // mouth keeps every frame r.fillPoly(state.outer[mf], IDX.dark); // mouth keeps every frame
if (!state.hidden[mf]) { if (!state.hidden[mf]) {
@ -295,13 +452,34 @@ function renderFrame(f, mode = plateMode()) {
return r; return r;
} }
// Lash ring, then sclera, then iris - the same three-layer structure the mouth
// has, for the same reason: the dark ring outside the pale interior is what
// makes a flat shape read as an opening rather than a blob.
//
// The iris is stencilled to the sclera it was just drawn over, so the lid crops
// it automatically. Nothing needs to clamp the gaze to keep the iris inside the
// eye, which matters because a clamp would flatten the performance at exactly
// the extremes that carry it.
function drawEyes(r, e) {
for (const s of [e.r, e.l]) {
r.fillPoly(s.lash, IDX.dark);
if (s.shut) continue; // a shut eye IS the lash line, alone
r.fillPoly(s.lid, IDX.white);
r.fillDisc(s.iris.x, s.iris.y, s.iris.r, IDX.iris, IDX.white);
// Stencilled to the iris, which is itself stencilled to the sclera - so the
// pupil is cropped by the lid transitively, and a blink or an extreme gaze
// takes the right bite out of it without anything having to compute where.
if (s.pupil) r.fillRect(s.iris.x, s.iris.y, s.pupil, IDX.pupil, IDX.iris);
}
}
const plateMode = () => el('plateMode').value; const plateMode = () => el('plateMode').value;
// Photo modes composite under the indexed layer, so the flat shapes stay exactly // Photo modes composite under the indexed layer, so the flat shapes stay exactly
// as they render while the reference sits behind them. // as they render while the reference sits behind them.
function compositeRender(canvas, f, zoom) { function compositeRender(canvas, f, zoom) {
const mode = plateMode(); const mode = plateMode();
const pf = heldFrame(keptSorted(), f); const pf = plateIndex(f);
const img = state.images[pf]; const img = state.images[pf];
const showPhoto = img && (mode === 'photo' || mode === 'photo-dim' || mode === 'oval+photo'); const showPhoto = img && (mode === 'photo' || mode === 'photo-dim' || mode === 'oval+photo');
@ -347,11 +525,22 @@ const keptSorted = () => [...state.keep].sort((a, b) => a - b);
// Positive lead = the mouth arrives earlier. Only performance parts shift; the // Positive lead = the mouth arrives earlier. Only performance parts shift; the
// head stays with the audio, because it is the mouth that should anticipate. // head stays with the audio, because it is the mouth that should anticipate.
function leadIndex(f) { function leadIndex(f) {
// Reads a cached scalar, not opts(): this runs once per strip thumbnail, and // Reads cached scalars, not opts(): this runs once per strip thumbnail, and
// calling opts() here meant ~14 DOM reads x 74 frames on every redraw. // calling opts() here meant ~14 DOM reads x 74 frames on every redraw.
return shiftIndex(f, state.lead, state.dense.length); //
// Exposure first, then lead. The grid decides WHICH frames get a new drawing;
// the lead then shifts which pose that drawing carries, by whole frames of the
// original track. Applying them the other way round would put the changes on
// the wrong beats - the picture would update on the odd frames instead of
// holding on the twos.
return shiftIndex(exposeIndex(f, state.exposure), state.lead, state.dense.length);
} }
// The plate rides the same grid, so the whole picture updates together. On 2s
// means on 2s - a head that cut on the odd frames while the mouth cut on the
// even ones would read as two performances laid over each other.
const plateIndex = (f) => heldFrame(keptSorted(), exposeIndex(f, state.exposure));
function blit(canvas, raster, zoom) { function blit(canvas, raster, zoom) {
canvas.width = RW * zoom; canvas.height = RH * zoom; canvas.width = RW * zoom; canvas.height = RH * zoom;
canvas.getContext('2d').putImageData(raster.toImageData(PALETTE.map((p) => p.hex), zoom), 0, 0); canvas.getContext('2d').putImageData(raster.toImageData(PALETTE.map((p) => p.hex), zoom), 0, 0);
@ -372,21 +561,42 @@ function drawReadout() {
el('readout').textContent = el('readout').textContent =
`${state.dense.length} frames → ${kept.length} drawings · ` + `${state.dense.length} frames → ${kept.length} drawings · ` +
`teeth on ${teethFrames}f · ` + `teeth on ${teethFrames}f · ` +
`${blinkRuns(state.eyes.shutR).length}/${blinkRuns(state.eyes.shutL).length} blinks R/L · ` +
`${gazeCells(state.eyes.gaze)} gaze cells · ` +
(state.exposure > 1
? `on ${state.exposure}s = ${(state.fps / state.exposure).toFixed(4).replace(/\.?0+$/, '')}fps · `
: '') +
(lead ? `mouth leads ${lead}f (${(lead / state.fps * 1000).toFixed(0)}ms) · ` : '') + (lead ? `mouth leads ${lead}f (${(lead / state.fps * 1000).toFixed(0)}ms) · ` : '') +
`holds ${Math.min(...runs)}–${Math.max(...runs)} frames · ` + `holds ${Math.min(...runs)}–${Math.max(...runs)} frames · ` +
`neutral f${state.neutral} · residual ` + `neutral f${state.neutral} · residual ` +
`${(state.stab.residual.reduce((a, b) => a + b, 0) / state.dense.length).toFixed(4)}`; `${(state.stab.residual.reduce((a, b) => a + b, 0) / state.dense.length).toFixed(4)}`;
} }
// Blinks as RUNS, not as shut frames: a three-frame blink is one blink, and the
// count is only useful as "did the performer blink six times or sixty".
function blinkRuns(shut) {
const runs = [];
for (let f = 0; f < shut.length; f++) {
if (shut[f] && !shut[f - 1]) runs.push(f);
}
return runs;
}
// How many distinct positions the iris ever occupies. This is the number the
// gaze knobs exist to control: two or three is a character who looks at things,
// forty is an unquantised iris sliding around, which is what the grid is for.
const gazeCells = (gaze) => new Set(gaze.map((g) => `${g.x},${g.y}`)).size;
function drawPanes() { function drawPanes() {
const f = state.frame, kept = keptSorted(); const f = state.frame, kept = keptSorted();
const pf = heldFrame(kept, f); const pf = plateIndex(f);
// The mouth frame is always shown, not only when shifted, so the number can be // The mouth frame is always shown, not only when shifted, so the number can be
// watched diverging from f rather than taken on trust. // watched diverging from f rather than taken on trust.
const lead = state.lead; const lead = state.lead;
el('framelabel').textContent = el('framelabel').textContent =
`f ${f} / ${state.dense.length - 1} · ${(f / state.fps).toFixed(2)}s · ` + `f ${f} / ${state.dense.length - 1} · ${(f / state.fps).toFixed(2)}s · ` +
`plate f${pf} · mouth f${leadIndex(f)}` + `plate f${pf} · mouth f${leadIndex(f)}` +
(state.exposure > 1 && f % state.exposure ? ' (held)' : '') +
(lead ? ` (${lead > 0 ? '+' : ''}${lead} = ${(lead / state.fps * 1000).toFixed(0)}ms)` : '') + (lead ? ` (${lead > 0 ? '+' : ''}${lead} = ${(lead / state.fps * 1000).toFixed(0)}ms)` : '') +
(state.keep.has(f) ? ' · KEPT' : ' · held'); (state.keep.has(f) ? ' · KEPT' : ' · held');
@ -404,12 +614,14 @@ function drawPanes() {
const map = (p) => ({ x: dx + (p.x * im.naturalWidth - sx) * s, y: dy + (p.y * im.naturalHeight - sy) * s }); const map = (p) => ({ x: dx + (p.x * im.naturalWidth - sx) * s, y: dy + (p.y * im.naturalHeight - sy) * s });
strokePts(g1, LIPS_OUTER.map((i) => map(state.dense[f][i])), '#4ade80'); strokePts(g1, LIPS_OUTER.map((i) => map(state.dense[f][i])), '#4ade80');
strokePts(g1, LIPS_INNER.map((i) => map(state.dense[f][i])), '#f87171'); strokePts(g1, LIPS_INNER.map((i) => map(state.dense[f][i])), '#f87171');
drawEyeOverlay(g1, map, f);
} else { } else {
g1.fillStyle = '#555'; g1.font = '13px system-ui'; g1.fillStyle = '#555'; g1.font = '13px system-ui';
g1.fillText('synthetic — no source frames', 14, 24); g1.fillText('synthetic — no source frames', 14, 24);
const sc = (p) => ({ x: p.x * c1.width, y: p.y * c1.height }); const sc = (p) => ({ x: p.x * c1.width, y: p.y * c1.height });
strokePts(g1, LIPS_OUTER.map((i) => sc(state.dense[f][i])), '#4ade80'); strokePts(g1, LIPS_OUTER.map((i) => sc(state.dense[f][i])), '#4ade80');
strokePts(g1, LIPS_INNER.map((i) => sc(state.dense[f][i])), '#f87171'); strokePts(g1, LIPS_INNER.map((i) => sc(state.dense[f][i])), '#f87171');
drawEyeOverlay(g1, sc, f);
} }
const c2 = el('cv-stab'), g2 = c2.getContext('2d'); const c2 = el('cv-stab'), g2 = c2.getContext('2d');
@ -426,9 +638,18 @@ function drawPanes() {
if (mf !== f) strokePts(g2, z(state.outer[f]), '#2f6b46'); if (mf !== f) strokePts(g2, z(state.outer[f]), '#2f6b46');
strokePts(g2, z(state.outer[mf]), '#4ade80'); strokePts(g2, z(state.outer[mf]), '#4ade80');
if (!state.hidden[mf]) strokePts(g2, z(state.inner[mf]), '#f87171'); if (!state.hidden[mf]) strokePts(g2, z(state.inner[mf]), '#f87171');
for (const e of [state.eyes.frames[f].r, state.eyes.frames[f].l]) {
strokePts(g2, z(e.lid), e.shut ? '#f87171' : '#60a5fa');
if (e.shut) continue;
g2.strokeStyle = '#fbbf24';
g2.beginPath();
g2.arc(e.iris.x * ZOOM, e.iris.y * ZOOM, e.iris.r * ZOOM, 0, Math.PI * 2);
g2.stroke();
}
compositeRender(el('cv-render'), f, ZOOM); compositeRender(el('cv-render'), f, ZOOM);
drawInteriorDebug(f); drawInteriorDebug(f);
drawGazeDebug(f);
} }
// What the teeth measurement actually saw: sampled region, pixels above // What the teeth measurement actually saw: sampled region, pixels above
@ -458,6 +679,78 @@ function drawInteriorDebug(fRaw) {
`area ${m.area}px · ${te.show ? 'SHOWN' : 'hidden'}`; `area ${m.area}px · ${te.show ? 'SHOWN' : 'hidden'}`;
} }
// Lid rings and the iris, on the raw frame. Landmark overlays are how you tell
// a tracking failure from a knob set wrong, and the eyes need it more than the
// mouth does: an iris that has latched onto an eyebrow looks, in the flat
// render alone, exactly like a gaze gain that is too high.
function drawEyeOverlay(g, map, f) {
const lm = state.dense[f];
for (const ring of [EYE_R_RING, EYE_L_RING]) {
strokePts(g, ring.map((i) => map(lm[i])), '#60a5fa');
}
if (!state.eyes.hasIris) return;
for (const iris of [IRIS_A, IRIS_B]) {
strokePts(g, iris.slice(1).map((i) => map(lm[i])), '#fbbf24');
}
}
// The gaze field: every position the iris takes over the whole take, plus where
// it is now. Tune against this, not against the numbers - "4 cells" tells you
// the quantisation is working, but only the picture tells you whether the four
// are the four looks the performance actually has.
function drawGazeDebug(f) {
const cv = el('cv-gaze'), S = 150;
cv.width = S; cv.height = S;
const g = cv.getContext('2d');
g.fillStyle = '#0d0f16'; g.fillRect(0, 0, S, S);
const ex = state.eyes;
// Scale so the widest excursion in the take fills the box, with a floor so a
// nearly-still gaze does not get magnified into a light show.
let m = 2;
for (const p of ex.gazePx) m = Math.max(m, Math.abs(p.x), Math.abs(p.y));
const k = (S / 2 - 8) / m;
const X = (v) => S / 2 + v * k, Y = (v) => S / 2 + v * k;
const o = opts();
if (o.gazeStep > 0) {
g.strokeStyle = '#1b2030'; g.lineWidth = 1;
for (let i = -20; i <= 20; i++) {
const v = i * o.gazeStep;
if (Math.abs(v) > m) continue;
g.beginPath(); g.moveTo(X(v), 0); g.lineTo(X(v), S); g.stroke();
g.beginPath(); g.moveTo(0, Y(v)); g.lineTo(S, Y(v)); g.stroke();
}
}
g.strokeStyle = '#2a2f3e';
g.beginPath(); g.moveTo(S / 2, 0); g.lineTo(S / 2, S);
g.moveTo(0, S / 2); g.lineTo(S, S / 2); g.stroke();
g.fillStyle = '#2f6b46';
for (const p of ex.gaze) g.fillRect(X(p.x) - 1.5, Y(p.y) - 1.5, 3, 3);
const raw = ex.gazePx[f], q = ex.gaze[f];
g.fillStyle = '#8891a5';
g.fillRect(X(raw.x) - 1, Y(raw.y) - 1, 2, 2);
g.fillStyle = '#fbbf24';
g.beginPath(); g.arc(X(q.x), Y(q.y), 4, 0, Math.PI * 2); g.fill();
const sig = state.eyeSig, fr = state.eyes.frames[f];
const og = state.gazeOriginValue;
// Per-eye raw gaze is the diagnostic for a wrong-looking eyeline. If the two
// agree and both point the wrong way, the ORIGIN is wrong. If they disagree in
// a sustained way, it is out-of-plane head rotation biasing the projection,
// which no 2D measurement can undo.
const sgn = (v) => `${v >= 0 ? '+' : ''}${v.toFixed(3)}`;
el('eyeinfo').textContent =
`open R ${sig.openR[f].toFixed(3)} L ${sig.openL[f].toFixed(3)} / cut ${o.blinkCut.toFixed(3)}\n` +
`${fr.r.shut ? 'R SHUT ' : ''}${fr.l.shut ? 'L SHUT' : ''}${!fr.r.shut && !fr.l.shut ? 'both open' : ''}\n` +
`gaze ${q.x >= 0 ? '+' : ''}${q.x.toFixed(1)}, ${q.y >= 0 ? '+' : ''}${q.y.toFixed(1)} px\n` +
`raw R ${sgn(sig.gazeR[f].x)} L ${sgn(sig.gazeL[f].x)} (x, eye widths)\n` +
`origin ${o.gazeOrigin} ${sgn(og.x)}, ${sgn(og.y)}` +
(ex.hasIris ? '' : ' — no iris landmarks');
}
function strokePts(g, pts, color, lw = 1) { function strokePts(g, pts, color, lw = 1) {
g.strokeStyle = color; g.lineWidth = lw; g.strokeStyle = color; g.lineWidth = lw;
g.beginPath(); g.beginPath();
@ -538,12 +831,57 @@ function drawWorksheet() {
/* ---------- export ---------- */ /* ---------- export ---------- */
// Six parts, three per eye, mirroring the mouth's lash/interior/content stack.
// `clip` is what tells the renderer the iris is stencilled by the sclera rather
// than merely drawn after it - without it an extreme gaze would put the iris on
// the cheek.
function eyeParts(grid) {
const out = [];
// Eyes ride the exposure grid but NOT the mouth lead: the lead is a lip-sync
// device and nothing about a blink is tied to the audio.
const src = (f) => exposeIndex(f, state.exposure);
[['r', 20], ['l', 23]].forEach(([side, z]) => {
const at = (f) => state.eyes.frames[src(f)][side];
out.push(
{ name: `eye_${side}`, kind: 'poly', z, color: 'skin_dark', interp: 'hold',
keys: grid.map((f) => ({ f, src: src(f), pts: at(f).lash })) },
{ name: `eye_${side}_in`, kind: 'poly', z: z + 1, color: 'eye_white', interp: 'hold',
parent: `eye_${side}`,
keys: grid.map((f) =>
(at(f).shut ? { f, hidden: true } : { f, src: src(f), pts: at(f).lid })) },
{ name: `iris_${side}`, kind: 'disc', z: z + 2, color: 'iris', interp: 'hold',
parent: `eye_${side}_in`, clip: `eye_${side}_in`,
keys: grid.map((f) => {
const e = at(f);
return e.shut ? { f, hidden: true }
: { f, src: src(f), c: { x: e.iris.x, y: e.iris.y }, r: e.iris.r };
}) },
);
if (!state.eyes.frames[0].r.pupil) return;
out.push(
{ name: `pupil_${side}`, kind: 'rect', z: z + 3, color: 'pupil', interp: 'hold',
parent: `iris_${side}`, clip: `iris_${side}`,
keys: grid.map((f) => {
const e = at(f);
return e.shut ? { f, hidden: true }
: { f, src: src(f), c: { x: e.iris.x, y: e.iris.y }, size: e.pupil };
}) },
);
});
return out;
}
function exportTake() { function exportTake() {
const kept = keptSorted(); const kept = keptSorted();
const N = state.dense.length; const N = state.dense.length;
// Output frames that actually carry a key. Everything between them is a hold,
// which the take format already expresses, so on 2s emits half the keys rather
// than emitting each pose twice.
const grid = [];
for (let f = 0; f < N; f += state.exposure) grid.push(f);
const take = { const take = {
name: el('takename').value || 'line_01', name: el('takename').value || 'line_01',
frames: N, width: RW, height: RH, exposure: 1, fps: state.fps, frames: N, width: RW, height: RH, exposure: state.exposure, fps: state.fps,
palette: PALETTE, palette: PALETTE,
slot: { x: RW / 2, y: RH / 2 }, slot: { x: RW / 2, y: RH / 2 },
parts: [ parts: [
@ -554,14 +892,19 @@ function exportTake() {
// key f carries the pose from source frame f+lead - so the renderer never // key f carries the pose from source frame f+lead - so the renderer never
// needs to know about it. // needs to know about it.
{ name: 'mouth', kind: 'poly', z: 30, color: 'skin_dark', interp: 'hold', { name: 'mouth', kind: 'poly', z: 30, color: 'skin_dark', interp: 'hold',
keys: state.outer.map((_, f) => ({ f, src: leadIndex(f), pts: state.outer[leadIndex(f)] })) }, keys: grid.map((f) => ({ f, src: leadIndex(f), pts: state.outer[leadIndex(f)] })) },
{ name: 'mouth_in', kind: 'poly', z: 31, color: 'mouth_dark', interp: 'hold', parent: 'mouth', { name: 'mouth_in', kind: 'poly', z: 31, color: 'mouth_dark', interp: 'hold', parent: 'mouth',
keys: state.inner.map((_, f) => { keys: grid.map((f) => {
const m = leadIndex(f); const m = leadIndex(f);
return state.hidden[m] ? { f, hidden: true } : { f, src: m, pts: state.inner[m] }; return state.hidden[m] ? { f, hidden: true } : { f, src: m, pts: state.inner[m] };
}) }, }) },
// Eyes. The lids are traced, so like the mouth they cost nothing and keep
// every frame. The iris is a primitive: its quantised position means the
// key stream is dense but the VALUES change only on saccades, so a
// hold-interpolating renderer cuts between fixations by itself.
...eyeParts(grid),
{ name: 'teeth', kind: 'poly', z: 32, color: 'teeth', interp: 'hold', parent: 'mouth_in', { name: 'teeth', kind: 'poly', z: 32, color: 'teeth', interp: 'hold', parent: 'mouth_in',
keys: state.teeth.map((_, f) => { keys: grid.map((f) => {
const m = leadIndex(f), te = state.teeth[m]; const m = leadIndex(f), te = state.teeth[m];
return te.show && te.pts ? { f, src: m, pts: te.pts } : { f, hidden: true }; return te.show && te.pts ? { f, src: m, pts: te.pts } : { f, hidden: true };
}) }, }) },
@ -577,7 +920,9 @@ function exportTake() {
a.href = URL.createObjectURL(new Blob([text], { type: 'text/plain' })); a.href = URL.createObjectURL(new Blob([text], { type: 'text/plain' }));
a.download = `${take.name}.take`; a.download = `${take.name}.take`;
a.click(); a.click();
status(`exported — ${kept.length} plate drawings, ${N} mouth frames`, 'ok'); status(`exported — ${kept.length} plate drawings, ${grid.length} mouth keys` +
(state.exposure > 1 ? ` on ${state.exposure}s` : '') + ', ' +
`${blinkRuns(state.eyes.shutR).length + blinkRuns(state.eyes.shutL).length} blinks`, 'ok');
} }
/* ---------- wiring ---------- */ /* ---------- wiring ---------- */
@ -605,6 +950,7 @@ async function runFrames() {
state.interior = measureAll(images, dense, opts()); state.interior = measureAll(images, dense, opts());
el('scrub').max = dense.length - 1; el('scrub').max = dense.length - 1;
state.frame = 0; state.frame = 0;
labelExposure();
rebuild(true); rebuild(true);
const dur = (dense.length / state.fps).toFixed(2); const dur = (dense.length / state.fps).toFixed(2);
status(`${images.length} frames · ${images[0].naturalWidth}x${images[0].naturalHeight} · ` + status(`${images.length} frames · ${images[0].naturalWidth}x${images[0].naturalHeight} · ` +
@ -626,24 +972,44 @@ function runSynthetic() {
state.dense = synthDense(72); state.dense = synthDense(72);
el('scrub').max = 71; el('scrub').max = 71;
state.frame = 0; state.frame = 0;
labelExposure();
rebuild(true); rebuild(true);
status('synthetic — exercises everything below detection', 'ok'); status('synthetic — exercises everything below detection', 'ok');
} }
// How each slider's raw value reads out. A table rather than the conditional
// chain this used to be: that chain grew a branch per knob and was one ternary
// away from being unreadable.
const FMT = {
apertureThresh: (v) => (v / 1000).toFixed(3),
tol: (v) => (v / 1000).toFixed(3),
blinkCut: (v) => (v / 1000).toFixed(3),
teethOn: (v) => (v / 100).toFixed(2),
teethErode: (v) => (v / 100).toFixed(2),
tongueReject: (v) => (v / 100).toFixed(2),
topBias: (v) => (v / 100).toFixed(2),
irisSize: (v) => `${v}%`,
gazeGain: (v) => (v / 100).toFixed(2),
gazeStep: (v) => (v ? `${v}px` : 'off'),
pupilPx: (v) => (v ? `${v}px` : 'off'),
lashPx: (v) => `${v}px`,
lead: (v) => (v > 0 ? `+${v}` : String(v)),
};
for (const id of ['verts', 'smoothWin', 'contourSmooth', 'apertureThresh', 'tol', for (const id of ['verts', 'smoothWin', 'contourSmooth', 'apertureThresh', 'tol',
'teethOn', 'teethDwell', 'teethErode', 'tongueReject', 'blobGrow', 'teethOn', 'teethDwell', 'teethErode', 'tongueReject', 'blobGrow',
'topBias', 'teethVerts', 'teethSmooth', 'lead']) { 'topBias', 'teethVerts', 'teethSmooth', 'lead',
'eyeVerts', 'lashPx', 'irisSize', 'pupilPx', 'gazeGain',
'gazeStep', 'gazeDwell', 'blinkCut', 'blinkHold', 'blinkDwell']) {
const show = () => {
el(id + 'v').textContent = FMT[id] ? FMT[id](+el(id).value) : el(id).value;
};
el(id).addEventListener('input', () => { el(id).addEventListener('input', () => {
el(id + 'v').textContent = id === 'apertureThresh' || id === 'tol' show();
? (+el(id).value / 1000).toFixed(3)
: ['teethOn', 'teethErode', 'tongueReject', 'topBias'].includes(id)
? (+el(id).value / 100).toFixed(2)
: id === 'lead' && +el(id).value > 0 ? `+${el(id).value}`
: el(id).value;
if (id === 'tol') return; // tol only matters when you ask for a suggestion if (id === 'tol') return; // tol only matters when you ask for a suggestion
rebuild(false); rebuild(false);
}); });
el(id + 'v').textContent = el(id).value; show();
} }
function seekTo(f) { function seekTo(f) {
@ -657,6 +1023,22 @@ el('btn-frames').onclick = runFrames;
el('btn-synth').onclick = runSynthetic; el('btn-synth').onclick = runSynthetic;
el('btn-export').onclick = exportTake; el('btn-export').onclick = exportTake;
el('plateMode').addEventListener('change', () => { if (state.dense) drawAll(); }); el('plateMode').addEventListener('change', () => { if (state.dense) drawAll(); });
el('exposure').addEventListener('change', () => { if (state.dense) rebuild(false); });
for (const id of ['irisAnchor', 'gazeOrigin']) {
el(id).addEventListener('change', () => { if (state.dense) rebuild(false); });
}
// Label the exposure options in the only units that mean anything here: the
// rate the picture actually changes at, which depends on the clip's own rate.
// "on 2s" is the animator's name for it and the number is what you hear against
// the audio, so the menu says both.
function labelExposure() {
for (const opt of el('exposure').options) {
const n = +opt.value;
const rate = (state.fps / n).toFixed(4).replace(/\.?0+$/, '');
opt.textContent = `${rate} fps · on ${n}s`;
}
}
el('btn-saveframe').onclick = () => { el('btn-saveframe').onclick = () => {
if (!state.dense) return; if (!state.dense) return;
const cv = document.createElement('canvas'); const cv = document.createElement('canvas');
@ -775,3 +1157,5 @@ else if (location.hash === '#frames') runFrames();
else status('ready — Load frames, then step with \u2190 \u2192 and delete with X'); else status('ready — Load frames, then step with \u2190 \u2192 and delete with X');
window.__roto = state; // headless smoke test reads this window.__roto = state; // headless smoke test reads this
window.__render = compositeRender; // ...and renders arbitrary frames off-screen
window.__lead = leadIndex; // ...and resolves the performance frame

View file

@ -52,3 +52,46 @@ export function subsampleSlots(len, n) {
export function subsampleRing(ring, n) { export function subsampleRing(ring, n) {
return subsampleSlots(ring.length, n).map((s) => ring[s]); return subsampleSlots(ring.length, n).map((s) => ring[s]);
} }
// ---- eyes ----
//
// Eyelid rings, under the same contract as the lip rings: ORDERED traversals
// where slot position IS vertex identity. Both eyes start at the OUTER corner
// and go over the UPPER lid first, so slot k means the same anatomy on both
// sides. On a 16-slot ring that puts the four cardinals exactly on the four
// quarter slots - 0 outer corner, 4 upper lid centre, 8 inner corner, 12 lower
// lid centre - so every even vertex budget lands on real landmarks.
//
// The two rings traverse opposite directions on screen, because they are
// mirrored anatomy described the same way. Nothing downstream cares: an
// even-odd fill has no winding, and ring SIMPLICITY is what is asserted.
export const EYE_R_RING = [
33, 246, 161, 160, 159, 158, 157, 173,
133, 155, 154, 153, 145, 144, 163, 7,
];
export const EYE_L_RING = [
263, 466, 388, 387, 386, 385, 384, 398,
362, 382, 381, 380, 374, 373, 390, 249,
];
// Outer, inner corner per eye. All four are also in RIGID, and that is the
// point: the eye's reference frame is built only from landmarks that do not
// move under performance, so a blink cannot be mistaken for a change of gaze.
export const EYE_R_CORNERS = [33, 133];
export const EYE_L_CORNERS = [263, 362];
// Upper and lower lid centres. Their separation over the corner distance is the
// openness signal that decides whether the eye is shut - the same shape of
// measurement as APERTURE is for the mouth, but normalised, so one threshold
// carries across takes and faces.
export const EYE_R_LIDS = [159, 145];
export const EYE_L_LIDS = [386, 374];
// The two iris blocks the refined mesh appends: centre first, then four ring
// points. WHICH BLOCK BELONGS TO WHICH EYE IS NOT DECLARED HERE - MediaPipe's
// own "left"/"right" is viewer-relative in some docs and subject-relative in
// others, and a swap looks almost right, so it would survive an eyeball and
// then read as a permanently wall-eyed character. pipeline.js resolves it from
// the geometry instead.
export const IRIS_A = [468, 469, 470, 471, 472];
export const IRIS_B = [473, 474, 475, 476, 477];

View file

@ -108,3 +108,27 @@ export function smoothTransforms(tfs, radius) {
theta: Math.atan2(sn[i], c[i]), s: s[i], tx: tx[i], ty: ty[i], theta: Math.atan2(sn[i], c[i]), s: s[i], tx: tx[i], ty: ty[i],
})); }));
} }
// Push a ring outward from its centroid by a FIXED distance, not by a scale
// factor.
//
// Scaling collapses with the shape: a shut eyelid scaled by 1.1 is still a shut
// eyelid, so the lash line - the only thing left to draw when the eye is closed
// - would vanish exactly on the frames where it is the whole drawing. A fixed
// radial offset gives a band of roughly constant thickness that survives the
// ring going degenerate, and it keeps a star-shaped ring simple, which
// docs/design.md requires of every cut part.
export function offsetRing(pts, d) {
if (!d) return pts;
let cx = 0, cy = 0;
for (const p of pts) { cx += p.x; cy += p.y; }
cx /= pts.length; cy /= pts.length;
return pts.map((p) => {
const dx = p.x - cx, dy = p.y - cy;
const m = Math.hypot(dx, dy);
// A vertex sitting exactly on the centroid has no outward direction. Leave
// it where it is rather than emitting NaN and poisoning the whole ring.
return m < 1e-9 ? { x: p.x, y: p.y }
: { x: p.x + (dx / m) * d, y: p.y + (dy / m) * d };
});
}

View file

@ -2,7 +2,9 @@
// All policy lives here, never in the renderer. See docs/design.md, // All policy lives here, never in the renderer. See docs/design.md,
// "The take is the contract". // "The take is the contract".
import { RIGID, LIPS_OUTER, LIPS_INNER, APERTURE, FACE_OVAL, EYE_INNER, subsampleSlots } from './landmarks.js'; import { RIGID, LIPS_OUTER, LIPS_INNER, APERTURE, FACE_OVAL, EYE_INNER,
EYE_R_RING, EYE_L_RING, EYE_R_CORNERS, EYE_L_CORNERS,
EYE_R_LIDS, EYE_L_LIDS, IRIS_A, IRIS_B, subsampleSlots } from './landmarks.js';
import { fitSimilarity, applySimAll, applySim, fitResidual, procrustesMean, smoothTransforms, movingAverage } from './mathutil.js'; import { fitSimilarity, applySimAll, applySim, fitResidual, procrustesMean, smoothTransforms, movingAverage } from './mathutil.js';
// MediaPipe normalises x by image WIDTH and y by image HEIGHT, so its normalised // MediaPipe normalises x by image WIDTH and y by image HEIGHT, so its normalised
@ -25,6 +27,13 @@ export function stabilize(dense, smoothRadius, aspect = 1) {
const raw = rigid.map((r) => fitSimilarity(r, ref)); const raw = rigid.map((r) => fitSimilarity(r, ref));
const tfs = smoothTransforms(raw, smoothRadius); const tfs = smoothTransforms(raw, smoothRadius);
// The refined mesh appends ten iris points to the 468 face points, but a
// plain mesh does not, and synthetic or hand-fed tracks need not. Checked
// rather than assumed: reading past the end would surface as NaN gaze deep
// downstream instead of as "this track carries no iris".
const hasIris = dense.every((f) => f && f.length > IRIS_B[IRIS_B.length - 1]);
const map = (table) => dense.map((f, i) => applySimAll(tfs[i], pick(f, table, aspect)));
return { return {
ref, ref,
transforms: tfs, transforms: tfs,
@ -43,9 +52,221 @@ export function stabilize(dense, smoothRadius, aspect = 1) {
const a = applySimAll(tfs[i], pick(f, APERTURE, aspect)); const a = applySimAll(tfs[i], pick(f, APERTURE, aspect));
return Math.hypot(a[0].x - a[1].x, a[0].y - a[1].y); return Math.hypot(a[0].x - a[1].x, a[0].y - a[1].y);
}), }),
// Eyes. Lid rings are a feature and get traced like the mouth; corners and
// lid centres are the measurement frame; the iris blocks are raw until
// pairIrises decides which is which.
lidR: map(EYE_R_RING), lidL: map(EYE_L_RING),
cornersR: map(EYE_R_CORNERS), cornersL: map(EYE_L_CORNERS),
lidsR: map(EYE_R_LIDS), lidsL: map(EYE_L_LIDS),
irisA: hasIris ? map(IRIS_A) : null,
irisB: hasIris ? map(IRIS_B) : null,
}; };
} }
/* ---------- eyes ---------- */
const mid = (a, b) => ({ x: (a.x + b.x) / 2, y: (a.y + b.y) / 2 });
const dist = (a, b) => Math.hypot(a.x - b.x, a.y - b.y);
// Which iris block belongs to which eye is RESOLVED FROM THE DATA, not declared
// in a table.
//
// The naming in MediaPipe's own material is viewer-relative in some places and
// subject-relative in others, and the two blocks are otherwise
// indistinguishable. Getting it backwards swaps the irises, which looks almost
// right - each eye still has a disc in roughly the right place - so it survives
// a casual eyeball and then reads as a subtly wall-eyed character for the rest
// of the project. Proximity to the eye's corner midpoint settles it in one
// comparison, is impossible to get wrong, and keeps working if the model is
// ever renumbered.
//
// Voted across every frame rather than read off frame zero: one bad detection
// should not decide the whole shot.
export function pairIrises(stab) {
if (!stab.irisA) return null;
let votes = 0;
for (let f = 0; f < stab.irisA.length; f++) {
const cR = mid(stab.cornersR[f][0], stab.cornersR[f][1]);
votes += dist(stab.irisA[f][0], cR) < dist(stab.irisB[f][0], cR) ? 1 : -1;
}
return votes > 0 ? { right: 'irisA', left: 'irisB' }
: { right: 'irisB', left: 'irisA' };
}
// Per-frame eye measurements, in units of eye width. Measurement only - every
// threshold and every stylisation is applied by the callers.
//
// Everything here stays in HEAD-LOCAL space, which is the same space the mouth
// lives in and the same space the registered photo underlay is drawn in. An
// earlier version pinned each eye into a fixed socket fitted to its corners'
// mean over the shot. That does remove the wobble, but it removes too much: the
// residual from out-of-plane rotation is real motion of the eye relative to the
// head, it is still there in the footage, and pinning it away leaves the drawn
// eyes hanging still over a photo whose eyes are moving. The eye has to track
// the face exactly as the mouth does.
//
// The wobble the socket was aimed at is dealt with the way docs/design.md deals
// with it everywhere else - the bounded contour average, the same knob and the
// same radius the mouth uses - and by placing the iris in the frame of the
// ALREADY-SMOOTHED lid ring, so the iris cannot jitter independently of the eye
// it sits in. See buildEyes in app.js.
export function eyeSignals(stab) {
const N = stab.transforms.length;
const pairing = pairIrises(stab);
const openR = [], openL = [], gazeRaw = [], gazeR = [], gazeL = [];
for (let f = 0; f < N; f++) {
const cR = mid(stab.cornersR[f][0], stab.cornersR[f][1]);
const cL = mid(stab.cornersL[f][0], stab.cornersL[f][1]);
const wR = dist(stab.cornersR[f][0], stab.cornersR[f][1]);
const wL = dist(stab.cornersL[f][0], stab.cornersL[f][1]);
// Openness is the lid gap over the CORNER distance. Normalising by the
// corners rather than by anything derived from the lids keeps the
// denominator rigid, so the ratio measures the lid and nothing else, and
// one threshold carries across takes, faces and framings.
openR.push(dist(stab.lidsR[f][0], stab.lidsR[f][1]) / wR);
openL.push(dist(stab.lidsL[f][0], stab.lidsL[f][1]) / wL);
if (!pairing) {
gazeRaw.push({ x: 0, y: 0 }); gazeR.push({ x: 0, y: 0 }); gazeL.push({ x: 0, y: 0 });
continue;
}
const iR = stab[pairing.right][f][0], iL = stab[pairing.left][f][0];
// Gaze is the iris centre relative to the CORNER MIDPOINT, in eye widths -
// a pure offset WITHIN the eye, with the eye's own position divided out, so
// that quantising it quantises the glance and not the head motion carrying
// it.
//
// Measuring against the lid ring's centroid instead would track the lid:
// every blink pulls that centroid down and would fake a glance at the
// floor, on precisely the frames where the eye is most conspicuous. The
// corners are in RIGID, so this origin and this denominator are both immune
// to the performance they are measuring.
const gR = { x: (iR.x - cR.x) / wR, y: (iR.y - cR.y) / wR };
const gL = { x: (iL.x - cL.x) / wL, y: (iL.y - cL.y) / wL };
// ONE gaze for both eyes, and deliberately so. At 320x200 an iris is a
// handful of pixels and its centre comes from five landmarks on an eye
// twenty pixels wide, so the difference between the two measurements is
// noise, not vergence - and independent per-eye noise reads as wall-eyed
// immediately, which is the most expensive artefact on a face. Openness
// stays per-eye, because a wink is real performance and should survive.
gazeRaw.push({ x: (gR.x + gL.x) / 2, y: (gR.y + gL.y) / 2 });
// Kept separately purely as a diagnostic. The two eyes should agree; when
// they disagree in a sustained way rather than frame to frame, that is not
// noise but out-of-plane head rotation biasing the projected iris offset,
// and no 2D measurement can undo it.
gazeR.push(gR); gazeL.push(gL);
}
return { openR, openL, gazeRaw, gazeR, gazeL, hasIris: !!pairing };
}
// Where "not looking anywhere in particular" sits on THIS face. Everything the
// character does is measured as a departure from it, so getting it wrong does
// not bias the gaze slightly - it re-points the whole performance.
//
// `median` is the default and the safe one: the middle of the take, per axis.
// docs/design.md already gives this rule for the anchor fit - the reference is
// the MEAN configuration over the shot, not one frame - and gaze needs it for
// the same reason. The median rather than the mean because a couple of frames
// of hard glance should not drag the rest-point after them.
//
// `neutral` reads the origin off the take's neutral frame instead, which is
// only correct when there genuinely is a held neutral to read. That frame is
// chosen by MINIMUM MOUTH APERTURE, and a closed mouth says nothing whatever
// about where the eyes are pointed - so on footage with no deliberate neutral
// at the top it is an arbitrary frame, and whichever way the performer happened
// to glance on it becomes "straight ahead" for the entire shot. It is kept
// because it is right when the take was shot for this tool, and because being
// able to switch is how you find out that it was not.
export function gazeOrigin(gazeRaw, mode = 'median', neutral = 0, radius = 2) {
if (mode === 'neutral') {
let sx = 0, sy = 0, n = 0;
// A window, not a single frame: one frame of a five-landmark iris centre is
// worth about a pixel of noise, and that pixel would become a permanent
// squint in the output.
for (let f = neutral - radius; f <= neutral + radius; f++) {
const k = Math.min(gazeRaw.length - 1, Math.max(0, f));
sx += gazeRaw[k].x; sy += gazeRaw[k].y; n++;
}
return { x: sx / n, y: sy / n };
}
const mid1 = (vals) => {
const v = vals.slice().sort((a, b) => a - b);
return v.length % 2 ? v[(v.length - 1) / 2]
: (v[v.length / 2 - 1] + v[v.length / 2]) / 2;
};
return { x: mid1(gazeRaw.map((g) => g.x)), y: mid1(gazeRaw.map((g) => g.y)) };
}
// Snap gaze onto a grid, then require a new cell to hold before it takes.
//
// This is the "Primitive - quantised" row of the part table in docs/design.md,
// and it is not a stylisation imposed on the truth: real eyes move in saccades,
// holding a fixation and then jumping. The smooth drift left in the measurement
// is tracker noise plus head-compensation error, so snapping to a grid and
// requiring a dwell removes the noise and recovers the saccade in the same
// operation - the rare case where the aesthetic rule and the physiology agree.
//
// The dwell is what stops a gaze parked on a cell boundary from chattering
// between two cells forever. It is meaningless without a grid, because
// continuous values never repeat, so step 0 short-circuits both.
export function quantizeGaze(gaze, step, dwell) {
if (!(step > 0)) return gaze.map((g) => ({ x: g.x, y: g.y }));
const q = gaze.map((g) => ({
x: Math.round(g.x / step) * step,
y: Math.round(g.y / step) * step,
}));
if (dwell <= 0 || !q.length) return q;
const out = [];
let live = q[0], pend = q[0], run = 0;
for (const g of q) {
if (g.x === pend.x && g.y === pend.y) run++;
else { pend = g; run = 1; }
if (run > dwell && (pend.x !== live.x || pend.y !== live.y)) live = pend;
out.push(live);
}
return out;
}
// Resolve openness into a shut/open decision per frame.
//
// `dwell` is the same guard the teeth get: a lid hovering at the threshold must
// commit before the state changes, so it cannot flicker.
//
// `hold` is the one that is NOT like the teeth, and it is the whole reason
// blinks are worth special-casing. A blink is 100-150ms, which at 12fps is one
// frame and at 24fps is two or three - and a single frame of closed eye reads
// as a dropped frame, not as a blink. Animators draw a blink over two or three
// drawings for exactly that reason. So once the eye shuts it stays shut for
// `hold` frames, which turns an unreadable flicker into a beat.
//
// The hysteresis runs the other way from the teeth: shutting needs a clear
// signal, and once shut the eye is given the benefit of the doubt on reopening,
// because the lid landmarks are least reliable mid-blink.
export function resolveBlink(open, { cut, dwell, hold }) {
const N = open.length;
const shut = new Array(N).fill(false);
let live = false; // current state
let run = 0; // frames the opposing reading has persisted
let held = 0; // frames spent in the current state
for (let f = 0; f < N; f++) {
const reading = live ? open[f] < cut * 1.35 : open[f] < cut;
if (reading === live) run = 0;
else {
run++;
// Leaving a blink additionally requires the blink to have been on screen
// long enough to be legible; entering one never waits.
if (run > dwell && (!live || held >= hold)) { live = reading; held = 0; run = 0; }
}
held++;
shut[f] = live;
}
return shut;
}
// Stage 4: fixed-index subsample of a stabilised ring, then map from normalised // Stage 4: fixed-index subsample of a stabilised ring, then map from normalised
// face space into character raster space. // face space into character raster space.
export function toRasterRing(stabRing, ringTable, n, xform) { export function toRasterRing(stabRing, ringTable, n, xform) {
@ -178,6 +399,24 @@ export function heldFrame(kept, f) {
return hit; return hit;
} }
// Hold every output frame back onto an exposure grid: 1 = on 1s, 2 = on 2s, and
// so on. Frame 5 at exposure 2 reads the pose from frame 4.
//
// This is where "aesthetic sparseness" belongs. docs/design.md used to put it at
// the extraction rate - pick 12fps and the timing is already chosen - but that
// makes the timing a property of a directory of PNGs, so auditioning 12 against
// 24 means re-ripping the clip and re-running detection over all of it. Rip
// dense once and quantise here instead: the dense track stays at the camera's
// rate, the decision stays reversible, and the audio clock is untouched, so
// sync cannot drift while you try timings.
//
// Floor, never round. Rounding would let an output frame read a pose from the
// FUTURE, which is a lead - a separate control, applied after this one, for a
// separate reason.
export function exposeIndex(f, exposure) {
return exposure > 1 ? Math.floor(f / exposure) * exposure : f;
}
// Shift a performance track against the clock, clamped at the ends. // Shift a performance track against the clock, clamped at the ends.
// //
// Pure and exported so the shift can actually be asserted: "the slider feels // Pure and exported so the shift can actually be asserted: "the slider feels

View file

@ -46,14 +46,47 @@ export class IndexedRaster {
} }
} }
fillDisc(cx, cy, r, index) { // `over` is an optional stencil: when given, only pixels that currently hold
// that index are written. The indexed buffer is its own clip mask, which is
// how Animator Pro would do it - and it is what keeps the iris inside the
// eye. A disc clipped by the sclera cannot spill past the lid at any gaze or
// any radius, including mid-blink when the opening is a two-pixel sliver, so
// the lid crops the iris for free instead of the gaze range needing a
// clamp that would flatten the performance at the extremes.
fillDisc(cx, cy, r, index, over = null) {
const rr = r * r; const rr = r * r;
const y0 = Math.max(0, Math.floor(cy - r)), y1 = Math.min(this.h - 1, Math.ceil(cy + r)); const y0 = Math.max(0, Math.floor(cy - r)), y1 = Math.min(this.h - 1, Math.ceil(cy + r));
const x0 = Math.max(0, Math.floor(cx - r)), x1 = Math.min(this.w - 1, Math.ceil(cx + r)); const x0 = Math.max(0, Math.floor(cx - r)), x1 = Math.min(this.w - 1, Math.ceil(cx + r));
for (let y = y0; y <= y1; y++) { for (let y = y0; y <= y1; y++) {
for (let x = x0; x <= x1; x++) { for (let x = x0; x <= x1; x++) {
const dx = x + 0.5 - cx, dy = y + 0.5 - cy; const dx = x + 0.5 - cx, dy = y + 0.5 - cy;
if (dx * dx + dy * dy <= rr) this.buf[y * this.w + x] = index; if (dx * dx + dy * dy > rr) continue;
const o = y * this.w + x;
if (over === null || this.buf[o] === over) this.buf[o] = index;
}
}
}
// An exactly size x size block of pixels, snapped to the pixel grid, with the
// same optional stencil as fillDisc.
//
// The pupil is a SQUARE because at 320x200 it is three pixels across, and a
// circle of radius 1.5 is not a circle - it is a plus sign with the corners
// gnawed off, and it changes shape as it moves. A square that size is a
// deliberate mark that stays the same mark wherever it lands, which is the
// whole argument for flat shapes at this resolution.
//
// The top-left is rounded rather than the centre, so the block is size x size
// on every frame. Round the extents instead and a fractional centre gives you
// three pixels on one frame and four on the next, which reads as the pupil
// breathing.
fillRect(cx, cy, size, index, over = null) {
if (size < 1) return;
const x0 = Math.round(cx - size / 2), y0 = Math.round(cy - size / 2);
for (let y = Math.max(0, y0); y < Math.min(this.h, y0 + size); y++) {
for (let x = Math.max(0, x0); x < Math.min(this.w, x0 + size); x++) {
const o = y * this.w + x;
if (over === null || this.buf[o] === over) this.buf[o] = index;
} }
} }
} }

View file

@ -7,9 +7,12 @@
// as blocks meeting at corners. It is invisible at some vertex counts and obvious // as blocks meeting at corners. It is invisible at some vertex counts and obvious
// at others, so it needs an assertion rather than an eyeball. // at others, so it needs an assertion rather than an eyeball.
import { LIPS_OUTER, LIPS_INNER, FACE_OVAL, RIGID, subsampleSlots, subsampleRing } from './landmarks.js'; import { LIPS_OUTER, LIPS_INNER, FACE_OVAL, RIGID, subsampleSlots, subsampleRing,
import { fitSimilarity, applySim, procrustesMean, smoothTransforms } from './mathutil.js'; EYE_R_RING, EYE_L_RING, EYE_R_CORNERS, EYE_L_CORNERS,
import { stabilize, toRasterRing, selectKeys, activeKey, shiftIndex } from './pipeline.js'; EYE_R_LIDS, EYE_L_LIDS } from './landmarks.js';
import { fitSimilarity, applySim, procrustesMean, smoothTransforms, offsetRing } from './mathutil.js';
import { stabilize, toRasterRing, smoothContours, selectKeys, activeKey, shiftIndex,
exposeIndex, eyeSignals, pairIrises, gazeOrigin, quantizeGaze, resolveBlink } from './pipeline.js';
import { IndexedRaster, hexToRgb } from './raster.js'; import { IndexedRaster, hexToRgb } from './raster.js';
import { writeTake } from './take.js'; import { writeTake } from './take.js';
import { otsuForTest, scaleRing } from './interior.js'; import { otsuForTest, scaleRing } from './interior.js';
@ -206,6 +209,25 @@ export function run() {
ok('activeKey holds between keys', ok('activeKey holds between keys',
activeKey(sel.keys, sel.keys[1].f - 1).f === sel.keys[0].f); activeKey(sel.keys, sel.keys[1].f - 1).f === sel.keys[0].f);
// exposure: rip dense, choose the timing here. On 2s every odd frame must
// reuse the even frame's pose, and the grid must never read from the future -
// that direction is the lead, which is a different control for a reason.
ok('exposure 1 is identity', [0, 1, 7, 71].every((f) => exposeIndex(f, 1) === f));
ok('on 2s holds each pose for two frames',
[0, 1, 2, 3, 4, 5].map((f) => exposeIndex(f, 2)).join(',') === '0,0,2,2,4,4');
ok('on 3s holds each pose for three frames',
[0, 1, 2, 3, 4, 5, 6].map((f) => exposeIndex(f, 3)).join(',') === '0,0,0,3,3,3,6');
ok('exposure never reads a pose from the future',
[0, 1, 2, 3, 4, 5, 6, 7].every((f) => exposeIndex(f, 3) <= f));
{
// Exposure then lead, in that order: the picture must change on the grid
// beats and carry a pose shifted by whole frames of the original track.
const N = 72, at = (f) => shiftIndex(exposeIndex(f, 2), 1, N);
ok('exposure and lead compose without moving the beats',
at(0) === 1 && at(1) === 1 && at(2) === 3 && at(3) === 3,
[0, 1, 2, 3].map(at).join(','));
}
// mouth lead: a shift that "feels like it does nothing" is indistinguishable // mouth lead: a shift that "feels like it does nothing" is indistinguishable
// from one that does nothing, so assert the arithmetic directly. // from one that does nothing, so assert the arithmetic directly.
ok('lead 0 is identity', [0, 5, 71].every((f) => shiftIndex(f, 0, 72) === f)); ok('lead 0 is identity', [0, 5, 71].every((f) => shiftIndex(f, 0, 72) === f));
@ -288,6 +310,241 @@ export function run() {
(tt.mBright - tt.mDark) / 255 > 0.4, `sep ${((tt.mBright - tt.mDark) / 255).toFixed(4)}`); (tt.mBright - tt.mDark) / 255 > 0.4, `sep ${((tt.mBright - tt.mDark) / 255).toFixed(4)}`);
} }
/* ---- eyes ---- */
ok('eye rings have 16 distinct ids each',
new Set(EYE_R_RING).size === 16 && new Set(EYE_L_RING).size === 16);
ok('the two eye rings share no landmark',
!EYE_R_RING.some((i) => EYE_L_RING.includes(i)));
// The cardinal contract, asserted rather than trusted: on a 16-slot ring the
// quarter slots must be the four anatomical cardinals, which is what makes
// every even vertex budget land on real landmarks instead of between them.
ok('eye ring slot 0/4/8/12 are outer, upper, inner, lower',
EYE_R_RING[0] === EYE_R_CORNERS[0] && EYE_R_RING[8] === EYE_R_CORNERS[1] &&
EYE_R_RING[4] === EYE_R_LIDS[0] && EYE_R_RING[12] === EYE_R_LIDS[1] &&
EYE_L_RING[0] === EYE_L_CORNERS[0] && EYE_L_RING[8] === EYE_L_CORNERS[1] &&
EYE_L_RING[4] === EYE_L_LIDS[0] && EYE_L_RING[12] === EYE_L_LIDS[1]);
// The gaze origin and denominator are built from the eye corners, so if a
// corner were not rigid a blink could move it and fake a glance.
ok('every eye corner is a rigid landmark',
[...EYE_R_CORNERS, ...EYE_L_CORNERS].every((i) => RIGID.includes(i)));
// Same simplicity requirement as the lips, and for the same reason: a cut
// part with a self-intersecting ring renders as blocks meeting at corners.
// Checked on blink frames too, where the ring is nearly degenerate.
for (const [label, table] of [['right', EYE_R_RING], ['left', EYE_L_RING]]) {
let worst = null;
for (let n = 4; n <= 12 && !worst; n += 2) {
const slots = subsampleSlots(table.length, n);
for (let f = 0; f < dense.length; f++) {
const hits = ringSelfIntersections(slots.map((sl) => dense[f][table[sl]]));
if (hits.length) { worst = `verts=${n} frame=${f} edges ${JSON.stringify(hits[0])}`; break; }
}
}
ok(`${label} eye ring is simple at every vertex budget`, !worst, worst || '');
}
// offsetRing must grow by a FIXED amount and survive a degenerate ring - the
// shut eyelid is exactly the degenerate case, and it is the frame where the
// lash line is the entire drawing.
{
const sq = [{ x: -1, y: 0 }, { x: 0, y: -1 }, { x: 1, y: 0 }, { x: 0, y: 1 }];
const g = offsetRing(sq, 2);
ok('offsetRing pushes every vertex out by exactly d',
g.every((p, i) => Math.abs(Math.hypot(p.x, p.y) - (Math.hypot(sq[i].x, sq[i].y) + 2)) < 1e-9));
ok('offsetRing(0) is identity', offsetRing(sq, 0) === sq);
// A shut lid: a flat sliver. The offset must still open it into a band.
const shutLid = [{ x: -10, y: 0 }, { x: 0, y: -0.02 }, { x: 10, y: 0 }, { x: 0, y: 0.02 }];
const band = offsetRing(shutLid, 1.5);
const h = Math.max(...band.map((p) => p.y)) - Math.min(...band.map((p) => p.y));
ok('offsetRing gives a shut lid a visible lash band', h > 2.9, `height ${h.toFixed(3)}`);
ok('offsetRing keeps the shut lid simple', ringSelfIntersections(band).length === 0);
}
{
const stE = stabilize(dense, 2);
const sig = eyeSignals(stE);
ok('synthetic track carries iris landmarks', sig.hasIris);
// THE load-bearing eye assertion. The pairing is resolved from geometry
// rather than declared, so the test feeds a track built the OTHER way round
// and demands the resolver follow the data. A resolver only ever checked
// against the convention it was written for is checking nothing.
const pairA = pairIrises(stE);
const pairB = pairIrises(stabilize(synthDense(72, { swapIris: true }), 2));
ok('iris pairing is resolved from the data, not assumed',
pairA.right === 'irisA' && pairB.right === 'irisB',
`normal ${pairA.right}, swapped ${pairB.right}`);
// The eye must TRACK the face, not sit in a fixed socket. An earlier
// version pinned each eye to its corners' mean over the shot, which does
// kill the wobble but leaves the drawn eyes hanging still over a registered
// photo whose eyes are moving. Head-local is the same space the mouth and
// the underlay live in, so the eye moves with the head exactly as they do.
{
const spread = (arr, sel) => {
const v = arr.map(sel);
return Math.max(...v) - Math.min(...v);
};
const w = Math.hypot(stE.cornersR[0][0].x - stE.cornersR[0][1].x,
stE.cornersR[0][0].y - stE.cornersR[0][1].y);
const moves = Math.max(spread(stE.lidR, (r) => r[0].x), spread(stE.lidR, (r) => r[0].y));
ok('the eye stays in head-local space and tracks the face', moves / w > 0.02,
`corner travels ${(moves / w * 100).toFixed(1)}% of an eye width`);
// Subsampling a 16-slot ring to any even budget must keep the two corners
// at output indices 0 and n/2. That is what lets the socket be read back
// off the drawn polygon instead of measured separately, which is what
// stops the iris drifting relative to the eye it sits in.
let bad = null;
for (let n = 4; n <= 12; n += 2) {
const sl = subsampleSlots(16, n);
if (sl[0] !== 0 || sl[n / 2] !== 8) bad = `n=${n} -> ${sl.join(',')}`;
}
ok('the drawn lid ring carries its own corners at 0 and n/2', !bad, bad || '');
// The contour average is what removes the jitter, and it is the mouth's
// knob doing the mouth's job - no second mechanism for the eyes.
const ring = (rad) => smoothContours(
stE.lidR.map((r) => toRasterRing(r, EYE_R_RING, 8, (p) => ({ x: p.x * 600, y: p.y * 600 }))), rad);
const jitter = (rings) => {
let acc = 0;
for (let f = 1; f < rings.length; f++) {
const a = rings[f], b = rings[f - 1];
acc += Math.hypot((a[0].x + a[4].x) / 2 - (b[0].x + b[4].x) / 2,
(a[0].y + a[4].y) / 2 - (b[0].y + b[4].y) / 2);
}
return acc / (rings.length - 1);
};
ok('contour averaging steadies the eye without pinning it',
jitter(ring(1)) < jitter(ring(0)) * 0.8,
`${jitter(ring(0)).toFixed(3)} -> ${jitter(ring(1)).toFixed(3)} px/frame`);
}
// Blink: synth shuts the lids for exactly one frame every 19.
const lo = Math.min(...sig.openR), hi = Math.max(...sig.openR);
ok('openness collapses on a blink and not otherwise', lo < hi * 0.2,
`${lo.toFixed(3)} .. ${hi.toFixed(3)}`);
const shut = resolveBlink(sig.openR, { cut: hi * 0.3, dwell: 0, hold: 3 });
const runs = [];
for (let f = 0; f < shut.length; f++) if (shut[f] && !shut[f - 1]) runs.push(f);
const lens = runs.map((a) => { let n = 0; while (shut[a + n]) n++; return n; });
ok('blinks are found', runs.length >= 3, `${runs.length} runs at ${runs.join(',')}`);
// The knob that is not like the teeth: a one-frame blink reads as a dropped
// frame, so `hold` must stretch it into something legible.
ok('a one-frame blink is held to the minimum length',
lens.every((n) => n >= 3), `run lengths ${lens.join(',')}`);
ok('a shorter hold leaves the blink shorter',
resolveBlink(sig.openR, { cut: hi * 0.3, dwell: 0, hold: 1 }).filter(Boolean).length <
shut.filter(Boolean).length);
// Gaze, against ground truth: synth commands +0.16 eye widths at f12 and
// -0.16 at f23, holding each for eleven frames.
const org = gazeOrigin(sig.gazeRaw, 'neutral', 0);
const gx = (f) => (sig.gazeRaw[f].x - org.x);
ok('gaze recovers the commanded direction',
gx(12) > 0.12 && gx(12) < 0.20 && gx(23) < -0.12 && gx(23) > -0.20,
`f12 ${gx(12).toFixed(3)}, f23 ${gx(23).toFixed(3)}`);
// Measuring gaze against the lid centroid instead of the corner midpoint
// would drag the iris down on every blink and fake a glance at the floor,
// on exactly the frames where the eye is most conspicuous.
const gy = (f) => (sig.gazeRaw[f].y - org.y);
ok('a blink does not fake a change of gaze',
Math.abs(gy(19) - gy(18)) < 0.02, `f18 ${gy(18).toFixed(4)} -> f19 ${gy(19).toFixed(4)}`);
// Quantisation is what turns drift into saccades: four commanded
// fixations must come back as a handful of cells, not one per frame.
// The origin re-points the whole performance, so a wrong one does not bias
// the gaze slightly - it makes the character look the other way. The median
// must sit inside the range it summarises; the neutral-frame origin need
// not, which is exactly the failure mode it has on footage with no
// deliberate neutral at the top.
{
const med = gazeOrigin(sig.gazeRaw, 'median');
const xs = sig.gazeRaw.map((g) => g.x);
ok('the median origin lies inside the take\'s own gaze range',
med.x > Math.min(...xs) && med.x < Math.max(...xs),
`${med.x.toFixed(3)} in ${Math.min(...xs).toFixed(3)}..${Math.max(...xs).toFixed(3)}`);
// Synth looks left as much as right, so the rest point is near zero.
ok('the median origin finds the rest point, not a glance',
Math.abs(med.x) < 0.08, `median x ${med.x.toFixed(3)}`);
ok('the two origins actually differ, so the toggle is a real A/B',
Math.abs(med.x - gazeOrigin(sig.gazeRaw, 'neutral', 12).x) > 0.02);
}
const px = sig.gazeRaw.map((g) => ({ x: (g.x - org.x) * 30, y: (g.y - org.y) * 30 }));
const cells = (a) => new Set(a.map((g) => `${g.x},${g.y}`)).size;
ok('quantisation collapses drift into a few fixations',
cells(quantizeGaze(px, 2, 2)) <= 6 && cells(px) > 40,
`${cells(px)} raw -> ${cells(quantizeGaze(px, 2, 2))} cells`);
ok('gaze step 0 leaves the track untouched',
quantizeGaze(px, 0, 2).every((g, i) => g.x === px[i].x && g.y === px[i].y));
ok('quantised values land on the grid',
quantizeGaze(px, 2, 0).every((g) => Math.abs(g.x % 2) < 1e-9 && Math.abs(g.y % 2) < 1e-9));
// A one-frame excursion is noise; the dwell must swallow it.
{
const spike = [{ x: 0, y: 0 }, { x: 0, y: 0 }, { x: 4, y: 0 }, { x: 0, y: 0 }, { x: 0, y: 0 }];
ok('the dwell suppresses a one-frame gaze spike',
quantizeGaze(spike, 2, 1).every((g) => g.x === 0));
ok('a sustained move still gets through',
quantizeGaze([...spike, { x: 4, y: 0 }, { x: 4, y: 0 }, { x: 4, y: 0 }], 2, 1).pop().x === 4);
}
}
// The iris is stencilled by the sclera and the pupil by the iris, which is
// what keeps both inside the lid at any gaze without clamping the gaze itself.
{
const rr = new IndexedRaster(40, 40);
rr.clear(0);
rr.fillPoly([{ x: 10, y: 10 }, { x: 30, y: 10 }, { x: 30, y: 20 }, { x: 10, y: 20 }], 1);
rr.fillDisc(28, 15, 9, 2, 1); // a disc reaching well past the "lid"
let spill = 0, inside = 0;
for (let y = 0; y < 40; y++) for (let x = 0; x < 40; x++) {
const v = rr.buf[y * 40 + x];
if (v !== 2) continue;
if (x >= 10 && x < 30 && y >= 10 && y < 20) inside++; else spill++;
}
ok('a stencilled disc cannot spill past its clip', spill === 0 && inside > 20,
`${inside} in, ${spill} out`);
rr.fillDisc(5, 35, 3, 3); // no stencil: writes freely
ok('an unstencilled disc still writes anywhere', rr.buf.includes(3));
// The stencil chain: pupil over iris over sclera. A pupil placed where the
// iris has already been cropped must be cropped the same way.
rr.fillRect(28, 15, 5, 4, 2);
let pSpill = 0;
for (let y = 0; y < 40; y++) for (let x = 0; x < 40; x++) {
if (rr.buf[y * 40 + x] === 4 && !(x >= 10 && x < 30 && y >= 10 && y < 20)) pSpill++;
}
ok('the pupil inherits the iris clip transitively', pSpill === 0);
}
// A square pupil is only worth having if it is the SAME square every frame:
// exactly its nominal size at any centre, or it breathes as the gaze moves.
{
const sizes = [];
for (const [cx, cy] of [[20, 20], [20.5, 20.5], [20.49, 19.51], [21, 20]]) {
const rr = new IndexedRaster(40, 40);
rr.clear(0);
rr.fillRect(cx, cy, 3, 1);
let n = 0, minX = 99, maxX = -1, minY = 99, maxY = -1;
for (let y = 0; y < 40; y++) for (let x = 0; x < 40; x++) {
if (rr.buf[y * 40 + x] !== 1) continue;
n++; minX = Math.min(minX, x); maxX = Math.max(maxX, x);
minY = Math.min(minY, y); maxY = Math.max(maxY, y);
}
sizes.push(`${maxX - minX + 1}x${maxY - minY + 1}:${n}`);
}
ok('a 3px pupil is 3x3 at every centre', sizes.every((v) => v === '3x3:9'), sizes.join(' '));
const rr = new IndexedRaster(40, 40);
rr.clear(0); rr.fillRect(20, 20, 0, 1);
ok('pupil size 0 draws nothing', !rr.buf.includes(1));
}
// take writer round-trip // take writer round-trip
const take = { const take = {
name: 'test', frames: 72, width: 320, height: 200, exposure: 2, name: 'test', frames: 72, width: 320, height: 200, exposure: 2,
@ -297,6 +554,12 @@ export function run() {
{ name: 'head', kind: 'plate', z: 0, interp: 'hold', keys: [{ f: 0, plate: 0 }] }, { name: 'head', kind: 'plate', z: 0, interp: 'hold', keys: [{ f: 0, plate: 0 }] },
{ name: 'mouth', kind: 'poly', z: 30, color: 'skin', interp: 'hold', { name: 'mouth', kind: 'poly', z: 30, color: 'skin', interp: 'hold',
keys: sel.keys.map((k) => ({ f: k.f, src: k.src, pts: shapes[k.src] })) }, keys: sel.keys.map((k) => ({ f: k.f, src: k.src, pts: shapes[k.src] })) },
{ name: 'iris_r', kind: 'disc', z: 22, color: 'iris', interp: 'hold',
parent: 'eye_r_in', clip: 'eye_r_in',
keys: [{ f: 0, src: 0, c: { x: 120.4, y: 88.7 }, r: 5.5 }, { f: 2, hidden: true }] },
{ name: 'pupil_r', kind: 'rect', z: 23, color: 'pupil', interp: 'hold',
parent: 'iris_r', clip: 'iris_r',
keys: [{ f: 0, src: 0, c: { x: 120, y: 89 }, size: 3 }] },
], ],
}; };
const text = writeTake(take); const text = writeTake(take);
@ -309,6 +572,13 @@ export function run() {
}), `${keyLines.length} key lines`); }), `${keyLines.length} key lines`);
ok('take declares a plate and a part table', ok('take declares a plate and a part table',
/^plate\s+0/m.test(text) && /^part\s+mouth/m.test(text)); /^plate\s+0/m.test(text) && /^part\s+mouth/m.test(text));
ok('a disc part declares its clip', /^part\s+iris_r.*clip=eye_r_in/m.test(text));
ok('a disc key is three integers', /^key\s+iris_r\s+f=0\s+src=0\s+disc=120,89,6$/m.test(text),
(text.split('\n').find((l) => l.startsWith('key iris_r')) || '').trim());
ok('a hidden disc key emits hidden', /^key\s+iris_r\s+f=2\s+hidden$/m.test(text));
ok('a pupil key is a square, not a tessellated polygon',
/^key\s+pupil_r\s+f=0\s+src=0\s+rect=120,89,3$/m.test(text) &&
/^part\s+pupil_r.*clip=iris_r/m.test(text));
ok('coordinates are integers', !/-?\d+\.\d/.test(text.split('\n').filter((l) => l.startsWith('key')).join(''))); ok('coordinates are integers', !/-?\d+\.\d/.test(text.split('\n').filter((l) => l.startsWith('key')).join('')));
return results; return results;

View file

@ -5,11 +5,16 @@
// verified without a video file. A synthetic face is also the only way to test // verified without a video file. A synthetic face is also the only way to test
// stabilisation against a KNOWN head motion, since real footage gives no ground // stabilisation against a KNOWN head motion, since real footage gives no ground
// truth to compare against. // truth to compare against.
import { LIPS_OUTER, LIPS_INNER, FACE_OVAL, RIGID, EYE_INNER } from './landmarks.js'; import { LIPS_OUTER, LIPS_INNER, FACE_OVAL, RIGID,
EYE_R_RING, EYE_L_RING, IRIS_A, IRIS_B } from './landmarks.js';
const NUM = 478; const NUM = 478;
export function synthDense(nFrames = 72) { // `swapIris` places the two iris blocks on the opposite eyes. It exists so the
// pairing resolver can be tested against a track it actually disagrees with:
// a resolver checked only against the convention it was written for is checking
// nothing at all.
export function synthDense(nFrames = 72, { swapIris = false } = {}) {
const frames = []; const frames = [];
for (let t = 0; t < nFrames; t++) { for (let t = 0; t < nFrames; t++) {
const pts = new Array(NUM); const pts = new Array(NUM);
@ -35,11 +40,49 @@ export function synthDense(nFrames = 72) {
const openAmt = target[beat]; const openAmt = target[beat];
const wide = 0.10 + (beat === 1 ? 0.012 : beat === 3 ? -0.008 : 0); const wide = 0.10 + (beat === 1 ? 0.012 : beat === 3 ? -0.008 : 0);
place(RIGID[0], -0.075, -0.045); place(RIGID[1], -0.028, -0.043);
place(RIGID[2], 0.028, -0.043); place(RIGID[3], 0.075, -0.045);
place(RIGID[4], 0.000, -0.050); place(RIGID[5], 0.000, -0.020); place(RIGID[4], 0.000, -0.050); place(RIGID[5], 0.000, -0.020);
place(RIGID[6], 0.000, 0.012); place(RIGID[6], 0.000, 0.012);
place(EYE_INNER[0], -0.028, -0.043); place(EYE_INNER[1], 0.028, -0.043);
// Eyes. The corners (RIGID[0..3]) are placed BY the lid rings rather than
// separately, because they are slots 0 and 8 of those rings: writing them
// twice is how the mouth grew a bowtie, and a corner that disagrees with
// its own ring would make the eye self-intersect at some vertex budgets
// and not others.
//
// A blink is ONE frame, which is the honest hard case: at 12fps that is
// what a real blink costs, and it is exactly the length that reads as a
// dropped frame rather than as a blink unless `hold` extends it.
const blink = t > 5 && t % 19 === 0;
const openness = blink ? 0.05 : 1;
// Gaze holds and then jumps, the way gaze actually behaves, with a little
// jitter on top so quantisation has noise to remove and the dwell has
// something to suppress.
const LOOK = [[0, 0], [0.16, 0.0], [-0.16, 0.05], [0.0, -0.09]];
const [gx, gy] = LOOK[Math.floor(t / 11) % LOOK.length];
const jit = () => (Math.random() - 0.5) * 0.012;
// Half the corner separation, and the lid half-height at full open.
const EYE_RX = 0.0235, EYE_RY = 0.011, EYE_Y = -0.044;
const eye = (ring, cx, dir, iris) => {
const n = ring.length;
for (let k = 0; k < n; k++) {
// dir flips the traversal so each ring runs the direction its real
// table does: slot 0 outer corner, 4 upper lid, 8 inner, 12 lower.
const a = dir > 0 ? Math.PI + (k / n) * Math.PI * 2 : -(k / n) * Math.PI * 2;
place(ring[k], cx + EYE_RX * Math.cos(a),
EYE_Y + EYE_RY * openness * Math.sin(a));
}
// Iris: centre first, then four ring points, as the refined mesh emits.
const ix = cx + (gx + jit()) * EYE_RX * 2, iy = EYE_Y + (gy + jit()) * EYE_RX * 2;
place(iris[0], ix, iy);
for (let k = 1; k < iris.length; k++) {
const a = ((k - 1) / (iris.length - 1)) * Math.PI * 2;
place(iris[k], ix + 0.008 * Math.cos(a), iy + 0.008 * Math.sin(a));
}
};
eye(EYE_R_RING, -0.0515, 1, swapIris ? IRIS_B : IRIS_A);
eye(EYE_L_RING, 0.0515, -1, swapIris ? IRIS_A : IRIS_B);
// Lip rings as ellipse arcs, traversed so ring ORDER matches the tables: // Lip rings as ellipse arcs, traversed so ring ORDER matches the tables:
// slot 0 = right corner, 5 = top centre, 10 = left corner, 15 = bottom // slot 0 = right corner, 5 = top centre, 10 = left corner, 15 = bottom

View file

@ -15,7 +15,12 @@ export function writeTake(take) {
// v1 emits a single frozen plate derived from the face oval. A real project // v1 emits a single frozen plate derived from the face oval. A real project
// replaces this with hand-drawn angles referenced by cel frame; the record // replaces this with hand-drawn angles referenced by cel frame; the record
// shape is the same either way. // shape is the same either way.
L.push(`plate 0 kind=poly slot_mouth=${r(take.slot.x)},${r(take.slot.y)} scale=1.00 rot=0 squash=1.00`); L.push(`plate 0 kind=poly slot_mouth=${r(take.slot.x)},${r(take.slot.y)}` +
(take.eyeSlots
? ` slot_eye_r=${r(take.eyeSlots.r.cx)},${r(take.eyeSlots.r.cy)},${r(take.eyeSlots.r.w)}` +
` slot_eye_l=${r(take.eyeSlots.l.cx)},${r(take.eyeSlots.l.cy)},${r(take.eyeSlots.l.w)}`
: '') +
` scale=1.00 rot=0 squash=1.00`);
L.push(''); L.push('');
for (const part of take.parts) { for (const part of take.parts) {
const bits = [`part ${part.name.padEnd(9)} kind=${part.kind} z=${part.z}`]; const bits = [`part ${part.name.padEnd(9)} kind=${part.kind} z=${part.z}`];
@ -23,6 +28,10 @@ export function writeTake(take) {
if (part.kind === 'poly') bits.push('closed=1 fill=1'); if (part.kind === 'poly') bits.push('closed=1 fill=1');
bits.push(`interp=${part.interp}`); bits.push(`interp=${part.interp}`);
if (part.parent) bits.push(`parent=${part.parent}`); if (part.parent) bits.push(`parent=${part.parent}`);
// `clip` names a part this one is stencilled by, not merely drawn after.
// The iris needs it: at an extreme gaze the disc reaches past the lid, and
// ordering alone would put it on the cheek.
if (part.clip) bits.push(`clip=${part.clip}`);
L.push(bits.join(' ')); L.push(bits.join(' '));
} }
L.push(''); L.push('');
@ -30,6 +39,20 @@ export function writeTake(take) {
for (const k of part.keys) { for (const k of part.keys) {
if (k.hidden) { L.push(`key ${part.name.padEnd(9)} f=${k.f} hidden`); continue; } if (k.hidden) { L.push(`key ${part.name.padEnd(9)} f=${k.f} hidden`); continue; }
if (part.kind === 'plate') { L.push(`key ${part.name.padEnd(9)} f=${k.f} plate=${k.plate}`); continue; } if (part.kind === 'plate') { L.push(`key ${part.name.padEnd(9)} f=${k.f} plate=${k.plate}`); continue; }
// A disc is three numbers, so it gets its own key shape rather than being
// pre-tessellated into a polygon here: the renderer draws a real circle
// with hard edges, and a five-pixel iris approximated by a polygon would
// lose a pixel off its silhouette on some frames and not others.
// A square, in whole pixels, at an integer centre. Same argument as the
// disc: the renderer is told the shape, not a polygon approximating it.
if (part.kind === 'rect') {
L.push(`key ${part.name.padEnd(9)} f=${k.f} src=${k.src} rect=${r(k.c.x)},${r(k.c.y)},${k.size}`);
continue;
}
if (part.kind === 'disc') {
L.push(`key ${part.name.padEnd(9)} f=${k.f} src=${k.src} disc=${r(k.c.x)},${r(k.c.y)},${r(k.r)}`);
continue;
}
const pts = k.pts.map((p) => `${r(p.x)},${r(p.y)}`).join(' '); const pts = k.pts.map((p) => `${r(p.x)},${r(p.y)}`).join(' ');
// f is authoritative (what renders); src is the pre-snap extreme frame, // f is authoritative (what renders); src is the pre-snap extreme frame,
// kept only as a tuning signal - a key dragged far means the minimum-hold // kept only as a tuning signal - a key dragged far means the minimum-hold